Amazon FSx for NetApp ONTAP's NVMe/TCP procedure assumes RHEL 9.3: why multipathing is unavailable on Amazon Linux 2023, and how to check before you are billed
Followed the AWS procedure and run cat /sys/module/nvme_core/parameters/multipath only to find
the file isn't there?
Moving your target Linux distribution to Amazon Linux 2023 can stop a documented command from
working exactly as written. You are reading the procedure correctly, yet your environment alone
gives a different result — and tracking the cause down to a kernel build setting takes time.
This article is for anyone who got stuck at that point.
What this article does and does not cover: it covers how to confirm multipathing is unusable on
AL2023, and how to count whether it is actually using the second path. It does not cover
performance comparisons or a generalized conclusion for distributions other than AL2023 (details
under "What this article does not cover" below). All figures are from a Single-AZ configuration;
Multi-AZ is unverified.
dev.to series: FSx for ONTAP Block Protocols, Measured.
What was checked
Assuming no prior knowledge of FSx for ONTAP, here's what this article is confirming.
- Location: a single AWS VPC, single AZ
- Client: one EC2 instance (Amazon Linux 2023)
-
Target: one Amazon FSx for NetApp ONTAP file system (second-generation SINGLE_AZ_2). The HA
pair's two controllers hold two paths to one namespace
- Path 1: to the
optimizedcontroller (shortest path) - Path 2: to the
non-optimizedcontroller. Only when multipathing is available can the host tell that path 2 is a standby
- Path 1: to the
Figure: an FSx for ONTAP HA pair has two paths to one namespace, and **Asymmetric Namespace
Access is the mechanism that tells the host which one is the shortest path* (referred to as
multipathing throughout this article). Without this mechanism in the kernel, those two paths
don't appear as "one shortest path plus one standby": they appear as two separate devices
pointing at the same data.*
| Item | Role in this article |
|---|---|
| VPC / AZ | Single AZ, so both paths to the HA pair's controllers sit on the same network |
| EC2 client | Checks for multipathing and the kernel configuration behind it |
| optimized path | The shortest path as determined by multipathing. Real data is expected to flow here |
| non-optimized path | The failover path. Without multipathing it appears as an ordinary path too |
AWS's
Provisioning NVMe/TCP for Linux
includes a step that confirms multipath has come up: the fourth item under "To discover the target
NVMe nodes". It says that a returned Y means success (wording as of 2026-09-17; every
reference to that page below was checked on the same day).
cat /sys/module/nvme_core/parameters/multipath
# Y means success
On the Amazon Linux 2023 AMIs I used, that file did not exist. The kernel does not have
Asymmetric Namespace Access (a mechanism that tells the host which of several paths to the same
namespace is the shortest one) in it. The range I checked is the three kernel series returned by
al2023-ami-kernel-default-x86_64, so I cannot write "AL2023 always behaves this way". Checking
takes one command, and a later section gives it.
The procedure is written for RHEL 9.3. The same page's "Before you begin" says
Create an EC2 instance running Red Hat Enterprise Linux (RHEL) 9.3. Immediately after that there
is a statement about other distributions. In summary: on AMIs other than RHEL 9.3 some of the
utilities may already be present or the install command may differ, but apart from installing
packages, the commands in this section are valid on other EC2 Linux AMIs too (the original wording
is at the end of "Before you begin" on that page).
What this article reports is two counterexamples to that statement. Neither of them is
explained by package installation.
| Command from the documentation | Result on AL2023 | A packaging problem? |
|---|---|---|
cat /etc/nvme/hostnqn |
File does not exist | No. nvme-cli is installed; whether it creates this file differs by distribution |
cat /sys/module/nvme_core/parameters/multipath |
File does not exist | No. This is a kernel build option (CONFIG_NVME_MULTIPATH) and a package cannot change it |
You can check this before any billing starts. That is the most practical part of this article,
I think.
This is not a claim that "the documentation is wrong". It is a report that these two do not
fall inside the stated range (apart from installing packages, valid on other AMIs too).
The state of my feedback to AWS is at the end of this article.
Then, after bringing in a kernel that does have multipathing, I found that the optimized /
non-optimized output of nvme list-subsys says "there are two paths", not "two paths are in
use". Whether they are in use has to be counted separately, and there is a trap in choosing
which table to count.
What this article does not cover. It does not compare performance. The figures here exist to
show that two clients agreed under the same conditions, and they cannot be used for sizing. What
moves the numbers is in a separate article. All figures are from a
Single-AZ configuration; Multi-AZ is unverified.
This is one of three, but there is no reading order. The block-protocol measurements are split
by reader.
Article Reader A Session and queue counts Anyone about to mount B Checking whether multipathing is available (this article) Anyone who followed the procedure and got stuck C What moves the numbers Anyone measuring or citing figures None of them assumes the other two. Read only the part you need.
The environment I measured
| Item | Value |
|---|---|
| File system | Amazon FSx for NetApp ONTAP, second generation SINGLE_AZ_2, 6,144 MBps × 1 HA pair, 4,096 GiB SSD, 200,000 provisioned IOPS |
| Region | ap-northeast-1, single AZ |
| ONTAP |
9.18.1 (the reads and writes in the agreement table). For the 12 sequential-write points I failed to record the version. FileSystemTypeVersion returned null and I tore the environment down before re-reading it from the cluster API. The kernel investigation also spans deployments on other versions. The cross-article version mapping is in the ONTAP version matrix
|
| NVMe read cache | Disabled (confirmed on both nodes) |
| LUN / namespace | 600 GiB, on a 900 GiB volume |
| Before measuring | Wrote the full 600 GiB once. Unwritten blocks on a thin-provisioned volume return zeros, so reading without writing first measures nothing |
| Client | c5n.9xlarge; 50 Gbps network is a guaranteed figure |
| Instrument | VDBENCH 5.04.07 directly, iorate=max (unlimited), 512 threads, 60 s warm-up + 300 s measurement |
| Test data / efficiency settings |
Not recorded whether the VDBENCH payload was compressible, or whether volume inline efficiency (StorageEfficiencyEnabled) was enabled. On the file-protocol side, efficiency settings measurably affected figures; whether the same applies to block is unconfirmed |
The window is 300 seconds, so the figures in this article include burst. They cannot be cited as a
baseline.
The figures in this article are
iorate=maxsaturation points, not operating points. No target
IOPS was given; they are what the system took when everything available was pushed at it. On the
file-protocol side of the same repository, giving a target of 4,400 IOPS produced 4,406 MB/s at
8.86 ms, while unlimited produced 4,204 MB/s at 121.79 ms. Unlimited was 5% lower in
throughput and 14 times higher in response time.With 512 threads fixed and no rate limit, response time becomes the identity
512 ÷ IOPS
(derived 216.6 ms against measured 216.588 ms in a separate run). The figures below can be read
through that identity: 1,277.48 MB/s is about 401 ms, and 131,736.5 IOPS is about
3.9 ms. Both are derived values and were not recorded as measurements.For this article's purpose, being a saturation point is not a disadvantage. The figures are
here to show that Rocky and RHEL agreed under the same conditions, and as long as both are
saturated the same way, the agreement test holds. Those values still cannot be used for
sizing.
Behaviour on this AMI
On this AMI the file named at the top does not exist. Reading the kernel configuration gives
# CONFIG_NVME_MULTIPATH is not set, and modinfo nvme_core has no multipath parameter either.
So multipathing is not in the kernel.
As a result the two controllers are not merged into one device; they appear as separate devices
pointing at the same namespace UUID (/dev/nvme2n1 and /dev/nvme3n1). nvme list-subsys shows
no optimized / non-optimized either.
The AWS procedure is written for RHEL 9.3, where it is present.
Please do not generalise this to "multipathing is not available on AL2023". What I observed were the
AMIs that theal2023-ami-kernel-default-x86_64SSM parameter returned on 2026-09-12 and
09-13, and I did not record the kernel versions. The template said "record the kernel version
at measurement time" in place of pinning the AMI, and the block phase did not implement that.
That is a gap in my records.Checking in your own environment is more reliable, and it is one command.
grep -i NVME_MULTIPATH /boot/config-$(uname -r)
I later widened the observation to three kernel series. All three returned by
al2023-ami-kernel-default-x86_64 (6.1.186 / 6.12.103 / 6.18.48) gave
# CONFIG_NVME_MULTIPATH is not set. You can check this before creating a file system, by
extracting the config from the kernel package in the repository. It is known before billing
starts, so checking in that order is cheaper.
To measure multipathing I then brought in another distribution. The AWS procedure is written for RHEL 9.3,
so I put both Rocky Linux 9.7, a rebuild, and RHEL 9.7 itself on the same namespace and compared
them.
| Workload | RHEL 9.7 | Rocky 9.7 |
|---|---|---|
| 1 MiB sequential read | 1,277.48 MB/s | 1,276.60–1,279.22 MB/s |
| 4 KiB random read | 131,736.5 IOPS | 131,550.6 / 131,749.4 IOPS |
| 4 KiB random write | 195,439.8 IOPS | 194,960.6 IOPS |
Only the reads agreed in the same state (within 0.1%). Same minor version, same nvme-cli
(2.16-1.el9), same cold state. The kernels differ in z-stream (5.14.0-611.55.1 against 611.5.1).
For reads, there is now one piece of grounds for reading a rebuild's figure as the figure for the
product itself.
Both write workloads had the Rocky side measured warm, so the states are not aligned. Counting
the 0.3% difference on 4 KiB random write as "agreement" was my error, and the note in that same
table contradicted my own conclusion (corrected 2026-09-20). For that shape I can say neither
that it holds nor that it does not.Sequential write was re-measured later, in the same state, for that shape alone. It does not
agree (2026-09-17). On a different file system (second generation, 1,536 MBps, 50,000 provisioned
IOPS), same namespace, same state, same parameters, six runs each.
Client Min Median Max Spread Rocky 9.7 842.70 1,040.26 1,092.50 29.6% RHEL 9.7 1,105.75 1,112.65 1,129.30 2.1% Units MB/s (1 MiB sequential write, 8 threads,
o_direct, 120-second window). These 12
points were taken under conditions that differ from the environment described at the top: a
different file system (second generation, 1,536 MBps, 50,000 provisioned IOPS), 8 threads rather
than 512, and 120 seconds rather than 300. The512 ÷ IOPSreading given at the top cannot be
applied to these 12 points. RHEL's minimum is above Rocky's maximum: six against six, with no
overlap in range. RHEL's median is 7.0% above Rocky's, and the asymmetry in spread is larger than
the difference in medians.So "a rebuild's figure can be read as the figure for the product itself" is limited to the
reads. When citing a sequential-write figure, say which of the two it was measured on.
4 KiB random write was not re-measured in the same state, so for that shape I can say neither
that it holds nor that it does not. I cannot state a cause. The only difference on record is
the kernel z-stream, and I did not vary it on its own. During the runs the file system side was at
72–81% disk throughput, 16–19% IOPS and 31–38% network, none of them saturated (so this is not a
case of the two meeting at a ceiling). For these 12 points I failed to record the ONTAP version.
FileSystemTypeVersionreturnednulland I tore the environment down before re-reading it from
the cluster API.
iopolicy on a kernel with multipathing, and the value the procedure states
This is the third counterexample. iopolicy (Linux's multipath setting that decides how I/O
is distributed across multiple paths to an NVMe device: round-robin cycles through them,
queue-depth sends to whichever has the shortest queue) is what step 5 of the same procedure
says to confirm as round-robin (and the example nvme list-subsys output at step 7 also shows
iopolicy=round-robin).
| Value | |
|---|---|
| Value the documentation says to confirm | round-robin |
| Measured on Rocky Linux 9.7 / RHEL 9.7 | queue-depth |
| What set it | the udev rule 71-nvmf-netapp.rules shipped by nvme-cli (2.16-1.el9) |
The kernel default is numa, neither round-robin nor queue-depth. It was queue-depth in
this environment because the NetApp-oriented udev rule bundled with the nvme-cli package set it.
Following the confirmation step returns a value other than the expected one.
I measured the effect on performance. It is negligible: switching policy mid-run produced no
step change, and the difference was 0.5%. The figures and the method are in
a separate article (C). So this is not a story about round-robin being
required for speed. The only part worth reporting is that the confirmation step returns a value
that differs from the documentation.
This can change with the
nvme-cliversion. What I observed here is one version, 2.16-1.el9,
and I did not confirm whether other versions ship the same udev rule.
How to check whether multipathing is using the second path
This section was added afterwards. It became necessary in order to do the analysis above, and it
is also the reason I once published a wrong conclusion.
nvme list-subsys shows the two controllers as optimized / non-optimized, and prints iopolicy
as well. That is a display of "there are two paths", not of "two paths are in use".
At first I varied iopolicy three ways, compared total throughput, saw that only queue-depth was
faster, and wrote that "only queue-depth uses the second path". That was wrong. I had not
counted bytes per path. A client-side total that does not increase is consistent with two different
facts: "the second path is not being used", and "it is being used but the ceiling is nearer than
the path".
Bytes per path can be read from two places
On the ONTAP side they can be read from a counter table. There is a trap here.
| Table | Result in this configuration |
|---|---|
nvmf_lif |
Zero rows. Still empty after writing 600 GiB over NVMe/TCP |
lif |
Reports 0 bytes for the LIF in question. It does not count NVMe-oF |
nvmf_tcp_port |
Has the real data. read_data / write_data / total_ops per LIF |
"The table exists but has no rows" and "this version does not have that table" add up to the same
appearance. Return 0 and move on, and you reach the correct conclusion, "the second path saw
0 bytes", for the wrong reason.
It can also be counted on the client side. NVMe/TCP controllers have different destination
addresses, so aggregating established sockets by destination separates the paths. That is a second
source, independent of the ONTAP side.
ss -tin state established | awk '/:4420/{...}' # sum bytes_sent / bytes_received per destination
Counted properly, the second path was not running
I ran four workloads under queue-depth and took before/after differences.
| Path | Read bytes | Write bytes | Total ops |
|---|---|---|---|
| optimized | +about 1,012 GiB | +about 730 GiB | 641,908 → 120,008,866 |
| non-optimized | ±0 | ±0 | 4 → 4 |
Neither reads nor writes went to the second path. The ONTAP side and the client side agreed, and
it was the same on Rocky and on RHEL (the second path's increment across all workloads was about
20 KB).
Further, switching iopolicy within a single run produces no step change. With a sequential read
running I changed queue-depth → numa → queue-depth, and the values sampled at one-second
intervals stayed inside the same 1,150–1,400 MB/s range both before and after each switch.
Summary
-
All three kernel series returned by
al2023-ami-kernel-default-x86_64(6.1.186 / 6.12.103 / 6.18.48) hadCONFIG_NVME_MULTIPATHunset. Measuring multipathing requires a different distribution. Asking AWS why AL2023 is excluded from the procedure (reply 2026-09-21) surfaced a likely reason:nvme connect-allestablishes both paths and auto-discovers the alternate LIF even when only one is given, but withoutCONFIG_NVME_MULTIPATHthe kernel never merges them into one namespace device. The procedure's later steps assume that merge already happened, so proceeding unmerged leaves the other path writable against the same namespace with no multipath software mediating it: a data-corruption risk -
You can check before creating a file system. Extract the config from the kernel package in the
repository and you know before billing starts. On a running instance it is the one command
grep -i NVME_MULTIPATH /boot/config-$(uname -r) - Rocky Linux 9.7's figures can be read as RHEL 9.7's figures, but only for the reads. On the same namespace in the same state they agreed within 0.1%. Both write workloads had the Rocky side measured warm, so the states are not aligned. Counting the 0.3% difference on 4 KiB random write as agreement was my error, and for that shape I can say neither that it holds nor that it does not (corrected 2026-09-20). Sequential write does not agree: re-measured later in the same state, six runs each, RHEL's median was 7.0% higher and RHEL's minimum was above Rocky's maximum (spread 29.6% for Rocky against 2.1% for RHEL). When citing, say which of the two it was measured on
-
optimized/non-optimizedinnvme list-subsysis a display of path existence. For whether they are in use, count with thenvmf_tcp_portcounters.nvmf_lifreturns no rows, andlifdoes not count NVMe-oF -
In this configuration, not one byte went to the second path even under
queue-depth. Two places agreed (the ONTAP-side counters and the client-side sockets), and it was the same on Rocky and on RHEL
AWS documentation notes and support inquiries
This article touches three places in the documentation. All three were submitted on 2026-09-17.
The source for all of them is
Provisioning NVMe/TCP for Linux.
| Remark | State |
|---|---|
Against "apart from installing packages, valid on other EC2 Linux AMIs too": cat /etc/nvme/hostnqn finds no file on AL2023 |
Submitted (2026-09-17). Reproduced in the same environment, confirming the wording doesn't hold as-is for the AL2023 environment checked. The two proposed fixes (scoping the wording, adding a generation step) have been shared as an improvement request. Adoption, content, and timing are unconfirmed |
Against the same statement: cat /sys/module/nvme_core/parameters/multipath finds no file on AL2023 (kernel build option) |
Submitted (2026-09-17). The observation was confirmed and filed as documentation feedback. Fix timing undecided. A follow-up question ("why is AL2023 excluded from scope?") also got a technical explanation: proceeding with the paths unmerged leaves the other path writable against the same namespace with no multipath software mediating it, a data-corruption risk, so the exclusion has a technical rationale |
Step 5 directs confirming that iopolicy is round-robin, but the measurement on Rocky 9.7 / RHEL 9.7 is queue-depth. nvme-cli's version boundary is 2.11 -> 2.12 (see below) |
Submitted (2026-09-17). On 2026-09-30 it was acknowledged that a note is needed stating whether iopolicy resolves to round-robin or queue-depth depends on the environment, and this was fed to the responsible team (with a note that NetApp's own published documentation carries the same point). The content and timing of any change are not disclosed in advance, so they are unconfirmed |
The form I submitted was this. Verbatim quotation (which line of which step), reproduction steps,
expected and measured values, and the range I confirmed across three kernel series, written as
such. Not generalising to "multipathing is not available on AL2023" is the same in the feedback: the
observation is limited to the three series returned by al2023-ami-kernel-default-x86_64.
For the second one I also included the control under which the procedure does hold (on a kernel
with CONFIG_NVME_MULTIPATH=y, Y is returned). It is not a report that the procedure itself is
wrong.
The third is not submitted as a performance problem. The difference between policies is 0.5% in
this configuration, and the report is the single point that the confirmation step returns a value
differing from the documentation.
The third changed shape before submission, after I looked at the package side. Reading
71-nvmf-netapp.rules out of four builds of the public nvme-cli package (the SRPMs for 2.11 /
2.13 / 2.16, plus the installed 2.16-1.el9) showed that 2.11 has round-robin and 2.13 changed it
to queue-depth (confirmed 2026-09-17). At that point 2.12 had not been examined, so the boundary
could only be written as "somewhere between 2.11 and 2.13".
Submission status across the three block-protocol articles and S3 Burst Part 2 is collected in
AWS documentation notes and inquiries across the block-protocol articles.
A pointer to the upstream v2.11 and v2.12 sources followed, showing the boundary is v2.12.
Fetching and checking that commit (2026-09-19, 71-nvmf-netapp.rules.in) confirmed that
queue-depth was already in place as of v2.12. So the version boundary is settled at
2.11 -> 2.12 (the original "somewhere between 2.11 and 2.13" narrowed to 2.12, from both that
pointer and my own re-check). This is agreed to be a version drift rather than an error in the
documentation.
On 2026-09-30, it was acknowledged for this point that a note is needed stating that whether
iopolicy resolves to round-robin or queue-depth depends on the environment, and it was fed to
the responsible team.
NetApp's NVMe-oF procedure for RHEL 9.x
carries the same point (from RHEL 9.7 the default is queue-depth; on 9.6 round-robin is the
default with queue-depth selectable), consistent with this understanding. The content and timing
of any change are not disclosed in advance, so its reflection in the documentation is unconfirmed.
Primary record of the figures
The primary record of the figures and their conditions is in
the protocol measurement results,
and the reproduction steps are in
the block measurement runbook.
What is unconfirmed is written as unconfirmed there, so please do not drop that when citing.
How to check whether multipathing is available, counter-table traps, and iopolicy version differences are
collected in
Considerations when measuring FSx for ONTAP's block protocols.

Top comments (0)