DEV Community

Amazon S3 Files vs. Amazon FSx for NetApp ONTAP: two designs for using S3 as a file (S3 Burst Part 2)

Amazon S3 Files vs. Amazon FSx for NetApp ONTAP: two designs for using S3 as a file, measured on the same host (S3 Burst Part 2)

Keeping the primary copy of your data in one place while reaching it both as a file (NFS/SMB)
and over the S3 API: have you had to weigh which design to pick before committing to one?

There is more than one way on AWS to combine the two. Worth pausing on this, though: every
design has a point where "fast enough" turns into a ceiling, and that point sits in a different
place for each one. It helps to know where each one bottlenecks before choosing. This article
measures three of them from the same host.

The three-path configuration and where each bottlenecks

The same configuration as the image above, reproduced so it reads in environments where images
don't render.

flowchart LR
    subgraph AWS["AWS Cloud"]
        subgraph VPC["VPC, single AZ"]
            EC2["EC2 client"]
            subgraph A["A: This architecture"]
                AP["S3 Access Point"]
                FSXN[("Amazon FSx for<br/>NetApp ONTAP")]
                AP --> FSXN
            end
            S3["Amazon S3<br/>(bucket)"]
            S3FILES["Amazon S3 Files<br/>(via efs-proxy)"]
            EC2 -->|"write over S3 API"| AP
            FSXN -->|"read over NFS / SMB"| EC2
            EC2 <-->|"read/write over S3 API"| S3
            EC2 <-->|"read/write over NFS"| S3FILES
        end
    end

The figure states the same thing as the table below. Only A uses a different protocol for
writes versus reads; B and C have the same client reading and writing over the same protocol.
The figure's three panels map directly to the three below.

Path Write Read
A FSx for ONTAP's S3 Access Point (this architecture, below) S3 API NFS / SMB
B Amazon S3 S3 API S3 API
C Amazon S3 Files (an S3 bucket mounted over NFS) NFS NFS

Only A uses a different protocol and client for writes versus reads. That's the substance of
this architecture: something written over the S3 API is read over NFS/SMB with no copy job in
between. That's why the figure shows two arrow types for A, while B and C are a single
bidirectional arrow — the same client reads and writes over the same protocol in B and C.

Measure throughput and lay it out in a table, and the three numbers sit side by side. Sitting side
by side reads as "this one is faster," but each of the three was bottlenecked for a different
reason. Under each panel of the figure is where that path was actually bottlenecked. A throughput
value provisioned on the file system, client-side network bandwidth, the CPU of a process running
on the client. What the bottleneck is matters more to design than what the ceiling number is.

The measuring side has its own prerequisite. There are five settings that, left at default,
measure something other than what you intended.
That's covered in the
previous article (Part 5, "five defaults that can invalidate a measurement").
This article's figures were measured after changing all five from default.

Scope of this article

What it covers

  • The three ceilings that bind on the FSx for ONTAP side, and which one actually applied
  • What bottlenecks each of the three paths
  • What FlexCache does for reads
  • What becomes the ceiling as client count scales to 8
  • Whether a 15-minute sustained write decays
  • The bottleneck shift with small objects
  • What each tool can and can't observe
  • What hasn't been measured yet

What it does not cover

  • Which configuration is faster. The three figures in this article mix a provisioned ceiling with an unprovisioned elastic ceiling. They aren't laid out to be read side by side for a winner, so wherever a number appears, the type of ceiling it's hitting is stated alongside it.
  • How to choose a configuration itself. Where to place the read side, which side holds the primary copy, when you want writes visible on the bucket side — these are decided by other factors and throughput numbers alone don't decide them.
  • The semantics of Amazon S3 Files propagation. That file-to-bucket propagation converges within roughly a 60-second window, and the reverse direction is on the order of seconds, is a different subject and belongs in a separate article.
  • IAM policy and the authorization model. The AWS side of an S3 Access Point (IAM identity policy, access point policy, VPC endpoint policy, SCP) and ONTAP's two-layer authorization are a different axis from this article's subject (throughput). Minimum-privilege policy examples are in the S3 AP design guide's authorization design section.

Measurement environment

Figures can't be reproduced without the environment they came from, so it's placed first.

Item Value
Measured on 2026-09-01 and 2026-09-02
Region ap-northeast-1
ONTAP 9.18.1P3D1 (the second-generation NFS 2-host measurement used 9.18.1P6. The cross-article version mapping is in the ONTAP version matrix)
File system SINGLE_AZ_1 (first-generation Single-AZ). Measured at two throughput capacities, 128 MBps and 2048 MBps
SSD 1024 GiB, DiskIopsConfiguration at two points, 3,072 and 40,000
Client c5n.9xlarge (single-client measurements), c5n.2xlarge x8 (client-count scaling test — adding one client at a time and observing how the combined throughput grows)
Volume UNIX security style, NFSv3, actimeo=0
Test data Uncompressible (from /dev/urandom). Zero-filled data was not used. These are figures from a world with efficiency features off
AZ configuration All figures are from a Single-AZ configuration. Multi-AZ is unverified

Both NFS and SMB were measured, but the scope differs. The per-protocol comparison table and
per-path bottlenecks are measured over NFS. What was measured over SMB is the growth pattern with
increasing client count, the 15-minute sustained write, and Multichannel channel count. The
seven scenarios under cache control were not measured over SMB.
This architecture can be read
over either NFS or SMB, but this article's figures are all NFS. If you plan to read over SMB,
the figures below don't directly apply. What carries over is the measurement approach and the
procedure for isolating which type of ceiling is binding.

Reproducibility under identical conditions was 0.02-0.14%. With that baseline established, a
"moved 30%" statement later in this article can be said not to be run-to-run noise.

Unless otherwise noted, figures in the body are single measurements. Where something was
measured twice, both values are shown as 415.1 / 415.7 rather than averaged or taking a median.
Totals in the client-count tables are the sum of each host's value.

And every figure in this article is a saturation point — however much could be pushed through at
a fixed concurrency. No point gives a target throughput or target IOPS. This distinction
matters as much as stating which type of ceiling applies — a saturation point answers "where does
it choke," not "what response time does it deliver at that load." Measuring both on a separate file
system (second-generation, 6,144 MBps), a target of 4,400 IOPS produced 4,406 MB/s at 8.86 ms,
while unlimited produced 4,204 MB/s at 121.79 ms. Unlimited is 5% lower throughput and 14x
the response time.
Source:
the series with a varied target value
(Japanese).

Response time exists for the S3 API path and not for the file path. The S3 API side's body
notes p50 alongside throughput (363.7 ms at concurrency 1, 1054.9 ms at concurrency 64, and so on).
The file path's load generator didn't retain the values it output, and the environment has since
been torn down. Because concurrency was fixed, it can be re-derived through the identity
concurrency / IOPS, but that's a derived value, not a measurement.

A separate article pushes further on the same property on the block protocol side
(separate article). It
covers a case where quadrupling thread count moves throughput by only 0.07% while response time
alone quadruples, and the same 2,364 MB/s appears at both a 216 ms point and an 860 ms point.
That same caution applies when citing this article's figures.

The block protocol side is a set of three.
Session and queue counts
(for anyone about to mount),
checking whether ANA is available
(for anyone who followed the procedure and got stuck), and
what moves the numbers
(for anyone measuring or citing figures). There's no order to read them in.

The measurement host was c5n.9xlarge. The reason is isolation. With an "Up to"-labeled instance,
once a measured value plateaus, you can't tell whether you hit the storage-side ceiling or the
client-side bandwidth ceiling.
c5n.9xlarge's 50 Gbps is a guaranteed value. Every figure below
sits well under 50 Gbps, so the possibility that the client-side bandwidth was the cause of a
plateau can be ruled out.

Only the client-count scaling test used eight c5n.2xlarge instances instead, since that test wants
more clients rather than a bigger single client.

The client-count scaling test and the sustained write used a separate file system

The table above is first-generation. Only the client-count scaling test and the 15-minute
sustained write were measured on a separate, second-generation file system. The first-generation
one is a shared environment, and raising or lowering its provisioned throughput capacity would
affect other measurements.

Item Value
Measured on 2026-09-05
File system SINGLE_AZ_2 (second-generation Single-AZ), 6,144 MBps x1 HA pair
SSD 4,096 GiB, provisioned 200,000 IOPS
Client c5n.2xlarge at 1 / 2 / 4 / 6 / 8 hosts. 64 threads and nconnect=16 fixed per host

Because the environment differs, these figures can't be placed in the same column as the table
above. Load per host was fixed and only the host count was varied, because fixing total thread
count and dividing by host count would change two things at once per data point, making it
impossible to separate which one caused the effect.

Six settings to change from default before measuring throughput (Part 5)

Each one, left at default, measures something other than what you intended. And the number
that comes back doesn't look unnatural, so measuring alone won't reveal it. The table and how each
was noticed are in Part 5; only the names are listed here.

  1. Linux's NFS mount opens only one TCP connection by default (plateaus around 590 MB/s)
  2. Data created with dd if=/dev/zero (all-zero blocks never reach disk)
  3. Volume inline efficiency (compression/deduplication) enabled
  4. DiskIopsConfiguration at AUTOMATIC
  5. Reads slightly exceeding the read cache's capacity

More on item 5 — the read cache's two-tier structure

Item 5 is covered in detail because it bears directly on this article's figures.

FSx for ONTAP's file server has two tiers of read cache: in-memory and NVMe
(performance specs). For
ap-northeast-1's first-generation Single-AZ, what's published differs by tier.

Tier Capacity at 2048 MBps Source
In-memory 256 GB Performance specs, the "other regions" table
NVMe read cache Not stated No column exists in that same table

For NVMe, only a condition is stated in a separate section: "A Single-AZ 1 created on or after
2022-11-28 with a throughput capacity of 2 GBps or more gets a read cache." This file system meets
that condition. Checking with the ONTAP CLI directly confirmed it was enabled.

128 MBps configuration : system node external-cache -> 0 records
2048 MBps configuration: is_enabled: true (both nodes)
Enter fullscreen mode Exit fullscreen mode

And how SSD IOPS gets used is stated explicitly in the same documentation. SSD IOPS is used only
when reading data that's in neither the in-memory cache nor the NVMe cache.

With that in mind, the measurement can be read this way. A 280 GiB read exceeds the in-memory
cache's 256 GB (= 238 GiB) by 18%, but there's an NVMe read cache behind it. This
configuration's NVMe capacity isn't published (in the four regions with a column, it's 1,900 GB at
the same 2048 MBps). The measured 98.5-99.9% of read bytes not going to disk is consistent with
coming back from these two tiers
(breakdown under "Per-tool visibility" below).

That's why the observation "raising SSD IOPS from 3,072 to 40,000 made reads 7.14x faster"
didn't square with the documentation. If it's coming back from cache, SSD IOPS shouldn't matter.

That mismatch went away on re-measurement (2026-09-17). Setting the working set to 512 GiB —
2.15x the in-memory cache's 238 GiB — and disabling the NVMe cache, re-measuring showed
96-126% of read bytes coming from disk, with the low end saturating at DiskIopsUtilization
112%. "SSD IOPS determines the disk throughput level" is consistent with measurement under
this condition.
So the mismatch's cause wasn't the documentation — it was my measurement:
a 280 GiB working set only exceeded the cache by 18%.

The ratio itself changed, though. 256.47 -> 1,246.67 MB/s, 4.86x. 7.14x doesn't
reproduce. And what's capping the high side, around 1,250 MB/s, is none of the published
utilization metrics (disk throughput 67% / IOPS 43% / network 43%). It isn't concurrency
either — going from 8 to 32 to 64 threads drops it from 1,071 to 999 to 958, with response time
alone scaling proportionally (6.9 -> 31.0 -> 59.3 ms). This is left unresolved.

To check in your own environment, the order is two steps. Check for an NVMe cache with
system node external-cache, and then read a volume that reliably exceeds the sum of both
tiers in one pass. A first-generation Single-AZ at 2 GBps or more gets an NVMe cache by
default — there's no way to prevent it from being created, only to create it and then disable
it.
I forgot this once and measured four points with it left enabled.

Item 6 — SMB Multichannel disabled by default

ONTAP's SMB Multichannel is disabled by default. This is a separate setting from SMB 3.x
itself, so a channel can stay at one even while negotiated at dialect=3.1.1. That state held in
this measurement environment: while one host was reading, only one established TCP connection to
the SVM existed.

Enabling it doesn't reach connections already established. Enabling only affects new
connections. Removing and re-mapping the drive kept the same local port, and it only became four
channels after restarting the SMB client service.

max_connections_per_session is a separate value. With Multichannel disabled, setting it to 32
still yields one channel, and you can misread "raised the limit but it didn't increase" as a
result.

The mechanism, sources, and where to read channel count are in
SMB Multichannel is disabled by default, and enabling it doesn't reach already-established connections
(Japanese).

Three ceilings that limit writes on the FSx for ONTAP side

Before getting into the three-path comparison, here's what limits this architecture's own side.
There's a ceiling that doesn't move even when throughput capacity is raised, and not knowing it
leads to a sizing miss.

One term first. Throughput capacity is a value — 128 / 256 / 512 / 1024 / 2048 MBps and so on —
that you set on the file system. It isn't metered usage; you're billed by time against the
value you set. This article measures at two points, 128 MBps and 2048 MBps.

For ap-northeast-1's first-generation Single-AZ, three ceilings bind on writes, and the one that
actually applies is whichever is smallest.

Ceiling This environment's value How it's determined
Provisioned throughput capacity 128 / 2048 MBps Set by you
Disk throughput out of SSD 768 MBps (768 MBps/TiB x 1 TiB, when SSD IOPS is AUTOMATIC) Driven by the provisioned SSD IOPS
Per-HA-pair ceiling Write 750 MBps (read is 2,048 MBps) Determined by the file system's generation and configuration

Source: performance specs.

The second row isn't a fixed value determined by SSD capacity alone. 768 MBps/TiB is what you
get when SSD IOPS is set to AUTOMATIC (3 IOPS/GiB). Raising SSD IOPS with USER_PROVISIONED
raises that 768 MBps/TiB figure right along with it. Conversely, raising only throughput
capacity to 2048 MBps while leaving SSD IOPS at AUTOMATIC leaves disk throughput capped at 768
MBps/TiB. There are two levers, and raising only one doesn't move the ceiling.

Laying out which configuration hits which ceiling:

Throughput capacity Ceiling that applied Measured (uncompressible, 8 MiB, concurrency 64) Against that ceiling
128 MBps The provisioned 128 MBps 133.4 MB/s 104%
2048 MBps The per-HA-pair write ceiling, 750 MBps 415 MB/s (measured twice: 415.1 / 415.7) 55%

Landing at 104% of the 128 MBps configuration — slightly over the provisioned value — is a
burst effect.
This configuration's disk throughput can sustain 128 MBps continuously, or 600
MBps briefly. Bursting was available but the measurement stopped at 133.4, so the provisioned
throughput capacity was the binding ceiling in this configuration.

On second-generation, the meaning of the provisioned value itself is different. In a separate
measurement, a second-generation Single-AZ's provisioned value points to a disk throughput
baseline, and the network-side baseline is roughly twice that. On a configuration provisioned
at 1,536 MBps, 1 MiB sequential reads produced 2,882 MB/s for 27 minutes, then 1,439 MB/s.
Carrying this table's first-generation reading over to second-generation misses a read sizing
estimate by 2x

(measurement record)
(Japanese).

At the 2048 MBps configuration, the provisioned value stopped being the ceiling. The denominator
for percentage changes here, so both are given. 415 MB/s is 20% against the provisioned 2048
MBps. Against the per-HA-pair write ceiling of 750 MBps that actually applied, it's 55%. The
surprise of "provisioned 2048 MBps, only got 415" is explained by the provisioned value no longer
being the applicable ceiling.

What's consuming the remaining 45% — the unused portion of the 750 MBps — hasn't been isolated.
Two candidates come from the published specs.

  • Writes use twice the network bandwidth of reads. This is documented explicitly, because a write replicates to the secondary file server, so a single write turns into twice the network throughput.
  • A fixed cost per request. Raising concurrency only stretched p50 (median latency) out to 1054.9 ms (at concurrency 64). If there's a fixed cost per single request, raising concurrency doesn't raise requests-per-second. If request rate doesn't rise, total throughput doesn't rise either.

Which one is doing how much hasn't been isolated. This isn't a case where "the reason it doesn't
reach 750 is this one thing" can be stated.

Bottlenecks sitting in different places per path

This is the main point. Three numbers that come out in the same unit, MB/s, were each plateauing
somewhere different. The figure's A / B / C map in that order.

Path Where the bottleneck sits What increases it
This architecture (FSx for ONTAP S3 Access Point) Over the S3 API path, the provisioned throughput capacity on the file system, shared across all clients Raise the provisioned value
Amazon S3 Not the service side — client-side network bandwidth Add more clients
Amazon S3 Files The CPU of the proxy process running on the client Use a bigger client

Note: the row above is scoped to the S3 API path. It's based on measurement (1.01x with 4
clients). On the file path (NFS/SMB) of the same file system, reads exceeded the provisioned
value of 6,144 MBps by more than 2x.
Reads served from cache don't go through disk, so stating
"the provisioned value is the ceiling" without naming the path would contradict that
measurement. Details in
per-protocol measurement results
(Japanese).

Here's the decision flow for telling which of the above a measurement is hitting, in your own
environment.

flowchart TD
    START["A measured value plateaus"] --> Q1{"Does the total grow<br/>when adding clients?"}
    Q1 -->|"No"| C1["Provisioned throughput capacity<br/>on the file system<br/>(this architecture's S3 API path)"]
    Q1 -->|"Yes"| Q2{"Is the per-client value<br/>near the client's network bandwidth?"}
    Q2 -->|"Yes"| C2["Client-side network bandwidth<br/>(Amazon S3)"]
    Q2 -->|"No"| Q3{"Is CPU high on a process<br/>running on the client?"}
    Q3 -->|"Yes"| C3["CPU of a proxy process<br/>on the client<br/>(Amazon S3 Files)"]
    Q3 -->|"No"| C4["A factor outside these three.<br/>Isolate it separately"]

The figure states the same thing as the table below. A single-host measurement alone can't
determine the type
— telling them apart needs either adding clients, or holding transfer size
and concurrency fixed while varying something else at one host (details under
a single-client figure that was actually the client's own limit
above).

FSx for ONTAP — a shared ceiling that wasn't one thing

This is the left side of the figure's panel A, the S3 API write side. Adding clients doesn't
raise the total.
With the same setting and varying client count, 4 clients produced 1.01x.
Since the total doesn't change, each client's share is roughly 1/n of it.

The reason is that the ceiling is a value provisioned on the file system. This is a
configuration where capacity planning is done per file system, not per client.
If more clients
are expected, plan for it on the throughput capacity side.

Note that reads came out higher than the provisioned value. On a 128 MBps configuration, reads
reached 579.3 MB/s (8 MiB, concurrency 16), 4.5x the provisioned value. Using the write-side
ceiling to size reads will miss.

How the ceiling splits as client count grows to 8

The result above, up to 4 clients, is first-generation with the provisioned value as the ceiling.
Scaling to 8 clients / 128 connections on second-generation, with the same file system, same
setting, and same connection count, the total swung 5.5x.
The only thing changed was "where each
host reads from."

The two ceilings a read can hit

With every host reading the same file:

Hosts Connections Client total Port-measured Per host
1 16 2,664.58 MB/s 2,698.7 MiB/s 2,664.58
4 64 8,562.94 MB/s 8,635.2 MiB/s 2,140.74
8 128 11,916.29 MB/s 12,173.0 MiB/s 1,489.54

With the working set held at a fixed 600 GiB, split into 75 GiB per host at 8 hosts, so no block
is shared:

Hosts Connections Client total Port-measured Per host
2 32 2,299.05 MB/s 2,323.4 MiB/s 1,149.53
8 128 2,173.37 MB/s 2,191.7 MiB/s 271.67

Quadrupling connections from 32 to 128 doesn't grow the total. It's already saturated at 32.

The difference comes down to how much is served from the file server's read cache. SSD IOPS is
used only when reading data in neither the in-memory nor the NVMe cache, per the
performance specs. Overlap
in read ranges pulls toward the former; no overlap pulls toward the latter. There isn't one
ceiling — there are two, and which one a workload hits depends on the workload.

A hypothesis that turned out wrong

Before this measurement, I had assumed the SVM's single data LIF was a roughly 4,400 MB/s
ceiling. That was wrong.

Since client-reported totals aren't evidence on their own, I cross-checked against ONTAP's
physical port cumulative counters (nic_common's transmit_bytes), differenced. The
port-measured value exceeded the client total by within 1.3% at every point, and all 8 points went
through the same single port. That means a single LIF was carrying 12,173 MiB/s — 102 Gbps.

One thing tripped up this cross-check. The instantaneous-value counter returned 0 even while
traffic was actually flowing.
Only the differenced cumulative value was usable.

This changes how capacity planning should work. Rather than counting hosts, decide first whether
each host's read range overlaps.
A non-overlapping design pulls toward the disk path's ceiling,
so adding hosts doesn't grow the total — that's what the 32-connection and 128-connection rows
above show.

A single-client figure that was actually the client's own limit

A second, similarly shaped mistake, this time by me. In a separate round, on a first-generation
2,048 MBps configuration, I got roughly 1,250 MB/s and treated it as "this configuration's read
ceiling."
I checked with two clients (2026-09-19).

First-generation 2,048 MBps, SSD 1,024 GiB, SSD IOPS 40,000 (USER_PROVISIONED), NVMe read
cache disabled, volume 800 GiB, client c5n.9xlarge x2, each reading a non-overlapping 300
GiB
, 1 MiB sequential, 8 threads, o_direct, ONTAP 9.18.1P6, ap-northeast-1.

Condition Client total Disk throughput utilization
Single host only 1,195.27 MB/s 63% (SSD IOPS also 40%)
Two hosts simultaneously 981.90 + 919.58 = 1,901.48 MB/s (1.59x) 102-103% (burst)

With a single host, the disk path was only at 63% — the file system side had headroom. Only
with two hosts did the disk side saturate.

From here I chased "so what's the per-host ceiling?" The answer was two of my own settings, not
a resource.
On the same file system, I varied rsize to 1 MiB and thread count (the next day).

Condition (1 host) Measured Response Disk throughput utilization
Transfer 64 KiB, 8 threads 1,195.27 MB/s 6.69 ms 63%
Transfer 1 MiB, 8 threads 1,446.66 MB/s 5.443 ms 77.2%
Transfer 1 MiB, 16 threads 1,646.17 MB/s 8.208 ms 102.6%
Transfer 1 MiB, 32 threads 1,647.42 MB/s 16.420 ms 102.6%
Transfer 1 MiB, 64 threads 1,648.71 MB/s 32.838 ms 102.6%

Transfer size added +21%, concurrency added another +14%, and then it stopped. Past 16 threads,
only response time keeps doubling while throughput doesn't move. It's queueing, not gaining.

And a single host reaches the same 102.6% that two hosts reached. IOPS sat at 50.9% and network
at 66.3%, neither saturated. I could not name a specific client-side device that was saturated.

This "can't name it" was traced to a different tier and resolved separately (2026-09-20). On
second-generation, 6,144 MBps / 200,000 IOPS, measuring five ways of splitting the working set
shows adding volumes moves it only +1.0%, while adding SVMs moves it +26.3% and adding
clients moves it +52.7%.
So what matters is the number of distinct IPs the client talks to
and the client count
, not thread count (+8.9% going from 512 to 1,024). With two clients, the
disk side reaches 96.6% of the provisioned value, and only then does the file system side become
the actual limiter.
No single device can be named because the limiter isn't one counter — it's
how receive processing is distributed
(measurement)
(Japanese).

This reinforces this article's claim. The table at the top says to confirm the type of ceiling;
a single-host measurement can't determine that type. Using a value measured alone as a ceiling
for capacity planning under-buys.
There are two ways to check — adding hosts and fixing transfer
size/concurrency — and both arrive at the same conclusion.

I also chased the discrepancy between them, without using any additional resources. Both file
systems had already been deleted, but CloudWatch retains data for 63 days after deletion.

  • This architecture doesn't publish a single burst-remaining metric (zero out of 77 metrics contain Balance). My earlier explanation, "because burst remainders weren't aligned," doesn't hold, because nothing to align exists.
  • The denominator for utilization was the provisioned value itself (2,048 MB/s). DiskReadBytes / (utilization / 100) converges to 2,048 across every minute. On first-generation, over 100% just means "more than the provisioned value came out," not evidence of spending a burst allowance. Second-generation's 1,536 config at 203.5% is against a baseline, so the same metric has a different denominator by generation.
  • And both sessions were hitting the same server-side ceiling of about 2,040 MB/s. The 15% client-reported difference wasn't a difference in what the file system delivered.

As one discrepancy closed, a second one — between two counter systems — opened, and it closed
the same day too. At saturation, the server reported about 2,043 MB/s while vdbench reported
1,646-1,649 (19% short). The cause was how the measurement was taken. Aligning the window's
start to a minute boundary and comparing only the minutes fully contained in the window, the
two systems agree within 0.2% (300s window: 271.04 vs 271.57; 900s window: 326.70 vs 326.07).
The 19% arose from taking a 60-second average of a minute where load was only partially applied
as the server-side value. Against a steady-state 271 MB/s, the start minute was 216.85 and the end
minute 58.47. There's nothing to restate about FSx for ONTAP here. Making the two agree
requires both aligning the window and discarding the partial minutes
(measurement)
(Japanese).

The shared ceiling split two ways over SMB too

The same shape showed up over SMB. Varying only client count across 8 clients didn't produce
one single answer.

The SMB host-count test's configuration

Everything below the share is identical between the two. Same single CIFS share, same SMB 3.1.1
with 4 Multichannel channels, same physical port on the same node of the same SVM. Only what each
host reads differs.

Total at 8 hosts Read range
10,721.19 MB/s (port-measured 11,194.7 MiB/s) Every host reads the same range of the same file
4,225.31 MB/s Each host reads its own 1/8 of the range

The former was still climbing at 8 hosts; the latter plateaued between 4 and 8. The former's
growth is the result of ONTAP's memory serving the overlap. Citing only one of them leads to the
opposite conclusion.

Client-side totals aren't evidence on their own here either. A single read can satisfy multiple
clients, so bytes never sent over the wire can get counted. Every point here is cross-checked
against ONTAP's physical port cumulative counters.

Per-point figures, and what this measurement doesn't answer (the 180-second window doesn't
separate burst from baseline), are in the
SMB host-count test
(Japanese).

Amazon S3 — a total proportional to client count

The figure's panel B is made of just this one bidirectional arrow.

Hosts Write total Per host Read total Per host
1 581.5 MB/s 581.5 496.3 MB/s 496.3
2 1158.0 MB/s 579.0 977.5 MB/s 488.8
4 2362.5 MB/s 590.6 1974.7 MB/s 493.7
6 3601.4 MB/s 600.2 3038.5 MB/s 506.4
8 4783.2 MB/s 597.9 3992.2 MB/s 499.0

Going from 1 to 8 clients, write totals grew 8.23x and reads 8.04x. The per-host value doesn't
drop.
The total reached 38.3 Gbps at 8 hosts, and this range never showed a point where it
stopped being proportional to host count. Each host used a distinct prefix, so per-prefix request
ceilings weren't tested.

The roughly 500 MB/s visible with a single host was not Amazon S3's ceiling — it was the
client's.
Since adding hosts grows the total, capacity planning here comes down to counting
client-side bandwidth.

Amazon S3 Files — the CPU of a proxy process on the client

This path's structure differs from the other two, so the structure is described first. The
figure's panel C shows the client and efs-proxy sharing the same box — that's it.
Mounting S3
Files starts a transfer process called efs-proxy on the client, and the NFS client connects to
that process. It's visible in the mount info.

127.0.0.1:/ nfs4 rw,vers=4.2,rsize=1048576,wsize=1048576,proto=tcp,port=20812,addr=127.0.0.1
Enter fullscreen mode Exit fullscreen mode

The mount target is 127.0.0.1 — itself. Data is then carried from efs-proxy to the service
side. FSx for ONTAP's mount target is the file system itself, so this is structurally different.

With that structure in mind, I measured what
not supporting nconnect
means in practice. nconnect is an option that increases the number of TCP connections an NFS
mount opens.

Stream count Read Write
1 67.9 MB/s 34.6 MB/s
4 264.3 MB/s 140.5 MB/s
8 451.3 MB/s 281.0 MB/s
16 404.8 / 450.0 MB/s 531.8 MB/s

Reads reach about 450 MB/s at 8 streams and don't grow further at 16 (writes are still growing
at 16, 531.8 MB/s). What follows covers the read plateau. 450 MB/s is 3.6 Gbps, under the
single-flow ceiling
(roughly 5 Gbps per network flow).
So the single-flow ceiling isn't what's binding.

The plateau sits at efs-proxy's CPU. While reading at 16 streams, CPU usage on the
efs-proxy process was 67.3% + 18.1% + 14.9% (an 8-vCPU host).

This is where the meaning of not supporting nconnect becomes clear. What nconnect increases
is the connection count of the NFS mount — between the NFS client and efs-proxy. Since the
plateau is efs-proxy's own CPU, even if it were supported, it wouldn't reach where the plateau
actually is.

So it isn't "slow because nconnect isn't supported." Since the bottleneck is a process's CPU on
the client, the fix is a bigger client, not a mount option.
efs-proxy --tls being on this path
is the cost of a configuration where in-transit encryption is the default, not a defect.

One practical note. Specifying nconnect hung the mount (90-second timeout). Not a rejection,
not silent ignoring. The first measurement was entirely lost to this. When trying an unsupported
option, wrap it with timeout.

Absence of decay over a 15-minute sustained write

Writes had been measured in a 120-second window, so I re-measured at 900 seconds, on the
second-generation environment.

Item Value
Steady-state window average 2,063.00 MB/s (900 seconds after a 60-second warmup)
10-second interval min / max 1,749.0 / 2,128.4 MB/s
Cumulative written Roughly 1.77 TiB
Decay Not observed. The first minute and the last minute sit in the same range

Absorption somewhere can be ruled out. In-memory cache is 256 GB and NVMe cache is disabled;
the amount written is 7x that. Volume efficiency was disabled, with 0 savings.

Stating the relationship to the published value directly: the documentation says for
second-generation, "reads get the full throughput capacity, writes get one-third of it," and
it separately names 6,144 MBps as an exception to that general rule, listing it in a separate
table with Single-AZ writes at 1,024 MBps. So the published value that should apply to a 6,144
configuration is 1,024 MBps.
The measurement was 2.05x that.

This single point reproduced on a separate file system later. Same 6,144 MBps provisioned
value, 1,800 seconds (twice this measurement's window), 2,097 MB/s, a 0.3% difference between
the first 300 seconds and the last 300. Different day, different file system, different window
length: 2,063 and 2,097, a 1.6% difference.
The no-decay conclusion is supported by two
observations.

Lowering the provisioned value and measuring again still exceeds it in the same direction.

Provisioned Published value that should apply Measured (sustained, median) Against the published value
1,536 One-third of the general rule -> 512 MB/s 1,333 MB/s 2.60x
6,144 Exception table's 1,024 MB/s 2,097 MB/s 2.05x

Both points exceed the published write ceiling. Same direction, ratios of 2.60 and 2.05.

This used to be described a different way. It previously read "the general rule holds at
6,144 and misses at 1,536," but 6,144 / 3 = 2,048 being close to the measured 2,097 is a
coincidence.
The documentation names 6,144 as an exception to the general rule, so applying
the general rule there doesn't hold in the first place.
Corrected.

Which figure refers to what still isn't determined. The table header reads "Maximum
throughput from SSD storage," which may refer to a ceiling scoped to the SSD path only.
But this measurement closed off the likely absorption points — the cumulative 1.77 TiB is
7x the 256 GB in-memory cache, the NVMe cache was confirmed disabled on both nodes, efficiency
was disabled with 0 savings. Why it still comes out roughly 2x can't be explained from this
side.

The design-level effect is over-, not under-provisioning. Working backward from the general
rule to hit a write of 1,333 MB/s looks like it needs a 4,000 MBps-class provisioned value, but
1,536 was measured to be enough. The second-generation provisioned rate is $2.013/MBps-month
(ap-northeast-1, fetched 2026-09-04), so this misreading translates directly into overbuying.

Only two points were measured: Single-AZ, 1 HA pair. Multi-AZ, 2+ pairs, and provisioned values
other than 1,536 and 6,144 weren't measured
(measurement record)
(Japanese).

Capacity growth from continued writing

Trying to continue into a read-side measurement after this, it didn't complete. The cause is a
documented specification — a gap in my own design, not a documentation error.

Before starting, snapshot count was 0 and the snapshot policy was none. 30 minutes into the run, a
snapshot named starting with backup- appeared, holding 739 GiB, and the write rate dropped from
2,200 MB/s to 267 MB/s. Volume usage went from 88% to 100%, with 9.9 GiB free. Deleting files
with rm didn't return the free space.

The cause is the automatic daily backup.
Volume backups state that a
backup first takes a volume snapshot, and that snapshot sits inside the volume, consuming
capacity, and stays until the next backup.
For an overwrite-heavy workload, the pre-overwrite
blocks are retained, so the amount written becomes capacity as-is.

Why it doesn't return on deletion is
also documented.
Data deleted from the active file system is not released as long as a snapshot still references
it.

Checking for snapshots before starting doesn't prevent this. Before the run it reads normally,
because no snapshot exists yet. There was a recovery mechanism, though.
It's documented
that volume autosize and Snapshot autodelete are designed to work together. Neither was active
by default in this setup. The actual situation was that the default was in effect because
neither was specified in the template. "Not specified" isn't "disabled."

15 minutes over SMB

No decay over SMB either. The 900-second steady-state window averaged 1,488.03 MB/s, with a
10-second-interval min of 1,452.70 and max of 1,514.40; the first-third average was 1,487.55
against a last-third of 1,486.57.

Changing the client from an 8-vCPU type to a 36-vCPU type only moved it 0.6% (1,478.95 versus
1,488.03). This value isn't determined on the client side.

The initial observation that a 300-second measurement (1,698.42 MB/s) was 12.4% higher than the
900-second one doesn't reproduce (re-measured 2026-09-20). Varying only window length at 6,144
MBps, the longer window actually came out 4.0% higher (1,375.4 versus 1,430.3 MB/s). Channel count
stayed constant regardless of measurement duration, and even writes moved only 1.9% across window
lengths. The original 12.4% gap looks like a one-off value, not a property of window length.
That said, the longer window's vdbench format was interrupted partway through, so it wasn't a
comparable pair
(re-measurement record)
(Japanese).

Against the same physical port on the same node of the same file system, NFS measured 2,063.00
MB/s and SMB 1,488.03 MB/s. Since the source of concurrency differs (16 connections specified by
the client versus 4 negotiated channels), this can't be attributed to a difference in the
protocols themselves.

Conditions and the 10-second-interval progression are in
the SMB sustained write
(Japanese).

Additional read capacity via FlexCache

One more thing was measured on the read side. This section covers the figure's panel A, where
Origin and Cache sit as separate file systems. FlexCache isn't a mechanism for sharing the
origin's (the replication source volume's) performance — it's a mechanism that adds separate
capacity on the read side.

Measurement point Measured, uncompressible
FlexCache first read (cache not yet filled; read while filling) 306.8 MB/s
FlexCache resident (second read onward; cache already filled) 882.7 / 907.1 MB/s
Direct origin read (control — the comparison baseline; already read once and cached) 383.3 MB/s

Resident is 2.31x a direct origin read. The cache is a separate file system from the origin,
with its own throughput capacity, its own memory, and its own burst allowance. So while reading
from the cache, the origin's own ceiling doesn't apply.

The cost is the first read. Data that's never been read comes in at 306.8 the first time — 1/2.88
of resident.
That's the figure to design around.

One qualification. This was measured with both origin and cache at 128 MBps. Raising the
origin to 2048 MBps puts a direct origin read into the 2000 MB/s class, so depending on
configuration, the 2.31x could reverse. Unmeasured.

Bottleneck shift with small objects

Raising throughput capacity 16x and SSD IOPS 13x didn't raise request rate. The six measured
points were 4 KiB reads and writes at concurrency 64 and 256, and 64 KiB reads and writes at
concurrency 64. None of the six grew, and writes dropped 25%. Against 40,000 IOPS, 279-605
req/s — 0.7-1.5% of provisioned IOPS.

The bottleneck sits on the S3 Access Point's request-handling side. For designs that handle a
large volume of small objects, raising throughput capacity and SSD IOPS doesn't help.

Per-tool visibility

What's available to check "what is this actually limited by" while measuring differs by path.
Not knowing this means missing the numbers needed to identify a cause.

What you want to see Reaches it Doesn't reach it
Per-volume throughput ONTAP's volume counters Client-side measurement (turns into a single-TCP-flow ceiling)
Node CPU CloudWatch (AWS/FSx) fsxadmin's REST doesn't reach it
Fraction that went to disk CloudWatch's DiskReadBytes / DataReadBytes Can't be determined from the client side
S3 Files propagation lag CloudWatch (AWS/S3/Files's ExportAge) —

Measuring from the client side turns into roughly 590 MB/s. Linux's NFS mount opens only one
TCP connection per server by default, and that's what limits it first. If you want to measure
the storage-side ceiling, read the storage-side counters.

CPU on both paths was around 22%. The NFS path delivers twice this architecture's throughput at
the same CPU level.
So this architecture's cost isn't inside ONTAP — it's upstream, on the S3
translation side.

And on reads, DiskReadBytes / DataReadBytes was 1.4 / 0.9 / 1.5%. 98.5-99.9% of read bytes
didn't go to disk.
That means it came back from cache, which links directly to Part 5's fifth
default.

AWS documentation notes and support inquiries

This article touches three points in the published specs. Two were filed on 2026-09-17. The
source for all three is
the performance specs.

Content Type Status
Second-generation writes exceed the published ceiling at both points measured (2.60x the general rule's 512 at a 1,536 provisioned value; 2.05x the exception table's 1,024 at a 6,144 provisioned value) Gap between published value and measurement. Likely absorption points (two-tier cache, efficiency) closed off Filed (2026-09-17). AWS's reply (2026-09-24): the "up to a third" figure is not a hard limit but a sizing indicator. A write uses bandwidth twice over — once to the active server, once to sync the standby server — but network burst capacity can push measured throughput above a third of the provisioned value, and AWS considers that expected behavior, not an anomaly. The table heading's "Maximum" is described as the ceiling to expect once burst is not a factor and the same file system is shared with other workloads. For sizing, AWS pointed back to the original rule of thumb: provision three times the required write throughput (e.g., a 1,333 MB/s requirement calls for a 6,144 MBps configuration)
ap-northeast-1's NVMe read cache capacity is missing from the table (the four regions with a column have it listed). The management page is also written scoped to "second-generation" only, so a first-generation user with a qualifying config can't tell they have a cache Missing content. Presence was confirmable via ONTAP; capacity was not Filed (2026-09-17). AWS's reply (2026-09-21): no documentation change. AWS's position is that a first-generation Single-AZ configuration at 2,048 MBps lacking a cache is not something the documentation promises against, and that enumerating every such corner case isn't practical, so the existing wording stands. Pointed instead to system node external-cache show for checking presence and capacity per file system. Both the capacity-publication ask and the management-page scope ask were effectively declined
The observation that raising SSD IOPS made reads 7.14x faster didn't square with "SSD IOPS is used only for reads not served from cache" My own measurement gap. Hadn't read a volume that reliably exceeded the sum of both cache tiers Not filed. Re-measuring 2026-09-17 with the working set at 2.15x the cache size squared with the documentation (ratio came out 4.86x; 7.14x doesn't reproduce). Not a documentation error

Why the third isn't filed. The mismatch's cause was on my side. Re-measuring settled it
(2026-09-17). With the working set at 512 GiB — 2.15x the in-memory cache — and the NVMe cache
disabled, 96-126% of read bytes came from disk, with the low end saturating at
DiskIopsUtilization 112%. This matches the documentation. Not a documentation error, so not
filed.
Only the ratio changed — 4.86x, and 7.14x doesn't reproduce. And what's capping the
high side, roughly 1,250 MB/s, remains unresolved — that isn't a finding against the published
spec, it's a range I haven't measured.

The first and second were filed. Both were filed without me deciding "the correct value is
this"
— as a gap or omission for whoever holds the values to judge. The first was framed as the
question "what does 'Maximum throughput from SSD storage (write)' actually bound?" and the second
as two points: publishing the capacity, and the management page's stated scope.

The first used to be described a different way in this article. It previously read "which of
the general rule and the exception table applies flips depending on the value," but the
documentation names 6,144 as an exception to the general rule
, so applying the general rule
there doesn't hold in the first place. Corrected before filing.

Both first replies came in on 2026-09-21. The first stopped at confirming AWS's own measurement
clears a third of the provisioned value; the second landed on a decision not to change the
documentation. The underlying fact — a first-generation Single-AZ configuration at 2,048 MBps can
lack a cache — is something the documentation genuinely never promised against, and not
enumerating every corner case is a reasonable line to draw. What didn't go through, though, is
separate from that: asking for the capacity to be published, and asking the management page's
stated scope to cover first generation too, are both requests that a documented gap doesn't
answer by itself.
Recording that they didn't land. The per-file-system way to check
(system node external-cache show) is documented, so checking it yourself is unaffected.

The first one got a follow-up on 2026-09-24. The "up to a third" figure turned out to be a
sizing indicator rather than a hard limit — a write spends bandwidth twice over (once to the
active server, once to sync the standby), and network burst capacity can push it above a third of
the provisioned value, which AWS considers expected rather than anomalous. The table heading's
"Maximum" is described as the ceiling to expect once burst isn't a factor and the file system is
shared with other workloads. For sizing, AWS pointed back to the original rule of thumb: provision
three times the required write throughput. This explanation doesn't dispute that the
measurement exceeded the published value — it only says how to read that published value.

The block protocol side's findings and filing status can be traced from the
separate article.

Submission status across the three block-protocol articles and this article is collected in
AWS documentation notes and inquiries across the block-protocol articles.

What hasn't been measured yet

Meaning not yet measured, not "can't be."

  • Reading more than 2x the cache size in one pass. At 512 GiB (2.15x the cache), 4.86x (256.47 -> 1,246.67 MB/s). The low side saturating on SSD IOPS is confirmed. What was capping the high side, around 1,250 MB/s, turned out to be transfer size and concurrency, not a resource (details under a single-client figure that was actually the client's own limit above)
  • Effect of measurement duration. Measurement duration itself is not the cause. What carries the gap is provisioned IOPS, payload compressibility, and unit-to-unit variance (measurement) (Japanese)
  • Measuring FlexCache uncompressed on a 2048 MBps configuration. The "2.31x" could reverse depending on configuration
  • What's capping the roughly 300 MB/s on the 128 MBps configuration. Reads on the same configuration converge to 297-317 MB/s for both warm and cold (values from the same measurement set as this article, not shown in the body table). This used to be described as "consistent with plateauing on a burst," but reading burst remainders ruled that out. The remainder plateaued at 99 -> 95% and returned to 99% after the load stopped — not exhausted. Against a disk burst ceiling of 600 MBps, the measured 254 MB/s is 42%; against a network burst ceiling of 1,250 MBps, roughly 285 MB/s is 23% — neither ceiling was reached. It isn't the roughly 590 MB/s single-flow ceiling either. Where it sits is unresolved (measurement) (Japanese). Note that DiskReadBytes / DataReadBytes in the same window was 87-97%, so reads on this configuration are coming from disk — a contrast with the 2048 MBps configuration's 1.4%, consistent with an in-memory cache of only 16 GB
  • The host count at which Amazon S3's total stops being proportional to host count. It's still proportional at 8 hosts / 38.3 Gbps
  • Adding S3 Files proxies. A configuration with separate mounts running more efs-proxy instances hasn't been tested
  • More than 8 hosts. The side sharing the same file dropped to a 1.12x increment at 8 hosts, so where it stops hasn't been observed
  • A point below 32 connections on non-overlapping ranges. It was already saturated at 32, so the connection count where saturation begins hasn't been bracketed
  • A sustained write longer than 900 seconds. No decay over 15 minutes, but nothing beyond that has been checked
  • Reads with the fraction served from cache varied. Capacity filled up and the run was aborted, so none of the 7 planned points were captured
  • Time-of-day effects. Two different days were checked, but not different times of day
  • SMB's 7 scenarios under cache control. Stopped by fill, not capacity. format=yes ignores the run block's xfersize and threads, writing at 128 KiB and queue depth 2, so overwriting the same file that otherwise reaches 1,488 MB/s only produced 165 MB/s (315 MB/s with formatxfersize=1m; queue depth stayed at 2). The parameter file points 8 SDs at a single file; splitting each range into its own file looks likely to solve both the speed and the design problem at once. This is a prediction, not a result. Background in two places stuck under cache control (Japanese)
  • Why SMB's 300-second and 900-second values differ. Doesn't reproduce. Re-measuring at 6,144 MBps removes the window-length difference; the longer window's vdbench run had also been interrupted, so it wasn't a comparable pair in the first place (measurement) (Japanese)
  • Conditions for raising Multichannel's channel count above 4. Setting max_connections_per_session to 32 still produced 4. The 4 comes from the client side's ConnectionCountPerRssNetworkInterface, which determines channel count per RSS-capable NIC. The server side's 32 is a per-session ceiling, not the limiter on the increasing side. Adding NICs or raising this value hasn't been tested
  • The discrepancy between server-side counters and the client-side measurement tool. The cause was measurement method — aligning the window's start to a minute boundary and comparing only the minutes fully contained in the window, the two systems agree within 0.2% (measurement) (Japanese)
  • What the 4 KiB random read ceiling actually is. It's on the client side. What matters is the number of distinct IPs the client talks to and the client count, not thread count. When citing this figure, write "per client" — it isn't the file system's ceiling (measurement) (Japanese)
  • Why 4 KiB random reads plateau at 47.5% of provisioned SSD IOPS. The ratio doesn't reproduce. Smaller configurations are disk-IOPS-limited; larger ones follow a different formula. So "47.5% of provisioned" is a value specific to that one configuration, not an invariant with scale. When citing it, always state the 6,144 MBps / 200,000 IOPS condition alongside it (measurement) (Japanese)

Closing

Measuring three paths from the same host, the bottleneck sat in three different places for all
three. A throughput capacity provisioned on the file system, client-side network bandwidth, and
the CPU of a process on the client. This difference doesn't close as you add more clients.

And on the FSx for ONTAP side, there wasn't one ceiling. With the same file system and the same
connection count, the total swings 5.5x depending only on whether each EC2 client's read range overlaps.

Which one to choose comes down to where you place the read side, which side holds the primary
copy, and when you want a write visible on the bucket side. Throughput numbers are one input into
that decision, and they come later in the order of things to check.

What's worth taking away is a procedure: check the type of ceiling before comparing
throughput.
Whether it's a provisioned value, client-side bandwidth, or a process's CPU on the
client. Line up numbers of different types in the same column, and you get a table shaped like a
comparison without being one.

And a value measured with a single EC2 client can't determine that type. On the same file system,
adding a second EC2 client grew the total 1.59x, and fixing transfer size and concurrency while
staying at one client grew it 1.38x. Either way arrives at the same ceiling. There are two ways
to check the type, and neither costs much extra.

A measurement that reproduces within 0.02-0.14% under identical conditions can still swing 30%
from a single misunderstanding about conditions. "Stable" isn't "correct." And this article's
figures are for one specific throughput capacity and one specific payload. They don't transfer
directly as your environment's numbers. What transfers is the measurement approach.

Hope this helps anyone running the same kind of measurement.

Defaults to check before measuring and items to always record are collected in
Performance testing considerations.

Until next time.

Top comments (0)