Benchmarks

Strata Linux Multi-GPU: 8 GiB Host-Registration Cap Cut Prefill 3×

A controlled RTX 4090 + RTX 3060 experiment traced a severe Strata Q2_0 multi-GPU prefill slowdown to the Linux expert-copy path. Allowing full CUDA host registration raised 32K prefill from 665 to 2,110 tok/s.

Approximately 11 min read

Inside this experiment

The result, up front.

What we found
On this Linux RTX 3060 + RTX 4090 system, removing the multi-GPU 8 GiB CUDA host-registration cap raised measured 32K Q2_0 prefill from 664.9 to 2109.8 tok/s.
Test conditions
Strata 0.1.28 at bbaaabb; Qwen3.8-Flash-Next Q2_0; 32K context; int8 KV; GPU order 3060 -> 4090; auto K=2; same benchmark before and after the platform-guarded source change.
What this cannot prove
The primary measurements are one Linux machine and one Q2_0 model. The maintainer has confirmed the WDDM-cap diagnosis and announced a 0.1.31 fix, but that release was not yet public when this note was updated.

The A/B changed only the Linux host-registration cap after separate controls for memlock, PCIe bandwidth, prompt chunk size and GPU ordering.

The Linux full-registration A/BHigher is faster · Scale starts at zero
  1. ~1K prompt
    Stock multi-GPU597.8 tok/s
    Full registration785.6 tok/s
  2. ~4K prompt
    Stock multi-GPU612.8 tok/s
    Full registration1581.2 tok/s
  3. ~32K prompt
    Stock multi-GPU664.9 tok/s
    Full registration2109.8 tok/s

Single local runs using the same model, GPU order and auto K=2 split. The source change only altered the Linux host-registration cap. Original table & context ↓

Drawn directly from this article’s published tables at build time. No new benchmark run or synthetic data.

A mixed RTX 4090 + RTX 3060 setup gave me a result that initially looked like a simple heterogeneous-GPU penalty.

It was not.

With Strata 0.1.28 and Qwen3.8-Flash-Next Q2_0, a 32K prompt processed at only about 665 tok/s in the two-GPU configuration. The same model on the RTX 4090 alone, with the prompt chunk forced to the same 2,048-token size, reached 2,814 tok/s.

After tracing the prompt path with Strata’s own timing instrumentation, the bottleneck turned out to be much more specific: the downstream RTX 4090 was spending most of each prompt chunk waiting for expert copies.

A small Linux-only source change that allowed the full expert arena to use CUDA host registration raised 32K prefill to 2,110 tok/s and cut the full ~28K prompt wall time from 42.7 seconds to 13.4 seconds.

I filed the full reproduction upstream as Strata issue #253.

This is not an upstream-confirmed fix yet. It is a controlled local A/B with a strong source-level hypothesis.

Test system

The machine used for this investigation:

Component Configuration
CPU AMD Ryzen 5 7600
RAM ~128 GB
Primary GPU NVIDIA GeForce RTX 4090 24 GB
Secondary GPU NVIDIA GeForce RTX 3060 12 GB
OS Linux
NVIDIA driver 590.48.01
CUDA toolkit 13.1
Strata engine 0.1.28
Strata revision bbaaabb4643bb7873d4cef9d48b5dcf96e6cbff4
Model Qwen3.8-Flash-Next
Quantization Q2_0
Context 32,768
KV int8
Vision off

The key multi-GPU run used the device order RTX 3060 -> RTX 4090.

Strata’s automatic layer split selected:

RTX 3060: layers 0-1
RTX 4090: layers 2-47
K=2

That ordering matters because the last GPU also carries the output head and draft/MTP work. An earlier 4090 -> 3060 test was slower on decode, so I reversed the order before investigating the prompt regression.

The symptom

The first surprising result was how little the second GPU helped prompt processing.

The two-GPU setup had a decode expert-cache hit rate around 97-98%, and the RTX 3060 was responsible for only two layers. Yet prompt processing at 32K was roughly one quarter of the 4090-only result when both used a 2,048-token prompt chunk.

That suggested something more severe than “the 3060 is slower.”

The A/B result

The cleanest comparison uses the same model, same GPU order, same K=2 split, same context, and the same benchmark script.

Context Stock multi-GPU Linux full registration
1K 597.8 tok/s 785.6 tok/s
4K 612.8 tok/s 1581.2 tok/s
32K 664.9 tok/s 2109.8 tok/s

At 32K, the patched run was about 3.17x faster.

Decode in the same run

Decode also improved modestly in the same run:

Context Stock decode Patched decode
1K 123.6 tok/s 120.5 tok/s
4K 111.0 tok/s 125.3 tok/s
32K 111.5 tok/s 126.9 tok/s

The main effect is clearly prefill rather than decode.

The full prompt wall time moved with it

For the ~28K prompt used by the profiler:

stock multi-GPU:   42,674 ms
patched multi-GPU: 13,359 ms

That is almost exactly the same story as the throughput numbers.

The next question was where those extra 29 seconds were going.

Strata’s profiler showed the copy stall

Strata already exposes STRATA_PREFILL_TIMING=1, so no instrumentation patch was required.

In the stock multi-GPU run, representative 2,048-token chunks on the RTX 4090 stage looked like this:

GPU timeline 3374 ms ... wait copy 2644 ms (78.4%)
GPU timeline 2946 ms ... wait copy 2208 ms (74.9%)
GPU timeline 3030 ms ... wait copy 2004 ms (66.1%)
GPU timeline 3237 ms ... wait copy 2448 ms (75.6%)

The GPU was not spending most of the time doing dequantization or GEMMs. It was waiting for expert copies.

After the Linux full-registration change, representative chunks became:

GPU timeline 803 ms ... wait copy 35 ms (4.3%)
GPU timeline 822 ms ... wait copy 40 ms (4.9%)
GPU timeline 876 ms ... wait copy 35 ms (4.0%)
GPU timeline 827 ms ... wait copy 34 ms (4.1%)

Some patched chunks still showed higher copy waits, including roughly 100-215 ms, but the multi-second stalls disappeared.

This is the strongest evidence in the experiment because it links the throughput change to a specific internal phase rather than only comparing end-to-end numbers.

Why the 8 GiB cap became the main suspect

At this revision, Strata’s multi-GPU setup applies a CUDA host-registration limit:

const uint64_t pin_limit =
    (o.expert_cache_remote[0] > 0 || multi_gpu) ? (8ull << 30) : 0;

The surrounding source comment describes the motivation as a multi-GPU Windows/WDDM experiment where registering the full host expert arena into multiple CUDA contexts caused later allocations to fail.

On this Linux Q2_0 run, startup reported:

cudaHostRegister limited to 8 GiB for CUDA1
12 slices pinned (7 GiB)

The Q2_0 expert arena is much larger than that.

Inside the prefill path, Strata distinguishes between expert blobs that can be copied directly from pinned host memory and blobs that require the staging path. That made the registration cap a plausible explanation for the enormous wait copy term.

The local patch

For the A/B, I kept the existing cap on Windows and removed it on Linux:

#ifdef _WIN32
const uint64_t pin_limit =
    (o.expert_cache_remote[0] > 0 || multi_gpu) ? (8ull << 30) : 0;
#else
const uint64_t pin_limit = 0;
#endif

Nothing else in the benchmark configuration changed.

On Linux, startup could then use unrestricted CUDA host registration for the arena.

The result was the 3.17x 32K prefill increase shown above.

This does not prove that this exact conditional is the right upstream fix. It does show that the registration policy is causally important on this machine.

Control 1: raising memlock alone did not fix the capped path

One obvious theory was the process’s locked-memory limit.

The original shell had a finite RLIMIT_MEMLOCK and Strata reported an mlock failed fallback.

I raised the shell limit to unlimited:

Max locked memory    unlimited    unlimited    bytes

The startup message changed from failed mlock to successful mlock.

Performance did not:

before unlimited memlock:
32K prefill 654.2 tok/s
decode      124.6 tok/s

after unlimited memlock:
32K prefill 652.5 tok/s
decode      124.8 tok/s

That difference is noise.

So a successful mlock() fallback was not enough to remove the stock 8 GiB-cap regression. The independent 2x5060 Ti reproduction adds an important nuance: a high locked-memory limit may still be required for the full cudaHostRegisterPortable path itself on Linux. In other words, “memlock alone does not fix the capped path” is different from “memlock limits never matter.”

The RTX 3060 is electrically limited to x4 in this machine, so PCIe was another obvious suspect.

Topology:

GPU0 <-> GPU1: PHB
P2P: unavailable

Measured pinned-memory bandwidth:

RTX 4090 H2D 26.90 GB/s
RTX 4090 D2H 26.40 GB/s

RTX 3060 H2D  6.40 GB/s
RTX 3060 D2H  6.60 GB/s

That is normal for this secondary Gen4 x4 path.

It is also in the same broad range as the x4 secondary-GPU link documented in Strata’s own layer-split benchmark, where multi-GPU prompt processing can still outperform one GPU.

So x4 by itself does not explain a 4x prompt collapse.

Control 3: the 2,048-token chunk explains only part of the loss

Strata changes prompt-buffer behavior in multi-GPU mode and ends up using a 2,048-token prompt chunk in this configuration.

Large prompt chunks can be much faster because an expert streamed for a chunk can serve more tokens before the next chunk boundary.

To isolate that effect, I forced the RTX 4090-only run to the same 2,048-token chunk:

Context RTX 4090 only, prefill 2048
1K 1104.9 tok/s
4K 1988.0 tok/s
32K 2814.4 tok/s

So the smaller chunk is a real penalty, but it does not explain the stock dual-GPU result of about 665 tok/s.

After the registration patch, the dual-GPU result rose to 2,110 tok/s, much closer to the 2,814 tok/s single-4090 control.

Why the patched dual-GPU run is still slower than one RTX 4090

The patch removes a pathological copy stall. It does not make an RTX 3060 equivalent to an RTX 4090.

Strata estimated the two GPUs roughly as:

RTX 3060: 28 SMs @ 1.84 GHz
RTX 4090: 128 SMs @ 2.52 GHz

Even with only two layers on the 3060, the heterogeneous pipeline still has:

That is why 2,110 tok/s patched dual-GPU still trails 2,814 tok/s single-4090 at the same 2,048-token chunk.

For this particular model and hardware pair, the RTX 4090 alone remains the fastest configuration.

The multi-GPU experiment is valuable because it exposed a Linux-specific performance interaction, not because adding the 3060 became the best deployment choice.

What I think the evidence supports

The evidence supports four fairly narrow claims:

  1. The stock Linux multi-GPU run had a severe Q2_0 prefill regression on this machine.
  2. Profiling localized most of the lost time to expert-copy waits on the downstream RTX 4090 stage.
  3. Raising RLIMIT_MEMLOCK and validating PCIe bandwidth did not remove the regression.
  4. Allowing unrestricted CUDA host registration on Linux removed most of the copy wait and increased 32K prefill by about 3.17x.

What it does not yet establish:

That is why the upstream report is framed as a reproducible performance issue with a tested candidate change rather than a finished diagnosis.

Independent reproduction on 2x RTX 5060 Ti 16GB

After issue #253 was filed, GitHub user cha0yang reported an independent Linux reproduction on a very different machine:

Their 32K prompt result moved from 821 tok/s with the 8 GiB cap to 2,141 tok/s after the same Linux-only registration change, a reported +161% gain.

Their 8K result also improved, although the post-patch run required a manually selected prompt chunk, so it is not as clean an A/B for pinning alone.

Decode stayed around 70 tok/s and the reported expert-cache hit rate stayed at 87.8%.

The most interesting part is that their patched 32K result, 2,141 tok/s, is within about 1.5% of the 2,110 tok/s measured on the RTX 3060 + RTX 4090 system here despite completely different GPUs and a different quantization format.

That makes the original result much harder to explain as a quirk of one heterogeneous GPU pair.

The 16GB-card caveat

The independent reproduction also exposed a limitation that my 24GB + 12GB setup did not.

Full cudaHostRegisterPortable registration consumes GPU address-space / VRAM resources. On the two 16GB cards, that reduced the space available for the expert cache and prompt buffers enough that Strata’s automatic 2,048-token prefill choice no longer fit.

Their workaround was to select the chunk manually:

--prefill 8K prompt
512 1,110 tok/s
1024 1,576 tok/s
1536 1,694 tok/s
2048 startup failure

So a platform guard alone may not be the whole upstream solution. On lower-VRAM cards, Strata’s automatic prefill sizing may also need to account for the VRAM cost of fully mapping the host arena.

This is useful evidence in both directions: the Linux performance win reproduced independently, and the reproduction found a real resource tradeoff that a simple “remove the cap on Linux” patch could otherwise hide.

Maintainer confirmation and the 0.1.31 fix

Niko1221 has now confirmed the diagnosis in issue #253.

According to the maintainer, the 8 GiB cap is a Windows/WDDM workaround that was also being applied on Linux. Strata 0.1.31 will:

The maintainer also said they are measuring the change on a two-GPU AMD system.

That directly addresses both pieces exposed by the community testing: the Linux performance regression itself and the 16GB-card startup failure found by the independent 2x RTX 5060 Ti reproduction.

At the time of this update, the latest public GitHub release is still v0.1.30, so I am treating the 0.1.31 behavior as maintainer-confirmed upcoming behavior rather than a released result. Once 0.1.31 is public, the remaining useful check is to rerun the same stock benchmark without the local patch.

Upstream report

The reproduction, controls, source snippet, and local patch are now public in:

If the project changes the registration policy or asks for additional A/B runs, I will update this article with the follow-up results.

Sources and further reading

Continue reading