Strata Linux Multi-GPU: 8 GiB Host-Registration Cap Cut Prefill 3×
A controlled RTX 4090 + RTX 3060 experiment traced a severe Strata Q2_0 multi-GPU prefill slowdown to the Linux expert-copy path. Allowing full CUDA host registration raised 32K prefill from 665 to 2,110 tok/s.
Approximately 11 min read
Inside this experiment
The result, up front.
- What we found
- On this Linux RTX 3060 + RTX 4090 system, removing the multi-GPU 8 GiB CUDA host-registration cap raised measured 32K Q2_0 prefill from 664.9 to 2109.8 tok/s.
- Test conditions
- Strata 0.1.28 at bbaaabb; Qwen3.8-Flash-Next Q2_0; 32K context; int8 KV; GPU order 3060 -> 4090; auto K=2; same benchmark before and after the platform-guarded source change.
- What this cannot prove
- The primary measurements are one Linux machine and one Q2_0 model. The maintainer has confirmed the WDDM-cap diagnosis and announced a 0.1.31 fix, but that release was not yet public when this note was updated.
The A/B changed only the Linux host-registration cap after separate controls for memlock, PCIe bandwidth, prompt chunk size and GPU ordering.
- ~1K promptStock multi-GPU597.8 tok/sFull registration785.6 tok/s
- ~4K promptStock multi-GPU612.8 tok/sFull registration1581.2 tok/s
- ~32K promptStock multi-GPU664.9 tok/sFull registration2109.8 tok/s
Single local runs using the same model, GPU order and auto K=2 split. The source change only altered the Linux host-registration cap. Original table & context ↓
Drawn directly from this article’s published tables at build time. No new benchmark run or synthetic data.
A mixed RTX 4090 + RTX 3060 setup gave me a result that initially looked like a simple heterogeneous-GPU penalty.
It was not.
With Strata 0.1.28 and Qwen3.8-Flash-Next Q2_0, a 32K prompt processed at only about 665 tok/s in the two-GPU configuration. The same model on the RTX 4090 alone, with the prompt chunk forced to the same 2,048-token size, reached 2,814 tok/s.
After tracing the prompt path with Strata’s own timing instrumentation, the bottleneck turned out to be much more specific: the downstream RTX 4090 was spending most of each prompt chunk waiting for expert copies.
A small Linux-only source change that allowed the full expert arena to use CUDA host registration raised 32K prefill to 2,110 tok/s and cut the full ~28K prompt wall time from 42.7 seconds to 13.4 seconds.
I filed the full reproduction upstream as Strata issue #253.
This is not an upstream-confirmed fix yet. It is a controlled local A/B with a strong source-level hypothesis.
Test system
The machine used for this investigation:
| Component | Configuration |
|---|---|
| CPU | AMD Ryzen 5 7600 |
| RAM | ~128 GB |
| Primary GPU | NVIDIA GeForce RTX 4090 24 GB |
| Secondary GPU | NVIDIA GeForce RTX 3060 12 GB |
| OS | Linux |
| NVIDIA driver | 590.48.01 |
| CUDA toolkit | 13.1 |
| Strata engine | 0.1.28 |
| Strata revision | bbaaabb4643bb7873d4cef9d48b5dcf96e6cbff4 |
| Model | Qwen3.8-Flash-Next |
| Quantization | Q2_0 |
| Context | 32,768 |
| KV | int8 |
| Vision | off |
The key multi-GPU run used the device order RTX 3060 -> RTX 4090.
Strata’s automatic layer split selected:
RTX 3060: layers 0-1
RTX 4090: layers 2-47
K=2
That ordering matters because the last GPU also carries the output head and draft/MTP work. An earlier 4090 -> 3060 test was slower on decode, so I reversed the order before investigating the prompt regression.
The symptom
The first surprising result was how little the second GPU helped prompt processing.
The two-GPU setup had a decode expert-cache hit rate around 97-98%, and the RTX 3060 was responsible for only two layers. Yet prompt processing at 32K was roughly one quarter of the 4090-only result when both used a 2,048-token prompt chunk.
That suggested something more severe than “the 3060 is slower.”
The A/B result
The cleanest comparison uses the same model, same GPU order, same K=2 split, same context, and the same benchmark script.
| Context | Stock multi-GPU | Linux full registration |
|---|---|---|
| 1K | 597.8 tok/s | 785.6 tok/s |
| 4K | 612.8 tok/s | 1581.2 tok/s |
| 32K | 664.9 tok/s | 2109.8 tok/s |
At 32K, the patched run was about 3.17x faster.
Decode in the same run
Decode also improved modestly in the same run:
| Context | Stock decode | Patched decode |
|---|---|---|
| 1K | 123.6 tok/s | 120.5 tok/s |
| 4K | 111.0 tok/s | 125.3 tok/s |
| 32K | 111.5 tok/s | 126.9 tok/s |
The main effect is clearly prefill rather than decode.
The full prompt wall time moved with it
For the ~28K prompt used by the profiler:
stock multi-GPU: 42,674 ms
patched multi-GPU: 13,359 ms
That is almost exactly the same story as the throughput numbers.
The next question was where those extra 29 seconds were going.
Strata’s profiler showed the copy stall
Strata already exposes STRATA_PREFILL_TIMING=1, so no instrumentation patch was required.
In the stock multi-GPU run, representative 2,048-token chunks on the RTX 4090 stage looked like this:
GPU timeline 3374 ms ... wait copy 2644 ms (78.4%)
GPU timeline 2946 ms ... wait copy 2208 ms (74.9%)
GPU timeline 3030 ms ... wait copy 2004 ms (66.1%)
GPU timeline 3237 ms ... wait copy 2448 ms (75.6%)
The GPU was not spending most of the time doing dequantization or GEMMs. It was waiting for expert copies.
After the Linux full-registration change, representative chunks became:
GPU timeline 803 ms ... wait copy 35 ms (4.3%)
GPU timeline 822 ms ... wait copy 40 ms (4.9%)
GPU timeline 876 ms ... wait copy 35 ms (4.0%)
GPU timeline 827 ms ... wait copy 34 ms (4.1%)
Some patched chunks still showed higher copy waits, including roughly 100-215 ms, but the multi-second stalls disappeared.
This is the strongest evidence in the experiment because it links the throughput change to a specific internal phase rather than only comparing end-to-end numbers.
Why the 8 GiB cap became the main suspect
At this revision, Strata’s multi-GPU setup applies a CUDA host-registration limit:
const uint64_t pin_limit =
(o.expert_cache_remote[0] > 0 || multi_gpu) ? (8ull << 30) : 0;
The surrounding source comment describes the motivation as a multi-GPU Windows/WDDM experiment where registering the full host expert arena into multiple CUDA contexts caused later allocations to fail.
On this Linux Q2_0 run, startup reported:
cudaHostRegister limited to 8 GiB for CUDA1
12 slices pinned (7 GiB)
The Q2_0 expert arena is much larger than that.
Inside the prefill path, Strata distinguishes between expert blobs that can be copied directly from pinned host memory and blobs that require the staging path. That made the registration cap a plausible explanation for the enormous wait copy term.
The local patch
For the A/B, I kept the existing cap on Windows and removed it on Linux:
#ifdef _WIN32
const uint64_t pin_limit =
(o.expert_cache_remote[0] > 0 || multi_gpu) ? (8ull << 30) : 0;
#else
const uint64_t pin_limit = 0;
#endif
Nothing else in the benchmark configuration changed.
On Linux, startup could then use unrestricted CUDA host registration for the arena.
The result was the 3.17x 32K prefill increase shown above.
This does not prove that this exact conditional is the right upstream fix. It does show that the registration policy is causally important on this machine.
Control 1: raising memlock alone did not fix the capped path
One obvious theory was the process’s locked-memory limit.
The original shell had a finite RLIMIT_MEMLOCK and Strata reported an mlock failed fallback.
I raised the shell limit to unlimited:
Max locked memory unlimited unlimited bytes
The startup message changed from failed mlock to successful mlock.
Performance did not:
before unlimited memlock:
32K prefill 654.2 tok/s
decode 124.6 tok/s
after unlimited memlock:
32K prefill 652.5 tok/s
decode 124.8 tok/s
That difference is noise.
So a successful mlock() fallback was not enough to remove the stock 8 GiB-cap regression. The independent 2x5060 Ti reproduction adds an important nuance: a high locked-memory limit may still be required for the full cudaHostRegisterPortable path itself on Linux. In other words, “memlock alone does not fix the capped path” is different from “memlock limits never matter.”
Control 2: the RTX 3060 PCIe link was healthy
The RTX 3060 is electrically limited to x4 in this machine, so PCIe was another obvious suspect.
Topology:
GPU0 <-> GPU1: PHB
P2P: unavailable
Measured pinned-memory bandwidth:
RTX 4090 H2D 26.90 GB/s
RTX 4090 D2H 26.40 GB/s
RTX 3060 H2D 6.40 GB/s
RTX 3060 D2H 6.60 GB/s
That is normal for this secondary Gen4 x4 path.
It is also in the same broad range as the x4 secondary-GPU link documented in Strata’s own layer-split benchmark, where multi-GPU prompt processing can still outperform one GPU.
So x4 by itself does not explain a 4x prompt collapse.
Control 3: the 2,048-token chunk explains only part of the loss
Strata changes prompt-buffer behavior in multi-GPU mode and ends up using a 2,048-token prompt chunk in this configuration.
Large prompt chunks can be much faster because an expert streamed for a chunk can serve more tokens before the next chunk boundary.
To isolate that effect, I forced the RTX 4090-only run to the same 2,048-token chunk:
| Context | RTX 4090 only, prefill 2048 |
|---|---|
| 1K | 1104.9 tok/s |
| 4K | 1988.0 tok/s |
| 32K | 2814.4 tok/s |
So the smaller chunk is a real penalty, but it does not explain the stock dual-GPU result of about 665 tok/s.
After the registration patch, the dual-GPU result rose to 2,110 tok/s, much closer to the 2,814 tok/s single-4090 control.
Why the patched dual-GPU run is still slower than one RTX 4090
The patch removes a pathological copy stall. It does not make an RTX 3060 equivalent to an RTX 4090.
Strata estimated the two GPUs roughly as:
RTX 3060: 28 SMs @ 1.84 GHz
RTX 4090: 128 SMs @ 2.52 GHz
Even with only two layers on the 3060, the heterogeneous pipeline still has:
- a slower first stage,
- a host-mediated stage handoff,
- per-stage prompt buffers,
- less freedom to borrow expert-cache VRAM for larger prompt chunks.
That is why 2,110 tok/s patched dual-GPU still trails 2,814 tok/s single-4090 at the same 2,048-token chunk.
For this particular model and hardware pair, the RTX 4090 alone remains the fastest configuration.
The multi-GPU experiment is valuable because it exposed a Linux-specific performance interaction, not because adding the 3060 became the best deployment choice.
What I think the evidence supports
The evidence supports four fairly narrow claims:
- The stock Linux multi-GPU run had a severe Q2_0 prefill regression on this machine.
- Profiling localized most of the lost time to expert-copy waits on the downstream RTX 4090 stage.
- Raising RLIMIT_MEMLOCK and validating PCIe bandwidth did not remove the regression.
- Allowing unrestricted CUDA host registration on Linux removed most of the copy wait and increased 32K prefill by about 3.17x.
What it does not yet establish:
- that every Linux multi-GPU system is affected,
- that every Strata quantization format behaves the same way,
- that unrestricted registration is always safe,
- that the Windows cap should be removed,
- or that the maintainer will choose this exact fix.
That is why the upstream report is framed as a reproducible performance issue with a tested candidate change rather than a finished diagnosis.
Independent reproduction on 2x RTX 5060 Ti 16GB
After issue #253 was filed, GitHub user cha0yang reported an independent Linux reproduction on a very different machine:
- Strata 0.1.29,
- Qwen3.8-Flash-Next IQ3_XXS,
- two RTX 5060 Ti 16GB cards,
- Ryzen 7 9700X,
- 64 GB DDR5,
- auto layer split across both GPUs.
Their 32K prompt result moved from 821 tok/s with the 8 GiB cap to 2,141 tok/s after the same Linux-only registration change, a reported +161% gain.
Their 8K result also improved, although the post-patch run required a manually selected prompt chunk, so it is not as clean an A/B for pinning alone.
Decode stayed around 70 tok/s and the reported expert-cache hit rate stayed at 87.8%.
The most interesting part is that their patched 32K result, 2,141 tok/s, is within about 1.5% of the 2,110 tok/s measured on the RTX 3060 + RTX 4090 system here despite completely different GPUs and a different quantization format.
That makes the original result much harder to explain as a quirk of one heterogeneous GPU pair.
The 16GB-card caveat
The independent reproduction also exposed a limitation that my 24GB + 12GB setup did not.
Full cudaHostRegisterPortable registration consumes GPU address-space / VRAM resources. On the two 16GB cards, that reduced the space available for the expert cache and prompt buffers enough that Strata’s automatic 2,048-token prefill choice no longer fit.
Their workaround was to select the chunk manually:
--prefill |
8K prompt |
|---|---|
| 512 | 1,110 tok/s |
| 1024 | 1,576 tok/s |
| 1536 | 1,694 tok/s |
| 2048 | startup failure |
So a platform guard alone may not be the whole upstream solution. On lower-VRAM cards, Strata’s automatic prefill sizing may also need to account for the VRAM cost of fully mapping the host arena.
This is useful evidence in both directions: the Linux performance win reproduced independently, and the reproduction found a real resource tradeoff that a simple “remove the cap on Linux” patch could otherwise hide.
Maintainer confirmation and the 0.1.31 fix
Niko1221 has now confirmed the diagnosis in issue #253.
According to the maintainer, the 8 GiB cap is a Windows/WDDM workaround that was also being applied on Linux. Strata 0.1.31 will:
- apply the cap only under WDDM (Windows and WSL2),
- add
STRATA_ARENA_PIN_GIBto explicitly choose how much of the arena to pin (NGiB, or0for the whole arena), - and make the prompt path step down the chunk size instead of exiting when full arena pinning leaves too little VRAM for the originally selected chunk.
The maintainer also said they are measuring the change on a two-GPU AMD system.
That directly addresses both pieces exposed by the community testing: the Linux performance regression itself and the 16GB-card startup failure found by the independent 2x RTX 5060 Ti reproduction.
At the time of this update, the latest public GitHub release is still v0.1.30, so I am treating the 0.1.31 behavior as maintainer-confirmed upcoming behavior rather than a released result. Once 0.1.31 is public, the remaining useful check is to rerun the same stock benchmark without the local patch.
Upstream report
The reproduction, controls, source snippet, and local patch are now public in:
If the project changes the registration policy or asks for additional A/B runs, I will update this article with the follow-up results.