Local AI

Adding a Second GPU Made My 4090 Slower on Qwen3.8

An independent Qwen3.8 Flash-Next benchmark on RTX 4090 + RTX 3060 showing how auto-fit, manual MoE placement, quantization, and MTP changed decode speed.

Approximately 8 min read

I started this experiment with a simple assumption:

If a large MoE model does not fit in VRAM, adding another 12 GB GPU should help.

The machine was deliberately mismatched:

Component Hardware
Primary GPU RTX 4090 24 GB
Secondary GPU RTX 3060 12 GB
CPU Ryzen 5 7600
System RAM 128 GB
Runtime llama.cpp
OS Ubuntu

The model was Qwen3.8 Flash-Next, first using Unsloth’s UD-Q4_K_XL GGUF.

That quant is roughly 111 GB, so even with 36 GB of combined VRAM, a large fraction of the weights still has to live outside GPU memory.

I expected the 3060 to reduce CPU offload and give at least a small speedup.

It reduced CPU offload.

It also made inference slower.

The first surprise: 36 GB of VRAM was slower than 24 GB

I started with llama.cpp auto-fit and compared the 4090 alone with the 4090 + 3060.

For batch-1 decode:

Configuration Decode throughput
RTX 4090 only 26.88 tok/s
RTX 4090 + RTX 3060, auto-fit 22.18 tok/s

Adding the 3060 reduced decode throughput by about 17.5%.

Prefill also dropped:

Configuration pp512
RTX 4090 only ~347 tok/s
RTX 4090 + RTX 3060, auto-fit ~318 tok/s

This was not a one-run fluke. A five-run tg256 check produced tightly separated samples.

4090 only:

26.7558
26.9572
26.8872
26.9213
26.8802

Dual-GPU auto-fit:

22.2219
22.2269
22.2169
22.1496
22.0643

The slower second GPU was not merely failing to help. It was on the critical path.

What auto-fit was doing

The interesting part came from looking at the actual placement.

With the dual-GPU Q4 configuration, auto-fit assigned roughly:

RTX 4090: blocks 0-13
RTX 3060: blocks 14-48

That was the clue.

The 3060 was not acting as a small overflow device. It was handling a large part of the model.

For two similar GPUs, aggressively filling both devices can be reasonable. For a 4090 + 3060 pair, the optimization problem is different.

The useful objective is not simply:

maximize VRAM usage

It is closer to:

minimize work on the slow GPU unless the saved CPU/RAM traffic is worth the extra device cost.

Those are not the same objective.

Manual placement recovered the loss

I then experimented with llama.cpp’s placement controls, especially:

-ngl
-ncmoe
-sm layer
-ts

A much more asymmetric configuration worked better:

-ngl 999
-ncmoe 32
-sm layer
-ts 8/1

That shifted most regular layer work back toward the 4090.

The resulting decode speed was:

27.32 tok/s

So the Q4 comparison became:

Configuration Decode
4090 only 26.88 tok/s
Dual GPU auto-fit 22.18 tok/s
Dual GPU manual placement 27.32 tok/s

Manual placement was about 23.2% faster than dual-GPU auto-fit.

The five-run manual result was also stable:

27.3461
27.3088
27.3482
27.2700
27.3306

The result does not prove that one specific llama.cpp heuristic is universally wrong. It does show that on this heterogeneous pair, the memory-fit solution was not the latency-optimal solution.

ncmoe alone was not the magic knob

One thing that surprised me was how little ncmoe mattered while auto-fit remained enabled.

I swept:

ncmoe = 0
12
24
36
48

Decode stayed around 22 tok/s.

Once fit was disabled and placement became explicit, the behavior became much easier to interpret.

With all routed experts kept on CPU, I changed the layer split between the two GPUs:

4090:3060 split Decode
8:1 20.49 tok/s
6:1 19.61 tok/s
4:1 19.87 tok/s
2:1 19.12 tok/s
1:1 18.49 tok/s

The broad trend was clear: giving the 3060 more whole-layer work hurt batch-1 decode.

For this machine, the 3060 behaved more like a capacity device than an equal compute partner.

Then I changed the quant

The next experiment mattered more than the placement tuning.

I switched from:

UD-Q4_K_XL

to:

UD-IQ3_XXS

The GGUF size dropped from roughly 111 GB to about 82 GB.

On the RTX 4090 alone, the change was dramatic:

Quant 4090-only decode
UD-Q4_K_XL ~26.88 tok/s
UD-IQ3_XXS 37.89 tok/s

That is roughly a 41% decode improvement.

Prefill also improved substantially:

Q4 pp512:  ~347 tok/s
IQ3 pp512: ~529 tok/s

This was the biggest practical win in the entire experiment.

Not MTP.

Not tensor splitting.

Not Flash Attention.

Reducing the amount of model traffic that had to be serviced outside the fastest memory path mattered more.

IQ3 changed the value of the second GPU

I repeated the dual-GPU tests with IQ3.

The five-run tg256 results were:

Configuration Mean decode
4090-only auto-fit 37.89 ± 0.19 tok/s
4090 + 3060 auto-fit 32.17 ± 0.56 tok/s
Manual dual, ncmoe=24, ts=8/1 37.40 ± 0.53 tok/s

This time manual tuning recovered almost all of the dual-GPU loss, but it still did not beat the 4090 alone.

That was an important change.

As the target model became faster and more of its useful working set fit efficiently around the 4090, the value of assigning target-model work to the 3060 shrank further.

Single-GPU manual placement also lost to auto-fit

I also tried forcing whole expert blocks onto the 4090 instead of letting single-GPU auto-fit choose its own partial overflow.

The usable points looked like this:

ncmoe=31   34.23 tok/s
ncmoe=30   34.60 tok/s
ncmoe=29   34.77 tok/s
ncmoe=28   load failure

auto-fit   37.89 tok/s

So on a single 4090, auto-fit was actually doing something useful that the coarse whole-block placement could not reproduce.

That is an important nuance: the same auto-fit machinery that was poor for the heterogeneous two-GPU latency objective was still the best option I found for the single-GPU IQ3 case.

What about MTP?

I also tested Qwen3.8’s MTP speculative decoding using an Unsloth llama.cpp build and a separate MTP sidecar.

The obvious heterogeneous-GPU layout was:

RTX 4090 -> target model
RTX 3060 -> MTP drafter

For the slower Q4 target this worked reasonably well.

Using a Q4_K_M MTP sidecar with a short draft length:

Q4 target baseline:          25.6 tok/s
Q4 target + MTP + ngram:     30.7 tok/s

That was about a 20% improvement for the tested prompt.

The draft acceptance rate was around 75%, with a mean accepted length a little above three tokens.

So in the Q4 case, the 3060 finally had a useful job: instead of slowing the target model down, it could act as a dedicated speculative drafter.

Then IQ3 changed the economics again.

On the newer Unsloth build:

IQ3 target baseline:         36.1 tok/s
IQ3 + 3060 MTP:              32.0 tok/s

The drafter became too expensive relative to the faster target path.

Speculative decoding only helps when the draft path is cheap enough compared with target verification. A second GPU is not automatically a useful draft GPU.

The 26.8 GiB PLE table was not the decode bottleneck

Qwen3.8 Flash-Next also contains a very large PLE / n-gram embedding table.

In this GGUF:

per_layer_token_embd.weight
physical size: ~26.82 GiB

llama.cpp can lazy-load this structure, so I tested:

lazy-mode auto
lazy-mode on
lazy-mode off

The Q4 decode measurements were:

Lazy mode Decode
auto 26.794 tok/s
on 26.774 tok/s
off 26.855 tok/s

Essentially identical.

But process memory behavior changed dramatically.

Roughly:

lazy auto: ~75 GiB RSS
lazy off:  ~103 GiB RSS

Keeping the huge PLE table resident increased mapped memory by roughly the size of the table without producing a meaningful decode gain.

For this workload, the default lazy behavior was the sensible choice.

Flash Attention was not the answer either

On IQ3 I explicitly compared default Flash Attention behavior with FA disabled.

Five-run decode averages:

FA auto: 37.892 ± 0.192 tok/s
FA off:  37.780 ± 0.155 tok/s

The difference was noise-level.

Again, the dominant bottleneck was elsewhere.

The practical result

The final picture looked like this:

Configuration Decode
Q4_K_XL, dual auto-fit 22.18 tok/s
Q4_K_XL, 4090-only 26.88 tok/s
Q4_K_XL, manual dual 27.32 tok/s
Q4_K_XL + 3060 MTP drafter 30.7 tok/s
IQ3_XXS, dual auto-fit 32.17 tok/s
IQ3_XXS, manual dual 37.40 tok/s
IQ3_XXS, 4090-only 37.89 tok/s

The biggest improvement did not come from using more hardware.

It came from moving less data through the wrong hardware.

What this changed for me

It is tempting to think about local LLM hardware as a simple sum of VRAM:

24 GB + 12 GB = 36 GB

For large MoE inference, that mental model is too simple.

At least four things mattered in this experiment:

  1. Where the routed expert weights live
  2. How much weight traffic still comes from system RAM
  3. How much compute gets assigned to the slower GPU
  4. Whether the secondary GPU actually shortens the critical path

A memory placement that uses more VRAM can still be a worse latency configuration.

Likewise, a second GPU can be useful for one quant and useless for another.

The best Q4 setup I found could justify the 3060 as a speculative drafter.

The best IQ3 setup I found was simpler:

4090 -> target
3060 -> mostly idle

That looks wasteful.

It was also faster.

The broader lesson

The interesting result is not that an RTX 3060 is slower than an RTX 4090.

That part is obvious.

The useful result is that maximizing GPU utilization and maximizing inference speed were different optimization problems on this heterogeneous system.

Auto-fit solved a capacity problem.

It did not always solve the latency problem I actually cared about.

For mixed consumer GPUs running large MoE models, the next optimization target should probably not be “how do I fill every last gigabyte of VRAM?”

A better question is:

Which tensors reduce the critical path when they move to the second GPU, and which ones simply move the bottleneck there?

That is the experiment I want to continue next.

Sources and further reading

Continue reading