Local AI

MiMo 9B EXL3: 145 tok/s on RTX 4090, 56 tok/s on RTX 3060

Independent EXL3 4.0 bpw measurements for Xiaomi MiMo-V2.6-Distill-Qwen-9B across RTX 4090, RTX 3060 12GB, mixed-GPU layer split, long context, TabbyAPI, vision, and tool calling.

Approximately 8 min read

I converted Xiaomi’s MiMo-V2.6-Distill-Qwen-9B to EXL3 4.0 bpw because this is exactly the size where EXL3 should be practical: small enough to fit comfortably on consumer GPUs, large enough to be useful, and interesting enough that serving behavior matters more than whether the conversion merely completes.

The result is a roughly 6.4 GiB artifact that loads in ExLlamaV3 and serves successfully through TabbyAPI.

The headline result is simple:

On this machine, the same MiMo 9B EXL3 artifact generated about 145 tok/s on an RTX 4090, about 50-56 tok/s on an RTX 3060 12GB, and about 100-108 tok/s when split across the 4090 and 3060.

The mixed result is slower than 4090-only because this model already fits on the faster card. The more useful finding is that the 3060 result is not a barely-working fallback: the model remains comfortably interactive and passed 16K retrieval on a 12GB GPU without CPU model offload.

The quantized model and raw benchmark JSON are published on Hugging Face:

Test environment

The main serving tests used:

Component Configuration
Model XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B
Quantization EXL3 4.0 bpw
Runtime ExLlamaV3 1.5.2
API server TabbyAPI
PyTorch 2.10.0 + cu128
Primary GPU RTX 4090 24GB
Secondary GPU RTX 3060 12GB
KV cache FP16
Sampling for context tests temperature 0
Tool format Qwen3.5
API OpenAI-compatible chat completions

These are local smoke measurements, not a standardized benchmark suite. The goal was to answer practical deployment questions: does it load, how fast does it decode at different context sizes, can a 12GB card run it, does heterogeneous splitting work, and do important model features survive quantization?

RTX 4090 context scaling

The first sweep kept the model on the RTX 4090 and increased prompt length from roughly 1K to 64K tokens.

Each prompt contained a unique needle near the end. The response had to begin with the exact hidden code, which gives a basic retrieval/correctness check rather than timing an arbitrary long prompt with no verification.

Target context Actual prompt Prefill Decode Total time Needle
~1K 1,038 4,943 tok/s 145.2 tok/s 0.91 s PASS
~4K 4,104 8,922 tok/s 144.1 tok/s 1.09 s PASS
~8K 8,199 9,212 tok/s 140.5 tok/s 1.45 s PASS
~16K 16,392 8,957 tok/s 137.3 tok/s 2.37 s PASS
~32K 32,541 4,828 tok/s 127.2 tok/s 7.31 s PASS
~64K 64,587 6,577 tok/s 111.3 tok/s 10.48 s PASS

All six needle checks passed.

The decode curve is more useful than a single short-context number. On this setup, generation stayed near 145 tok/s at short context, remained above 137 tok/s around 16K, and was still above 111 tok/s around 64K.

Prefill throughput is more variable. It should not be treated as a stable hardware score from this one pass because kernel compilation, autotuning, cache state, and prompt shape can all move it.

RTX 3060 12GB: fully usable, not just loadable

I then excluded the RTX 4090 from model placement and forced the model onto the RTX 3060 12GB.

Vision was disabled for this accessibility test. There was no CPU model offload and no model layer placed on the 4090.

Context Prefill Decode Needle
~1K 1,479 tok/s 53.5 tok/s PASS
~8K 1,511 tok/s 49.0 tok/s PASS
~16K 1,440 tok/s 50.3 tok/s PASS

A separate normal-generation request completed naturally at 55.8 tok/s.

Observed RTX 3060 memory during the tests was about 6.4-6.6 GiB total GPU usage. That is the total observed device memory at measurement time, not an isolated model-only allocation.

The practical result is the important one:

A 12GB RTX 3060 can run this MiMo 9B EXL3 quant entirely on-GPU at roughly 50 tok/s while still handling a ~16K retrieval prompt.

That is fast enough to be useful for interactive local inference.

What happens if I split it across the 4090 and 3060?

The machine also has both GPUs, so I tested a conventional ExLlamaV3 layer split:

gpu_split: [4, 8]
tensor_parallel: false

This is deliberately an awkward heterogeneous configuration: a fast RTX 4090 plus a substantially slower RTX 3060 connected over PCIe.

Context 4090 only 4090 + 3060 3060 only
~1K 145.2 tok/s 106.2 tok/s 53.5 tok/s
~8K 140.5 tok/s 102.4 tok/s 49.0 tok/s
~16K 137.3 tok/s 100.7 tok/s 50.3 tok/s

A normal short generation on the mixed split reached 108.1 tok/s.

After model load, total observed memory was about 4.74 GiB on the RTX 4090 and 2.65 GiB on the RTX 3060.

For this specific model, using both GPUs is a bad performance optimization. The model already fits on the 4090, so putting part of the critical path on a slower GPU introduces an unnecessary penalty.

The result is still useful for a different reason: heterogeneous layer splitting works, and the resulting speed is still practical. That matters more when a future model is too large for the 4090 alone.

I would summarize the placement choices this way:

Model fits on 4090:
    use 4090 only

Only a 3060 12GB available:
    3060 only is surprisingly viable

Model no longer fits on 4090:
    heterogeneous layer split is a workable fallback

TabbyAPI validation

I did not stop at a direct ExLlamaV3 load.

The artifact was served through TabbyAPI, and the following checks passed:

The tool-call test asked the model to call a synthetic get_weather function for Toronto. TabbyAPI parsed a structured function call with the expected city argument.

For vision, I generated a simple white image with a large red square and asked for the square’s color. The model returned red.

Those tests are small, but they answer a useful distribution question: this is not just a set of quantized tensors that load in a custom script. The artifact works through an API server and preserves tool and vision paths.

Repetition and long-generation smoke

Quantized models occasionally get reported with repetition or looping behavior, so I also ran a small long-generation stability test.

The test used six requests:

Decode throughput stayed around 145-146 tok/s on the RTX 4090.

I checked repeated 4-grams and repeated 12-word windows. Under the published heuristic:

obvious loop =
    repeated 4-gram ratio > 0.20
    OR
    any 12-word window repeated >= 3 times

the result was:

obvious loops: 0 / 6

Four of the six requests reached the 1,024-token ceiling, so this is not an EOS-behavior benchmark and it is not proof that the model cannot loop under other prompts or sampling settings.

It is simply a useful smoke result: I did not observe a clear repeated n-gram degeneration in these six long outputs.

Two conversion quirks were worth documenting

The conversion exposed two source-metadata details that are useful if someone tries to reproduce the EXL3 build.

Processor metadata

The source processor metadata declared:

Qwen2VLImageProcessor

The current ExLlamaV3 Qwen3.5 loader expected:

Qwen2VLImageProcessorFast

I normalized that metadata in the local conversion copy. The source weights were not changed.

MTP is declared, but the tensors are absent

The source configuration declared:

text_config.mtp_num_hidden_layers = 1

but the source safetensors index contained no MTP/NextN tensors.

If the declared MTP side model is left enabled, the conversion eventually tries to load an MTP tensor that is not present.

The local conversion copy therefore disabled the absent MTP side model before compiling the final EXL3 artifact.

Again, this is a metadata normalization for conversion; it does not alter the source model weights.

What I would actually run

For this exact 9B model:

RTX 4090: use the 4090 alone. It is the fastest configuration by a large margin.

RTX 3060 12GB: use it directly. Roughly 50-56 tok/s is much better than I expected from a 12GB Ampere card for this model.

4090 + 3060: do not split a model that already fits on the 4090 just because both GPUs are available. The mixed result is slower. Save heterogeneous splitting for models that need the extra memory.

The larger point is why EXL3 is interesting for local inference.

A single quant artifact can cover a surprisingly wide hardware range: fast 24GB enthusiast hardware, an older 12GB consumer card, and a mixed-GPU fallback. The runtime behavior is different on each, but all three configurations were usable.

The raw measurement files, model card, and quantized artifact are available in the Hugging Face repository linked above.

Sources and further reading

Continue reading