MiMo-V2.6 QuantBench: Public Q4 and IQ4_XS Compared
We tested public MiMo-V2.6-Distill-Qwen-9B GGUFs on one RTX 4090, comparing loadability, perplexity, speed, exact-task accuracy, and BF16 fidelity.
Approximately 14 min read
MiMo-V2.6-Distill-Qwen-9B appeared on Hugging Face with multiple GGUF conversions within hours. That created an unusually useful quantization experiment: several public artifacts with the same model name, similar advertised quant types, and different conversion or calibration choices were available at nearly the same time.
Instead of assuming that all files called Q4_K_M or IQ4_XS are interchangeable, I downloaded the public artifacts and tested them on one machine under the same llama.cpp build.
The result was more interesting than a simple leaderboard.
One public Q4 artifact did not load at all because its metadata declared an MTP/NextN layer that was not present in the released checkpoint. Two other Q4 files that appeared to be separate releases were almost byte-for-byte identical in behavior and structure. A larger Q4 used a different tensor mix and shifted which reasoning tasks failed.
The IQ4_XS comparison was even better. RAMGPT used a small 64-chunk task-mixed iMatrix. Bartowski published a much larger 518-chunk calibration. The larger calibration won the cleaned exact-task pack, while perplexity, decode throughput, and BF16 behavior agreement stayed much closer.
The main lesson is not that one quant is universally best.
It is this:
Quantization changes decision boundaries in non-monotonic ways. Raw benchmark score, language-model loss, runtime speed, and behavioral fidelity to BF16 are different measurements and should not be collapsed into one number.
This note documents the artifact audit, the matched RTX 4090 tests, and the parts of the benchmark that broke before the models did.
Test environment
All local measurements in this article used the same machine and runtime:
| Component | Configuration |
|---|---|
| GPU | NVIDIA GeForce RTX 4090 24 GB |
| Secondary GPU | RTX 3060 present but excluded from these tests |
| CPU | AMD Ryzen 5 7600 |
| llama.cpp | commit ce8caa6, version 0.4.1-dev |
| Backend | CUDA |
| Flash attention | on |
| GPU layers | full offload to the 4090 |
| Main GPU | 0 |
| Chat template | original Xiaomi MiMo Jinja forced for matched chat tests |
| Sampling | temperature 0 for deterministic task probes |
The BF16 reference was converted from upstream revision:
f2773fb482ac3dd047a4af4003b86e56b7225d0d
The upstream config declares mtp_num_hidden_layers: 1, but the released safetensors contain no MTP/NextN tensors. The working BF16 reference therefore used the converter’s --no-nextn path and reports 32 blocks.
The original chat-template SHA256 used throughout the matched chat tests was:
59a64ebb4df6d1489d09a91267cf3ceb106162d4a893c4f84833cfb8c897ff63
First: audit the public artifacts before benchmarking them
A benchmark is pointless if two candidate files are actually the same artifact or if a model cannot load.
I therefore started with a small market census and inspected file sizes, LFS hashes, block metadata, NextN metadata, embedded templates, and tensor-type distributions before running any quality test.
The Q4 files found
The main public Q4 candidates tested were:
| Publisher | Artifact | File size |
|---|---|---|
| RAMGPT | MiMo-V2.6-Distill-Qwen-9B-Q4_K_M.gguf |
5,629,105,344 bytes |
| prithivMLmods | MiMo-V2.6-Distill-Qwen-9B.Q4_K_M.gguf |
5,629,105,344 bytes |
| CNWPlayer | mimo-v2.6-distill-qwen-9b-q4_k_m.gguf |
5,629,101,344 bytes |
| sweetprincessluna | mimo-v2.6-distill-qwen-9b-q4_k_m.gguf |
5,629,101,344 bytes |
| bartowski | MiMo-V2.6-Distill-Qwen-9B-Q4_K_M.gguf |
5,841,049,120 bytes |
The CNWPlayer and sweetprincessluna Q4 files had the exact same SHA256:
00c05c4b98dc927802993aca5a908a14f2f803a825567abc6d5f7cedc8aef521
They are therefore one artifact for benchmarking purposes, not two independent quants.
One Q4 failed before the benchmark started
The CNWPlayer/sweetprincessluna artifact reported:
qwen35.block_count = 33
qwen35.nextn_predict_layers = 1
It also did not contain the upstream chat template.
On the tested llama.cpp build, loading failed with:
check_tensor_dims: tensor 'blk.32.attn_norm.weight' not found
This is the same failure mode encountered during RAMGPT’s first conversion attempt. The config claims one MTP layer, but the released model weights do not include the corresponding block.
This is an artifact-specific compatibility result on the tested llama.cpp commit. It should not be generalized into a claim about the publishers or about future runtime versions.
RAMGPT and prithivMLmods were effectively the same Q4 path
The RAMGPT and prithivMLmods files had:
- identical file size
- 32 blocks
- no NextN metadata
- exact upstream chat-template match
- identical tensor-type counts
A binary comparison found only seven byte differences across the entire 5.63 GB files. The first 64 MiB were identical.
Their behavioral results also tracked each other almost perfectly.
For practical purposes, these two files represent the same standard Q4_K_M conversion path in this experiment.
Bartowski’s Q4 was genuinely different: it was about 212 MB larger and used a different tensor-type distribution.
Q4 perplexity and speed
Perplexity used wiki.test.raw from the public llama.cpp validation dataset.
The 128-chunk run used:
context per chunk: 512
chunks: 128
batch: 512
ubatch: 512
GPU: RTX 4090 only
The wiki.test.raw SHA256 used here was:
173c87a53759e0201f33e0ccf978e510c2042d7f2cb78229d9a50d79b9e7dd08
This test set was not part of RAMGPT’s own iMatrix calibration. I did not attempt a full overlap audit against every third-party calibration corpus, so the perplexity numbers should be treated as a matched measurement rather than a perfectly controlled held-out study for every publisher.
Q4 PPL-128
| Model | PPL |
|---|---|
| BF16 | 9.0972 ± 0.1350 |
| RAMGPT Q4_K_M | 9.0548 ± 0.1321 |
| bartowski Q4_K_M | 9.3122 ± 0.1391 |
The obvious trap is to say that the RAMGPT Q4 is “better than BF16” because its measured PPL is slightly lower.
That would be an overclaim.
The confidence ranges overlap, and low-bit noise can move local predictions in either direction. A quantized model can occasionally score better on a finite sample without becoming a better model in any general sense.
Q4 RTX 4090 throughput
llama-bench used prompt length 512, generation length 128, three repetitions, full GPU offload, and flash attention.
| Model | pp512 | tg128 |
|---|---|---|
| BF16 | 7,553.8 tok/s | 55.32 tok/s |
| RAMGPT Q4_K_M | 9,377.8 tok/s | 143.46 tok/s |
| prithiv Q4_K_M | 9,418.3 tok/s | 143.49 tok/s |
| bartowski Q4_K_M | 9,696.7 tok/s | 139.67 tok/s |
Bartowski’s larger Q4 prefills about 3.4% faster than RAMGPT’s Q4 in this run, but decodes about 2.6% slower.
RAMGPT and prithivMLmods are again effectively tied.
A small exact-task pack showed why raw accuracy is not fidelity
I also ran a deterministic 29-interaction probe pack covering:
- arithmetic
- modular arithmetic
- fractions and combinatorics
- exact TSP
- set cover
- formal-logic validity
- proof verification
- Python semantics
- simple network-security knowledge
- JSON
- structured tool calling
- multi-turn state
This is an exploratory diagnostic pack, not a general-purpose model leaderboard.
The pure final-answer accuracy after separating answer scoring from repetition hygiene was:
| Model | Correct |
|---|---|
| BF16 | 23 / 29 |
| RAMGPT Q4_K_M | 25 / 29 |
| prithiv Q4_K_M | 25 / 29 |
| bartowski Q4_K_M | 24 / 29 |
Again, the Q4 files scoring above BF16 does not mean quantization improved the model.
The error sets moved.
For example, the standard RAMGPT/prithiv Q4 path missed the modular-arithmetic and TSP probes that BF16 answered correctly. Bartowski’s Q4 preserved those but failed a GCD item and a proof-verification item that the standard Q4 answered correctly.
If BF16’s pass/fail pattern is used only as a behavioral reference, not as a truth oracle, the picture changes:
| Q4 | Same pass/fail outcome as BF16 |
|---|---|
| RAMGPT Q4_K_M | 21 / 29 |
| bartowski Q4_K_M | 26 / 29 |
Bartowski’s Q4 behaved more like BF16 even though its raw exact-task score was one point lower.
That distinction is important enough to state explicitly:
A quant can gain benchmark points by moving a BF16 error across the decision boundary. That is higher task accuracy on that sample, but it is not higher fidelity to the original model.
The more interesting comparison: IQ4_XS versus IQ4_XS
The Q4 market audit was useful, but the cleanest comparison was between two public IQ4_XS files.
Both used iMatrix calibration. Both loaded correctly. Both preserved the upstream template. Their file sizes differed by only about 31 MB.
| IQ4_XS | File size |
|---|---|
| RAMGPT | 5,196,437,024 bytes |
| bartowski | 5,227,304,480 bytes |
The difference is only about 0.6% of the file size.
Yet the calibration strategies were very different.
The two iMatrices were built very differently
Both importance matrices covered:
- 248 tensors
- all 32 model layers
- 512-token chunks
RAMGPT’s iMatrix used 64 chunks.
Bartowski’s published iMatrix metadata reports 518 chunks.
That is more than eight times as many calibration chunks.
RAMGPT calibration
RAMGPT deliberately used a task-mixed corpus rather than Wiki alone. It interleaved:
- general natural text
- real llama.cpp C/C++ and Python
- shell/config/JSON
- OWASP security documentation
- MiMo-style chat and tool-call examples
The calibration-text SHA256 was:
f4ce34541473d074e9e355772dee137dafe59330c6edb2c94f2397a19babfa11
The resulting iMatrix had:
| Statistic | RAMGPT |
|---|---|
| chunks | 64 |
| tensors | 248 |
| layers | 32 |
| minimum active rate | 97.78% |
| median active rate | 100% |
| median normalized entropy | 81.82% |
Bartowski calibration
Bartowski publishes the calibration file directly as MiMo-V2.6-Distill-Qwen-9B-calibration-v6.txt.
The file is about 1.11 MB and includes long-form natural and academic prose. I did not attempt to classify every section of the corpus, so I will not reduce its distribution to a single label.
Its iMatrix statistics were:
| Statistic | bartowski |
|---|---|
| chunks | 518 |
| tensors | 248 |
| layers | 32 |
| minimum active rate | 98.63% |
| median active rate | 100% |
| median normalized entropy | 81.80% |
The median entropy was nearly identical despite the large difference in chunk count.
That alone is a useful warning against treating “more calibration tokens” as automatically more diverse activation coverage.
The IQ4 tensor policies were also different
RAMGPT’s final IQ4_XS contained:
| Tensor type | Count |
|---|---|
| F32 | 177 |
| IQ4_XS | 217 |
| Q5_K | 32 |
| Q6_K | 1 |
Bartowski’s file contained:
| Tensor type | Count |
|---|---|
| F32 | 225 |
| IQ4_XS | 167 |
| Q5_K | 5 |
| Q6_K | 19 |
| Q8_0 | 11 |
Tensor counts are not proportional to byte size because many F32 tensors are small norms, biases, or state parameters.
Still, the policy difference is obvious. Bartowski preserves more tensors at higher precision, while RAMGPT pushes more of the large model body into IQ4_XS and relies on the iMatrix plus a smaller set of promoted tensors.
That makes the head-to-head more than a calibration-corpus comparison. It is a comparison of two complete IQ4 quantization policies.
IQ4_XS perplexity was effectively tied
The same PPL-128 test produced:
| Model | PPL |
|---|---|
| BF16 | 9.0972 ± 0.1350 |
| RAMGPT IQ4_XS | 9.0966 ± 0.1344 |
| bartowski IQ4_XS | 9.0520 ± 0.1333 |
These numbers are too close to support a strong quality ranking.
All three intervals overlap heavily.
The practical result is simpler:
Neither IQ4_XS showed a clear language-model-loss collapse on this 128-chunk sample.
IQ4_XS speed was also almost a tie
The same RTX 4090 benchmark gave:
| IQ4_XS | pp512 | tg128 |
|---|---|---|
| RAMGPT | 10,088.1 tok/s | 153.52 tok/s |
| bartowski | 10,252.4 tok/s | 153.47 tok/s |
Bartowski prefills about 1.6% faster in this run.
Decode throughput is effectively identical.
The smaller RAMGPT file therefore did not buy a meaningful decode-speed advantage here.
The cleaned hard pack favored the 518-chunk iMatrix
The first exploratory task pack was not sensitive enough, so I built a harder deterministic pack spanning:
- integer arithmetic and modular arithmetic
- fractions and GCD/LCM
- Python semantics
- formal logic
- networking and protocol knowledge
- six exact four-node TSP instances
- local rule switching
- structured tool calling
- multi-turn state
During audit I removed two prompts from the score because their wording could reasonably induce the model to repeat the expression rather than evaluate it. This left 48 scored interactions.
The cleaned final-answer accuracy was:
| Model | Correct |
|---|---|
| BF16 | 39 / 48 |
| RAMGPT IQ4_XS | 36 / 48 |
| bartowski IQ4_XS | 43 / 48 |
This was the clearest result in the IQ4 comparison.
Bartowski’s larger calibration and more conservative tensor policy won the exact-task pack by seven answers over the RAMGPT IQ4.
The category-level pattern also showed that the difference was not isolated to one task family.
Bartowski was especially strong on the arithmetic portion. RAMGPT lost more points on logic and TSP.
But behavioral fidelity to BF16 was nearly tied
Using only the pass/fail pattern of the BF16 reference:
| IQ4_XS | Same pass/fail outcome as BF16 |
|---|---|
| RAMGPT | 41 / 48 |
| bartowski | 40 / 48 |
That is essentially a tie on this small pack.
This produces a useful three-way distinction:
- raw exact-task accuracy: bartowski clearly ahead
- BF16 pass/fail fidelity: effectively tied
- PPL and decode speed: effectively tied
A single “quality” number would hide all of this.
What the 64-versus-518 result says
RAMGPT’s original hypothesis was that a smaller task-aware calibration corpus might outperform a much larger generic-style calibration by exercising code, shell, security, structured output, and tool-call behavior directly.
This experiment does not support that hypothesis strongly enough.
The 64-chunk iMatrix produced excellent activation coverage and nearly the same PPL and speed as the 518-chunk alternative, but it lost the cleaned hard exact-task pack.
The honest result is:
For this MiMo-V2.6 9B IQ4_XS experiment, the larger 518-chunk calibration plus more conservative tensor policy produced better raw exact-task accuracy at almost no runtime cost.
That does not prove that 518 chunks are intrinsically required.
The comparison is confounded by the tensor policy. Bartowski did not only use a larger iMatrix; the final GGUF also preserves more tensors in Q6_K, Q8_0, or F32.
A proper ablation would need at least four controlled variants:
64-chunk RAMGPT iMatrix + RAMGPT tensor policy
518-chunk bartowski iMatrix + RAMGPT tensor policy
64-chunk RAMGPT iMatrix + bartowski tensor policy
518-chunk bartowski iMatrix + bartowski tensor policy
That would separate calibration-data effects from tensor-policy effects.
This market comparison tells us that the complete bartowski recipe did better on the hard pack. It does not tell us which part of the recipe deserves the credit.
The reasoning-boundary leak was not a quantization failure
A separate probe repeatedly exposed a literal </think> boundary in message.content for some prompts.
This occurred in:
- BF16
- Q4_K_M
- RAMGPT IQ4_XS
- bartowski IQ4_XS
Other prompts correctly separated reasoning_content and final content.
The behavior therefore cannot be attributed to IQ4 quantization.
The Xiaomi upstream quickstart explicitly uses a model-specific mimo reasoning parser in SGLang. RAMGPT reported the reproducible BF16 behavior to the upstream model discussion separately.
For QuantBench, the boundary leak was recorded as parser/runtime hygiene and not counted as an IQ4-specific quality loss.
What I would take away from this market snapshot
Several conclusions survived the artifact audit.
1. Validate the file before benchmarking it
A model named Q4_K_M can still carry invalid architecture metadata.
The CNWPlayer/sweetprincessluna artifact tested here failed before inference because the GGUF declared a nonexistent 33rd block.
A benchmark table that never attempts a real model load can miss the most important failure.
2. Deduplicate by hash and structure
Two different repository names do not guarantee two independent quantizations.
The CNWPlayer and sweetprincessluna files were byte-identical.
RAMGPT and prithivMLmods were not hash-identical, but differed by only seven bytes and produced essentially the same behavior.
3. More bits do not guarantee lower PPL on every finite sample
Bartowski’s Q4 file was larger, yet its measured PPL on this sample was higher.
RAMGPT’s Q4 measured slightly below BF16.
Neither observation should be converted into a universal model-quality claim.
4. Raw accuracy and BF16 fidelity answer different questions
A quantization can correct a BF16 error by moving a decision boundary.
It can also break a BF16 success.
That is why an evaluation should report both:
task accuracy against the oracle
behavioral agreement with the BF16 reference
5. For IQ4_XS, bartowski won this exact-task round
The strongest result in this note is the cleaned 48-interaction hard pack:
BF16 39 / 48
RAMGPT IQ4_XS 36 / 48
bartowski IQ4_XS 43 / 48
The difference is large enough that it should not be hidden behind a story about RAMGPT’s task-aware calibration.
The larger published calibration plus more conservative tensor mix did better on this pack.
6. But the next experiment should be an ablation, not a rematch
The interesting research question is no longer “whose IQ4 is better?”
It is:
How much of the result comes from calibration data, and how much comes from tensor precision allocation?
That can be tested directly by crossing the two iMatrices with the two tensor policies under the same BF16 source.
That experiment would turn this market snapshot into a controlled quantization study.
Reproducibility notes
The RAMGPT GGUF and iMatrix used in this article are published at:
https://huggingface.co/ramgpt/MiMo-V2.6-Distill-Qwen-9B-GGUF
The tested RAMGPT IQ4_XS SHA256 was:
bc2d7e526df4bcc24fd5b1554c1f7c7062569706e2ce7835a8ebbfb6087e00e4
The tested RAMGPT iMatrix SHA256 was:
a9163b33e811a68f077a916d6b1c067e06de3590f075808dd60ce950cdaece66
The tested bartowski IQ4_XS LFS SHA256 was:
eccfbc188e71dec8350fbdd5af898d2a50691ac67e077327c52baa9da1e91d90
The tested bartowski iMatrix LFS SHA256 was:
946b69b405f8b1efb291fe9caacf4207d7d83fc36e9901478b67edb9f95cc735
The public wiki.test.raw input used for the matched PPL runs was SHA256:
173c87a53759e0201f33e0ccf978e510c2042d7f2cb78229d9a50d79b9e7dd08
All measurements in this article are specific to the named artifacts, the listed llama.cpp commit, and the tested RTX 4090 configuration.
They are not intended as permanent rankings of the publishers or of future revisions.
That is exactly why the artifact hashes are included.