Benchmarks

MiMo-V2.6 QuantBench: Public Q4 and IQ4_XS Compared

We tested public MiMo-V2.6-Distill-Qwen-9B GGUFs on one RTX 4090, comparing loadability, perplexity, speed, exact-task accuracy, and BF16 fidelity.

Approximately 14 min read

MiMo-V2.6-Distill-Qwen-9B appeared on Hugging Face with multiple GGUF conversions within hours. That created an unusually useful quantization experiment: several public artifacts with the same model name, similar advertised quant types, and different conversion or calibration choices were available at nearly the same time.

Instead of assuming that all files called Q4_K_M or IQ4_XS are interchangeable, I downloaded the public artifacts and tested them on one machine under the same llama.cpp build.

The result was more interesting than a simple leaderboard.

One public Q4 artifact did not load at all because its metadata declared an MTP/NextN layer that was not present in the released checkpoint. Two other Q4 files that appeared to be separate releases were almost byte-for-byte identical in behavior and structure. A larger Q4 used a different tensor mix and shifted which reasoning tasks failed.

The IQ4_XS comparison was even better. RAMGPT used a small 64-chunk task-mixed iMatrix. Bartowski published a much larger 518-chunk calibration. The larger calibration won the cleaned exact-task pack, while perplexity, decode throughput, and BF16 behavior agreement stayed much closer.

The main lesson is not that one quant is universally best.

It is this:

Quantization changes decision boundaries in non-monotonic ways. Raw benchmark score, language-model loss, runtime speed, and behavioral fidelity to BF16 are different measurements and should not be collapsed into one number.

This note documents the artifact audit, the matched RTX 4090 tests, and the parts of the benchmark that broke before the models did.

Test environment

All local measurements in this article used the same machine and runtime:

Component Configuration
GPU NVIDIA GeForce RTX 4090 24 GB
Secondary GPU RTX 3060 present but excluded from these tests
CPU AMD Ryzen 5 7600
llama.cpp commit ce8caa6, version 0.4.1-dev
Backend CUDA
Flash attention on
GPU layers full offload to the 4090
Main GPU 0
Chat template original Xiaomi MiMo Jinja forced for matched chat tests
Sampling temperature 0 for deterministic task probes

The BF16 reference was converted from upstream revision:

f2773fb482ac3dd047a4af4003b86e56b7225d0d

The upstream config declares mtp_num_hidden_layers: 1, but the released safetensors contain no MTP/NextN tensors. The working BF16 reference therefore used the converter’s --no-nextn path and reports 32 blocks.

The original chat-template SHA256 used throughout the matched chat tests was:

59a64ebb4df6d1489d09a91267cf3ceb106162d4a893c4f84833cfb8c897ff63

First: audit the public artifacts before benchmarking them

A benchmark is pointless if two candidate files are actually the same artifact or if a model cannot load.

I therefore started with a small market census and inspected file sizes, LFS hashes, block metadata, NextN metadata, embedded templates, and tensor-type distributions before running any quality test.

The Q4 files found

The main public Q4 candidates tested were:

Publisher Artifact File size
RAMGPT MiMo-V2.6-Distill-Qwen-9B-Q4_K_M.gguf 5,629,105,344 bytes
prithivMLmods MiMo-V2.6-Distill-Qwen-9B.Q4_K_M.gguf 5,629,105,344 bytes
CNWPlayer mimo-v2.6-distill-qwen-9b-q4_k_m.gguf 5,629,101,344 bytes
sweetprincessluna mimo-v2.6-distill-qwen-9b-q4_k_m.gguf 5,629,101,344 bytes
bartowski MiMo-V2.6-Distill-Qwen-9B-Q4_K_M.gguf 5,841,049,120 bytes

The CNWPlayer and sweetprincessluna Q4 files had the exact same SHA256:

00c05c4b98dc927802993aca5a908a14f2f803a825567abc6d5f7cedc8aef521

They are therefore one artifact for benchmarking purposes, not two independent quants.

One Q4 failed before the benchmark started

The CNWPlayer/sweetprincessluna artifact reported:

qwen35.block_count = 33
qwen35.nextn_predict_layers = 1

It also did not contain the upstream chat template.

On the tested llama.cpp build, loading failed with:

check_tensor_dims: tensor 'blk.32.attn_norm.weight' not found

This is the same failure mode encountered during RAMGPT’s first conversion attempt. The config claims one MTP layer, but the released model weights do not include the corresponding block.

This is an artifact-specific compatibility result on the tested llama.cpp commit. It should not be generalized into a claim about the publishers or about future runtime versions.

RAMGPT and prithivMLmods were effectively the same Q4 path

The RAMGPT and prithivMLmods files had:

A binary comparison found only seven byte differences across the entire 5.63 GB files. The first 64 MiB were identical.

Their behavioral results also tracked each other almost perfectly.

For practical purposes, these two files represent the same standard Q4_K_M conversion path in this experiment.

Bartowski’s Q4 was genuinely different: it was about 212 MB larger and used a different tensor-type distribution.

Q4 perplexity and speed

Perplexity used wiki.test.raw from the public llama.cpp validation dataset.

The 128-chunk run used:

context per chunk: 512
chunks: 128
batch: 512
ubatch: 512
GPU: RTX 4090 only

The wiki.test.raw SHA256 used here was:

173c87a53759e0201f33e0ccf978e510c2042d7f2cb78229d9a50d79b9e7dd08

This test set was not part of RAMGPT’s own iMatrix calibration. I did not attempt a full overlap audit against every third-party calibration corpus, so the perplexity numbers should be treated as a matched measurement rather than a perfectly controlled held-out study for every publisher.

Q4 PPL-128

Model PPL
BF16 9.0972 ± 0.1350
RAMGPT Q4_K_M 9.0548 ± 0.1321
bartowski Q4_K_M 9.3122 ± 0.1391

The obvious trap is to say that the RAMGPT Q4 is “better than BF16” because its measured PPL is slightly lower.

That would be an overclaim.

The confidence ranges overlap, and low-bit noise can move local predictions in either direction. A quantized model can occasionally score better on a finite sample without becoming a better model in any general sense.

Q4 RTX 4090 throughput

llama-bench used prompt length 512, generation length 128, three repetitions, full GPU offload, and flash attention.

Model pp512 tg128
BF16 7,553.8 tok/s 55.32 tok/s
RAMGPT Q4_K_M 9,377.8 tok/s 143.46 tok/s
prithiv Q4_K_M 9,418.3 tok/s 143.49 tok/s
bartowski Q4_K_M 9,696.7 tok/s 139.67 tok/s

Bartowski’s larger Q4 prefills about 3.4% faster than RAMGPT’s Q4 in this run, but decodes about 2.6% slower.

RAMGPT and prithivMLmods are again effectively tied.

A small exact-task pack showed why raw accuracy is not fidelity

I also ran a deterministic 29-interaction probe pack covering:

This is an exploratory diagnostic pack, not a general-purpose model leaderboard.

The pure final-answer accuracy after separating answer scoring from repetition hygiene was:

Model Correct
BF16 23 / 29
RAMGPT Q4_K_M 25 / 29
prithiv Q4_K_M 25 / 29
bartowski Q4_K_M 24 / 29

Again, the Q4 files scoring above BF16 does not mean quantization improved the model.

The error sets moved.

For example, the standard RAMGPT/prithiv Q4 path missed the modular-arithmetic and TSP probes that BF16 answered correctly. Bartowski’s Q4 preserved those but failed a GCD item and a proof-verification item that the standard Q4 answered correctly.

If BF16’s pass/fail pattern is used only as a behavioral reference, not as a truth oracle, the picture changes:

Q4 Same pass/fail outcome as BF16
RAMGPT Q4_K_M 21 / 29
bartowski Q4_K_M 26 / 29

Bartowski’s Q4 behaved more like BF16 even though its raw exact-task score was one point lower.

That distinction is important enough to state explicitly:

A quant can gain benchmark points by moving a BF16 error across the decision boundary. That is higher task accuracy on that sample, but it is not higher fidelity to the original model.

The more interesting comparison: IQ4_XS versus IQ4_XS

The Q4 market audit was useful, but the cleanest comparison was between two public IQ4_XS files.

Both used iMatrix calibration. Both loaded correctly. Both preserved the upstream template. Their file sizes differed by only about 31 MB.

IQ4_XS File size
RAMGPT 5,196,437,024 bytes
bartowski 5,227,304,480 bytes

The difference is only about 0.6% of the file size.

Yet the calibration strategies were very different.

The two iMatrices were built very differently

Both importance matrices covered:

RAMGPT’s iMatrix used 64 chunks.

Bartowski’s published iMatrix metadata reports 518 chunks.

That is more than eight times as many calibration chunks.

RAMGPT calibration

RAMGPT deliberately used a task-mixed corpus rather than Wiki alone. It interleaved:

The calibration-text SHA256 was:

f4ce34541473d074e9e355772dee137dafe59330c6edb2c94f2397a19babfa11

The resulting iMatrix had:

Statistic RAMGPT
chunks 64
tensors 248
layers 32
minimum active rate 97.78%
median active rate 100%
median normalized entropy 81.82%

Bartowski calibration

Bartowski publishes the calibration file directly as MiMo-V2.6-Distill-Qwen-9B-calibration-v6.txt.

The file is about 1.11 MB and includes long-form natural and academic prose. I did not attempt to classify every section of the corpus, so I will not reduce its distribution to a single label.

Its iMatrix statistics were:

Statistic bartowski
chunks 518
tensors 248
layers 32
minimum active rate 98.63%
median active rate 100%
median normalized entropy 81.80%

The median entropy was nearly identical despite the large difference in chunk count.

That alone is a useful warning against treating “more calibration tokens” as automatically more diverse activation coverage.

The IQ4 tensor policies were also different

RAMGPT’s final IQ4_XS contained:

Tensor type Count
F32 177
IQ4_XS 217
Q5_K 32
Q6_K 1

Bartowski’s file contained:

Tensor type Count
F32 225
IQ4_XS 167
Q5_K 5
Q6_K 19
Q8_0 11

Tensor counts are not proportional to byte size because many F32 tensors are small norms, biases, or state parameters.

Still, the policy difference is obvious. Bartowski preserves more tensors at higher precision, while RAMGPT pushes more of the large model body into IQ4_XS and relies on the iMatrix plus a smaller set of promoted tensors.

That makes the head-to-head more than a calibration-corpus comparison. It is a comparison of two complete IQ4 quantization policies.

IQ4_XS perplexity was effectively tied

The same PPL-128 test produced:

Model PPL
BF16 9.0972 ± 0.1350
RAMGPT IQ4_XS 9.0966 ± 0.1344
bartowski IQ4_XS 9.0520 ± 0.1333

These numbers are too close to support a strong quality ranking.

All three intervals overlap heavily.

The practical result is simpler:

Neither IQ4_XS showed a clear language-model-loss collapse on this 128-chunk sample.

IQ4_XS speed was also almost a tie

The same RTX 4090 benchmark gave:

IQ4_XS pp512 tg128
RAMGPT 10,088.1 tok/s 153.52 tok/s
bartowski 10,252.4 tok/s 153.47 tok/s

Bartowski prefills about 1.6% faster in this run.

Decode throughput is effectively identical.

The smaller RAMGPT file therefore did not buy a meaningful decode-speed advantage here.

The cleaned hard pack favored the 518-chunk iMatrix

The first exploratory task pack was not sensitive enough, so I built a harder deterministic pack spanning:

During audit I removed two prompts from the score because their wording could reasonably induce the model to repeat the expression rather than evaluate it. This left 48 scored interactions.

The cleaned final-answer accuracy was:

Model Correct
BF16 39 / 48
RAMGPT IQ4_XS 36 / 48
bartowski IQ4_XS 43 / 48

This was the clearest result in the IQ4 comparison.

Bartowski’s larger calibration and more conservative tensor policy won the exact-task pack by seven answers over the RAMGPT IQ4.

The category-level pattern also showed that the difference was not isolated to one task family.

Bartowski was especially strong on the arithmetic portion. RAMGPT lost more points on logic and TSP.

But behavioral fidelity to BF16 was nearly tied

Using only the pass/fail pattern of the BF16 reference:

IQ4_XS Same pass/fail outcome as BF16
RAMGPT 41 / 48
bartowski 40 / 48

That is essentially a tie on this small pack.

This produces a useful three-way distinction:

A single “quality” number would hide all of this.

What the 64-versus-518 result says

RAMGPT’s original hypothesis was that a smaller task-aware calibration corpus might outperform a much larger generic-style calibration by exercising code, shell, security, structured output, and tool-call behavior directly.

This experiment does not support that hypothesis strongly enough.

The 64-chunk iMatrix produced excellent activation coverage and nearly the same PPL and speed as the 518-chunk alternative, but it lost the cleaned hard exact-task pack.

The honest result is:

For this MiMo-V2.6 9B IQ4_XS experiment, the larger 518-chunk calibration plus more conservative tensor policy produced better raw exact-task accuracy at almost no runtime cost.

That does not prove that 518 chunks are intrinsically required.

The comparison is confounded by the tensor policy. Bartowski did not only use a larger iMatrix; the final GGUF also preserves more tensors in Q6_K, Q8_0, or F32.

A proper ablation would need at least four controlled variants:

64-chunk RAMGPT iMatrix + RAMGPT tensor policy
518-chunk bartowski iMatrix + RAMGPT tensor policy
64-chunk RAMGPT iMatrix + bartowski tensor policy
518-chunk bartowski iMatrix + bartowski tensor policy

That would separate calibration-data effects from tensor-policy effects.

This market comparison tells us that the complete bartowski recipe did better on the hard pack. It does not tell us which part of the recipe deserves the credit.

The reasoning-boundary leak was not a quantization failure

A separate probe repeatedly exposed a literal </think> boundary in message.content for some prompts.

This occurred in:

Other prompts correctly separated reasoning_content and final content.

The behavior therefore cannot be attributed to IQ4 quantization.

The Xiaomi upstream quickstart explicitly uses a model-specific mimo reasoning parser in SGLang. RAMGPT reported the reproducible BF16 behavior to the upstream model discussion separately.

For QuantBench, the boundary leak was recorded as parser/runtime hygiene and not counted as an IQ4-specific quality loss.

What I would take away from this market snapshot

Several conclusions survived the artifact audit.

1. Validate the file before benchmarking it

A model named Q4_K_M can still carry invalid architecture metadata.

The CNWPlayer/sweetprincessluna artifact tested here failed before inference because the GGUF declared a nonexistent 33rd block.

A benchmark table that never attempts a real model load can miss the most important failure.

2. Deduplicate by hash and structure

Two different repository names do not guarantee two independent quantizations.

The CNWPlayer and sweetprincessluna files were byte-identical.

RAMGPT and prithivMLmods were not hash-identical, but differed by only seven bytes and produced essentially the same behavior.

3. More bits do not guarantee lower PPL on every finite sample

Bartowski’s Q4 file was larger, yet its measured PPL on this sample was higher.

RAMGPT’s Q4 measured slightly below BF16.

Neither observation should be converted into a universal model-quality claim.

4. Raw accuracy and BF16 fidelity answer different questions

A quantization can correct a BF16 error by moving a decision boundary.

It can also break a BF16 success.

That is why an evaluation should report both:

task accuracy against the oracle
behavioral agreement with the BF16 reference

5. For IQ4_XS, bartowski won this exact-task round

The strongest result in this note is the cleaned 48-interaction hard pack:

BF16              39 / 48
RAMGPT IQ4_XS     36 / 48
bartowski IQ4_XS  43 / 48

The difference is large enough that it should not be hidden behind a story about RAMGPT’s task-aware calibration.

The larger published calibration plus more conservative tensor mix did better on this pack.

6. But the next experiment should be an ablation, not a rematch

The interesting research question is no longer “whose IQ4 is better?”

It is:

How much of the result comes from calibration data, and how much comes from tensor precision allocation?

That can be tested directly by crossing the two iMatrices with the two tensor policies under the same BF16 source.

That experiment would turn this market snapshot into a controlled quantization study.

Reproducibility notes

The RAMGPT GGUF and iMatrix used in this article are published at:

https://huggingface.co/ramgpt/MiMo-V2.6-Distill-Qwen-9B-GGUF

The tested RAMGPT IQ4_XS SHA256 was:

bc2d7e526df4bcc24fd5b1554c1f7c7062569706e2ce7835a8ebbfb6087e00e4

The tested RAMGPT iMatrix SHA256 was:

a9163b33e811a68f077a916d6b1c067e06de3590f075808dd60ce950cdaece66

The tested bartowski IQ4_XS LFS SHA256 was:

eccfbc188e71dec8350fbdd5af898d2a50691ac67e077327c52baa9da1e91d90

The tested bartowski iMatrix LFS SHA256 was:

946b69b405f8b1efb291fe9caacf4207d7d83fc36e9901478b67edb9f95cc735

The public wiki.test.raw input used for the matched PPL runs was SHA256:

173c87a53759e0201f33e0ccf978e510c2042d7f2cb78229d9a50d79b9e7dd08

All measurements in this article are specific to the named artifacts, the listed llama.cpp commit, and the tested RTX 4090 configuration.

They are not intended as permanent rankings of the publishers or of future revisions.

That is exactly why the artifact hashes are included.

Sources and further reading

Continue reading