Benchmarks

Qwen3.8-27B on RTX 4090: 5,000 tok/s Prefill Reproduced

An independent Ubuntu RTX 4090 reproduction of NInfer's Qwen3.8-27B A8 claims reaches 4,999.5 tok/s pp2048, 500.9 tok/s edit decode, and exact retrieval at 260K context.

Approximately 9 min read

A LocalLLM post claimed something unusually aggressive for a dense 27B-class model on one RTX 4090: roughly 5,000 tokens per second of prompt prefill, more than 100 tokens per second of speculative decode, and usable context extending to roughly 262K tokens.

The repository behind the post exposes enough implementation detail to test the claim directly. I reproduced it on Ubuntu instead of Windows, using the exact public branch and the author’s A8 NInfer artifact.

The short version: the 5,000 tok/s prefill claim reproduces almost exactly, the long-context timings reproduce through 128K, and a 260,096-token needle test returns the correct value on a 24GB RTX 4090.

The important caveat is that the largest decode numbers are workload-dependent. N-gram drafting is spectacular when the answer mostly copies the prompt, such as code or document editing. It is not a 500 tok/s general-chat mode.

Test system and exact artifacts

Component Configuration
GPU NVIDIA GeForce RTX 4090 24GB
Secondary GPU RTX 3060 12GB, not used for NInfer
CPU AMD Ryzen 5 7600
System RAM 124 GiB visible to Linux
OS Ubuntu, Linux 6.8.0-94-generic
NVIDIA driver 590.48.01
CUDA 13.1
NInfer branch JGamboa/ninfer-4090-windows, feat/bonsai-ternary
Commit f20ee462ee1734cc374b5e69b23ae37a1277ef37
Model artifact qwen3_8_27b_a8.ninfer
Artifact size 20,437,521,664 bytes
Artifact SHA-256 49bf76388e139defbe416c8cc0079abc1c78b59be2a41ff56ae960aeacf5fc19

The repository describes this branch as Windows-developed and says the inherited Linux build had not been revalidated for the newer work. Linux compatibility itself was therefore part of the reproduction.

I built it natively for Ada:

cmake -S . -B build-sm89 -G Ninja \
  -DCMAKE_BUILD_TYPE=Release \
  -DCMAKE_CUDA_COMPILER=/usr/local/cuda-13.1/bin/nvcc \
  -DCMAKE_CUDA_ARCHITECTURES=89 \
  -DNINFER_BUILD_APPS=ON \
  -DNINFER_BUILD_BENCHMARKS=ON
cmake --build build-sm89 --target ninfer ninfer_bench --parallel 4

The build completed successfully on Linux. The most expensive compilation unit was the large Ada attention small_t.cu path, which spent several minutes in cicc and ptxas, but no source patch was required.

The 5,000 tok/s prefill claim reproduces

The author’s README publishes approximately 4,436 tok/s at pp512 and 5,008 tok/s at pp2048 for the A8 artifact.

I ran the repository’s own public benchmark path:

build-sm89/bench/ninfer_bench \
  --weights qwen3_8_27b_a8.ninfer \
  -p 512,2048 -n 128 -r 3 \
  --kv-dtype int8
Test RAMGPT Ubuntu RTX 4090 Published result
pp512 4,462.39 ± 6.24 tok/s 4,436 tok/s
pp2048 4,999.51 ± 36.97 tok/s 5,008 tok/s
tg128, no speculation 53.54 tok/s 47.0 tok/s

The pp2048 result is only about 0.17% below the published 5,008 tok/s figure.

That is close enough that the short-prompt A8 prefill result should be considered independently reproduced on this Linux system.

Why the A8 artifact matters

This performance does not come from the ordinary Qwen3.8 NInfer artifact alone.

The A8 model card describes a separate prefill execution route. For sufficiently large prompt steps, activations are quantized per token and per 64 channels and fed into int8 tensor-core matrix multiplication against the model’s groupwise weight representation. Decode remains a different problem. Autoregressive decode is bandwidth-sensitive and processes very few new tokens per step, while prompt prefill exposes much larger matrix operations.

pp2048 prefill: ~5,000 tok/s
tg128 baseline:    53.5 tok/s

A single “tokens per second” number is therefore not a useful description of the runtime.

MTP helps ordinary prose, but acceptance controls the gain

For a 512-token technical prose response with thinking disabled and greedy decoding:

prompt tokens:          39
generated tokens:      512
decode speed:        117.0 tok/s
MTP drafted tokens:    575
MTP accepted tokens:   319
MTP acceptance:       55.5%
accepted length:       2.66 tok/round

This is a more representative free-form generation number than the largest edit benchmark below.

MTP proposes multiple future tokens and wins only when the target model accepts them. In this prose run, 55.5% draft acceptance produced an average accepted length of 2.66 tokens per speculative round.

N-gram drafting is extremely fast on edit workloads

The most dramatic result came from a deliberately edit-shaped task. I constructed a 2,469-token configuration file, asked the model to return the entire file while changing only one line, and limited generation to 512 tokens.

Mode Prefill Decode Acceptance Accepted length
MTP3 4.51k tok/s 172.2 tok/s 100% 3.99 tok/round
MTP3 + n-gram chain 4.52k tok/s 500.9 tok/s 100% 15.48 tok/round
The n-gram verify window was 15 tokens and n-gram acceptance was 100%.

The 500.9 tok/s result is real, but it needs the right interpretation. The answer mostly reproduces text that already exists in the context. The n-gram path extends an MTP proposal with long spans copied from that context, then lets the target verify them.

That is fundamentally different from asking the model to invent 500 tokens of new prose.

For local coding agents, patch application, structured-document rewriting, JSON regeneration, and configuration edits, this behavior is still highly relevant. It just should not be advertised as a universal chat throughput number.

DFlash2 reached 282 tok/s on a code prompt

I also tested the branch’s DFlash2 route with a 12-token draft window on a code-generation prompt:

generated tokens:           512
decode speed:             282.4 tok/s
DFlash2 draft window:        12
drafted tokens:             738
accepted tokens:            449
acceptance rate:           60.8%
acceptance length:          8.24 tok/round

The prompt asked for a complete Python LRU cache implementation using a dictionary and doubly linked list.

This is a single workload, not a general DFlash2 average, and it should not be divided by the tg128 result as though the workloads were matched. It does show that the optimized branch can turn a predictable code-generation stream into long accepted speculative runs on the same GPU.

Long-context prefill also reproduces

The repository includes needle-in-a-haystack fixtures containing the target:

ORCHID=493817; COLOR=COBALT

I ran the A8 artifact at four depths using the repository’s fixtures, max context 262144, prefill chunk 1024, rk4v4-e8 KV, thinking disabled, greedy decoding, and MTP3.

All four runs returned the exact needle.

Fixture Actual prompt tokens Model elapsed Prefill speed Decode speed Retrieval
8K 7,680 1.7 s 4.74k tok/s 159.4 tok/s PASS
64K 64,512 16.9 s 3.84k tok/s 146.4 tok/s PASS
128K 130,048 41.9 s 3.11k tok/s 137.0 tok/s PASS
256K 260,096 115.2 s 2.26k tok/s 119.7 tok/s PASS

The published A8 figures were approximately 1.6 seconds at 8K, 17.2 seconds at 64K, and 43.0 seconds at 128K.

The Ubuntu reproduction landed at 1.7, 16.9, and 41.9 seconds respectively.

The 256K fixture goes beyond the headline table and is the more useful stress test. At 260,096 prompt tokens, the engine reported:

GPU weights used:       16.7 GiB
GPU sequence used:       4.68 GiB
free after startup:      1.14 GiB
model elapsed:         115.2 s
prefill speed:          2.26k tok/s
decode speed:          119.7 tok/s
retrieval:              PASS

During the run, the RTX 4090 was at 100% utilization with roughly 22.9GB of its 24.6GB reported memory in use.

So this is not merely a configuration that accepts a 262144 max-context flag. The model actually processed roughly 260K prompt tokens and retrieved the correct key near the card’s practical memory limit.

Model time and wall time are not the same thing

Each long-context test launched a fresh CLI process. Total wall time therefore includes model loading, weight upload, graph setup, and process startup.

Prompt Model elapsed End-to-end wall time
8K 1.7 s 6.68 s
64K 16.9 s 21.87 s
128K 41.9 s 46.94 s
256K 115.2 s 120.31 s

A persistent server amortizes most of that fixed startup cost across requests.

The Linux result matters

The fork’s README explicitly says the newer branch is developed and validated on Windows and that the Linux build is inherited rather than revalidated.

On this Ubuntu machine, the branch:

That does not prove every server, vision, concurrency, or Windows-specific path is equally portable. It establishes that the main optimized text-inference path works on this Linux RTX 4090 setup without a source patch.

What this reproduction establishes

It does show

It does not show

Those require matched model, quantization, prompt, sampling, KV, context, and concurrency controls.

Why this matters beyond a benchmark screenshot

On one 24GB consumer GPU, the same runtime can occupy several very different operating points:

short-prompt ingestion
    -> ~5,000 tok/s pp2048

free-form prose with MTP
    -> ~117 tok/s in this run

predictable code with DFlash2
    -> ~282 tok/s in this run

small edits to large existing text
    -> ~501 tok/s with MTP + n-gram

very long context
    -> exact retrieval at 260K
    -> ~2.26k tok/s prefill
    -> ~120 tok/s decode after the prompt

That combination is unusually well aligned with local-agent workloads.

Agents spend a lot of time ingesting repositories, logs, generated files, previous tool output, and structured context. They also frequently return modified versions of material already present in the prompt. Those are exactly the cases where fast prefill, large KV capacity, speculative verification, and context-copy drafting can matter more than a conventional one-number decode benchmark.

The result I would carry forward is narrower than “NInfer does 500 tok/s.”

On an RTX 4090, the optimized Qwen3.8-27B A8 path independently reproduced ~5k tok/s short-prompt prefill, remained usable at ~260K context with exact needle retrieval, and showed very large decode gains when the workload gives speculation something predictable to verify.

That is a more useful system property than the peak screenshot.

Sources and further reading

Continue reading