Qwen3.8-27B on RTX 4090: 5,000 tok/s Prefill Reproduced
An independent Ubuntu RTX 4090 reproduction of NInfer's Qwen3.8-27B A8 claims reaches 4,999.5 tok/s pp2048, 500.9 tok/s edit decode, and exact retrieval at 260K context.
Approximately 9 min read
A LocalLLM post claimed something unusually aggressive for a dense 27B-class model on one RTX 4090: roughly 5,000 tokens per second of prompt prefill, more than 100 tokens per second of speculative decode, and usable context extending to roughly 262K tokens.
The repository behind the post exposes enough implementation detail to test the claim directly. I reproduced it on Ubuntu instead of Windows, using the exact public branch and the author’s A8 NInfer artifact.
The short version: the 5,000 tok/s prefill claim reproduces almost exactly, the long-context timings reproduce through 128K, and a 260,096-token needle test returns the correct value on a 24GB RTX 4090.
The important caveat is that the largest decode numbers are workload-dependent. N-gram drafting is spectacular when the answer mostly copies the prompt, such as code or document editing. It is not a 500 tok/s general-chat mode.
Test system and exact artifacts
| Component | Configuration |
|---|---|
| GPU | NVIDIA GeForce RTX 4090 24GB |
| Secondary GPU | RTX 3060 12GB, not used for NInfer |
| CPU | AMD Ryzen 5 7600 |
| System RAM | 124 GiB visible to Linux |
| OS | Ubuntu, Linux 6.8.0-94-generic |
| NVIDIA driver | 590.48.01 |
| CUDA | 13.1 |
| NInfer branch | JGamboa/ninfer-4090-windows, feat/bonsai-ternary |
| Commit | f20ee462ee1734cc374b5e69b23ae37a1277ef37 |
| Model artifact | qwen3_8_27b_a8.ninfer |
| Artifact size | 20,437,521,664 bytes |
| Artifact SHA-256 | 49bf76388e139defbe416c8cc0079abc1c78b59be2a41ff56ae960aeacf5fc19 |
The repository describes this branch as Windows-developed and says the inherited Linux build had not been revalidated for the newer work. Linux compatibility itself was therefore part of the reproduction.
I built it natively for Ada:
cmake -S . -B build-sm89 -G Ninja \
-DCMAKE_BUILD_TYPE=Release \
-DCMAKE_CUDA_COMPILER=/usr/local/cuda-13.1/bin/nvcc \
-DCMAKE_CUDA_ARCHITECTURES=89 \
-DNINFER_BUILD_APPS=ON \
-DNINFER_BUILD_BENCHMARKS=ON
cmake --build build-sm89 --target ninfer ninfer_bench --parallel 4
The build completed successfully on Linux. The most expensive compilation unit was the large Ada attention small_t.cu path, which spent several minutes in cicc and ptxas, but no source patch was required.
The 5,000 tok/s prefill claim reproduces
The author’s README publishes approximately 4,436 tok/s at pp512 and 5,008 tok/s at pp2048 for the A8 artifact.
I ran the repository’s own public benchmark path:
build-sm89/bench/ninfer_bench \
--weights qwen3_8_27b_a8.ninfer \
-p 512,2048 -n 128 -r 3 \
--kv-dtype int8
| Test | RAMGPT Ubuntu RTX 4090 | Published result |
|---|---|---|
| pp512 | 4,462.39 ± 6.24 tok/s | 4,436 tok/s |
| pp2048 | 4,999.51 ± 36.97 tok/s | 5,008 tok/s |
| tg128, no speculation | 53.54 tok/s | 47.0 tok/s |
The pp2048 result is only about 0.17% below the published 5,008 tok/s figure.
That is close enough that the short-prompt A8 prefill result should be considered independently reproduced on this Linux system.
Why the A8 artifact matters
This performance does not come from the ordinary Qwen3.8 NInfer artifact alone.
The A8 model card describes a separate prefill execution route. For sufficiently large prompt steps, activations are quantized per token and per 64 channels and fed into int8 tensor-core matrix multiplication against the model’s groupwise weight representation. Decode remains a different problem. Autoregressive decode is bandwidth-sensitive and processes very few new tokens per step, while prompt prefill exposes much larger matrix operations.
pp2048 prefill: ~5,000 tok/s
tg128 baseline: 53.5 tok/s
A single “tokens per second” number is therefore not a useful description of the runtime.
MTP helps ordinary prose, but acceptance controls the gain
For a 512-token technical prose response with thinking disabled and greedy decoding:
prompt tokens: 39
generated tokens: 512
decode speed: 117.0 tok/s
MTP drafted tokens: 575
MTP accepted tokens: 319
MTP acceptance: 55.5%
accepted length: 2.66 tok/round
This is a more representative free-form generation number than the largest edit benchmark below.
MTP proposes multiple future tokens and wins only when the target model accepts them. In this prose run, 55.5% draft acceptance produced an average accepted length of 2.66 tokens per speculative round.
N-gram drafting is extremely fast on edit workloads
The most dramatic result came from a deliberately edit-shaped task. I constructed a 2,469-token configuration file, asked the model to return the entire file while changing only one line, and limited generation to 512 tokens.
| Mode | Prefill | Decode | Acceptance | Accepted length |
|---|---|---|---|---|
| MTP3 | 4.51k tok/s | 172.2 tok/s | 100% | 3.99 tok/round |
| MTP3 + n-gram chain | 4.52k tok/s | 500.9 tok/s | 100% | 15.48 tok/round |
| The n-gram verify window was 15 tokens and n-gram acceptance was 100%. |
The 500.9 tok/s result is real, but it needs the right interpretation. The answer mostly reproduces text that already exists in the context. The n-gram path extends an MTP proposal with long spans copied from that context, then lets the target verify them.
That is fundamentally different from asking the model to invent 500 tokens of new prose.
For local coding agents, patch application, structured-document rewriting, JSON regeneration, and configuration edits, this behavior is still highly relevant. It just should not be advertised as a universal chat throughput number.
DFlash2 reached 282 tok/s on a code prompt
I also tested the branch’s DFlash2 route with a 12-token draft window on a code-generation prompt:
generated tokens: 512
decode speed: 282.4 tok/s
DFlash2 draft window: 12
drafted tokens: 738
accepted tokens: 449
acceptance rate: 60.8%
acceptance length: 8.24 tok/round
The prompt asked for a complete Python LRU cache implementation using a dictionary and doubly linked list.
This is a single workload, not a general DFlash2 average, and it should not be divided by the tg128 result as though the workloads were matched. It does show that the optimized branch can turn a predictable code-generation stream into long accepted speculative runs on the same GPU.
Long-context prefill also reproduces
The repository includes needle-in-a-haystack fixtures containing the target:
ORCHID=493817; COLOR=COBALT
I ran the A8 artifact at four depths using the repository’s fixtures, max context 262144, prefill chunk 1024, rk4v4-e8 KV, thinking disabled, greedy decoding, and MTP3.
All four runs returned the exact needle.
| Fixture | Actual prompt tokens | Model elapsed | Prefill speed | Decode speed | Retrieval |
|---|---|---|---|---|---|
| 8K | 7,680 | 1.7 s | 4.74k tok/s | 159.4 tok/s | PASS |
| 64K | 64,512 | 16.9 s | 3.84k tok/s | 146.4 tok/s | PASS |
| 128K | 130,048 | 41.9 s | 3.11k tok/s | 137.0 tok/s | PASS |
| 256K | 260,096 | 115.2 s | 2.26k tok/s | 119.7 tok/s | PASS |
The published A8 figures were approximately 1.6 seconds at 8K, 17.2 seconds at 64K, and 43.0 seconds at 128K.
The Ubuntu reproduction landed at 1.7, 16.9, and 41.9 seconds respectively.
The 256K fixture goes beyond the headline table and is the more useful stress test. At 260,096 prompt tokens, the engine reported:
GPU weights used: 16.7 GiB
GPU sequence used: 4.68 GiB
free after startup: 1.14 GiB
model elapsed: 115.2 s
prefill speed: 2.26k tok/s
decode speed: 119.7 tok/s
retrieval: PASS
During the run, the RTX 4090 was at 100% utilization with roughly 22.9GB of its 24.6GB reported memory in use.
So this is not merely a configuration that accepts a 262144 max-context flag. The model actually processed roughly 260K prompt tokens and retrieved the correct key near the card’s practical memory limit.
Model time and wall time are not the same thing
Each long-context test launched a fresh CLI process. Total wall time therefore includes model loading, weight upload, graph setup, and process startup.
| Prompt | Model elapsed | End-to-end wall time |
|---|---|---|
| 8K | 1.7 s | 6.68 s |
| 64K | 16.9 s | 21.87 s |
| 128K | 41.9 s | 46.94 s |
| 256K | 115.2 s | 120.31 s |
A persistent server amortizes most of that fixed startup cost across requests.
The Linux result matters
The fork’s README explicitly says the newer branch is developed and validated on Windows and that the Linux build is inherited rather than revalidated.
On this Ubuntu machine, the branch:
- configured with CUDA 13.1 and sm_89;
- compiled the q8/A8 prefill kernels;
- linked ninfer and ninfer_bench;
- loaded the published A8 artifact with the exact expected SHA-256;
- reproduced the short-prompt benchmark;
- ran MTP, n-gram, and DFlash2 paths;
- completed 260K-context inference; and
- returned the correct needle at every tested depth.
That does not prove every server, vision, concurrency, or Windows-specific path is equally portable. It establishes that the main optimized text-inference path works on this Linux RTX 4090 setup without a source patch.
What this reproduction establishes
It does show
- The published A8 artifact and exact branch run on Ubuntu with an RTX 4090.
- The pp2048 result reproduces at 4,999.5 tok/s, essentially matching the published 5,008 tok/s.
- MTP can substantially improve practical free-form decode when draft acceptance is high enough.
- DFlash2 can produce long accepted draft sequences on code-shaped output.
- N-gram chaining can exceed 500 tok/s when the task mostly reproduces existing context.
- The engine can process 260,096 prompt tokens on a 24GB 4090 using rk4v4-e8 KV and still retrieve the exact needle.
It does not show
- that normal chat runs at 500 tok/s;
- that every prompt gets high MTP or DFlash2 acceptance;
- that n-gram speculation helps prose with little reusable context;
- that A8 prefill is quality-identical under every downstream task;
- that one 260K needle test proves perfect long-context reasoning; or
- that NInfer is universally faster than llama.cpp, vLLM, or other runtimes across all workloads.
Those require matched model, quantization, prompt, sampling, KV, context, and concurrency controls.
Why this matters beyond a benchmark screenshot
On one 24GB consumer GPU, the same runtime can occupy several very different operating points:
short-prompt ingestion
-> ~5,000 tok/s pp2048
free-form prose with MTP
-> ~117 tok/s in this run
predictable code with DFlash2
-> ~282 tok/s in this run
small edits to large existing text
-> ~501 tok/s with MTP + n-gram
very long context
-> exact retrieval at 260K
-> ~2.26k tok/s prefill
-> ~120 tok/s decode after the prompt
That combination is unusually well aligned with local-agent workloads.
Agents spend a lot of time ingesting repositories, logs, generated files, previous tool output, and structured context. They also frequently return modified versions of material already present in the prompt. Those are exactly the cases where fast prefill, large KV capacity, speculative verification, and context-copy drafting can matter more than a conventional one-number decode benchmark.
The result I would carry forward is narrower than “NInfer does 500 tok/s.”
On an RTX 4090, the optimized Qwen3.8-27B A8 path independently reproduced ~5k tok/s short-prompt prefill, remained usable at ~260K context with exact needle retrieval, and showed very large decode gains when the workload gives speculation something predictable to verify.
That is a more useful system property than the peak screenshot.