AI Hardware

llama.cpp Q4_K Prefill Fusion: Gate, Up, and SwiGLU in One CUDA Path

A new llama.cpp CUDA PR targets duplicate activation quantization and intermediate FP32 traffic in dense Q4_K MLP prefill, with contributor-reported gains from 4.1% to 14.7%.

Approximately 6 min read

Quantized weights do not automatically make every part of inference cheap.

A fresh llama.cpp CUDA pull request is a good example. PR #28702 targets a very specific dense MLP path: separate Q4_K Gate and Up projections followed by split SwiGLU during prompt processing. The proposal is not a new quantization format and it does not change model outputs. It tries to remove work that exists between already-quantized weights and the next projection.

The contributor’s summary is unusually concrete: quantize the shared activation to Q8_1 once, compute Gate and Up in one CUDA kernel, and apply SwiGLU after the full K reduction. The claimed benefit comes from eliminating duplicated activation quantization, intermediate FP32 projection round trips, and a separate SwiGLU kernel launch.

That makes this PR interesting beyond its benchmark table. It shows where quantized inference can still pay a large orchestration and memory-traffic tax even when the matrix weights themselves are only four bits.

PR #28702 is still a draft and open as of September 11. All performance numbers below are contributor measurements, not RAMGPT benchmarks.

What the existing path is paying for

A dense gated MLP commonly has two projections consuming the same input activation:

                 ┌─> Gate projection ─┐
input activation ┤                    ├─> SwiGLU ─> Down projection
                 └─> Up projection ───┘

With Q4_K weights on CUDA, the activation side still needs the representation expected by the quantized matrix path. If Gate and Up are processed independently, shared input work can be repeated. Their projection results also have to exist somewhere before the elementwise SwiGLU operation consumes them.

PR #28702 attacks those boundaries rather than changing the mathematical graph.

The proposed fused path does three things together:

  1. quantizes the shared F32 input activation to Q8_1 once;
  2. consumes separate dense Q4_K Gate and Up tensors inside one CUDA kernel;
  3. applies the split SwiGLU after the complete K reduction instead of launching another kernel over intermediate projection outputs.

The Down projection remains unchanged. The PR also leaves the FP32 output after SwiGLU in place, so this is not an attempt to fuse the entire feed-forward block.

Why one Q8_1 conversion matters

The Gate and Up projections read the same activation. Treating them as unrelated operations means the runtime can repeat preparation work for identical input data.

The PR reuses llama.cpp’s existing Q8_1 quantization machinery but shares that quantized activation across both projections. This is the sort of optimization that is easy to miss if performance analysis stops at the weight format.

Q4_K describes how the weights are stored and processed. It does not mean the surrounding activation conversion, temporary storage, synchronization, or kernel launches disappear.

For local inference users, that distinction matters. A model can fit comfortably in VRAM and use an efficient weight quantization while prompt processing still leaves performance on the table in the plumbing around GEMM.

The bigger target is intermediate traffic

The PR description explicitly calls out FP32 projection-output round trips.

Without fusion, Gate and Up produce intermediate results that are materialized before SwiGLU combines them. Materialization means writes, later reads, and another kernel boundary. On a modern GPU, avoiding those transfers can be as important as reducing arithmetic.

The fused kernel keeps the relationship between the two projections closer to where the values are produced. Only after the K reduction is complete does it apply SwiGLU.

This is a useful constraint. Applying the nonlinear operation too early would change the computation. Fusion is valuable only when it preserves the original reduction semantics.

What hardware and model paths are actually covered

This is not a universal Q4_K acceleration patch.

The PR currently scopes the fused path to separate dense Q4_K Gate/Up weights, F32 activations, split SwiGLU, and supported NVIDIA Ampere-or-newer configurations. HIP and MUSA keep their existing implementations.

There is also an escape hatch:

GGML_CUDA_DISABLE_FUSION=1

That matters for a draft optimization because fusion changes execution strategy even when the intended math is identical. A runtime-level disable switch makes regression isolation substantially easier.

The contributor says the kernel uses a fusion-specific configuration of 64 rows and 128 threads per CTA, targeting two resident blocks per SM. The PR correctly qualifies that actual residency depends on hardware and compiled resource use.

That last point is important when interpreting cross-GPU results. A kernel configuration that improves scheduling and traffic on one architecture does not imply an identical percentage gain on another.

The reported numbers are already showing that asymmetry

The contributor reports the following pp16K improvements with Q4_K_M models on a DGX Spark:

Model Fused layers Contributor-reported pp16K gain
Qwen3.8-27B 64/64 +11.1%
Qwen3.6-27B 64/64 +10.5%
Llama-3.1-8B-Instruct 32/32 +14.7%
Qwen3-8B 36/36 +13.5%

The PR reports identical perplexity for those comparisons.

On an RTX 5090, the submitted result for Qwen3.6-27B is +4.1% at pp16K, again with identical reported perplexity.

The obvious lesson is not that one GPU is “better” for fusion. It is that the benefit is sensitive to the surrounding execution bottleneck. Saving launches and memory traffic matters most when those costs occupy a meaningful fraction of total prompt-processing time.

The same optimization can therefore produce very different percentage gains depending on GPU architecture, model shape, matrix dimensions, occupancy, and how efficiently the unfused kernels already run.

Why this is specifically a prefill story

The measurements in the PR are prompt-processing (pp) numbers, not token-generation (tg) numbers.

That distinction matters.

Prefill processes many input tokens in parallel and exposes large matrix operations. Decode usually operates at much smaller effective batch dimensions per generated token and can be dominated by a different balance of memory bandwidth, kernel overhead, KV-cache access, and synchronization.

A double-digit pp16K improvement does not imply a double-digit end-to-end chat-speed improvement, and it definitely should not be translated into an equivalent tg claim.

For workloads with long prompts, document ingestion, RAG context, or repeated large-prefix processing, prompt throughput may matter a lot. For a short interactive chat dominated by long generation, the visible benefit can be much smaller.

This PR also illustrates a useful benchmarking rule

When an inference optimization is described as “faster,” the first questions should be:

PR #28702 provides enough of those details to make its early measurements interpretable, but the results are still contributor measurements from a draft PR.

RAMGPT has not reproduced these numbers independently.

Until the patch is merged and tested across more hardware, it should be treated as an implementation result under active development rather than a guaranteed speedup for every Q4_K model.

What to watch before calling this a llama.cpp feature

The PR is still open and marked draft. That means stock llama.cpp users should not assume the fused path is present in their current build.

The important follow-ups are straightforward:

Until it merges, stock llama.cpp users should not assume the optimization exists in a release build.

The strongest takeaway today is narrower and more useful: Q4_K weight quantization does not eliminate duplicated activation work or intermediate FP32 traffic. PR #28702 shows that fusing Gate, Up, and SwiGLU can attack those costs directly, and the contributor’s early results suggest the payoff is strongly hardware-dependent.

Sources and further reading

Continue reading