Troubleshooting

llama.cpp Metal Tensor Bug: The 2 GiB Slice Boundary Explained

A fresh llama.cpp Metal PR isolates a large-tensor addressing failure at the 2 GiB slice offset and replaces Tensor API slicing with explicit 64-bit address calculation.

Approximately 6 min read

A GPU backend bug becomes much easier to reason about when the failure has a sharp boundary.

A fresh llama.cpp pull request, #28748, reports exactly that kind of boundary in the Apple Metal Tensor path. The contributor found that kernel_mul_mm can address the wrong memory when the offset from a tensor’s base address to a slice reaches 0x80000000, or 2 GiB. Instead of landing at the requested slice, the Tensor API path was observed writing 4 GiB earlier than expected.

The proposed fix is small but technically revealing: stop asking the Metal Tensor API to derive the slice from the whole output tensor for this path. Compute the destination address explicitly with 64-bit arithmetic and construct the tile tensor from that address.

PR #28748 is open as of September 12, 2026. The measurements and reproduction results below are the contributor’s, not RAMGPT benchmarks.

The failure is about an offset, not a 2 GiB allocation limit

The important detail is easy to misread. The report does not say that Metal cannot allocate a tensor larger than 2 GiB. It says the problematic condition appears when the slice start offset relative to the tensor base reaches 2 GiB.

The contributor compared two ways of storing the same matrix-multiplication tile. One path first calculates a destination pointer and then creates a tensor at that pointer. The other wraps the complete output tensor and calls slice(row, column).

Below the boundary, both approaches agreed. At and above 0x80000000, the slice-derived address shifted by exactly -4 GiB in the experiment.

That discontinuity matters. A gradual numerical error would suggest arithmetic precision or accumulation behavior. A clean address jump at a power-of-two boundary instead points toward address/index representation.

Why a large tensor can expose a bug that small tests miss

Consider an output buffer whose base address is B. A tile begins at byte offset O, so the expected address is conceptually:

expected = B + O

For small outputs, O never approaches the reported boundary. Unit tests and ordinary inference workloads can therefore exercise the same kernel for a long time without seeing the failure.

Once a workload produces a sufficiently large tensor, the slice offset can cross 2 GiB even though the underlying GPU allocation itself is valid. That is why a backend may appear stable on smaller models, shorter requests, or smaller intermediate tensors and then fail abruptly when tensor dimensions grow.

This is also a useful troubleshooting distinction: capacity failure and addressing failure are different classes of problem. An out-of-memory condition usually fails allocation or reports memory pressure. A bad slice address can instead produce invalid values or corrupted results while memory is technically available.

The experiment isolates the slice operation

The strongest part of PR #28748 is not merely that the author’s original workload stopped failing. The contributor constructed a controlled experiment around the boundary.

Each dispatch performs the same matrix multiplication. The difference is only how the destination tile is represented.

The reported cases include slice starts 1024 bytes below the boundary, 64 bytes below it, a tile that crosses the boundary, exactly at the boundary, and offsets above it. The direct-address path produced the expected 256 of 256 values in every listed case.

The Tensor API slice path also behaved correctly while the start remained below 2 GiB, including cases where part of the tile crossed the boundary. But when the start became exactly 0x80000000, the expected address received none of the 256 values while the location 4 GiB earlier received all 256. The same result was reported at +64 bytes and +1024 bytes.

That narrows the hypothesis considerably: crossing the boundary somewhere inside a tile was not sufficient. The slice’s starting offset was the trigger in the submitted experiment.

What the patch changes in kernel_mul_mm

The affected code is the GGML_METAL_HAS_TENSOR == true implementation in:

ggml/src/ggml-metal/kernels/mul_mm.metal

The old strategy conceptually wraps the complete output memory as a two-dimensional tensor and stores into a derived slice:

whole tensor -> slice(row, column) -> store tile

The proposed path instead calculates the tile pointer directly using 64-bit address arithmetic and creates the destination tensor at that location:

base pointer + 64-bit offset -> tile tensor -> store tile

This is not a new matrix multiplication algorithm. The multiplication result is still the same cooperative tensor. The change is in how the destination memory for each tile is addressed.

Why the exact -4 GiB displacement is interesting

The contributor reports a step from correct addressing below 0x80000000 to an address 4 GiB early starting exactly at 0x80000000.

RAMGPT should not overstate what that proves. The PR demonstrates the observed behavior in the tested Metal environment; it does not by itself establish the internal implementation of Apple’s Tensor API or prove a specific signed-integer bug inside it.

Still, the boundary is diagnostically valuable. 0x80000000 is where the high bit of a 32-bit signed integer becomes set. A 4 GiB displacement is 0x100000000, the full range of a 32-bit unsigned integer. Those values make integer-width or signed-offset handling a natural hypothesis, but that remains an inference unless the underlying API implementation confirms it.

That distinction is worth preserving: the experiment establishes the boundary and displacement; the mechanism behind the API behavior is not yet proven.

What users might see

The related reproduction was built around a long embedding request. After the patch, the contributor reports an output size of 1024 with zero invalid values for the long request, and the related issue was no longer reproducible in that environment.

If this bug is involved in a real workload, the useful clues are therefore not simply “Metal crashed”. More specific signals would be:

Those clues can separate this problem from VRAM exhaustion, malformed GGUF files, quantization errors, or model-specific architecture support.

Hardware scope is still narrow

The PR’s reported environment is a 13-inch MacBook Air with M5 and 16 GB memory, macOS 26.6.2, and the listed Apple Metal compiler version. The contributor also supplied a targeted probe rather than claiming universal reproduction across Apple GPUs.

That means it would be premature to say every Metal-capable Mac has the same bug. It would also be premature to claim the patch is already an upstream fix: PR #28748 remains open.

For users tracking a production regression, the relevant questions are therefore:

Am I on the Metal Tensor path?
Does failure correlate with tensor/request size?
Can I reproduce below and above the boundary?
Does a non-Metal backend provide a clean control?
Has PR #28748 merged into the llama.cpp build I am using?

The broader engineering lesson

Large-model inference bugs are not always caused by the model.

A GGUF can load correctly. The quantization can be valid. The matrix multiplication can be mathematically correct. Available memory can be sufficient. Yet an intermediate tensor can still fail because a backend API represents or derives a large offset incorrectly.

That is why controlled boundary testing is so useful. Instead of treating “large request fails” as one opaque symptom, vary the dimension that controls the address and look for a transition. A reproducible threshold can turn a vague backend failure into a testable systems hypothesis.

PR #28748 is a particularly clean example: the reported transition is not merely “large versus small.” It is a specific slice-start boundary, a specific displacement, and a patch that bypasses the suspect address derivation while leaving the actual matrix multiplication intact.

Until the PR is reviewed and merged, treat it as an upstream investigation and proposed fix rather than released llama.cpp behavior.

Sources and further reading

Continue reading