Local AI

llama.cpp v0.6.0 Extended Batch API: A Migration Checklist

llama.cpp v0.6.0 adds llama_batch_ext and llama_process for mixed token and embedding input. Here is what C API integrators need to test before upgrading.

Approximately 5 min read

llama.cpp v0.6.0 changes a meaningful part of the public inference interface. The October 5 release introduces an extended batching API, llama_batch_ext, together with llama_process(). It allows inputs that are not adequately described by a flat array of token IDs.

If you run llama-server as a normal OpenAI-compatible endpoint, you may not need any application change. If you maintain an inference wrapper, multimodal adapter, embedded application, or speculative-decoding integration that calls the C API directly, this release warrants an explicit compatibility review.

What changed in the input model?

The familiar text-only path starts with token IDs, positions and sequence metadata. For a plain chat completion, that is often sufficient. An encoder or multi-token prediction component may instead supply embeddings or model-specific state alongside ordinary tokens.

The new batch representation is designed for those combinations. The v0.6.0 release describes mixed token/embedding batches and per-token state embeddings for MTP and deepstack use cases. Those are input representation capabilities, not an unconditional promise that any model can consume any embedding.

A conceptual comparison:

Earlier mental model:
token IDs + positions + sequence IDs -> inference

Extended path:
tokens or embeddings + positions + sequence/state data
    -> llama_batch_ext
    -> llama_process()
    -> model execution

This diagram deliberately omits C function arguments. Use the tagged public header for exact signatures, rather than copying an example from an earlier review revision.

Why downstream compatibility needs deliberate testing

The upstream implementation discussion covers the shape of the new interface and its semantics. The v0.6.0 release notes also say that the examples, speculative-decoding components, multimodal tooling and server were migrated to the extended path. That is evidence of an upstream transition, not proof that all external language bindings were migrated simultaneously.

A binding can compile and still have important unsupported cases. For instance, it may:

The last case is particularly worth isolating. v0.6.0 bumps LLAMA_SESSION_VERSION to 11 and LLAMA_STATE_SEQ_VERSION to 4. These are format-version changes. They do not establish that old state snapshots are interchangeable with new ones.

The five tests I would require before upgrading a wrapper

1. A known-good text-only prompt

Start with the exact model artifact, chat template, tokenizer and generation settings from your earlier build. Confirm that the new version loads it and gives a coherent deterministic response. This establishes a baseline but does not test the extended API’s distinguishing feature.

2. More than one sequence

Test multiple sequences and position boundaries. A wrapper that confuses sequence metadata may return well-formed output assigned to the wrong request. Record the sequence and token positions at the API boundary, not merely the final answer.

3. Mixed input on a model that supports it

For an appropriate multimodal or MTP model, exercise a batch containing supported token and embedding entries. Do not synthesize arbitrary embeddings and assume all text models will accept them. Check the actual model-specific requirements first.

4. A saved-state round trip

Save and restore state using the same tagged version before attempting cross-version tests. Record the producing version with each persisted snapshot. If a previous snapshot is rejected, the application should report that clearly and reconstruct it from its source prompt rather than guessing at binary compatibility.

5. Recovery after a rejected request

Force a normal input-validation error, then send a valid request on the same service. A healthy wrapper should not retain stale request metadata or continue with a partially populated batch. Observe whether the application restores a usable state without leaking memory or retaining abandoned work.

These tests are a proposed integration plan; they are not test results from RAMGPT.

Performance is a separate claim

The extended API changes what a batch can represent. It does not, by itself, imply a universal decode-speed improvement.

If you benchmark the upgrade, keep hardware, model weights, quantization and workload matched. Report prompt processing and token generation separately. Multiple changes in the v0.6.0 release affect kernels and model support, so a before/after comparison of entire releases would not isolate the effect of this API.

The release also ships other significant features, including support for additional model architectures, a decision-model server endpoint, typed multimodal embedding requests and kernel work. They are separate from the batch interface and should be assessed separately.

When should you migrate?

Ordinary server user: pin a known-good version and verify the endpoints your client actually uses. No need to rewrite code merely to adopt the new batch type.

C/C++ wrapper author: add a release-pinned test against the tagged header, model the token/embedding alternatives explicitly, and define an error path for unsupported input.

Multimodal or speculative integration: treat input ownership, per-item positions and state boundaries as correctness requirements. Validate mixed inputs before optimizing throughput.

The engineering lesson is simple: the word batch now describes a richer unit of inference input. A migration succeeds when its meaning is preserved, not just when the program links.

Primary sources

Sources and further reading

Continue reading