Troubleshooting

vLLM Mistral Tool Call IDs Can Collide After Truncation

A Sep 29 vLLM bug shows that adapting OpenAI-style tool call IDs to Mistral by keeping only the last nine characters can reject valid requests or create duplicate IDs.

Approximately 4 min read

A tool call ID looks like plumbing until two different calls become the same ID.

That is the failure documented in vLLM issue #59212 on September 29, 2026. With --tokenizer-mode mistral, vLLM adapts conversation history to Mistral’s stricter tool-call format. The existing helper keeps the last nine characters of IDs that are too long.

Two distinct failures follow.

First, some OpenAI-style IDs are still invalid after the rewrite:

call_1

is too short, while:

functions.get_weather:0

can become:

weather:0

which still contains punctuation. Mistral then rejects the request with an error requiring nine alphanumeric characters.

Second, suffix truncation can destroy uniqueness:

call_A_000000001
call_B_000000001

both become:

000000001

Newer Mistral tokenizer validation can then fail with:

Duplicate tool call id

This is not a GPU, quantization, KV-cache, or model-weight problem. It happens in request normalization and chat-template rendering before inference.

Why the ID matters

Tool call IDs are correlation keys. The assistant emits an action with an ID; a later tool result carries the same ID so the conversation can associate the result with the correct action.

Conceptually:

assistant tool call A
id = call_A_000000001
        |
        v
tool executes
        |
        v
tool result
tool_call_id = call_A_000000001

An adapter can change the external representation, but it must preserve that one-to-one relationship.

Suffix truncation does not.

Why this can fail on the next turn

The awkward operational detail is that the first tool execution may work. The error can appear when the client sends the prior assistant tool calls and tool results back as history on the next request.

turn 1
assistant emits tool calls
-> client executes them

turn 2
client replays calls + results
-> vLLM normalizes history
-> Mistral validates IDs
-> request fails

That can make the incident look like an agent framework problem even though the tool itself executed correctly.

The proposed upstream fix

Open PR #59213 replaces truncate_tool_call_ids with a normalization step.

If an ID is already exactly nine ASCII alphanumeric characters, the proposed code leaves it alone.

Otherwise it derives a deterministic nine-character base62 value from a SHA-256 digest of the complete original ID.

The important properties are:

same original ID
-> same normalized ID

different full IDs
-> no deterministic shared-suffix collision

output
-> nine ASCII alphanumeric characters

This is a better adapter contract than simply trimming until the parser accepts a value.

The mapping is finite, so hashing does not make collisions mathematically impossible. It does remove the guaranteed collision pattern created by IDs that share the same last nine characters.

What the upstream tests report

PR #59213 adds cases covering long IDs, short IDs, non-alphanumeric IDs, and IDs with a shared suffix, across multiple Mistral tokenizer generations.

The contributor reports:

new tests on main: 6 / 9 fail
with PR #59213: all pass

The tests also check that call IDs and result IDs remain paired, normalized IDs are unique within the tested request, and already-valid IDs stay unchanged.

These are upstream contributor results, not RAMGPT measurements.

As of September 29, PR #59213 is open. Do not assume a released vLLM build already contains it.

How to diagnose this quickly

If a Mistral-backed vLLM request fails with:

Tool call id was ... but must be a-z, A-Z, 0-9, with a length of 9.

or:

Duplicate tool call id

collect:

vLLM version or commit
mistral_common version
tokenizer mode
model/tokenizer name
original tool call IDs
first renderer/tokenizer exception

Do not start by changing CUDA settings or model quantization. Issue #59212 reproduces the problem in a tokenizer-only CPU path.

Safe short-term choices

Until the fix is present in the vLLM build you actually deploy:

  1. If you control the client, generate IDs that already satisfy the Mistral format.
  2. If you carry a local adapter, transform assistant call IDs and tool-result IDs with the same deterministic mapping.
  3. Do not merely truncate IDs in a gateway; that reproduces the same collision class outside vLLM.
  4. Do not drop IDs. They are what bind results to requested actions.

Verify the entire history round trip

A good smoke test includes IDs such as:

abcDEF123
call_1
functions.get_weather:0
call_A_000000001
call_B_000000001

Verify that the request renders, every Mistral-facing ID satisfies the required format, calls and results still match, distinct calls remain distinct, and a second conversation turn can replay the history.

That last check matters because this is a history-round-trip compatibility bug, not just a one-shot generation bug.

The broader lesson

“OpenAI-compatible” does not mean every downstream model family accepts every value carried inside the same JSON schema.

Production inference gateways need explicit compatibility adapters at model-specific boundaries:

external protocol
-> validate
-> normalize while preserving semantics
-> model-specific renderer
-> model

Tool call IDs are not decorative strings. They are execution-graph identifiers. An adapter that makes them invalid or non-unique can break an otherwise correct agent workflow before the model runs.

Sources and further reading

Continue reading