Troubleshooting

SGLang Incremental Streaming Can Duplicate Text on Graceful Abort

SGLang issue #40901 shows that incremental streaming can resend the full accumulated text on graceful abort; PR #40902 fixes the text and token delta bookkeeping.

Approximately 7 min read

A streaming API is only predictable if every chunk obeys the same contract, including the last one.

SGLang issue #40901, filed on September 23, 2026, describes a narrow bug with a very production-shaped failure mode: when --incremental-streaming-output is enabled, a graceful abort can send the entire accumulated text again instead of sending only the text that has not yet been streamed.

The request does not need to crash. The model does not need to produce wrong logits. The network can be healthy. A client that correctly concatenates incremental deltas can still end up with duplicated output.

That makes this a useful serving bug to study because the failure is at the boundary between internal request state and an external API contract.

The contract is supposed to be disjoint deltas

The upstream report points to two places in SGLang that describe incremental streaming as a sequence of non-overlapping segments. In that mode, a client should be able to reconstruct the response by concatenating chunks in order:

chunk 1: "hello "
chunk 2: "worl"
chunk 3: "d"

client result:
"hello world"

The important property is not the size of each chunk. It is that a piece already emitted should not appear again in a later chunk.

That contract becomes especially important around termination paths. Normal decode chunks may be correct for thousands of requests while a timeout, explicit abort, disconnect, or session cleanup takes a different code path and violates the same protocol.

What the abort path does on current main

Issue #40901 traces the problem to TokenizerManager._handle_abort_req.

SGLang already tracks last_output_offset for output token IDs. But on the affected abort path, text is obtained from state.get_text(), which represents the full accumulated text, not only the portion that remains unstreamed.

So the two fields in the same terminal message can describe different slices of the response:

text       -> full accumulated text
output_ids -> attempted token delta

That is already an inconsistent state boundary.

The upstream headless reproduction makes the mismatch concrete. It sets up a request where most output has already been streamed and then calls the abort handler. The expected terminal delta is reported as:

{"text": "d", "output_ids": [6, 7], ...}

Current main instead returns:

{"text": "hello world", "output_ids": [7], ...}

Those values are from the upstream reproducer, not a RAMGPT measurement.

There are actually two bugs visible in that one message.

First, the text is replayed from the beginning. Second, the token-ID logic keeps only output_ids[-1], which can drop pending token IDs when more than one token arrived after the last emitted incremental chunk. If no new token is pending, taking the last element can also resend an ID that was already accounted for.

Why a client can duplicate the response

The serving layer treats incremental output as deltas. If the client follows the contract and appends each text segment, an abort can produce something like:

already received:
"hello " + "worl"

abort chunk:
"hello world"

concatenated client text:
"hello worlhello world"

The client is not necessarily doing anything wrong. Under an incremental protocol, concatenation is exactly what it is expected to do.

This distinction matters operationally because duplicate text is easy to misdiagnose as a retry problem. Teams often investigate reverse proxies, SSE reconnect behavior, frontend state, idempotency keys, or duplicated HTTP requests first.

Here, the upstream report says a single request can produce the duplication because the terminal abort chunk does not use the same delta semantics as ordinary streaming chunks.

Graceful abort is not an obscure edge case

The issue lists several ways the abort path can be reached, including client disconnects, /abort_request, and session cleanup.

That means the trigger can appear during normal operations:

user closes a tab
client cancels generation
application enforces its own timeout
session cleanup aborts unfinished work

A bug that exists only on abort may therefore remain invisible in happy-path load tests while still corrupting transcripts in real traffic.

This is also why I would classify the impact as a serving correctness problem, not a model-quality problem. The generated content can be fine internally; the API representation of that content is wrong at termination.

PR #40902 makes text and token cursors explicit

The proposed fix is SGLang PR #40902, also opened on September 23.

The PR adds a text-side cursor, last_streamed_text_len, to ReqState. It is updated beside the existing token output offset when incremental BatchStrOutput chunks are emitted.

Then the abort path can slice both representations from their last known streamed positions:

text cursor      -> emit only unseen text
output-id cursor -> emit only unseen token IDs

That is the important design change. The fix is not trying to infer whether a string “looks duplicated.” It records how much was already delivered and uses that state to construct the terminal delta.

The PR says non-incremental behavior remains unchanged. The IDs-only BatchTokenIDOutput path also does not need text tracking because it never streams text.

The regression tests target the state transition, not the model

This bug does not require a GPU or a model forward pass to reproduce.

The issue includes a headless Python reproducer, and PR #40902 adds tests around three cases:

The PR author reports that the new regression tests fail on main and pass with the change, and that ruff, isort, and the repository pre-commit hooks pass. Those are upstream author-reported validation results. PR #40902 is still open as of this article, so this should not be read as a statement that every released SGLang build already contains the fix.

There is no performance benchmark to report here. The PR describes the runtime overhead as one integer update per incremental text chunk, and it does not change model forward or kernel code.

How I would diagnose this in a serving stack

If an SGLang-backed client sometimes shows repeated text after cancellation, I would first separate three failure classes:

1. request-level retry
2. transport-level replay/reconnect
3. one request emitting overlapping incremental chunks

For the third case, capture the raw streaming events before the application concatenates them.

With --incremental-streaming-output enabled, verify whether every event contains a disjoint text segment. Then intentionally trigger a graceful abort after several chunks have arrived.

If the final event suddenly contains the full prefix again, the symptom matches issue #40901 much more closely than a duplicate HTTP request.

I would also inspect token IDs if the client exposes them. The upstream bug is not just a text replay: the old abort slicing can lose pending IDs as well. A mismatch between the text delta and token-ID delta is a strong clue that the error is inside output-state bookkeeping rather than generation itself.

Do not paper over this with string de-duplication

A tempting client workaround is:

if new chunk starts with text I already have:
    remove the repeated prefix

That is risky.

Generated text can legitimately repeat itself. A model may intentionally emit the same phrase, code block, JSON fragment, or line twice. Content-based de-duplication cannot reliably distinguish a protocol replay from valid generation.

The safer place to fix this class of bug is where the server already knows the delivery cursor.

That is exactly what PR #40902 proposes: track the stream boundary as state, then slice from that boundary.

The broader platform lesson: terminal paths are part of the protocol

Streaming implementations often get most attention on the hot path:

decode token
-> emit chunk
-> decode token
-> emit chunk

But the protocol also includes exits:

normal finish
abort
client disconnect
validation failure
server-side cancellation
exception

If those paths construct output differently, they need the same contract tests as the steady-state loop.

For a production inference platform, I would want a regression matrix that asks more than “did the request finish?” It should verify invariants such as:

incremental chunks are disjoint
concatenated chunks equal final text
token IDs and text advance consistently
abort never replays an acknowledged prefix
non-incremental mode preserves its documented semantics

Those checks are cheap compared with GPU inference tests, and bugs in them can be just as visible to users.

SGLang #40901 is a small example of that principle. The model can be correct, the scheduler can still be alive, and every GPU kernel can pass — yet one missing piece of stream-state bookkeeping can turn a clean cancellation into a corrupted response.

Sources and further reading

Continue reading