bugfix: defer streaming chat role chunk until generation starts - #1710
Merged
valarLip merged 1 commit intoJul 28, 2026
Merged
Conversation
Contributor
🏷️ CI GuideRuns automatically on every eligible PR before approval:
Heavy model tests:
|
yitingw1
approved these changes
Jul 28, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Motivation
The streaming chat handler sent the
role: "assistant"chunk as soon as therequest was accepted, before the engine produced anything.
That is not what OpenAI does -- the first chunk is supposed to mean generation
has actually started (vLLM emits it on the first iteration of the result
generator). It also skews client-side metrics: a chat client that times the
first SSE chunk without checking for content records a near-zero TTFT and
counts the real first token as an inter-token latency. The
/v1/completionsbenchmark path has no such chunk, so existing numbers are unaffected, but the
chat path is a trap for external clients.
This PR defers the role chunk until the first engine chunk arrives, for both
the single-sequence and n>1 fanout paths. It still precedes any content or
tool_calls delta, so client-visible ordering is unchanged.