Don't salvage an aborted prefill as a continuation endpoint under MTP (fixes HTTP 500 'published MTP checkpoint is not materializable') - #4
Draft
hermes-pimentel wants to merge 1 commit into
Conversation
The Prefilling branch of salvage_continuation checked the DFlash context but not the MTP state, so the next request of the conversation planned a PrivateEndpoint reuse its MTP head could not materialize and failed with HTTP 500 "published MTP checkpoint is not materializable" on every retry.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Issues and Discussions are disabled on this repository, so I'm opening this as a Draft PR that carries the bug
report together with the smallest fix I found. Happy to change the approach, or to close this in favor of your own
fix.
Summary
With
--spec mtp, when a client disconnects while a streaming Chat Completions request is still prefilling,salvage_continuationpublishes the partial prefill as the conversation's continuation endpoint. The next request ofthat conversation plans a
PrivateEndpointreuse that the MTP head cannot materialize, and fails with HTTP 500published MTP checkpoint is not materializable. Every retry takes the same plan and fails the same way, until theentry is evicted or the server restarts. An agent session on the server stops there.
Environment
masteratf118551fb401de073555807a48c50238e180e3b8, built withCMAKE_CUDA_ARCHITECTURES=120a(this repository's Dockerfile pinned to that commit),
nvidia/cuda:13.1.2images, no local patchesWaveCut/Ternary-Bonsai-2-27B-NInfer-v3atb85b33627b27b9757a5a094fc74785a53e99f8eeWaveCut/Qwen3.8-27B-GSQ-RCO-IQ3_S-NInfer-v3at1daef874cd622636e122de5dd0b2a40727bebb2eReproduction
A ~20K-token prompt, a streaming request cut by the client while it prefills, then the same messages again (set
modelto the served--model-id):The second request returns:
Server log, Bonsai:
Server log, Qwen3.8 GSQ:
In real use the disconnect came from an agent client (Hermes Agent), which aborts its background request when the
main turn needs the single lane. The main turn then failed three times in a row and the session ended.
Root cause
ProgramImpl::salvage_continuation(src/models/qwen3_5/program/transactions/commit.cpp):Lifecycle::Active/Finishablebranch refuses MTP state that lags the frontier (state.mtp_kv_valid + 1 < frontier);Lifecycle::Prefillingbranch only checks the DFlash context (state.dflash_context_frontier < frontier) andhas no MTP check.
The next request then plans
PrivateEndpointreuse from that endpoint. Inrequest_plan.cpp,append_readyneedstail_hidden_validandmtp_kv_valid >= reuse_base - 1, which an endpoint cut mid-prefill does not have, so theplanner throws
std::logic_error("published MTP checkpoint is not materializable"), reported as HTTP 500.Fix
Refuse to publish an aborted prefill under MTP, next to the DFlash guard, so the conversation resumes from its last
captured checkpoint instead.
With it, on the Bonsai setup, the same two requests return HTTP 200, and the second one re-prefills from the shared
prefix:
Normal requests, tool calls, vision and MTP decoding were unchanged in a smoke test afterwards.
Alternatives
Active branch expects. That keeps the optimization, but I don't know the MTP state layout well enough to do it safely.
stale entry would then cost one re-prefill instead of failing every retry of the conversation.
Testing
Reproduced on both artifacts without the change. Verified before and after the change on the Bonsai setup; the fix
is in the shared
qwen3_5program, so it covers both. I did not add a regression test. If you point me to where one would fit (a program-level test that aborts a prefill under MTPand plans the next request), I'll add it.