Skip to content

Don't salvage an aborted prefill as a continuation endpoint under MTP (fixes HTTP 500 'published MTP checkpoint is not materializable') - #4

Draft
hermes-pimentel wants to merge 1 commit into
iamwavecut:masterfrom
hermes-pimentel:fix/mtp-salvage-aborted-prefill
Draft

hermes-pimentel wants to merge 1 commit into
iamwavecut:masterfrom
hermes-pimentel:fix/mtp-salvage-aborted-prefill

Conversation

@hermes-pimentel

Copy link
Copy Markdown

Issues and Discussions are disabled on this repository, so I'm opening this as a Draft PR that carries the bug
report together with the smallest fix I found. Happy to change the approach, or to close this in favor of your own
fix.

Summary

With --spec mtp, when a client disconnects while a streaming Chat Completions request is still prefilling,
salvage_continuation publishes the partial prefill as the conversation's continuation endpoint. The next request of
that conversation plans a PrivateEndpoint reuse that the MTP head cannot materialize, and fails with HTTP 500
published MTP checkpoint is not materializable. Every retry takes the same plan and fails the same way, until the
entry is evicted or the server restarts. An agent session on the server stops there.

Environment

  • ninfer-all master at f118551fb401de073555807a48c50238e180e3b8, built with CMAKE_CUDA_ARCHITECTURES=120a
    (this repository's Dockerfile pinned to that commit), nvidia/cuda:13.1.2 images, no local patches
  • RTX 5070 Ti 16 GB (sm_120, calibrated device profile), driver 616.92, Docker Desktop on WSL2
  • Reproduced with two artifacts:
    • WaveCut/Ternary-Bonsai-2-27B-NInfer-v3 at b85b33627b27b9757a5a094fc74785a53e99f8ee
    • WaveCut/Qwen3.8-27B-GSQ-RCO-IQ3_S-NInfer-v3 at 1daef874cd622636e122de5dd0b2a40727bebb2e
ninfer-serve /models/Ternary-Bonsai-2-27B-ninfer-v3.ninfer --model-id bonsai2-27b --host 0.0.0.0 --port 8080 \
  --max-context 73728 --kv-capacity 73728 --kv-dtype bf16 --gdn-state-fp16 --max-concurrency 1 \
  --default-reasoning-effort medium --default-thinking-budget 32768 --default-max-tokens 49152 \
  --vision --vision-residency resident --derive-session-keys --auto-long-anchors \
  --spec mtp --draft-tokens 3 --lm-head-draft

ninfer-serve /models/Qwen3.8-27B-GSQ-RCO-IQ3_S-ninfer-v3.ninfer --model-id qwen3.8-27b-gsq --host 0.0.0.0 --port 8080 \
  --max-context 73728 --kv-capacity 73728 --kv-dtype rk4v4-e8 --gdn-state-fp16 --max-concurrency 1 \
  --default-max-tokens 0 --vision --vision-residency overlay --vision-max-merged 4096 --kv-lease-growth \
  --derive-session-keys --auto-long-anchors --spec mtp --draft-tokens 3 --lm-head-draft \
  --structured-output --ngram-draft-tokens 0

Reproduction

A ~20K-token prompt, a streaming request cut by the client while it prefills, then the same messages again (set
model to the served --model-id):

body() { printf '{"model":"bonsai2-27b","stream":%s,"max_tokens":64,"messages":[{"role":"user","content":"' "$1"
  for i in $(seq 1 800); do printf 'Line %05d: the daily report shows stable temperature, normal pressure and no alarm in sector %d. ' $i $((i % 17)); done
  printf 'Summarize in one sentence."}]}'; }
body true > stream.json; body false > plain.json
curl -sN --max-time 3 -H 'Content-Type: application/json' --data-binary @stream.json http://127.0.0.1:8080/v1/chat/completions > /dev/null
curl -s -H 'Content-Type: application/json' --data-binary @plain.json http://127.0.0.1:8080/v1/chat/completions

The second request returns:

HTTP 500
{"error":{"code":null,"message":"published MTP checkpoint is not materializable","param":null,"type":"internal_error"}}

Server log, Bonsai:

INFO  req#432 done | openai-chat | cancelled | prompt 23,558 | output 0 | cache 0 (0.0%) | TTFT 139 ms | total 3.3s | queue 283 ms | thinking 0/32,768
INFO  req#432 response failed during transport | HTTP 499 | client disconnected
ERROR req#433 failed during generation | openai-chat | HTTP 500 | internal error

Server log, Qwen3.8 GSQ:

INFO  req#111 done | openai-chat | cancelled | prompt 20,388 | output 0 | cache 0 (0.0%) | TTFT 7.28 ms | total 3.3s | queue 10.2 ms
INFO  req#111 response failed during transport | HTTP 499 | client disconnected
ERROR req#112 failed during generation | openai-chat | HTTP 500 | internal error

In real use the disconnect came from an agent client (Hermes Agent), which aborts its background request when the
main turn needs the single lane. The main turn then failed three times in a row and the session ended.

Root cause

ProgramImpl::salvage_continuation (src/models/qwen3_5/program/transactions/commit.cpp):

  • the Lifecycle::Active / Finishable branch refuses MTP state that lags the frontier (state.mtp_kv_valid + 1 < frontier);
  • the Lifecycle::Prefilling branch only checks the DFlash context (state.dflash_context_frontier < frontier) and
    has no MTP check.

The next request then plans PrivateEndpoint reuse from that endpoint. In request_plan.cpp, append_ready needs
tail_hidden_valid and mtp_kv_valid >= reuse_base - 1, which an endpoint cut mid-prefill does not have, so the
planner throws std::logic_error("published MTP checkpoint is not materializable"), reported as HTTP 500.

Fix

Refuse to publish an aborted prefill under MTP, next to the DFlash guard, so the conversation resumes from its last
captured checkpoint instead.

With it, on the Bonsai setup, the same two requests return HTTP 200, and the second one re-prefills from the shared
prefix:

INFO  req#1 done | openai-chat | cancelled | prompt 23,558 | output 0 | cache 0 (0.0%) | total 3.1s | thinking 0/32,768
INFO  req#1 response failed during transport | HTTP 499 | client disconnected
INFO  req#2 done | openai-chat | output limit | prompt 23,558 | output 64 | cache 12 (0.1%, shared prefix) | total 8.1s | prefill 3.09k tok/s | decode 164.2 tok/s

Normal requests, tool calls, vision and MTP decoding were unchanged in a smoke test afterwards.

Alternatives

  • Keep salvaging under MTP by bringing the MTP KV and the tail hidden state up to the salvaged cursor, as the
    Active branch expects. That keeps the optimization, but I don't know the MTP state layout well enough to do it safely.
  • Have the planner fall back to a root prefill when a planned reuse cannot be materialized, instead of throwing. A
    stale entry would then cost one re-prefill instead of failing every retry of the conversation.

Testing

Reproduced on both artifacts without the change. Verified before and after the change on the Bonsai setup; the fix
is in the shared qwen3_5 program, so it covers both. I did not add a regression test. If you point me to where one would fit (a program-level test that aborts a prefill under MTP
and plans the next request), I'll add it.

The Prefilling branch of salvage_continuation checked the DFlash context but not
the MTP state, so the next request of the conversation planned a PrivateEndpoint
reuse its MTP head could not materialize and failed with HTTP 500 "published MTP
checkpoint is not materializable" on every retry.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant