fix: preserve structured draft row pitch across speculative widths - #5
Draft
wongkiller wants to merge 1 commit into
Draft
wongkiller wants to merge 1 commit into
wongkiller wants to merge 1 commit into
Conversation
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem and scope
Related Issue: none; Issues are disabled in this repository. The full bug report and reproducer are included below.
Draft for maintainer confirmation of scope/direction (not yet approved). This is a focused local crash fix, submitted with evidence under the draft-PR policy.
When DFlash2 uses seven drafts but the structured round reserves a wider ngram/copy window, compact device rows and fixed-pitch host rows have different layouts. A contiguous D2H copy leaves later callback rows reading the wrong tokens. With concurrent structured requests, incorrect masks can lead to
generated token ... violates the structured output grammar, HTTP 500, and a failed worker returning HTTP 503.Reproduced on base
f118551fb401de073555807a48c50238e180e3b8; the report below includes the artifact, server options and HTTP reproducer. Feature provenance: Neroued#294. No claim is made that upstream master reproduces this fork's integration failure. This is separate from #4 (aborted-prefill MTP continuation checkpoints); that change is not included here.Implementation
cudaMemcpy2DAsync, mapping the tensor's actual device row pitch into the callback's reserved host pitch.The change stays within Program-owned staging. It does not disable grammar checks, catch-and-ignore worker failures, change quantization, reduce concurrency, or change the public API.
Verification
Environment: RTX 5090 32 GB; Windows / WSL2 / Docker Desktop; Linux CUDA 13.1.2, Release
sm_120acompatibility build. The tested model was Ternary Bonsai 2 27B Heretic with DFlash2-7, default ngram proposals, two active lanes,rk4v4-e8KV and a 917,504-token configured pool. The live regression requests were short; this is not a full-context correctness claim.Previously executed local verification:
compact draft row used reserved-width stride; against the fixed implementation it exited 0.ninfer,ninfer-serve, and the focused regression target successfully.git diff --checkpassed.The regression was compiled as a standalone target linked to the existing built runtime so the old and fixed implementations could be tested with the identical test source. Local commands used:
The standalone target definition was:
Maintainers can instead build and run the repository's existing
ninfer_qwen3_5_structured_round_testtarget with the modified test source in this PR.Limitations: the full repository test suite, CUDA sanitizers, other GPU architectures, and an isolated transfer-performance benchmark were not run. This is a correctness fix; no speedup or performance-parity claim is made. Code and test preparation used AI assistance; the listed results are local test observations, not assumed CI results.
Full bug report and HTTP reproducer
Problem
At
f118551fb401de073555807a48c50238e180e3b8, concurrent structured-output requests can crash the worker withgenerated token ... violates the structured output grammar. Active requests return HTTP 500; subsequent requests return HTTP 503 until restart.This was reproduced on an RTX 5090 (32 GB), Windows / WSL2 / Docker Desktop, using a Linux CUDA 13.1.2
sm_120acompatibility build (NINFER_SM120_NATIVE=OFF) andemiltsoi/Ternary-Bonsai-2-27B-Uncensored-Heretic-NInfer. Artifact SHA-256:99f94f44bed48892b7c5e1c4dc8349e0db8b4b44d2bd8ad7cd438cfc181de2fd.Relevant server configuration:
The live failure occurred on the sixth pair of concurrent strict-JSON requests. The following is a minimal HTTP workload matching that test (adjust the URL to the server):
The exact failing round/token depends on sampling and batching. A deterministic GPU regression is included in the accompanying proposed fix.
Root cause and proposed scope
StructuredRound::enqueue_dflash()copies the compact device draft frame contiguously.StructuredRound::callback()indexes the host staging rows using the reserved maximum pitch,width_ - 1. With seven neural drafts and a larger ngram reservation, later compact rows are therefore read at the wrong offset. The grammar masks are conditioned on stale/wrong draft tokens.The proposed change is confined to Program-owned structured-round staging: use
cudaMemcpy2DAsyncwith the actual device row pitch and reserved host pitch, validate frame bounds, and add regression coverage. Grammar enforcement and worker invariant checks remain enabled. No sampling, model, KV-capacity, or public API changes are proposed.This is a locally reproduced bug in this consolidated fork. The related feature originates in Neroued#294; this report does not claim that upstream master or that PR independently reproduces the same mixed-width integration failure.
A local repair already passed focused tests. Issues are disabled in this repository, so this report accompanies the draft PR under
PR_POLICY.md; maintainer confirmation of scope/direction is still pending.