Skip to content

Fix vLLM routing capture and add GB10 validation - #139

Merged
WestWaters merged 1 commit into
WestWaters:mainfrom
SloptimistPrime:contrib/gb10-routing-capture
Oct 6, 2026
Merged

WestWaters merged 1 commit into
WestWaters:mainfrom
SloptimistPrime:contrib/gb10-routing-capture

Conversation

@SloptimistPrime

Copy link
Copy Markdown
Contributor

Fix FlashInfer routing accounting, including one-token prompt continuations that use a decode kernel. Keep the scheduler's phase flags without changing kernel dispatch. Unsupported observations stay in an unknown bucket.

Added code, reasoning and general-text workloads at two context lengths. Serial and batched runs on one GB10 matched 21,504 prompt tokens and 744 decode-feedback tokens in every router layer, with no recording errors or unknown tokens.

Serial outputs matched two runs after removing the recorder. Five batched profiles matched both controls. Short-code controls differed from each other, so batched output stability remains unresolved.

45 focused tests pass. Includes a CPU-only test job and aggregate results, not raw prompts, outputs or machine details. The runtime was a local vLLM development build, not a stock release. This is routing validation, not a quality, speed or multi-node claim.

Details: notes/vllm-routing-capture.md and notes/vllm-routing-profiles.md.

I used AI assistance and reviewed and tested the changes.

@WestWaters
WestWaters merged commit 099b4a7 into WestWaters:main Oct 6, 2026
2 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants