Skip to content

Add SGLang backend for native token logprob scoring - #12

Merged
MorrisZJ merged 1 commit into
nokia-applied-research:mainfrom
shentonyan:feat/sglang-backend
Oct 7, 2026
Merged

MorrisZJ merged 1 commit into
nokia-applied-research:mainfrom
shentonyan:feat/sglang-backend

Conversation

@shentonyan

Copy link
Copy Markdown
Contributor

Summary

  • Add SGLangBackend using SGLang's native /generate API.
  • Use token_ids_logprob and max_new_tokens=0 to score requested tokens without sampling.
  • Add protocol tests and an optional engine smoke test.
  • Add a parity script against HFBackend.

Validation

  • 113 passed, 1 skipped
  • Ruff passes
  • Real SGLang parity requires a running GPU-backed SGLang server.

@MorrisZJ

MorrisZJ commented Oct 7, 2026

Copy link
Copy Markdown
Collaborator

Thanks @shentonyan, and sorry for the wait; 0.3.0 landed in between. We ran this against a real server: SGLang 0.5.10, Qwen2.5-7B-Instruct, scripts/sglang_parity.py against HFBackend in bf16. As written, every request fails with "SGLang did not return input token log-probabilities". The server answers 200, but the scores aren't where the backend looks. Two things we found:

  1. The response puts them under meta_info, and at the placeholder position input_token_ids_logprobs is null. The output_token_ids_logprobs that do come back are the next token after the placeholder, so the placeholder approach scores the wrong position even once it is parsed.
  2. Dropping the placeholder fixes it: send the prompt ids alone with max_new_tokens=0, return_logprob=True, token_ids_logprob=ids, and read meta_info["output_token_ids_logprobs"][0] (entries are [logprob, token_id, text], so map them back by id). With that change your parity script gives 15/15 argmax agreement, and L0 decisions through Decider match the transformers backend on 24/24 choice, yes/no and score items (max |Δp| 0.04).

Could you:

  • (a) make that change;
  • (b) make the stub return the real response shape and drop the transformers dependency from the tests (CI installs numpy only, so they currently skip; see FakeTokenizer in anyjev/backends/fake.py);
  • (c) rebase onto main for CHANGELOG.md and pyproject.toml;
  • (d) note "checked on SGLang 0.5.10" in the docstring?

We'll rerun the same checks and merge.

Score the requested ids at the position after the prompt: send the prompt ids alone with max_new_tokens=0, return_logprob and token_ids_logprob, and read meta_info.output_token_ids_logprobs[0] by token id. The tests use the real response shape and need no transformers.
@shentonyan
shentonyan force-pushed the feat/sglang-backend branch from 961d18a to 63c0575 Compare October 7, 2026 14:11
@shentonyan

Copy link
Copy Markdown
Contributor Author

Thanks @MorrisZJ, and thanks for running it against a real server. I pushed an update: one commit on top of current main (f82fe03).

  • (a) Dropped the placeholder. The backend now sends the prompt ids alone with max_new_tokens=0, return_logprob=True and token_ids_logprob=ids, and reads meta_info["output_token_ids_logprobs"][0], mapping the [logprob, token_id, text] entries back by id.
  • (b) The stub returns the real response shape (including a null input_token_ids_logprobs, and entries in a different order than requested). The tests use FakeTokenizer and no longer import transformers, so they run in a numpy-only environment; the engine smoke test stays behind @pytest.mark.engine.
  • (c) Rebased onto main: the changelog entry is under Unreleased and pyproject.toml keeps the engine marker.
  • (d) The docstring notes "checked on SGLang 0.5.10". That comes from your run; I do not have a GPU server here, so I could not reproduce the parity numbers myself.

ruff check anyjev scripts space tests and pytest -q are clean on a numpy-only install.

@MorrisZJ
MorrisZJ merged commit 71d3270 into nokia-applied-research:main Oct 7, 2026
2 checks passed
MorrisZJ added a commit that referenced this pull request Oct 7, 2026
- CREDITS / CHANGELOG: @shentonyan (SGLang backend, #12), @tak-bro (vLLM missing-label report, #13)
- ROADMAP / README (EN, zh): SGLang done, off the help-wanted list
- VLLMBackend docs: vLLM's default raw logprobs can leave a label out (seen on 0.17.1 with
  Nemotron-H); recommend --logprobs-mode processed_logprobs (docs only)
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants