Repository navigation
Add llama.cpp backend (GGUF) for raw / L0 / L1 - #2
Conversation
|
Saw @MorrisZJ comment on the issue about the serving-path rework after I'd already opened this, sorry for the timing. I'll rebase once you confirm main has settled. |
@Tusm11 Thanks for the contribution, and no worries about the timing! We just landed the serving-path rework, so main should be stable now. Could you rebase onto the latest main and refactor this PR to follow the new method? Feel free to ping me if anything in the new interface is unclear. |
5fb55a8 to
c090a0d
Compare
Rebased onto the current main. The rework (transformers-5 compat fix in HFBackend, L2 support in VLLMBackend) doesn't touch the Backend protocol this PR implements, so no functional changes were needed. I did update the docstring to state the raw/L0/L1 vs. no-L2 boundary explicitly, matching the style in the updated vllm.py. Re-ran the full suite (103 passed, 1 skipped), the engine test on a real 7B model, and the parity script against transformers. same result as before, 15/15 argmax agreement with the F32 GGUF. Ready for review. |
|
Independent check from a Windows box — i7-14700F, RTX 4060 8 GB, Python 3.13,
One environment note that may be worth a line for Windows users: the cu124 wheel from the usual extra-index crashes here at context creation — FWIW the same GGUF reads correctly through a llama-server build (PrismML b10660) on this machine too, so it is not the model file. |
|
I made a small follow-up PR against your branch to support reusing an existing |
|
Thanks @Tusm11, and sorry for the wait — I wanted to run this rather than just read it. Thanks also to @monke-sniper for the independent reproduction and to @craftingmod for the follow-up. Everything below was run against this branch on Linux with What holds upThe hard part is right. I checked the chunked decode, the absolute positions across chunks, the single output flag on the last token of the last chunk, and And one concern of mine turned out to be wrong, which is worth reporting explicitly. I expected that on a SentencePiece vocabulary What I would like changed before merge
That artifact carries a frozen temperature (5.214 here) and a frozen prior, so it silently miscalibrates the other model. The inverse hurts the workflow this backend most plausibly exists for — calibrate on a GPU, ship the GGUF:
Use after Setting
The version gate is wrong in both directions. Running it as written:
The chat template cannot compile Performance — follow-ups, except one line
HousekeepingOne structural note: the default path — no Yes please to committing the parity JSONs — I will add you to Separately: yes, please open that one-line PR adding @craftingmod — reusing an externally created |
|
Thanks @monke-sniper for the independent reproduction and the Windows/cu124 report, and @MorrisZJ for actually running this. Pushed fixes for everything flagged:
Validation:
Intentionally not touched in this PR:
Thanks again for the reports and testing. |
|
Thanks @Tusm11, and sorry this sat for a while. 0.3.0 landed in between. I re-checked your branch on top of the current main (0.3.0 plus the evaluation commit): with your files applied, ruff is clean and the whole suite passes (100 passed, 9 skipped), including all 16 of your stub tests. Every point from the last review is addressed. Thank you, this is ready apart from the 0.3.0 changes. 0.3.0 removed the label-trained levels (L1 temperature artifacts, L2 heads, observe, level="auto"), so this backend is now simply raw / L0. Could you rebase and adjust a few things? Rebase onto main. pyproject.toml conflicts because the bench extra is gone; put llamacpp = ["llama-cpp-python>=0.3.16"] next to vllm / client / eval and keep your markers line. In CHANGELOG.md, move your line under ## Unreleased and say raw / L0. Once that's in, we'll merge. On our side we'll add you to CREDITS.md with #1 on the CHANGELOG line, and take llama.cpp off the help-wanted lists. |
a3128c0 to
f82fe03
Compare
|
Thanks @MorrisZJ , rebased onto the current main (0.3.0 plus the evaluation commit) as a single commit. One note: the PR was briefly closed by my own force-push while I was rebasing, and I've reopened it. Done as requested: -> pyproject.toml: llamacpp = ["llama-cpp-python>=0.3.16"] is next to vllm / client / eval, and the engine markers line is kept. I added jinja2 to the dev extra so CI exercises the default GGUF-template path. ruff check anyjev scripts space tests is clean, and pytest -q gives 109 passed, 1 skipped on Windows (the skip is the engine test). Separately, tests/test_readme_numbers.py calls read_text() without an encoding, so on Windows it fails with a UnicodeDecodeError unless PYTHONUTF8=1 is set. Adding encoding="utf-8" there would fix it. I left it out of this PR to keep it to one change. |
… README test on Windows - CREDITS / CHANGELOG: @Tusm11 (llama.cpp backend, #1 #2), @monke-sniper (independent Windows reproduction) - ROADMAP / README (EN, zh): llama.cpp done, off the help-wanted list - anyjev.truncate: load on the CPU without device_map (it needs accelerate, which the hf extra does not install) - tests/test_readme_numbers.py: read files as UTF-8 (reported by @Tusm11 in #2)
|
Thanks @MorrisZJ for the thorough reviews and for merging, and @monke-sniper and @craftingmod for testing it on your machines. Happy to help if anything turns up on other platforms. |
Closes Issue #1
Adds
LlamaCppBackend(anyjev/backends/llamacpp.py), the llama.cpp slot from the ROADMAP's help-wanted list. It runs GGUF models through llama-cpp-python at raw / L0 / L1.What it does
llama_memory_clear). Only the final token is flagged for output in thellama_batch, and its row comes fromllama_get_logits_ith(-1). There is nologits_all. The backend returns the full-vocabulary log-softmax of that row at the requested ids. Nothing is sampled.tokenizer=(a Hugging Face name or object), prompts are rendered exactly asHFBackendrenders them.TokenMismatchError.pip install "anyjev[llamacpp]"installsllama-cpp-python>=0.3.16. Older versions are refused with a clear message.hidden_states_to, so no L2, and noscore_shared. Both could be follow-ups.Files
anyjev/backends/llamacpp.py(new)scripts/llamacpp_parity.py(new)tests/test_llamacpp.py(new): 12 tests on a stubllama_cppmodule, which run in CI, plus one@pytest.mark.enginesmoke test that runs whenANYJEV_GGUFpoints to a GGUF filepyproject.toml: adds thellamacppextra and registers theenginemarkerCHANGELOG.md: one lineResults (all runs on my machine; the commands are below)
Parity against
HFBackend. Qwen2.5-0.5B-Instruct, transformers in float32 on CPU, 15 prompts (choice / noul / score), same rendered prompts and label ids on both sides:convert_hf_to_gguf.py --outtype f32)Qwen/Qwen2.5-0.5B-Instruct-GGUFfp16The F16 gap comes from llama.cpp computing with the F16 weights, not from the backend. On the same F16 file:
Llama.eval(logits_all=True)last row to max |Δ| 0.0002 (15/15)The parity script's docstring therefore recommends an F32 conversion.
Engine smoke test.
pytest -m enginepasses on Qwen2.5-7B-Instruct Q8_0.CI.
ruff check anyjev bench demo handoff scripts space tests && pytest -qpasses.Environment: Windows, CPU, Python 3.13, llama-cpp-python , torch , transformers . The versions are in
envin the parity JSON.I haven't committed the parity JSONs, since vLLM's parity script doesn't commit its output either. I can add them under
bench/if you'd like.Side note (not changed here)
HFBackendpassesdevice_map=, which needsaccelerate, but thehfextra doesn't list it. I had topip install accelerateto run the parity script. I can open a separate one-line PR if that's useful.