Skip to content

Reuse tokenizer counts and request length estimates - #4

Merged
hotchpotch merged 1 commit into
mainfrom
perf/reuse-token-counts
Sep 21, 2026
Merged

hotchpotch merged 1 commit into
mainfrom
perf/reuse-token-counts

Conversation

@hotchpotch

Copy link
Copy Markdown
Owner

Summary

Avoid repeated local tokenization when preparing documents, checking request budgets, and recording execution details. Reuse exact counts and estimates while preserving serialized-payload counting and scoring prompts.

  • Cache tokenizer counts across calls on a reranker instance with a lock-protected LRU, bounded to 2,048 strings and 4,000,000 retained Unicode code points. Custom length_fn calls are not memoized.
  • Reuse listwise group estimates with keys that preserve document order, and reuse pointwise and pairwise preflight estimates in request details.
  • Reuse the initial document encoding during truncation, while still checking the token count of decoded text.
  • Add regression coverage for cache eviction, concurrent access, estimate reuse, context-dependent subword tokenization, and truncation at partial UTF-8 boundaries. Record the changes in the development release notes.

Validation

  • uv run --locked tox: 222 passed, 52 live tests skipped; lint and type checks passed.
  • uv run --locked pytest tests/test_lengths.py -q: 26 passed.
  • git diff --check: passed before committing.

@hotchpotch
hotchpotch marked this pull request as ready for review September 21, 2026 06:23
@hotchpotch
hotchpotch merged commit 580b0ad into main Sep 21, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant