Conversation
fschlatt
force-pushed
the
fix-numba-recompile-variable-id-length
branch
from
September 23, 2026 13:55
a16c8b9 to
f7ef12f
Compare
`Run`/`Qrels` store document ids as a fixed-width unicode numpy array whose
dtype (`<U{max_len}`) is derived from the longest id in the input. numba
encodes that width in the type signature of `create_and_sort` and the helpers
it calls, so each distinct max id length triggers a fresh JIT compilation
(seconds to minutes). Workloads that build many `Run`/`Qrels` objects with
differing id lengths — e.g. computing IR metrics per batch during model
validation — recompile the numba kernels on nearly every call, which can turn
a few-minute evaluation into tens of minutes.
Round the fixed width up to the next power of two so the number of distinct
dtypes is bounded by ~log2(max_len). The kernels then compile a handful of
times and are reused. The width only ever grows, so ids are never truncated
and metric values are unchanged.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
fschlatt
force-pushed
the
fix-numba-recompile-variable-id-length
branch
from
September 23, 2026 13:56
f7ef12f to
59410e8
Compare
fschlatt
marked this pull request as ready for review
September 23, 2026 14:33
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
Run/Qrelsbuild their document-id arrays as a fixed-width unicode numpy array whose dtype is derived from the longest id in the input:numba encodes the charseq width in the type signature (
Array(UnicodeCharSeq(N), ...)) ofcreate_and_sortand every helper it calls (create_bulk,create_dict_from_lists,sort_dict_*,to_typed_list,typed_list_argosrt, …). So each distinctmax_lentriggers a fresh JIT compilation of the whole chain — seconds to minutes each.Any workload that constructs many
Run/Qrelsobjects with differing max id lengths therefore recompiles on nearly every call. This shows up badly when computing IR metrics per batch during neural-model validation: with document ids of varying length (e.g. 27–208 chars),evaluate()re-JITs ranx on every batch (~40 s each on a fresh interpreter), turning a few-minute evaluation into ~45 min, plus steadily growing memory from the accumulating compiled overloads.Measured (fresh process, real data, 8 queries/batch):
<U{max_len})(Identical behavior can also be forced by remapping ids to a constant length before calling ranx, which confirmed the cause; the last-compiled signature was literally
Array(UnicodeCharSeq(208), 1, 'C').)Fix
Round the fixed width up to the next power of two, so the number of distinct dtypes is bounded by
~log2(max_len)instead of "one per distinct id length". The numba kernels then compile a handful of times over a program's lifetime and are reused thereafter.The width only ever grows relative to
max_len, so ids are never truncated and metric values are unchanged. Applied symmetrically inRun.__init__andQrels.__init__.Opened as a draft for discussion — happy to adjust (e.g. a fixed cap, or a shared helper) or add a test.
🤖 Generated with Claude Code