Skip to content

Avoid numba recompilation for varying document-id lengths - #83

Open
fschlatt wants to merge 1 commit into
AmenRa:masterfrom
fschlatt:fix-numba-recompile-variable-id-length
Open

fschlatt wants to merge 1 commit into
AmenRa:masterfrom
fschlatt:fix-numba-recompile-variable-id-length

Conversation

@fschlatt

Copy link
Copy Markdown

Problem

Run/Qrels build their document-id arrays as a fixed-width unicode numpy array whose dtype is derived from the longest id in the input:

max_len = max(len(y) for x in doc_ids for y in x)
dtype = f"<U{max_len}"
doc_ids = TypedList([np.array(x, dtype=dtype) for x in doc_ids])
self.run = create_and_sort(q_ids, doc_ids, scores)

numba encodes the charseq width in the type signature (Array(UnicodeCharSeq(N), ...)) of create_and_sort and every helper it calls (create_bulk, create_dict_from_lists, sort_dict_*, to_typed_list, typed_list_argosrt, …). So each distinct max_len triggers a fresh JIT compilation of the whole chain — seconds to minutes each.

Any workload that constructs many Run/Qrels objects with differing max id lengths therefore recompiles on nearly every call. This shows up badly when computing IR metrics per batch during neural-model validation: with document ids of varying length (e.g. 27–208 chars), evaluate() re-JITs ranx on every batch (~40 s each on a fresh interpreter), turning a few-minute evaluation into ~45 min, plus steadily growing memory from the accumulating compiled overloads.

Measured (fresh process, real data, 8 queries/batch):

batch 0 batch 1 batch 2 batch 3
current (<U{max_len}) 125 s 55 s 72 s 49 s
this PR / constant width ~40 s (compile once) 0.0 s 0.0 s 0.0 s

(Identical behavior can also be forced by remapping ids to a constant length before calling ranx, which confirmed the cause; the last-compiled signature was literally Array(UnicodeCharSeq(208), 1, 'C').)

Fix

Round the fixed width up to the next power of two, so the number of distinct dtypes is bounded by ~log2(max_len) instead of "one per distinct id length". The numba kernels then compile a handful of times over a program's lifetime and are reused thereafter.

max_len = max(len(y) for x in doc_ids for y in x)
width = 1 << (max_len - 1).bit_length() if max_len > 1 else 1
dtype = f"<U{width}"

The width only ever grows relative to max_len, so ids are never truncated and metric values are unchanged. Applied symmetrically in Run.__init__ and Qrels.__init__.

Opened as a draft for discussion — happy to adjust (e.g. a fixed cap, or a shared helper) or add a test.

🤖 Generated with Claude Code

@fschlatt
fschlatt force-pushed the fix-numba-recompile-variable-id-length branch from a16c8b9 to f7ef12f Compare September 23, 2026 13:55
`Run`/`Qrels` store document ids as a fixed-width unicode numpy array whose
dtype (`<U{max_len}`) is derived from the longest id in the input. numba
encodes that width in the type signature of `create_and_sort` and the helpers
it calls, so each distinct max id length triggers a fresh JIT compilation
(seconds to minutes). Workloads that build many `Run`/`Qrels` objects with
differing id lengths — e.g. computing IR metrics per batch during model
validation — recompile the numba kernels on nearly every call, which can turn
a few-minute evaluation into tens of minutes.

Round the fixed width up to the next power of two so the number of distinct
dtypes is bounded by ~log2(max_len). The kernels then compile a handful of
times and are reused. The width only ever grows, so ids are never truncated
and metric values are unchanged.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
@fschlatt
fschlatt force-pushed the fix-numba-recompile-variable-id-length branch from f7ef12f to 59410e8 Compare September 23, 2026 13:56
@fschlatt
fschlatt marked this pull request as ready for review September 23, 2026 14:33

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant