Skip to content

Widen CSR nnz-offset to 64-bit (offset_t) so one segment can hold >2.1B nnz - #54

Merged
zirui-song-18 merged 1 commit into
opensearch-project:mainfrom
zirui-song-18:widen-csr-offset-64bit
Sep 10, 2026
Merged

zirui-song-18 merged 1 commit into
opensearch-project:mainfrom
zirui-song-18:widen-csr-offset-64bit

Conversation

@zirui-song-18

@zirui-song-18 zirui-song-18 commented Sep 10, 2026 •

Copy link
Copy Markdown
Collaborator

Description

The forward-index CSR indptr (cumulative-nnz prefix sum) was 32-bit (idx_t) everywhere, so it wraps past the ~2.1-billionth nnz — capping a single segment well below a full MS MARCO V2 corpus (138M docs / 28.7B nnz).

This adds a distinct offset_t = int64_t (in types.h, beside idx_t) for CSR nnz-offsets only, and widens every indptr writer/reader in lockstep: the SparseVectors storage root, the Index API and all overrides, the native .mcsr layout + reader/writer, the build-time corpus scans, the SIMD summary kernels, and the disk/mmap serialization. The two int32 nnz caps are removed.

Kept narrow on purpose (offsets-only): doc-ids, labels, term-ids, per-row nnz counts, cluster/centroid CSC, the per-block on-disk off[]/doc_id[] (uint32) and its INT32_MAX cap, and the id-map's int32 external ids. Only the O(N_docs) indptrgrows (4→8 B/entry, ~+553 MB over 138M docs); the O(nnz)indices/values` arrays are untouched.

GPU build (off by default) stays int32 — cuSPARSE's CSR API is itself 32-bit — but now fails closed with a clear error on a >INT32_MAX corpus instead of silently truncating.

Testing

  • All 650 unit tests pass (Release, AVX2), including the AVX2/AVX-512 kernel-equivalence tests — the widened kernels stay bit-equivalent.
  • Added offset_width_test.cpp: type invariants, an 8-byte-indptr native-layout regression guard, and a DISABLED_ >2³¹-nnz acceptance test (the real overflow gate; needs a large-RAM host).
  • Real search on base_small (100k docs, dim 30109) returns correct results — confirming no regression on the normal-scale path.

The Python bindings (SWIG) are widened too: the indptr typemap accepts a 32- or 64-bit int buffer and widens to offset_t, and the buffer views are zero-initialised so a wrong dtype raises TypeError instead of segfaulting. The Python bindings build and their pytest suite run in CI (green).

Issues Resolved

#53

By submitting this pull request, I confirm that my contribution is made under the terms of the Apache 2.0 license.
For more information on following Developer Certificate of Origin and signing off your commits, please check here.

@chishui

chishui commented Sep 10, 2026

Copy link
Copy Markdown
Collaborator

could you benchmark on 8.8M?

@chishui chishui left a comment •

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

  1. ENABLE_GPU=ON broken: gpu_cluster_assigner_test.cpp:34,167.
  2. ENABLE_BENCHMARKS=ON broken: benchmarks/*.
  3. Bug: class_wrappers.py:48,80 force int32, numpy wraps silently.
  4. Python bindings build in CI, green. fix description.

Comment thread nsparse/index.cpp Outdated
std::string("CSR row count exceeds 32-bit doc-id range: ") +
file_path);
}
if (nnz > std::numeric_limits<offset_t>::max()) {

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

nnz is already int64_t and offset_t is int64_t. remove

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Removed the guard (kept the meaningful num_rows > idx_t::max())

Comment thread nsparse/mmap_index.h Outdated
std::string("CSR row count exceeds 32-bit doc-id range: ") +
file_path);
}
if (nnz > std::numeric_limits<offset_t>::max()) {

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

same here

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Removed

Comment thread nsparse/utils/csr_layout.cpp Outdated
const auto indptr = narrow<idx_t>(wide_indptr, "indptr", interchange_path);
write_or_throw(out, indptr.data(), indptr.size() * sizeof(idx_t),
const auto indptr =
narrow<offset_t>(wide_indptr, "indptr", interchange_path);

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

narrow<offset_t> on an int64_t vector range-checks nothing and copies (rows+1)*8 bytes — ~1.1 GB at 138M rows. write wide_indptr directly?

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Write wide_indptr (already int64) directly; kept a cheap in-place negative-offset check (no allocation) so throws_on_negative_indptr still holds.

Comment thread nsparse/python/nsparse_typemaps.i Outdated
// on itemsize, not view.format (NumPy leaves format null for some int
// dtypes, and strcmp(nullptr, ...) would crash).
const char* fmt = view.format;
const bool int_fmt = fmt == nullptr || strcmp(fmt, "i") == 0 ||

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

a float64 buffer with a null format would then be accepted on itemsize alone and read as int64. Have you verified numpy leaves format null under PyBUF_FORMAT? reject null instead?

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Now rejects null/non-int format; dispatches on explicit 'i'/'l'/'q'

// Round-trips a small corpus whose terminal offset fits int32; the point is that
// indptr survives serialize/deserialize/mmap as 64-bit words, exercising the
// widened read_padded<offset_t>/read_array<offset_t> format path.
TEST(OffsetWidth, IndptrRoundTripsAsSixtyFourBit) {

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

comment says serialize/deserialize/mmap, but the test only calls add_vectors and reads indptr_data(). serialize it?

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Now actually serializes → deserialize and mmap_deserialize, asserting indptr matches as 64-bit across all three

Comment thread tests/offset_width_test.cpp Outdated
// must keep the full 64-bit count rather than wrapping negative. Disabled by
// default because the single indices/values buffer needs ~6.5 GB; run manually
// on a large host with --gtest_also_run_disabled_tests.
TEST(OffsetWidth, DISABLED_CrossesInt32Boundary) {

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

this only exercises map_vectors. With clusters/CSC/off[]/id-map all still int32, has a >2.1B-nnz corpus actually been built and searched end to end?

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Rewrote the DISABLED_ test to add → build → search a >2³¹-nnz corpus on a SeismicIndex (real clustering/forward-index/search path)

Signed-off-by: Zirui Song <zrsong@amazon.com>
@zirui-song-18

Copy link
Copy Markdown
Collaborator Author
  1. ENABLE_GPU=ON broken: gpu_cluster_assigner_test.cpp:34,167.
  2. ENABLE_BENCHMARKS=ON broken: benchmarks/*.
  3. Bug: class_wrappers.py:48,80 force int32, numpy wraps silently.
  4. Python bindings build in CI, green. fix description.
  1. Widened gpu_cluster_assigner_test.cpp indptr reads (pattern-matched to the .cu; not compilable here — no CUDA)
  2. Widened all 4 benchmark CSR structs to offset_t; verified they build clean.
  3. Coerce indptr to int64 (lossless widening); the typemap accepts int32 or int64
  4. Ack

@zirui-song-18

Copy link
Copy Markdown
Collaborator Author

could you benchmark on 8.8M?

Sure:
Metric │ int32 (main) │ int64 (new) │ Difference │
Disk Size(Single segment) │ 41.361 GB │ 41.406 GB │ +0.11%(+45MB) │
Peak JVM RSS │ 105,617 MB │ 105,616 MB │ −0.001%(noise) │
Construct time │ 2227 s │ 2223 s │ −0.18%(noise) │
warm p50 hf=1.0 │ 0.407 ms │ 0.422 ms │ +3.7% │
warm p50 hf=1.5 │ 1.388 ms │ 1.458 ms │ +5.0% │
p99 hf=1.0/1.5 │ 1.21/3.88 ms │ 1.30/4.11 ms │ +7.2% / +5.9% │

@zirui-song-18
zirui-song-18 merged commit da0d43c into opensearch-project:main Sep 10, 2026
11 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants