[Sparse ANN] [Native] Add support for neural-sparse-cpp's wider 64-bit nnz-offset - #2001
Conversation
PR Code Analyzer ❗AI-powered 'Code-Diff-Analyzer' found issues on commit d29f374. ⛔ Hard block: Issues at High severity or above will block this PR from merging.
The table above displays the top 10 most important findings. Pull Requests Author(s): Please update your Pull Request according to the report above. Repository Maintainer(s): You can Thanks. |
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## main #2001 +/- ##
=========================================
Coverage 84.22% 84.23%
- Complexity 4239 4241 +2
=========================================
Files 317 317
Lines 14877 14879 +2
Branches 2426 2425 -1
=========================================
+ Hits 12530 12533 +3
Misses 1484 1484
+ Partials 863 862 -1 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
275767c to
d29f374
Compare
Ack. The streaming path's indptr lives in |
Signed-off-by: Zirui Song <zrsong@amazon.com>
d29f374 to
4155a4d
Compare
PR Reviewer Guide 🔍Here are some key observations to aid the review process:
|
PR Code Suggestions ✨Explore these optional code suggestions:
|
Description
neural-sparse-cpp #54 widened the CSR forward-index offset from 32-bit to a new
offset_t = int64_tso a single segment can hold more than ~2.1B non-zeros. The 32-bit cap was hit by large corpora — e.g. MS MARCO V2 (138M docs, ~28.7B nnz) overflows a signed-int32indptrand cannot be built as a single force-merged segment on the native engine.This PR bumps the submodule to the merged nsparse commit and adapts the plugin's JNI + Java side so the wider offset flows through end to end. Doc-ids, labels, term-ids, and the per-block on-disk structures deliberately stay 32-bit (nsparse
idx_t) — only the CSR nnz-offset widens.Changes
jni/external/neural-sparse-cpp354c756→da0d43c("Widen the CSR offset to 64 bit", Widen CSR nnz-offset to 64-bit (offset_t) so one segment can hold >2.1B nnz neural-sparse-cpp#54, now merged to nsparsemain).jni/src/common.h) — the off-heap CSRindptraccumulator widens fromstd::vector<int32_t>tostd::vector<int64_t>, so a segment's cumulative nnz can exceedINT32_MAX. Each per-flush indptr is still a relative int32 array (reset to 0 per flush, bounded by the streaming flush limit) and is widened before the cumulative offset is added — so thetransferVectors/insertToIndexJNI signatures are unchanged.jni/src/nsparse_wrapper.cpp) — handsadd_with_idsandsearchtheindptrasnsparse::offset_t(wasidx_t) on both the insert and the single-query paths..csrwriter (CsrSparseVectorsFile.java) — emits int64indptrentries (8 bytes/row, matching nsparseread_csr); the id file'sexternal_idsstay int32 (idx_t). Removes the now-obsoleteINT_MAXnnz cap.CsrSparseVectorsFileTestsupdated for the 8-byte indptr (reads + byte-offset math); all 11 tests pass.Cost (measured on MS MARCO V1, 8.8M docs, native
seismic_sq, 1 segment)Widening the offset is effectively free — only the
indptrarray grows 4→8 bytes/row:Related Issues
Resolves #[Issue number to be closed when this PR is merged]
Check List
--signoff.By submitting this pull request, I confirm that my contribution is made under the terms of the Apache 2.0 license.
For more information on following Developer Certificate of Origin and signing off your commits, please check here.