Add 8-bit scalar-quantized GPU acceleration for the index build path - #58
Merged
Merged
Conversation
Signed-off-by: Zirui Song <zrsong@amazon.com>
zirui-song-18
requested review from
chishui,
model-collapse and
yuye-aws
as code owners
September 16, 2026 07:16
chishui
reviewed
Sep 16, 2026
| int32_t acc = 0; | ||
| for (int32_t p = start; p < end; ++p) { | ||
| const int32_t col = corpus_indices[p]; | ||
| acc += static_cast<int32_t>(corpus_values[p]) * |
Collaborator
There was a problem hiding this comment.
Issue: acc is int32, the CPU path uses int64. 255² × overlap overflows past ~33k terms; dim 30109 already hits 91% of INT32_MAX, so a bigger vocab silently diverges from the CPU argmax. int64_t acc, or reject dim > INT32_MAX/65025 in ensure_resident?
Collaborator
Author
There was a problem hiding this comment.
Good point. Now fixed.
Signed-off-by: Zirui Song <zrsong@amazon.com>
chishui
approved these changes
Sep 16, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Description
What
Extends the (opt-in, build-only) GPU acceleration from float-only to 8-bit scalar-quantized (SQ) indexes. Previously,
SeismicScalarQuantizedIndex/ DiskSeismicScalarQuantizedIndex` builds silently fell back to the CPU for both the cluster-assignment and summarization stages; only the unquantized (float) family used the GPU. This PR adds a GPU path for 8-bit codes so quantized index builds get the same acceleration.16-bit SQ intentionally stays on the CPU (see Scope).
Why not just reuse the float/cuSPARSE path
8-bit SQ codes are unsigned
uint8in[0, 255]. cuSPARSE has no unsigned int8 SpMM:CUDA_R_8Uis rejected bycusparseCreateCsr, and the signedCUDA_R_8Imisreads any code ≥ 128 (e.g. 200 → −56, 255 → −1), which would corrupt the cluster assignment. So the float path's cuSPARSE SpMM cannot be reused as-is.Instead the assignment uses a custom
uint8kernel that accumulates inint32. int32 is exact for byte-sized products (255² × overlap stays well underINT32_MAX), so the argmax matches the CPUint64reference bit-for-bit. The summarizer's max-pool is templated to readuint8(integeratomicMax) alongside the existing float bit-trick path.Scope / safety
NSPARSE_ENABLE_GPU=OFF); all changes are under#ifdef NSPARSE_WITH_GPU, so the default build is unaffected.DeviceCorpusresidency cache now also keys onelement_sizeanddim, so a reused corpus of a different code width can never be misinterpreted.Testing
Verified on a real NVIDIA L4 (CUDA 13.2 / cuSPARSE 12.7):
GpuClusterAssignerTest: 7/7 pass — 4 float (unchanged) + 3 new 8-bit tests that assert bit-exact parity vs a CPUint64reference, using codes that include values > 127 (which a signed int8 read would corrupt).Benchmark — build speedup
Real MS MARCO corpus (8.84M docs × 30,109 dims × 1.12B nnz) on the L4, 16-core host,
lambda=6000 beta=400 alpha=0.4, in-memory corpus, default 8-bit quantizer. GPU on = device visible +NSPARSE_GPU_SUMMARIZE=1; GPU off =CUDA_VISIBLE_DEVICES="". Times arebuild()only (corpus load and index write excluded).GPU-on runs held 80–100% utilization with ~7.6 GB resident and zero fallback warnings; GPU-off runs confirmed 0% utilization (pure CPU). The float control run on the same binary reproduces the expected ~1.7×, confirming the on/off toggle genuinely engages the GPU. 8-bit SQ now gets the same ~1.7× build acceleration as the float path.
Issues Resolved
List any issues this PR will resolve, e.g. Closes [...].
By submitting this pull request, I confirm that my contribution is made under the terms of the Apache 2.0 license.
For more information on following Developer Certificate of Origin and signing off your commits, please check here.