Skip to content

Add 8-bit scalar-quantized GPU acceleration for the index build path - #58

Merged
chishui merged 2 commits into
opensearch-project:mainfrom
zirui-song-18:gpu-sq-8bit
Sep 16, 2026
Merged

chishui merged 2 commits into
opensearch-project:mainfrom
zirui-song-18:gpu-sq-8bit

Conversation

@zirui-song-18

Copy link
Copy Markdown
Collaborator

Description

What

Extends the (opt-in, build-only) GPU acceleration from float-only to 8-bit scalar-quantized (SQ) indexes. Previously, SeismicScalarQuantizedIndex / DiskSeismicScalarQuantizedIndex` builds silently fell back to the CPU for both the cluster-assignment and summarization stages; only the unquantized (float) family used the GPU. This PR adds a GPU path for 8-bit codes so quantized index builds get the same acceleration.

16-bit SQ intentionally stays on the CPU (see Scope).

Why not just reuse the float/cuSPARSE path

8-bit SQ codes are unsigned uint8 in [0, 255]. cuSPARSE has no unsigned int8 SpMM: CUDA_R_8U is rejected by cusparseCreateCsr, and the signed CUDA_R_8I misreads any code ≥ 128 (e.g. 200 → −56, 255 → −1), which would corrupt the cluster assignment. So the float path's cuSPARSE SpMM cannot be reused as-is.

Instead the assignment uses a custom uint8 kernel that accumulates in int32. int32 is exact for byte-sized products (255² × overlap stays well under INT32_MAX), so the argmax matches the CPU int64 reference bit-for-bit. The summarizer's max-pool is templated to read uint8 (integer atomicMax) alongside the existing float bit-trick path.

Scope / safety

  • Float path unchanged — still cuSPARSE SpMM.
  • 8-bit (U8): new custom kernel (assignment) + uint8 max-pool (summarize).
  • 16-bit (U16): stays CPU-only — it needs int64 accumulation, which cuSPARSE can't do and is out of scope here.
  • GPU support remains off by default (NSPARSE_ENABLE_GPU=OFF); all changes are under #ifdef NSPARSE_WITH_GPU, so the default build is unaffected.
  • GPU acceleration covers index build only — query is always CPU/mmap.
  • DeviceCorpus residency cache now also keys on element_size and dim, so a reused corpus of a different code width can never be misinterpreted.

Testing

Verified on a real NVIDIA L4 (CUDA 13.2 / cuSPARSE 12.7):

  • GpuClusterAssignerTest: 7/7 pass — 4 float (unchanged) + 3 new 8-bit tests that assert bit-exact parity vs a CPU int64 reference, using codes that include values > 127 (which a signed int8 read would corrupt).
  • Full suite: 659/659 pass.
  • Local no-GPU build unaffected (guarded changes compile out).

Benchmark — build speedup

Real MS MARCO corpus (8.84M docs × 30,109 dims × 1.12B nnz) on the L4, 16-core host, lambda=6000 beta=400 alpha=0.4, in-memory corpus, default 8-bit quantizer. GPU on = device visible + NSPARSE_GPU_SUMMARIZE=1; GPU off = CUDA_VISIBLE_DEVICES="". Times are build() only (corpus load and index write excluded).

Config GPU on GPU off Speedup
8-bit SQ, batches = 1 69.5s 118.4s 1.70×
8-bit SQ, batches = 10 70.5s 120.4s 1.71×
float control, batches = 1 78.7s 133.2s 1.69×

GPU-on runs held 80–100% utilization with ~7.6 GB resident and zero fallback warnings; GPU-off runs confirmed 0% utilization (pure CPU). The float control run on the same binary reproduces the expected ~1.7×, confirming the on/off toggle genuinely engages the GPU. 8-bit SQ now gets the same ~1.7× build acceleration as the float path.

Issues Resolved

List any issues this PR will resolve, e.g. Closes [...].

By submitting this pull request, I confirm that my contribution is made under the terms of the Apache 2.0 license.
For more information on following Developer Certificate of Origin and signing off your commits, please check here.

Signed-off-by: Zirui Song <zrsong@amazon.com>
Comment thread nsparse/gpu/gpu_cluster_assigner.cu Outdated
int32_t acc = 0;
for (int32_t p = start; p < end; ++p) {
const int32_t col = corpus_indices[p];
acc += static_cast<int32_t>(corpus_values[p]) *

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Issue: acc is int32, the CPU path uses int64. 255² × overlap overflows past ~33k terms; dim 30109 already hits 91% of INT32_MAX, so a bigger vocab silently diverges from the CPU argmax. int64_t acc, or reject dim > INT32_MAX/65025 in ensure_resident?

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Good point. Now fixed.

Signed-off-by: Zirui Song <zrsong@amazon.com>
@chishui
chishui merged commit b5f53e4 into opensearch-project:main Sep 16, 2026
12 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants