Skip to content

DiskSeismic: Add disk_seismic_index - #32

Merged
chishui merged 2 commits into
opensearch-project:mainfrom
zirui-song-18:disk-seismic-index
Aug 26, 2026
Merged

chishui merged 2 commits into
opensearch-project:mainfrom
zirui-song-18:disk-seismic-index

Conversation

@zirui-song-18

Copy link
Copy Markdown
Collaborator

Description

Added the main classes for DiskSeismicIndex

Issues Resolved

By submitting this pull request, I confirm that my contribution is made under the terms of the Apache 2.0 license.
For more information on following Developer Certificate of Origin and signing off your commits, please check here.

Signed-off-by: Zirui Song <zrsong@amazon.com>
@chishui

chishui commented Aug 26, 2026 •

Copy link
Copy Markdown
Collaborator
  1. Could you post the benchmark of this index comparing with SeismicIndex?
  2. Do you plan to make it support quantization? I assume the quantized index could be using smaller disk side and faster query time.
  3. Do we support options to either store inline or as forward index

const IDSelector* id_selector, size_t element_size,
detail::TopKHolder<idx_t>& heap,
absl::flat_hash_set<idx_t>& visited) {
if (fwd != nullptr) {

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

what about throw exceptions if either fwd or vectors is null as I think you can do nothing without them

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It will throw, but in search() before entering OpenMP. If we throw in score_block(), it will go out of OMP parallel.

Comment thread nsparse/disk_seismic_index.cpp Outdated
Comment on lines +88 to +91
const auto data = vectors->get_all_data();
for (const idx_t doc_id : clusters[pl].get_docs(cid)) {
const idx_t start = data.indptr_data[doc_id];
const size_t len = data.indptr_data[doc_id + 1] - start;

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

let's assign variables outside of for loop

const idx_t* const indptr = data.indptr_data;
  const idx_t* const indices = data.indices_data;
  const float* const values = data.values_data;

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Good suggestions. Ack

throw_if_any_null(indptr, indices, values);
const size_t indptr_size = n + 1;
const size_t nnz = indptr[n];
if (vectors_ == nullptr) {

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

reset num_vectors_ to 0 here?

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ack.

Comment thread nsparse/disk_seismic_index.cpp Outdated
SearchParameters* search_parameters)
-> pair_of_score_id_vectors_t {
if (num_vectors_ == 0 || n == 0) {
return {std::vector<std::vector<float>>(n),

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

we use INVALID_IDX as not found doc, please check other indices

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ack.

dynamic_cast<const SeismicSearchParameters*>(
search_parameters)) {
cut = seismic_parameters->cut;
}

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

what about else situation?

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

cut/k_prime has been initialized with default. else should be fine with nothing.

search_parameters)) {
cut = seismic_parameters->cut;
}
if (k_prime <= 0) {

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

what if k_prime is smaller than k?

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

k_prime here refers to the number of blocks to be fully examined. It has nothing to do with k.

Comment thread nsparse/disk_seismic_index.cpp Outdated
}

std::vector<std::vector<float>> result_distances(n);
std::vector<std::vector<idx_t>> result_labels(n);

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

here set default labels to 0

@zirui-song-18 zirui-song-18 Aug 26, 2026 •

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Use INVALID_IDX now.

// score. score_summaries_transposed dots the full query with each cluster
// summary, so scores are comparable across posting lists for the global
// ranking.
std::vector<BlockCandidate> candidates;

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

are sizes of both variables known already? can you call .reserve()

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Now add a pre-scan:

    size_t total_clusters = 0;
    size_t max_clusters = 0;
    for (const term_t term : cuts) {
        if (term < clustered_inverted_lists.size()) {
            const size_t nc = clustered_inverted_lists[term].cluster_size();
            total_clusters += nc;
            max_clusters = std::max(max_clusters, nc);
        }
    }
    candidates.reserve(total_clusters);
    score_scratch.reserve(max_clusters);

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

no, no, I mean if you already know the size, if not, we don't do this just to calculate the size and reserve which could be even heavier.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Got it. That makes sense.

}

for (size_t i = 0; i < q_len; ++i) {
dense[q_indices[i]] = 0.0F;

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

this for loop may have multiple cache misses, is the performance of this for loop fine?

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We only reset the number of q_len position in query, which is quite small. This is much cheaper than memset the whole dim table.

@zirui-song-18

Copy link
Copy Markdown
Collaborator Author

Thanks for @chishui 's suggestions, a benchmark result can be seen here:
Host: g5.12xlarge (AMD EPYC 7R32 Zen2, 48 vCPU, 186.7 GiB, avx2+fma, noavx512f)
Data: 8.8M docs with 6,980 queries (MS MARCO); k 10

1. Iso-recall readout along the two Pareto frontiers

anchor recall paired cell recall Δrecall p50 ratio p99 ratio
Seis c3/hf=1.15 0.914011 Disk c3/k'=25 0.915014 +0.0010 1.44× 3.14×
Seis c3/hf=1.2 0.925817 Disk c3/k'=32 0.926963 +0.0011 1.44× 3.24×
Seis c5/hf=1.1 0.929943 Disk c5/k'=25 0.930960 +0.0010 1.39× 2.92×
Seis c3/hf=1.25 0.934212 Disk c3/k'=40 0.934943 +0.0007 1.48× 3.30×
Seis c5/hf=1.2 0.952521 Disk c5/k'=40 0.952765 +0.0002 1.44× 3.15×
Seis c8/hf=1.2 0.963639 Disk c8/k'=50 0.965143 +0.0015 1.33× 2.88×
Seis c5/hf=1.3 0.964269 Disk c8/k'=50 0.965143 +0.0009 1.36× 3.00×
Seis c8/hf=1.5 0.981232 Disk c8/k'=200 0.982106 +0.0009 1.13× 2.53×
Disk c3/k'=15 0.875917 Seis c3/hf=1.05 0.876218 +0.0003 1.43× 3.02×
Disk c3/k'=30 0.924298 Seis c3/hf=1.2 0.925817 +0.0015 1.50× 3.32×
Disk c5/k'=30 0.940745 Seis c8/hf=1.1 0.942507 +0.0018 1.58× 3.37×
Disk c8/k'=400 0.983625 Seis c8/hf=2 0.984112 +0.0005 1.57× 2.93×

Ratio > 1 = DiskSeismic faster. The p50 margin is 1.2–1.5× everywhere and narrows above recall 0.97;
the p99 margin is ~2.8–3.1× and does not narrow.

2. Footprint

variant file on disk RssAnon RssFile (pages this process touched) VmHWM
SeismicIndex in-memory c3/hf=1.0 13.35 GiB 13.371 GiB 0.004 GiB 13.376 GiB
SeismicIndex mmap c3/hf=1.0 13.35 GiB 0.011 GiB 13.359 GiB 13.369 GiB
DiskSeismic mmap c3/k'=50 114.65 GiB 0.332 GiB 35.828 GiB 36.160 GiB

Exact sizes: base_full.seismic.dat 14,339,730,216 B; base_full.disk_seismic.dat
123,106,695,136 B — 8.59× larger on disk.

3. Two secondary results

recall@10 p50 p99 RssAnon
SeismicIndex in-memory (documented default) 0.843367 0.1905 0.6325 13.371 GiB
SeismicIndex mmap, identical params 0.843367 0.1547 (1.23×) 0.4760 (1.33×) 0.011 GiB
DiskSeismic cut=3, k'=11 0.840358 (−0.0030) 0.1054 (1.81×) 0.1640 (3.86×) 0.332 GiB
DiskSeismic cut=3, k'=12 0.851619 (+0.0083) 0.1105 (1.72×) 0.1680 (3.76×) 0.332 GiB

@zirui-song-18

Copy link
Copy Markdown
Collaborator Author
  1. Could you post the benchmark of this index comparing with SeismicIndex?
  2. Do you plan to make it support quantization? I assume the quantized index could be using smaller disk side and faster query time.
  3. Do we support options to either store inline or as forward index
  1. The benchmark results can be seen in my above comment;
  2. Not yet — DiskSeismicIndex currently stores float32 values. But quantization is a natural fit and most of the plumbing already exists. The inline forward-index format (writer/reader/validator) already supports element_size ∈ {1,2,4}, and SeismicScalarQuantizedIndex (seismic_sq) is a complete reference for 8/16-bit scalar quantization: encode-on-add, element_size-dispatched scoring kernels (compute_similarity → uint8/uint16 dot products), query re-quantization, score decoding, and mmap of codes. The remaining work is contained within DiskSeismicIndex: drop the hardcoded float element size, quantize in add(), make score_block/score_doc element_size-aware (reuse the existing uint8 kernels instead of the float-only dot_product_float_dense), persist a quantizer header, and decode scores.
  3. Both layouts are available, but as two index types selected via the factory descriptor, not as a flag on a single index:
  • seismic (SeismicIndex) — the classic forward index: each document's vector is stored once in a global CSR (SparseVectors) and gathered by indptr[doc_id]. Loadable either in-memory or mmap (IndexIoFlag::kUseMmap).
  • disk_seismic (DiskSeismicIndex) — the inline layout: each document's vector is duplicated into every block it belongs to, laid out block-contiguously (InlineForwardIndex). mmap-only.

So you pick the storage by building "seismic,…" vs "disk_seismic,…" — there's no store=inline|forward parameter on one index.

void DiskSeismicIndex::write_index(IOWriter* io_writer) {
const uint64_t nv = num_vectors_;
io_writer->write(const_cast<uint64_t*>(&nv), sizeof(uint64_t), 1);
SeismicInvertedListsWriter inv_list_writer(clustered_inverted_lists);

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

check if clustered list are written twice

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Good catch, partially. The summaries are written once (only in the invlists section) and the per-doc vectors are written once (only in the inline forward section — DiskSeismic drops plain seismic's full-CSR forward write). What is written twice is the per-cluster doc-id membership: docs_/offsets_ in the invlists section and again as per-block doc_id[] in the inline forward. And on the mmap search path the invlists docs_/offsets_ are never read (membership comes from the inline doc_id[]; only the summaries + cluster counts are used) — so that copy is dead weight on disk. It's ~M×4 bytes (M = total doc-cluster memberships, ~0.4% of the file here), so it's a cleanliness/correctness fix rather than a size lever (quantization is the size lever). I can add a summaries-only serialize/deserialize path for the DiskSeismic invlists so docs_/offsets_ aren't persisted — docs_ is still needed in RAM at build time to lay out the blocks, just not on disk.

But the cost of that change is not small. Maybe we can leave it a separate commit later.

const detail::InlineForwardIndex* fwd =
fwd_.num_blocks() > 0 ? &fwd_ : nullptr;
const SparseVectors* vectors = fwd == nullptr ? vectors_.get() : nullptr;
const size_t element_size = fwd != nullptr

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

so here, can you throw when fwd and vectors are null?

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

oh, never mind

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

but can you quit early though?

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Done — switched from throwing to quitting early. The no-source case (fwd_.num_blocks()==0 && vectors_==nullptr) is now folded into the early-return at the top of search(), alongside num_vectors_==0 || n==0, returning the k-length padded empty result. A plain return is safe even though the scoring runs under OpenMP (only an escaping throw would be a problem), and it's checked once before the parallel region rather than per block.

Signed-off-by: Zirui Song <zrsong@amazon.com>

@chishui chishui left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Please update the duplicate serialization in next PR and also include a python integration tests.

@chishui
chishui merged commit 439b992 into opensearch-project:main Aug 26, 2026
9 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants