Skip to content

Pack the inline forward index's blocks instead of padding each to a page - #60

Merged
chishui merged 1 commit into
opensearch-project:mainfrom
chishui:pack-inline-forward-blocks
Sep 24, 2026
Merged

chishui merged 1 commit into
opensearch-project:mainfrom
chishui:pack-inline-forward-blocks

Conversation

@chishui

@chishui chishui commented Sep 23, 2026

Copy link
Copy Markdown
Collaborator

First step of #59.

Inline-forward blocks were page-aligned, so each block's tail page is padding no query reads:
24.3 of 74.2 GiB on base_full disk_seismic_sq, where blocks hold ~ten documents. Defaults to
InlineLayout::kPacked.

No format change — the header carries the alignment and readers honour either, so older files
still load and the A/B ran one binary over two files.

base_full, lambda=6000 beta=400 alpha=0.4, mmap, cut=3 k_prime=50 k=10, 6,980 queries,
1 thread, cold caches, aligned → packed: file 74.205 → 49.967 GiB (−32.7%), cold pass
982,994 → 888,829 ms (−9.6%), warm p50 0.221 → 0.217 ms, VmHWM 6.567 → 6.042 GiB (−8.0%),
recall@10 0.9252 → 0.9261 (seed noise). Padding cost storage and cache, not faults.

Adds index_size_stats; at one seed both layouts' labels match exactly.

  • New functionality includes testing
  • New functionality has been documented
  • Commits are signed per the DCO using --signoff

A block of the inline forward index was written page-aligned: each one starts
on a 4096 boundary, so the tail of its last page is padding. At the standard
lambda/beta a block holds about ten documents, which makes that padding a large
share of the file -- and none of it is ever read, since a query touches a
block's own payload. On MS MARCO base_full (lambda=6000 beta=400 alpha=0.4,
8-bit disk_seismic_sq) it was 24.3 GiB of a 74.2 GiB index.

Make InlineLayout::kPacked the default, which the format already supports: the
header carries the effective alignment and a reader honours whatever the file
declares. So this is not a format change -- an index written before this still
loads, it just has to be rebuilt to shrink -- and no kFormatVersion bump is
needed. base_full comes down to 50.0 GiB (-32.7%).

Nothing regresses. Padding cost storage and page cache rather than page faults,
so removing it leaves query work alone: on base_full, first (cold) query pass
-9.6%, warm p50 0.221 -> 0.217 ms, peak RSS 6.57 -> 6.04 GiB. Blocks now share
pages, which is where the resident-set gain comes from.

Also adds benchmarks/index_size_stats, which re-parses a serialized index and
prints where its bytes are plus what a narrower encoding of each array would
save -- how the padding was found, and what sizes the remaining levers
(delta-coded component ids, narrower off[], 4-bit values) before any of them is
built. And a labels_out.txt argument to sq_residency_bench, so a layout change
can be shown to be exact: dump one build's top-k, score the other against it.
With a fixed seed the two layouts' label files are byte-identical.

Signed-off-by: Liyun Xiu <xiliyun@amazon.com>

@zirui-song-18 zirui-song-18 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Basically LGTM

Comment on lines +183 to +186
if (layout.end <= kU16Max) {
stats->blocks_off_fits_u16 += 1;
stats->c_off_u16 += (n_docs + 1) * sizeof(uint16_t);
} else {

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

layout.end <= kU16Max tests the block's total byte length, but an off[] entry holds an element offset in [0, total_nnz] (off[n_docs] == total_nnz) — so a block's off[] fits a u16 table iff total_nnz <= 0xFFFF, not layout.end <= 0xFFFF. layout.end is always larger (it also counts doc_id[]/comps[]/vals[]), so this undercounts both blocks_off_fits_u16 and the u16 saving. For a ~4.5 KB block the two agree, but consider gating on total_nnz.

const std::vector<InvertedListClusters>* lists_ = nullptr;
const SparseVectors* vectors_ = nullptr;
uint64_t write_page_size_ = kDefaultPageSize;
InlineLayout write_layout_ = InlineLayout::kPageAligned;

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nit: Consider setting this to kPacked?

@zirui-song-18

Copy link
Copy Markdown
Collaborator

I also tested on 138M documents spreading on 3 data nodes:

  ┌──────────────────┬─────────────────────────┬──────────────────┬─────────────────┐                                        
  │                  │ baseline (kPageAligned) │ packed (kPacked) │      change     │                                        
  ├──────────────────┼─────────────────────────┼──────────────────┼─────────────────┤                                        
  │ store / shard    │ 760.2 GB                │ 617.8 GB         │ −18.7%(−142 GB) │                                        
  ├──────────────────┼─────────────────────────┼──────────────────┼─────────────────┤                                        
  │ store / 3 shards │ 2280.7 GB               │ 1853.5 GB        │ −18.7%          │                                        
  └──────────────────┴─────────────────────────┴──────────────────┴─────────────────┘
  ┌───────────────┬────────────────────┬──────────────────┬───────────────────┐                                              
  │    operating  │ baseline p50 / p99 │ packed p50 / p99 │ recall(base/pack) │                                              
  ├───────────────┼────────────────────┼──────────────────┼───────────────────┤                                              
  │ recall ≈ 0.90 │ 0.41 / 1.25        │ 0.39 / 1.17      │ 0.903 / 0.901     │                                              
  ├───────────────┼────────────────────┼──────────────────┼───────────────────┤                                              
  │ recall ≈ 0.95 │ 1.06 / 2.70        │ 1.07 / 2.70      │ 0.957 / 0.959     │                                              
  └───────────────┴────────────────────┴──────────────────┴───────────────────┘

@chishui
chishui merged commit 8e9d3e9 into opensearch-project:main Sep 24, 2026
12 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants