Skip to content

Bound the memory an index build needs, so a corpus larger than RAM can be indexed #44

Description

@chishui

Problem

A seismic build holds two whole-corpus intermediates — the inverted lists, then the clustered posting lists — so peak memory scales with non-zeros, with no way to trade time for memory. A corpus whose posting lists exceed RAM cannot be indexed at all.

On msmarco base_full (1.12B non-zeros), with the corpus mapped so the figure is the build's own, build() allocates 10.3 GB. piece1 (9.7B non-zeros) needs far more than a 68 GB host.

Proposal

Build one contiguous term window at a time — build, cluster, serialize, free — so peak memory tracks a window. Optionally stream the lists out and map them back, so build() still returns a usable index.

  • Output independent of window count: a memory knob only.
  • Serves all four index types.
  • Indifferent to corpus residency.
  • Windows cut by cost, not width.

Spill-then-concatenate breaks padding; term % batches cannot stream.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions