Skip to content

events: store dense terms' bitmaps in parts a window can read alone (format change) - #1043

Draft
tamirms wants to merge 1 commit into
tamirms/events-windowed-matchesfrom
tamirms/events-dense-term-parts
Draft

tamirms wants to merge 1 commit into
tamirms/events-windowed-matchesfrom
tamirms/events-dense-term-parts

Conversation

@tamirms

@tamirms tamirms commented Sep 25, 2026

Copy link
Copy Markdown
Contributor

What

  • When an index.pack bucket record would exceed 256 KiB, its largest terms of 16 KiB or more are demoted: the term's slot keeps only its fingerprint, and its bitmap is written again after the buckets as part records, each covering a fixed span of slabs.
  • A directory in index.pack's app data, keyed by the blinded term, records where each demoted term's parts start.
  • The cold reader reads only the parts the query's window reaches, and reports the range they cover.
  • Format change: the build stamp moves to version 2, and index.pack's format id from 0xFE1E000D to 0xFE1E0010, so an older binary refuses the new artifacts instead of misreading them.

Why

A popular term's bitmap sits whole in its bucket record, so a query that needs one or two slabs of it reads all of it. On pubnet chunk 6410 the fifteen most popular terms are 250 KiB to 1.08 MiB each, and a popular page reads about 9 MB of index to answer from one to three slabs. It also makes every other term in that bucket expensive, since a lookup reads the whole bucket record. With parts and the windowed lookup from the PR below, a page reads only the parts its slabs fall in. Parts and the directory add about 0.17% to the index.

Known limitations

  • Cold artifacts built before this must be rebuilt, and backfill won't do it because it skips chunks already marked frozen. That's acceptable only because nothing has shipped.
  • The streaming cold-index build on tamirms/full-history-p99-campaign will also need to write parts and the directory, to stay byte-identical with this build.
  • Not yet benchmarked end to end. The figures above come from the layout and the design doc.

🤖 Generated with Claude Code

A popular term's bitmap sits whole in its index.pack bucket record, so a
query that needs one or two slabs of it reads all of it: on pubnet chunk
6410 a popular page reads about 9 MB of index to answer from one to three
slabs. It also makes every other term in that bucket expensive to read.

When a bucket record would exceed 256 KiB, its largest terms (16 KiB or
more) are demoted: the term's slot keeps only its fingerprint, and its
bitmap is written again after the buckets as part records, each covering a
fixed span of slabs. A directory in index.pack's app data, keyed by the
blinded term, says where each demoted term's parts start. LookupKeys reads
only the parts the query's window reaches and reports the range they
cover, so the two-stage Matches reads a few parts for a typical page.

This changes the on-disk format: the stamp moves to version 2 and
index.pack's format id from 0xFE1E000D to 0xFE1E0010, so an older binary
refuses the new artifacts rather than misreading them. Parts and the
directory add about 0.17% to the index.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QHnF5BhsuoxpWxGmatQzmt
@tamirms
tamirms force-pushed the tamirms/events-dense-term-parts branch from 69dfba4 to a976f14 Compare September 25, 2026 20:35
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant