Skip to content

POC: Full-history RPC benchmarking — May 2026 session reference - #744

Merged
chowbao merged 8 commits into
stellar:rpc-hackfrom
chowbao:chowbao/full-history-benchmarks-poc
May 19, 2026
Merged

chowbao merged 8 commits into
stellar:rpc-hackfrom
chowbao:chowbao/full-history-benchmarks-poc

Conversation

@chowbao

@chowbao chowbao commented May 19, 2026 •

Copy link
Copy Markdown
Contributor

POC / reference PR. Not for merge — captures the end-state of the May 2026 full-history RPC benchmarking session so the team can see exactly what was built and measured. Numbers, methodology, raw CSVs, and tooling all in one tree. No tests; some files remain after merging upstream rpc-hack (which has since landed PR #740 and #743).

What's in this PR

  1. Merge of upstream/rpc-hack (d5bb683) — upstream landed PR Add eventstore: per-Chunk hot + cold event storage for full-history #740 (eventstore) and PR Add exploratory full-history-backfill script: BSB -> ColdStoreWriter #743 (full-history-backfill), superseding our original cherry-pick. Conflicts resolved in favor of upstream for the eventstore + events files; go.mod kept both indirect deps (streamhash + xxh3); go.sum reconciled via go mod tidy. Minor post-merge fixups: scripts updated for the new error-returning APIs (Reader.Offsets, HotStore.EventCount) and the removal of pkg/geometry (now pkg/chunk.LedgersPerChunk).
  2. Phase 5 — cold-cache + concurrency bench extensions (7757289):
    • cache.go — posix_fadvise(FADV_DONTNEED) eviction + mincore residency check (Linux only). No sudo required; sidecar-fd pattern.
    • bench_concurrent_runner.go — shared per-worker scaffold (evict → open ColdStoreReader → caller workload → close → record); used by both point and range concurrent benches.
    • bench_ledger_cold.go — ledger-point-cold-open (1 worker, fresh open per iter) and ledger-point-concurrent-cold (N workers).
    • bench_ledger_range_cold.go — ledger-range-concurrency-sweep: grid of (workers × page-size) cold-open range reads, per-cell p50/p99 + saturation summary.
  3. Phase 1–4 bench harness (d552a95) — 8-file harness under cmd/stellar-rpc/scripts/bench-fullhistory/ covering ledger-point/range, tx-page, tx-by-hash, events (4 filter scenarios × 2 tiers). Also: cold-read/, migrate-cold/, spot-check/, verify-pack/, upload-cold/ supporting tools.
  4. Results, learnings, raw CSVs (b526fa8, 7757289, 4908d62):
    • BENCHMARK_GOAL.md — full writeup of Phase 0–4 (plan, methodology, results, interpretation).
    • SESSION_LEARNINGS.md — state-of-the-world, decisions, surprises, gotchas, open follow-ups. Includes Phase 5 addendum with cold-cache methodology and getLedgers projections.
    • bench-out/ — 26 per-scenario CSVs + summary.csv from Phase 1–4; Phase 5 concurrency sweep (ledger-range-concurrency-sweep.csv) plus rep2/, rep3/ for run-to-run variance and post-merge/ confirming the harness behaves identically after the merge.

Headline results

Phase 1–4 (warm-cache, single-thread, chunk 5000)

Workload Hot vs Cold-NVMe
Ledger / tx-page / tx-by-hash ~tie; zstd decompression (~1.3 ms/ledger) dominates
Events sequential (no-filter) Cold 2.3× faster
Events high-cardinality filter (contract) p50 tie; cold tail much worse (p99 44 ms vs 12 ms)
Events low-cardinality filter (topic) Hot 1.8× faster

Full table (26 rows) in BENCHMARK_GOAL.md and bench-out/summary.csv.

Phase 5 — Cold-cache getLedgers concurrency sweep (chunks 4999–5050)

Each iteration: evict the chunk's packfile, open a fresh ColdStoreReader, read N consecutive ledgers, close. 100 iters/worker.

Page size Peak workers Peak ops/sec p50 @ peak p99 @ peak
1 48 5,721 8.0 ms 11.7 ms
10 24 1,262 17.9 ms 27.2 ms
20 24 687 33.2 ms 47.4 ms

Storage-layer getLedgers model: T(N) = cold-open (~1.3 ms) + N × per-ledger-decode (~1.3 ms). Default page size 50 → ~75 ms p50, ~150 req/s at the box's saturation knee. Theoretical aggregate ceiling ~7,000–10,000 cold ledgers/sec.

Below the knee, p50/ops are stable to ~2% across runs. Past the knee, p99 inflates 3–5× and peak ops/sec drifts 10–30% run-to-run. Variance + reproducibility analysis in SESSION_LEARNINGS.md Phase 5 addendum.

How to reproduce

export PATH="/tmp/go126/go/bin:$PATH"   # Go 1.26 needed for streamhash
go build -o /tmp/bench-fullhistory ./cmd/stellar-rpc/scripts/bench-fullhistory/...

# Seed hot store (only for hot-tier benches)
/tmp/bench-fullhistory seed-hot --chunk 5000 \
    --cold-dir /mnt/nvme/disk2/ledgers/cold \
    --hot-dir /mnt/nvme/disk2/ledgers/hot-5000

# Phase 1 warm-cache point lookup
/tmp/bench-fullhistory ledger-point --tier=cold --chunk=5000 --iters=2000 \
    --hot-dir=/mnt/nvme/disk2/ledgers/hot-5000 --out=bench-out

# Phase 5 cold-cache concurrency sweep
/tmp/bench-fullhistory ledger-range-concurrency-sweep \
    --workers 1,2,4,8,16,24,32,48,64 --page-sizes 1,10,20 --iters 100 \
    --chunk-lo 4999 --chunk-hi 5050 --out bench-out

Full reproduction recipes in SESSION_LEARNINGS.md.

Caveats / known not-production

  • Cold txhash uses a sorted-.bin shim, not the production RecSplit index. Sub-millisecond lookup either way; dwarfed by ~12 ms ledger-decode in the tx-hash path.
  • Warm-cache, single-threaded only. → Cold-cache + multi-worker now covered by Phase 5.
  • Single dataset (pubnet 50M–60M) for everything. Other chunks may differ in tx/event density.
  • EBS tier dropped from scope at user request.
  • No multi-chunk iterator — every bench stays within a single 10K-ledger chunk.
  • Nothing wired into methods.getLedger / getEvents — stand-alone bench drivers only. Production routing layer is separate work.
  • Concurrency curve untested past 64 workers. Page=1 still scaling at 48, collapses by 64 — real saturation likely 50–60.
  • All numbers are storage-layer only — no HTTP/TCP/TLS/JSON-RPC framing, no SQLite-first lookup, no XDR→JSON serialization. End-to-end RPC latency adds overhead on top.

Files NOT to merge

Most of this PR is reference material:

  • migrate-cold/, spot-check/, verify-pack/, upload-cold/ — session artifacts.
  • bench-out/, BENCHMARK_GOAL.md, SESSION_LEARNINGS.md — reference docs and raw data.

The bench harness in cmd/stellar-rpc/scripts/bench-fullhistory/ is the most reusable piece; productizing would require tests, configurable paths, and integration with the actual JSON-RPC handler.

After the merge, the cherry-pick of PR #740 (bfa3860) is content-redundant — the upstream version of those files is what's on this branch now.

Simon Chow and others added 3 commits May 19, 2026 03:02
…king)

Surgical pull from PR stellar#740 (stellar#740) — brings in the
new full-history eventstore (hot + cold) plus the chunk ID package
that the eventstore depends on. Skipped pr-740's ledger/ reverts and
its stores/errors.go change, both of which would have rolled back the
PR stellar#739 cold ledger store already on rpc-hack.

What's pulled:
  - cmd/stellar-rpc/internal/fullhistory/pkg/chunk/        (new)
  - cmd/stellar-rpc/internal/fullhistory/pkg/stores/eventstore/ (new)
  - cmd/stellar-rpc/internal/events/{index,ingest,payload}.go (new)
  - cmd/stellar-rpc/internal/events/{ledgeroffsets,membitmaps,benchmark_test}.go (modified to match new API)
  - cmd/stellar-rpc/internal/fullhistory/pkg/rocksdb/rocksdb.go (adds BatchMultiGet)
  - go.mod / go.sum (adds RoaringBitmap/roaring v2, mmap-go, streamhash)

What's deleted (replaced by new files in pr-740):
  - cmd/stellar-rpc/internal/events/eventindex.go
  - cmd/stellar-rpc/internal/events/termkey.go

What's NOT pulled (would conflict with PR stellar#739):
  - ledger/cold_store{,_test}.go  (deletes in pr-740; we keep)
  - ledger/hot_store{,_test}.go   (pr-740's older version)
  - stores/errors.go              (pr-740's older version w/o L2 sentinels)

Note: requires Go 1.26 (streamhash dependency). This commit will
become redundant once PR stellar#740 is rebased onto current rpc-hack.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
POC tooling produced during the May 2026 benchmarking session. Not
intended as production code — all under cmd/stellar-rpc/scripts/, no
tests, single-purpose drivers.

Bench harness (cmd/stellar-rpc/scripts/bench-fullhistory/):
  - main.go         sub-command dispatcher, latency stats, CSV writer
  - seed_hot.go     ColdStoreReader -> ledger.HotStore (one chunk)
  - seed_txhash.go  Seeds txhash.HotStore + a sorted .bin cold index
  - seed_events.go  Seeds eventstore.HotStore + eventstore.ColdWriter
                    via events.LCMToPayloads; finishes with WriteColdIndex
  - bench_ledger.go ledger-point + ledger-range scenarios
  - bench_tx_page.go     tx-page (page of N transactions) scenario
  - bench_tx_hash.go     tx-hash (hash -> seq -> ledger -> tx) scenario
  - bench_events.go events scenarios: no-filter / contract / topic / both

Numbers + interpretation are in BENCHMARK_GOAL.md (next commit).
Methodology: same seed across tiers, warm-cache, single-threaded,
1000+ iterations per latency measurement. CSVs in bench-out/.

Other exploratory scripts kept for completeness:
  - cold-read/    ColdStoreReader random-sample reader (used as a
                  getLedger shim while no production reader exists)
  - migrate-cold/ One-shot migration old-format pack -> cold-format
                  pack (obsolete now that full-history-backfill
                  writes cold-format directly; kept for posterity)
  - spot-check/   100-sample readback verifier for old-format packs
  - upload-cold/  GCS uploader using ADC (cloud.google.com/go/storage)
                  -- used to mirror the 1.4 TB cold tree to
                  gs://rpc-full-history/cold/
  - verify-pack/  Single-pack trailer + ledger sanity check

The full-history-backfill script (already on this branch from PR stellar#743)
writes cold-format packs directly, so migrate-cold is purely a session
artifact.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
BENCHMARK_GOAL.md (289 lines):
  - Goal, dataset, methodology, scope carve-outs
  - Phase 0 harness scaffolding
  - Phase 1 ledger benchmarks (point + range, hot vs cold-NVMe)
  - Phase 2 tx-page benchmarks (page sizes 5/20/100)
  - Phase 3 tx-by-hash benchmarks (sorted-.bin shim cold side)
  - Phase 4 events benchmarks (4 filter scenarios x 2 tiers)
  - Interpretation: zstd dominates ledger queries (hot~cold within
    10%); events split by workload shape (cold 2.3x faster sequential,
    hot 1.8x faster low-cardinality scattered).
  - Caveats, artifact pointers

SESSION_LEARNINGS.md (~300 lines):
  - Repo state, data inventory, machine layout
  - Tooling locations, build & run quick reference
  - Decisions with rationale, surprises, code gotchas
  - Open follow-ups, what an incoming engineer should read first
  - Reproducing a single bench from scratch

bench-out/summary.csv:
  - One row per scenario (26 total), all percentiles in ms,
    ops/sec, plus avg events/iter where applicable.
  - Ready to paste into a sheet.
  - Per-scenario raw per-iteration CSVs are reproducible locally
    by re-running the bench harness; not checked in.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
@chowbao
chowbao force-pushed the chowbao/full-history-benchmarks-poc branch from eab64ab to b526fa8 Compare May 19, 2026 03:07
Simon Chow and others added 4 commits May 19, 2026 17:12
Brings in the upstream version of PR stellar#740 (eventstore) and PR stellar#743
(full-history-backfill), which superseded our local cherry-pick.

Conflict resolution:
- All eventstore + events files: took upstream (canonical post-PR-stellar#740).
- go.mod: kept both indirect deps (streamhash from stellar#740 + xxh3 local).
- go.sum: took upstream; reconciled via go mod tidy.

Post-merge build fixes:
- scripts/full-history-backfill/main.go: pkg/geometry was removed by
  upstream PR stellar#740 but the script still imported it. Switched to
  pkg/chunk.LedgersPerChunk (same value, new home).
- scripts/bench-fullhistory/bench_events.go: eventstore.Reader.Offsets()
  now returns (Offsets, error) — added error handling.
- scripts/bench-fullhistory/seed_events.go: eventstore.HotStore.EventCount()
  now returns (uint32, error) — added error handling.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Cold-cache (FADV_DONTNEED + mincore) and multi-worker concurrency
benchmarks for the full-history ledger store. Replaces the prior
single-thread / warm-cache methodology with a more realistic SLA
picture for tail-history reads.

New files:
- cache.go: posix_fadvise eviction + mincore residency check (Linux).
- bench_concurrent_runner.go: shared per-worker scaffold (evict, open
  ColdStoreReader, run workload, close, record); used by both point
  and range concurrent benches.
- bench_ledger_cold.go: ledger-point-cold-open (single worker,
  cold cache, fresh open per iter) and ledger-point-concurrent-cold
  (N workers, same cold-open semantics).
- bench_ledger_range_cold.go: ledger-range-concurrency-sweep — grid of
  (workers x page-size) cold-open range reads, prints per-cell p50/p99
  and a saturation summary (peak ops/sec per page size).

Modified:
- main.go: register the new sub-commands; expand usage.
- SESSION_LEARNINGS.md: Phase 5 addendum (cold-cache methodology,
  measured results, projected getLedgers throughput, caveats).
- CLAUDE.md: branch-scoped context for full-history work.

Results:
- bench-out/ledger-range-concurrency-sweep.csv: primary sweep
  (10 worker levels x page sizes 1/10/20).
- bench-out/rep2,rep3/: repeat sweeps for run-to-run variance.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Branch-scoped Claude context; not intended for upstream.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Full ledger-range-concurrency-sweep run after merging upstream/rpc-hack,
to verify the harness still produces consistent numbers under the
shared runColdConcurrent helper.

Saturation knees unchanged: page=1 → 48 workers, page=10/20 → 24
workers. All per-cell p50/p99/ops-per-sec within ~1-2% of the
original sweep (well inside the ~5% run-to-run variance band).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
@chowbao
chowbao marked this pull request as ready for review May 19, 2026 17:36
Branch-scoped session notes; not intended for upstream.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
@chowbao
chowbao merged commit aa59a39 into stellar:rpc-hack May 19, 2026
5 of 7 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant