Skip to content

TxHash ingestion + query benchmarks (both backends) #730

Description

@karthikiyer56

TL;DR

Benchmark both ingestion and query paths for the txhash store across both backends — RecSplit-backed (cold: #696 writer + cold reader) and RocksDB-backed (hot store, read + write) — to confirm we hit ingestion-throughput and query-latency SLAs across the full history. Symmetric to Tamir's ledger benchmarks (#726).

What to measure

Ingestion throughput

  • RecSplit (cold): txhashes / sec via the four-phase pipeline (COUNT, ADD, BUILD, VERIFY) at production-shaped tx-index dataset sizes.
  • RocksDB (hot): txhashes / sec via PutTxHash calls from a simulated live-ingestion harness.

Query latency

  • Cold backend (RecSplit) only: P50 / P90 / P95 / P99 for Lookup(txhash) across the 16 CFs.
  • Hot backend (RocksDB) only: P50 / P90 / P95 / P99 for Lookup(txhash) across the 16 CFs.
  • Federated reader (hot + cold): P50 / P90 / P95 / P99 with a representative production hot/cold mix.

Resource footprint

  • Memory and disk footprint at representative dataset sizes (e.g., 1 yr, 5 yr, full chain).

Acceptance

  • Benchmark report comparing both backends across all metrics above, runnable from a documented harness in the repo.
  • Federated-reader latency reflects the hot/cold mix expected in production.
  • Reproducible: the harness seeds the data, runs the benchmarks, emits structured numbers (CSV or JSON) for plotting.
  • Report identifies any bottleneck above the agreed SLA, with a triage note (e.g., "needs index tuning", "needs CF tuning", "out of scope for this milestone").

Dependencies

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions