Skip to content

fullhistory: index fee-bump transactions by both hashes - #862

Merged
tamirms merged 1 commit into
feature/full-historyfrom
fullhistory-feebump-inner-hash
Jul 16, 2026
Merged

tamirms merged 1 commit into
feature/full-historyfrom
fullhistory-feebump-inner-hash

Conversation

@tamirms

@tamirms tamirms commented Jul 16, 2026 •

Copy link
Copy Markdown
Contributor

What

getTransaction(innerHash) for a fee-bump transaction currently returns not-found from the full-history index: both txhash writers key only the outer (result-pair) hash, while the SQLite path stores lookup entries for both hashes (cmd/stellar-rpc/internal/db/transaction.go), so this is a behavioral regression against the API being replaced.

This change indexes both hashes. The rule lives once in the txhash package (EntryHashes: the transaction hash, plus a fee-bump's inner hash), and the hot writer and the cold .bin writer both derive their entries from it, so the two tiers always index the same hash set. The read path needs no change: the SDK's LedgerTransactionViewByHash resolves either hash since stellar/go-stellar-sdk#5964, so the fetch-and-verify step confirms inner-hash candidates natively.

Depends on

stellar/go-stellar-sdk#5964 (merged), which exposes the inner hash on the extraction walk; go.mod pins its pseudo-version.

Tests

TestTxhashColdWriter_FeeBumpBothHashes covers the cold writer end to end (both prefixes land in the .bin, mapped to the same seq), and the shared-walk byte-identity guard's reference contract now includes inner-hash entries. Hot-side matching is covered by the SDK's own tests; an e2e fee-bump case through the daemon can follow separately.

🤖 Generated with Claude Code

@tamirms
tamirms force-pushed the fullhistory-feebump-inner-hash branch 2 times, most recently from 4bb8c20 to 81a9669 Compare July 16, 2026 15:04
@tamirms
tamirms marked this pull request as ready for review July 16, 2026 15:04
@socket-security

socket-security Bot commented Jul 16, 2026 •

Copy link
Copy Markdown

Review the following changes in direct dependencies. Learn more about Socket for GitHub.

Diff Package Supply Chain
Security
Vulnerability Quality Maintenance License
Updatedgolang/​github.com/​stellar/​go-stellar-sdk@​v0.6.1-0.20260618191317-308407eca8c6 ⏵ v0.6.1-0.20260716145807-2bfffb159f3675 +1100100100100

View full report

@tamirms
tamirms force-pushed the fullhistory-feebump-inner-hash branch from 81a9669 to b79c5e8 Compare July 16, 2026 15:51
@tamirms
tamirms requested review from a team and chowbao July 16, 2026 16:44
@tamirms
tamirms force-pushed the fullhistory-feebump-inner-hash branch from b79c5e8 to 0b1b982 Compare July 16, 2026 22:01
The v1 RPC stores getTransaction lookup entries for both a fee-bump's
hashes (outer and inner); full-history indexed only the outer, so
getTransaction(innerHash) regressed to not-found. Both txhash writers now
emit one entry per hash — two for a fee-bump — using the inner hash the
SDK's extraction walk exposes from the result pair it already visits. The
read path needs no change: LedgerTransactionViewByHash resolves either
hash since the same SDK change.

Requires the go-stellar-sdk fee-bump inner-hash extractors
(stellar/go-stellar-sdk#5964), pinned by the go.mod bump.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@tamirms
tamirms force-pushed the fullhistory-feebump-inner-hash branch from 0b1b982 to 24fd5c0 Compare July 16, 2026 22:02
@tamirms
tamirms merged commit cb59867 into feature/full-history Jul 16, 2026
8 of 15 checks passed
@tamirms
tamirms deleted the fullhistory-feebump-inner-hash branch July 16, 2026 22:07
tamirms pushed a commit that referenced this pull request Jul 30, 2026
The atomic flip onto the packed-row engine. Ingest writes ONE
seq-keyed row per ledger (hashes sorted in PhaseExtract, fee-bump
outer+inner per #862) instead of ~6,000 hash-keyed Puts — the measured
~10.9ms of per-ledger memtable-insert CPU those random 32B keys cost
(commit p50 20.5 -> 7.9ms in the 1k solo cell). Warmup replays the
dense row chain, trips loudly on the old format, and arms sealing only
after validation; the post-commit hook applies txhash before events
(independent, by convention). Read-only opens disable txhash queries
structurally — the freeze never queries. The freeze merges the sealed
runs with each tail row as its own sorted source (RAM flat at any tail
size) and emits duplicates verbatim: cold .bin bytes stay identical to
the walk path, gated by the composition test and freeze-at-crash-point
identity tests. Freeze txhash arm: 1m25 -> 17.9s at full chunk scale.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01C3xaaoD9LhzyKPdnZCJa2N
tamirms pushed a commit that referenced this pull request Jul 30, 2026
Both tiers take the ledger's raw []byte (the callers held it all along —
the stream loop, drain) instead of a view: IngestLedger/HotService.Ingest
and coldChunk.ingest wrap the view internally, and the ledgers writers
consume the bytes directly — never a re-size of the whole meta to get
them back.

The extracted per-tx slice is now consumed in ONE pass per tier: the hot
loop and coldChunk.walk accumulate the indexable tx hashes (outer, plus
inner for a fee-bump — #862, with fee-bump slack presizing) and feed
events.PayloadShaper together. Cold writer signatures take what the pass
produced — txhash.write([][32]byte), events.write([]Payload) — so
shaping leaves the events writer and its cost moves into the
ledger-scoped ColdExtract signal (comments and metric help updated); the
events failed-write latch now triggers on a garbage payload
(TermsForBytes) since shaping left write. The cold extract window reads
the close time first, so a garbage ledger now rejects there (one pinned
test string updated; the abort/no-artifact invariants are unchanged).

Byte identity across the reshape is pinned by the freeze-vs-walk and
shared-walk gates, which run against this exact accumulation order.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01C3xaaoD9LhzyKPdnZCJa2N
tamirms pushed a commit that referenced this pull request Jul 31, 2026
The atomic flip onto the packed-row engine. Ingest writes ONE
seq-keyed row per ledger (hashes sorted in PhaseExtract, fee-bump
outer+inner per #862) instead of ~6,000 hash-keyed Puts — the measured
~10.9ms of per-ledger memtable-insert CPU those random 32B keys cost
(commit p50 20.5 -> 7.9ms in the 1k solo cell). Warmup replays the
dense row chain, trips loudly on the old format, and arms sealing only
after validation; the post-commit hook applies txhash before events
(independent, by convention). Read-only opens disable txhash queries
structurally — the freeze never queries. The freeze merges the sealed
runs with each tail row as its own sorted source (RAM flat at any tail
size) and emits duplicates verbatim: cold .bin bytes stay identical to
the walk path, gated by the composition test and freeze-at-crash-point
identity tests. Freeze txhash arm: 1m25 -> 17.9s at full chunk scale.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01C3xaaoD9LhzyKPdnZCJa2N
tamirms pushed a commit that referenced this pull request Jul 31, 2026
Both tiers take the ledger's raw []byte (the callers held it all along —
the stream loop, drain) instead of a view: IngestLedger/HotService.Ingest
and coldChunk.ingest wrap the view internally, and the ledgers writers
consume the bytes directly — never a re-size of the whole meta to get
them back.

The extracted per-tx slice is now consumed in ONE pass per tier: the hot
loop and coldChunk.walk accumulate the indexable tx hashes (outer, plus
inner for a fee-bump — #862, with fee-bump slack presizing) and feed
events.PayloadShaper together. Cold writer signatures take what the pass
produced — txhash.write([][32]byte), events.write([]Payload) — so
shaping leaves the events writer and its cost moves into the
ledger-scoped ColdExtract signal (comments and metric help updated); the
events failed-write latch now triggers on a garbage payload
(TermsForBytes) since shaping left write. The cold extract window reads
the close time first, so a garbage ledger now rejects there (one pinned
test string updated; the abort/no-artifact invariants are unchanged).

Byte identity across the reshape is pinned by the freeze-vs-walk and
shared-walk gates, which run against this exact accumulation order.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01C3xaaoD9LhzyKPdnZCJa2N
tamirms pushed a commit that referenced this pull request Aug 3, 2026
The atomic flip onto the packed-row engine. Ingest writes ONE
seq-keyed row per ledger (hashes sorted in PhaseExtract, fee-bump
outer+inner per #862) instead of ~6,000 hash-keyed Puts — the measured
~10.9ms of per-ledger memtable-insert CPU those random 32B keys cost
(commit p50 20.5 -> 7.9ms in the 1k solo cell). Warmup replays the
dense row chain, trips loudly on the old format, and arms sealing only
after validation; the post-commit hook applies txhash before events
(independent, by convention). Read-only opens disable txhash queries
structurally — the freeze never queries. The freeze merges the sealed
runs with each tail row as its own sorted source (RAM flat at any tail
size) and emits duplicates verbatim: cold .bin bytes stay identical to
the walk path, gated by the composition test and freeze-at-crash-point
identity tests. Freeze txhash arm: 1m25 -> 17.9s at full chunk scale.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01C3xaaoD9LhzyKPdnZCJa2N
tamirms pushed a commit that referenced this pull request Aug 3, 2026
Both tiers take the ledger's raw []byte (the callers held it all along —
the stream loop, drain) instead of a view: IngestLedger/HotService.Ingest
and coldChunk.ingest wrap the view internally, and the ledgers writers
consume the bytes directly — never a re-size of the whole meta to get
them back.

The extracted per-tx slice is now consumed in ONE pass per tier: the hot
loop and coldChunk.walk accumulate the indexable tx hashes (outer, plus
inner for a fee-bump — #862, with fee-bump slack presizing) and feed
events.PayloadShaper together. Cold writer signatures take what the pass
produced — txhash.write([][32]byte), events.write([]Payload) — so
shaping leaves the events writer and its cost moves into the
ledger-scoped ColdExtract signal (comments and metric help updated); the
events failed-write latch now triggers on a garbage payload
(TermsForBytes) since shaping left write. The cold extract window reads
the close time first, so a garbage ledger now rejects there (one pinned
test string updated; the abort/no-artifact invariants are unchanged).

Byte identity across the reshape is pinned by the freeze-vs-walk and
shared-walk gates, which run against this exact accumulation order.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01C3xaaoD9LhzyKPdnZCJa2N
tamirms pushed a commit that referenced this pull request Aug 4, 2026
…s marshal, multithreaded zstd

The second latency pass: the commit phase's remaining floors — ~6,000 random txhash memtable keys, ~30k events write-path allocations, and the serial ledger encode — each go.

--- txhash: packed-row hot engine — window + sealed runs, point lookups
The hot tx-hash tier's storage engine, standalone (nothing wired): one
sorted 32B-hash row per ledger in the window, sealed every 256 ledgers
into hash-sorted 36B-record run files (TXHRUN01, CRC64, fsync-dir
ladder), routed by per-run blooms and a 16B-prefix page ladder — one
aligned pread per hit, full-hash verify, boundary-tie two-page rule.
Manifest-anchored recovery with drain-verified opens; the engine is
born DISARMED (sealing rights follow validation, the events-tier rule)
and enforces a dense per-ledger chain. No merge tier, no overlay, no
retired handles: a point lookup needs none of them.

--- txhash: one packed row per ledger — flip ingest, freeze, and reads
The atomic flip onto the packed-row engine. Ingest writes ONE
seq-keyed row per ledger (hashes sorted in PhaseExtract, fee-bump
outer+inner per #862) instead of ~6,000 hash-keyed Puts — the measured
~10.9ms of per-ledger memtable-insert CPU those random 32B keys cost
(commit p50 20.5 -> 7.9ms in the 1k solo cell). Warmup replays the
dense row chain, trips loudly on the old format, and arms sealing only
after validation; the post-commit hook applies txhash before events
(independent, by convention). Read-only opens disable txhash queries
structurally — the freeze never queries. The freeze merges the sealed
runs with each tail row as its own sorted source (RAM flat at any tail
size) and emits duplicates verbatim: cold .bin bytes stay identical to
the walk path, gated by the composition test and freeze-at-crash-point
identity tests. Freeze txhash arm: 1m25 -> 17.9s at full chunk scale.

--- events: flat-pair term accumulation, bucketed sort, derivation arenas
The events queue step drops its per-ledger map: the marshal loop
appends (term, id) pairs into writer-owned arenas (AppendTerms writes
term keys straight into a caller arena; topics walk via Count()+Raw()
with All()'s reject-on-truncation semantics preserved and golden-swept;
the per-event data key is one hoisted scratch), then a 256-bucket MSD
scatter + big-endian word-comparator sort over KEYS ONLY (sorting 20B
pairs measured 2-2.5x slower; stability = index tiebreak, property-
tested as the exact permutation of a stable bytes.Compare) and one
linear pass emits the exact-sized packed row. ApplyLedger takes the
sorted term-runs directly; promotion and overlay semantics are pinned
equivalent to the retired map path, which survives as test-only
reference code for the byte-identity differential gates. Phase p50
8.0 -> 5.9ms and ~30k write-path allocations/ledger -> ~1.

--- zstd: multithreaded ledger encode, config-selected (default 2 workers)
After the txhash/events/extract cuts, the ~21ms single-threaded zstd
encode became the binding floor of the pre-commit section (its join
wait absorbed further wins). WithWorkers enables libzstd's internal
multithreading — still ONE standard deterministic frame, so
FrameHeaderValid and the cold-inherits-hot verbatim-copy contract are
untouched, at a measured ~0.1% ratio cost. The workers count is a real
configuration field, not an env experiment: measurement settled on
default 2 (equal to 3 within noise — total p50 33.4 -> 29.2ms, join
7.35 -> 1.9ms — while claiming one fewer core). It is FORMAT-AFFECTING
for the stored ledger frames, so ONE resolved value
(storage.zstd_encode_workers in the daemon TOML, --zstd-workers on the
bench hot cell; 0 = explicit single-threaded, validated >= 0) feeds
BOTH encoders — hotchunk.Tuning.ZstdEncodeWorkers for hot ingest and
ingest.Config.ZstdEncodeWorkers for the walk/backfill cold writer —
because the freeze copies hot frames into the cold pack verbatim while
the walk re-encodes the same ledgers, and a chunk's pack must stay
byte-identical whichever materializer built it (the freeze-vs-walk
gates arbitrate, now exercising the MT default on both sides).
NewCompressor panics loudly on a non-MT libzstd rather than silently
degrading the format.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013zyTXU8wkocJgN6mBafnou
tamirms pushed a commit that referenced this pull request Aug 4, 2026
The redesigned views/streaming SDK consumed end to end; the pin makes this PR deliberately unmergeable until the SDK branch lands.

--- events: PayloadShaper — cursor-ordered shaping goes per-transaction
PayloadsFromLedgerEvents shaped a whole-ledger slice in three passes
(BeforeAllTxs / per-tx op+AfterTx / AfterAllTxs). PayloadShaper produces
the SAME cursor-ordered payload sequence from a per-transaction feed:
Add appends each tx's events into the three cursor-group buckets in
apply order, Finish concatenates them — the buckets ARE the old passes,
so event-ID assignment is unchanged (the shared-walk and freeze-vs-walk
byte-identity gates pin it). appendStageEventPayloads survives unchanged
as the shared per-event stage filter. countPayloads dies with the
three-pass shaping; both tiers drive the shaper over the extracted
per-tx slice (hot shapePayloads, cold events.write).

Also converges TermsForBytes onto AppendTerms — one term derivation for
the arena-owning hot path and the per-call cold path — with the golden
sweep now pinning against an independent test-local All()-slice
reference (the old TermsForBytes body).

This is groundwork for the streaming SDK extraction (next commits): a
per-tx consumer no longer needs the whole-ledger slice materialized.

--- ingest: thread raw ledger bytes; one pass feeds txhash and events
Both tiers take the ledger's raw []byte (the callers held it all along —
the stream loop, drain) instead of a view: IngestLedger/HotService.Ingest
and coldChunk.ingest wrap the view internally, and the ledgers writers
consume the bytes directly — never a re-size of the whole meta to get
them back.

The extracted per-tx slice is now consumed in ONE pass per tier: the hot
loop and coldChunk.walk accumulate the indexable tx hashes (outer, plus
inner for a fee-bump — #862, with fee-bump slack presizing) and feed
events.PayloadShaper together. Cold writer signatures take what the pass
produced — txhash.write([][32]byte), events.write([]Payload) — so
shaping leaves the events writer and its cost moves into the
ledger-scoped ColdExtract signal (comments and metric help updated); the
events failed-write latch now triggers on a garbage payload
(TermsForBytes) since shaping left write. The cold extract window reads
the close time first, so a garbage ledger now rejects there (one pinned
test string updated; the abort/no-artifact invariants are unchanged).

Byte identity across the reshape is pinned by the freeze-vs-walk and
shared-walk gates, which run against this exact accumulation order.

--- sdk: pin views-walk; migrate onto the redesigned views API
Pins github.com/stellar/go-stellar-sdk to
v0.6.1-0.20260727053836-9156a311aae9 — the pushed tip of the views-walk
branch, the redesigned views/streaming SDK this stack consumes (struct
views with NewXView constructors, Len()/ArmV0()/Scan(), and
ingest.StreamLedgerEvents replacing the deleted
ExtractLedgerEvents/ExtractTxHashes). views-walk is NOT yet merged into
the SDK's main branch, so this PR is deliberately unmergeable until it
lands — kept that way for review visibility of the real dependency.

The pin and the call-site swap are ONE commit by necessity, not haste:
the new SDK deletes the old API, so no tree exists that compiles with
the pin but without the swap (or vice versa) — an intermediate
pin-only or swap-only commit cannot build. The swap itself is
MECHANICAL: constructor/name substitutions (XView casts -> NewXView,
Count -> Len, V0 -> ArmV0, the error-yielding topics iteration ->
Scan), and the extracted-slice consumption becoming the
StreamLedgerEvents begin/per-tx hooks — the semantics were all landed
and reviewed in the preceding commits (PayloadShaper shaping, raw-bytes
threading, one-pass hash+payload accumulation), and the freeze-vs-walk
and shared-walk BYTE-IDENTITY gates plus the golden term sweep (its
reference moves All-slice -> Scan, the surviving error-returning form)
pin that no artifact byte moved across the swap.

One walk-visible behavior note, pinned by the gates: the stream
delivers per-transaction with no whole-ledger materialization, so
per-ledger extraction drops the []LedgerTransactionEvents slice + its
spines' growth entirely (A/B: extract p50 -46%, hot ingest_total
34.43/47.61 -> 28.74/39.58 p50/p99).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013zyTXU8wkocJgN6mBafnou
tamirms pushed a commit that referenced this pull request Aug 5, 2026
…s marshal, multithreaded zstd

The second latency pass: the commit phase's remaining floors — ~6,000 random txhash memtable keys, ~30k events write-path allocations, and the serial ledger encode — each go.

--- txhash: packed-row hot engine — window + sealed runs, point lookups
The hot tx-hash tier's storage engine, standalone (nothing wired): one
sorted 32B-hash row per ledger in the window, sealed every 256 ledgers
into hash-sorted 36B-record run files (TXHRUN01, CRC64, fsync-dir
ladder), routed by per-run blooms and a 16B-prefix page ladder — one
aligned pread per hit, full-hash verify, boundary-tie two-page rule.
Manifest-anchored recovery with drain-verified opens; the engine is
born DISARMED (sealing rights follow validation, the events-tier rule)
and enforces a dense per-ledger chain. No merge tier, no overlay, no
retired handles: a point lookup needs none of them.

--- txhash: one packed row per ledger — flip ingest, freeze, and reads
The atomic flip onto the packed-row engine. Ingest writes ONE
seq-keyed row per ledger (hashes sorted in PhaseExtract, fee-bump
outer+inner per #862) instead of ~6,000 hash-keyed Puts — the measured
~10.9ms of per-ledger memtable-insert CPU those random 32B keys cost
(commit p50 20.5 -> 7.9ms in the 1k solo cell). Warmup replays the
dense row chain, trips loudly on the old format, and arms sealing only
after validation; the post-commit hook applies txhash before events
(independent, by convention). Read-only opens disable txhash queries
structurally — the freeze never queries. The freeze merges the sealed
runs with each tail row as its own sorted source (RAM flat at any tail
size) and emits duplicates verbatim: cold .bin bytes stay identical to
the walk path, gated by the composition test and freeze-at-crash-point
identity tests. Freeze txhash arm: 1m25 -> 17.9s at full chunk scale.

--- events: flat-pair term accumulation, bucketed sort, derivation arenas
The events queue step drops its per-ledger map: the marshal loop
appends (term, id) pairs into writer-owned arenas (AppendTerms writes
term keys straight into a caller arena; topics walk via Count()+Raw()
with All()'s reject-on-truncation semantics preserved and golden-swept;
the per-event data key is one hoisted scratch), then a 256-bucket MSD
scatter + big-endian word-comparator sort over KEYS ONLY (sorting 20B
pairs measured 2-2.5x slower; stability = index tiebreak, property-
tested as the exact permutation of a stable bytes.Compare) and one
linear pass emits the exact-sized packed row. ApplyLedger takes the
sorted term-runs directly; promotion and overlay semantics are pinned
equivalent to the retired map path, which survives as test-only
reference code for the byte-identity differential gates. Phase p50
8.0 -> 5.9ms and ~30k write-path allocations/ledger -> ~1.

--- zstd: multithreaded ledger encode, config-selected (default 2 workers)
After the txhash/events/extract cuts, the ~21ms single-threaded zstd
encode became the binding floor of the pre-commit section (its join
wait absorbed further wins). WithWorkers enables libzstd's internal
multithreading — still ONE standard deterministic frame, so
FrameHeaderValid and the cold-inherits-hot verbatim-copy contract are
untouched, at a measured ~0.1% ratio cost. The workers count is a real
configuration field, not an env experiment: measurement settled on
default 2 (equal to 3 within noise — total p50 33.4 -> 29.2ms, join
7.35 -> 1.9ms — while claiming one fewer core). It is FORMAT-AFFECTING
for the stored ledger frames, so ONE resolved value
(storage.zstd_encode_workers in the daemon TOML, --zstd-workers on the
bench hot cell; 0 = explicit single-threaded, validated >= 0) feeds
BOTH encoders — hotchunk.Tuning.ZstdEncodeWorkers for hot ingest and
ingest.Config.ZstdEncodeWorkers for the walk/backfill cold writer —
because the freeze copies hot frames into the cold pack verbatim while
the walk re-encodes the same ledgers, and a chunk's pack must stay
byte-identical whichever materializer built it (the freeze-vs-walk
gates arbitrate, now exercising the MT default on both sides).
NewCompressor panics loudly on a non-MT libzstd rather than silently
degrading the format.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013zyTXU8wkocJgN6mBafnou
tamirms pushed a commit that referenced this pull request Aug 16, 2026
…s marshal, multithreaded zstd

The second latency pass: the commit phase's remaining floors — ~6,000 random txhash memtable keys, ~30k events write-path allocations, and the serial ledger encode — each go.

--- txhash: packed-row hot engine — window + sealed runs, point lookups
The hot tx-hash tier's storage engine, standalone (nothing wired): one
sorted 32B-hash row per ledger in the window, sealed every 256 ledgers
into hash-sorted 36B-record run files (TXHRUN01, CRC64, fsync-dir
ladder), routed by per-run blooms and a 16B-prefix page ladder — one
aligned pread per hit, full-hash verify, boundary-tie two-page rule.
Manifest-anchored recovery with drain-verified opens; the engine is
born DISARMED (sealing rights follow validation, the events-tier rule)
and enforces a dense per-ledger chain. No merge tier, no overlay, no
retired handles: a point lookup needs none of them.

--- txhash: one packed row per ledger — flip ingest, freeze, and reads
The atomic flip onto the packed-row engine. Ingest writes ONE
seq-keyed row per ledger (hashes sorted in PhaseExtract, fee-bump
outer+inner per #862) instead of ~6,000 hash-keyed Puts — the measured
~10.9ms of per-ledger memtable-insert CPU those random 32B keys cost
(commit p50 20.5 -> 7.9ms in the 1k solo cell). Warmup replays the
dense row chain, trips loudly on the old format, and arms sealing only
after validation; the post-commit hook applies txhash before events
(independent, by convention). Read-only opens disable txhash queries
structurally — the freeze never queries. The freeze merges the sealed
runs with each tail row as its own sorted source (RAM flat at any tail
size) and emits duplicates verbatim: cold .bin bytes stay identical to
the walk path, gated by the composition test and freeze-at-crash-point
identity tests. Freeze txhash arm: 1m25 -> 17.9s at full chunk scale.

--- events: flat-pair term accumulation, bucketed sort, derivation arenas
The events queue step drops its per-ledger map: the marshal loop
appends (term, id) pairs into writer-owned arenas (AppendTerms writes
term keys straight into a caller arena; topics walk via Count()+Raw()
with All()'s reject-on-truncation semantics preserved and golden-swept;
the per-event data key is one hoisted scratch), then a 256-bucket MSD
scatter + big-endian word-comparator sort over KEYS ONLY (sorting 20B
pairs measured 2-2.5x slower; stability = index tiebreak, property-
tested as the exact permutation of a stable bytes.Compare) and one
linear pass emits the exact-sized packed row. ApplyLedger takes the
sorted term-runs directly; promotion and overlay semantics are pinned
equivalent to the retired map path, which survives as test-only
reference code for the byte-identity differential gates. Phase p50
8.0 -> 5.9ms and ~30k write-path allocations/ledger -> ~1.

--- zstd: multithreaded ledger encode, config-selected (default 2 workers)
After the txhash/events/extract cuts, the ~21ms single-threaded zstd
encode became the binding floor of the pre-commit section (its join
wait absorbed further wins). WithWorkers enables libzstd's internal
multithreading — still ONE standard deterministic frame, so
FrameHeaderValid and the cold-inherits-hot verbatim-copy contract are
untouched, at a measured ~0.1% ratio cost. The workers count is a real
configuration field, not an env experiment: measurement settled on
default 2 (equal to 3 within noise — total p50 33.4 -> 29.2ms, join
7.35 -> 1.9ms — while claiming one fewer core). It is FORMAT-AFFECTING
for the stored ledger frames, so ONE resolved value
(storage.zstd_encode_workers in the daemon TOML, --zstd-workers on the
bench hot cell; 0 = explicit single-threaded, validated >= 0) feeds
BOTH encoders — hotchunk.Tuning.ZstdEncodeWorkers for hot ingest and
ingest.Config.ZstdEncodeWorkers for the walk/backfill cold writer —
because the freeze copies hot frames into the cold pack verbatim while
the walk re-encodes the same ledgers, and a chunk's pack must stay
byte-identical whichever materializer built it (the freeze-vs-walk
gates arbitrate, now exercising the MT default on both sides).
NewCompressor panics loudly on a non-MT libzstd rather than silently
degrading the format.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013zyTXU8wkocJgN6mBafnou
tamirms pushed a commit that referenced this pull request Aug 25, 2026
…s marshal, multithreaded zstd

The second latency pass: the commit phase's remaining floors — ~6,000 random txhash memtable keys, ~30k events write-path allocations, and the serial ledger encode — each go.

--- txhash: packed-row hot engine — window + sealed runs, point lookups
The hot tx-hash tier's storage engine, standalone (nothing wired): one
sorted 32B-hash row per ledger in the window, sealed every 256 ledgers
into hash-sorted 36B-record run files (TXHRUN01, CRC64, fsync-dir
ladder), routed by per-run blooms and a 16B-prefix page ladder — one
aligned pread per hit, full-hash verify, boundary-tie two-page rule.
Manifest-anchored recovery with drain-verified opens; the engine is
born DISARMED (sealing rights follow validation, the events-tier rule)
and enforces a dense per-ledger chain. No merge tier, no overlay, no
retired handles: a point lookup needs none of them.

--- txhash: one packed row per ledger — flip ingest, freeze, and reads
The atomic flip onto the packed-row engine. Ingest writes ONE
seq-keyed row per ledger (hashes sorted in PhaseExtract, fee-bump
outer+inner per #862) instead of ~6,000 hash-keyed Puts — the measured
~10.9ms of per-ledger memtable-insert CPU those random 32B keys cost
(commit p50 20.5 -> 7.9ms in the 1k solo cell). Warmup replays the
dense row chain, trips loudly on the old format, and arms sealing only
after validation; the post-commit hook applies txhash before events
(independent, by convention). Read-only opens disable txhash queries
structurally — the freeze never queries. The freeze merges the sealed
runs with each tail row as its own sorted source (RAM flat at any tail
size) and emits duplicates verbatim: cold .bin bytes stay identical to
the walk path, gated by the composition test and freeze-at-crash-point
identity tests. Freeze txhash arm: 1m25 -> 17.9s at full chunk scale.

--- events: flat-pair term accumulation, bucketed sort, derivation arenas
The events queue step drops its per-ledger map: the marshal loop
appends (term, id) pairs into writer-owned arenas (AppendTerms writes
term keys straight into a caller arena; topics walk via Count()+Raw()
with All()'s reject-on-truncation semantics preserved and golden-swept;
the per-event data key is one hoisted scratch), then a 256-bucket MSD
scatter + big-endian word-comparator sort over KEYS ONLY (sorting 20B
pairs measured 2-2.5x slower; stability = index tiebreak, property-
tested as the exact permutation of a stable bytes.Compare) and one
linear pass emits the exact-sized packed row. ApplyLedger takes the
sorted term-runs directly; promotion and overlay semantics are pinned
equivalent to the retired map path, which survives as test-only
reference code for the byte-identity differential gates. Phase p50
8.0 -> 5.9ms and ~30k write-path allocations/ledger -> ~1.

--- zstd: multithreaded ledger encode, config-selected (default 2 workers)
After the txhash/events/extract cuts, the ~21ms single-threaded zstd
encode became the binding floor of the pre-commit section (its join
wait absorbed further wins). WithWorkers enables libzstd's internal
multithreading — still ONE standard deterministic frame, so
FrameHeaderValid and the cold-inherits-hot verbatim-copy contract are
untouched, at a measured ~0.1% ratio cost. The workers count is a real
configuration field, not an env experiment: measurement settled on
default 2 (equal to 3 within noise — total p50 33.4 -> 29.2ms, join
7.35 -> 1.9ms — while claiming one fewer core). It is FORMAT-AFFECTING
for the stored ledger frames, so ONE resolved value
(storage.zstd_encode_workers in the daemon TOML, --zstd-workers on the
bench hot cell; 0 = explicit single-threaded, validated >= 0) feeds
BOTH encoders — hotchunk.Tuning.ZstdEncodeWorkers for hot ingest and
ingest.Config.ZstdEncodeWorkers for the walk/backfill cold writer —
because the freeze copies hot frames into the cold pack verbatim while
the walk re-encodes the same ledgers, and a chunk's pack must stay
byte-identical whichever materializer built it (the freeze-vs-walk
gates arbitrate, now exercising the MT default on both sides).
NewCompressor panics loudly on a non-MT libzstd rather than silently
degrading the format.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013zyTXU8wkocJgN6mBafnou
tamirms pushed a commit that referenced this pull request Aug 25, 2026
…s marshal, multithreaded zstd

The second latency pass: the commit phase's remaining floors — ~6,000 random txhash memtable keys, ~30k events write-path allocations, and the serial ledger encode — each go.

--- txhash: packed-row hot engine — window + sealed runs, point lookups
The hot tx-hash tier's storage engine, standalone (nothing wired): one
sorted 32B-hash row per ledger in the window, sealed every 256 ledgers
into hash-sorted 36B-record run files (TXHRUN01, CRC64, fsync-dir
ladder), routed by per-run blooms and a 16B-prefix page ladder — one
aligned pread per hit, full-hash verify, boundary-tie two-page rule.
Manifest-anchored recovery with drain-verified opens; the engine is
born DISARMED (sealing rights follow validation, the events-tier rule)
and enforces a dense per-ledger chain. No merge tier, no overlay, no
retired handles: a point lookup needs none of them.

--- txhash: one packed row per ledger — flip ingest, freeze, and reads
The atomic flip onto the packed-row engine. Ingest writes ONE
seq-keyed row per ledger (hashes sorted in PhaseExtract, fee-bump
outer+inner per #862) instead of ~6,000 hash-keyed Puts — the measured
~10.9ms of per-ledger memtable-insert CPU those random 32B keys cost
(commit p50 20.5 -> 7.9ms in the 1k solo cell). Warmup replays the
dense row chain, trips loudly on the old format, and arms sealing only
after validation; the post-commit hook applies txhash before events
(independent, by convention). Read-only opens disable txhash queries
structurally — the freeze never queries. The freeze merges the sealed
runs with each tail row as its own sorted source (RAM flat at any tail
size) and emits duplicates verbatim: cold .bin bytes stay identical to
the walk path, gated by the composition test and freeze-at-crash-point
identity tests. Freeze txhash arm: 1m25 -> 17.9s at full chunk scale.

--- events: flat-pair term accumulation, bucketed sort, derivation arenas
The events queue step drops its per-ledger map: the marshal loop
appends (term, id) pairs into writer-owned arenas (AppendTerms writes
term keys straight into a caller arena; topics walk via Count()+Raw()
with All()'s reject-on-truncation semantics preserved and golden-swept;
the per-event data key is one hoisted scratch), then a 256-bucket MSD
scatter + big-endian word-comparator sort over KEYS ONLY (sorting 20B
pairs measured 2-2.5x slower; stability = index tiebreak, property-
tested as the exact permutation of a stable bytes.Compare) and one
linear pass emits the exact-sized packed row. ApplyLedger takes the
sorted term-runs directly; promotion and overlay semantics are pinned
equivalent to the retired map path, which survives as test-only
reference code for the byte-identity differential gates. Phase p50
8.0 -> 5.9ms and ~30k write-path allocations/ledger -> ~1.

--- zstd: multithreaded ledger encode, config-selected (default 2 workers)
After the txhash/events/extract cuts, the ~21ms single-threaded zstd
encode became the binding floor of the pre-commit section (its join
wait absorbed further wins). WithWorkers enables libzstd's internal
multithreading — still ONE standard deterministic frame, so
FrameHeaderValid and the cold-inherits-hot verbatim-copy contract are
untouched, at a measured ~0.1% ratio cost. The workers count is a real
configuration field, not an env experiment: measurement settled on
default 2 (equal to 3 within noise — total p50 33.4 -> 29.2ms, join
7.35 -> 1.9ms — while claiming one fewer core). It is FORMAT-AFFECTING
for the stored ledger frames, so ONE resolved value
(storage.zstd_encode_workers in the daemon TOML, --zstd-workers on the
bench hot cell; 0 = explicit single-threaded, validated >= 0) feeds
BOTH encoders — hotchunk.Tuning.ZstdEncodeWorkers for hot ingest and
ingest.Config.ZstdEncodeWorkers for the walk/backfill cold writer —
because the freeze copies hot frames into the cold pack verbatim while
the walk re-encodes the same ledgers, and a chunk's pack must stay
byte-identical whichever materializer built it (the freeze-vs-walk
gates arbitrate, now exercising the MT default on both sides).
NewCompressor panics loudly on a non-MT libzstd rather than silently
degrading the format.

Route the two constant-key event terms (event type, topic count) around
the pair sort: every event emits the same few keys, so their pairs
collapse into one radix bucket and the comparison sort re-derives an
order the arena already has. termlanes.go collects their ids in per-key
ascending lanes and the run build merges them at their byte-order
positions; rows are byte-identical (differential-tested). Measured on
sac6000: -1.4ms/ledger in the events phase.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013zyTXU8wkocJgN6mBafnou
tamirms pushed a commit that referenced this pull request Aug 25, 2026
…s marshal, multithreaded zstd

The second latency pass: the commit phase's remaining floors — ~6,000 random txhash memtable keys, ~30k events write-path allocations, and the serial ledger encode — each go.

--- txhash: packed-row hot engine — window + sealed runs, point lookups
The hot tx-hash tier's storage engine, standalone (nothing wired): one
sorted 32B-hash row per ledger in the window, sealed every 256 ledgers
into hash-sorted 36B-record run files (TXHRUN01, CRC64, fsync-dir
ladder), routed by per-run blooms and a 16B-prefix page ladder — one
aligned pread per hit, full-hash verify, boundary-tie two-page rule.
Manifest-anchored recovery with drain-verified opens; the engine is
born DISARMED (sealing rights follow validation, the events-tier rule)
and enforces a dense per-ledger chain. No merge tier, no overlay, no
retired handles: a point lookup needs none of them.

--- txhash: one packed row per ledger — flip ingest, freeze, and reads
The atomic flip onto the packed-row engine. Ingest writes ONE
seq-keyed row per ledger (hashes sorted in PhaseExtract, fee-bump
outer+inner per #862) instead of ~6,000 hash-keyed Puts — the measured
~10.9ms of per-ledger memtable-insert CPU those random 32B keys cost
(commit p50 20.5 -> 7.9ms in the 1k solo cell). Warmup replays the
dense row chain, trips loudly on the old format, and arms sealing only
after validation; the post-commit hook applies txhash before events
(independent, by convention). Read-only opens disable txhash queries
structurally — the freeze never queries. The freeze merges the sealed
runs with each tail row as its own sorted source (RAM flat at any tail
size) and emits duplicates verbatim: cold .bin bytes stay identical to
the walk path, gated by the composition test and freeze-at-crash-point
identity tests. Freeze txhash arm: 1m25 -> 17.9s at full chunk scale.

--- events: flat-pair term accumulation, bucketed sort, derivation arenas
The events queue step drops its per-ledger map: the marshal loop
appends (term, id) pairs into writer-owned arenas (AppendTerms writes
term keys straight into a caller arena; topics walk via Count()+Raw()
with All()'s reject-on-truncation semantics preserved and golden-swept;
the per-event data key is one hoisted scratch), then a 256-bucket MSD
scatter + big-endian word-comparator sort over KEYS ONLY (sorting 20B
pairs measured 2-2.5x slower; stability = index tiebreak, property-
tested as the exact permutation of a stable bytes.Compare) and one
linear pass emits the exact-sized packed row. ApplyLedger takes the
sorted term-runs directly; promotion and overlay semantics are pinned
equivalent to the retired map path, which survives as test-only
reference code for the byte-identity differential gates. Phase p50
8.0 -> 5.9ms and ~30k write-path allocations/ledger -> ~1.

--- zstd: multithreaded ledger encode, config-selected (default 2 workers)
After the txhash/events/extract cuts, the ~21ms single-threaded zstd
encode became the binding floor of the pre-commit section (its join
wait absorbed further wins). WithWorkers enables libzstd's internal
multithreading — still ONE standard deterministic frame, so
FrameHeaderValid and the cold-inherits-hot verbatim-copy contract are
untouched, at a measured ~0.1% ratio cost. The workers count is a real
configuration field, not an env experiment: measurement settled on
default 2 (equal to 3 within noise — total p50 33.4 -> 29.2ms, join
7.35 -> 1.9ms — while claiming one fewer core). It is FORMAT-AFFECTING
for the stored ledger frames, so ONE resolved value
(storage.zstd_encode_workers in the daemon TOML, --zstd-workers on the
bench hot cell; 0 = explicit single-threaded, validated >= 0) feeds
BOTH encoders — hotchunk.Tuning.ZstdEncodeWorkers for hot ingest and
ingest.Config.ZstdEncodeWorkers for the walk/backfill cold writer —
because the freeze copies hot frames into the cold pack verbatim while
the walk re-encodes the same ledgers, and a chunk's pack must stay
byte-identical whichever materializer built it (the freeze-vs-walk
gates arbitrate, now exercising the MT default on both sides).
NewCompressor panics loudly on a non-MT libzstd rather than silently
degrading the format.

Route the two constant-key event terms (event type, topic count) around
the pair sort: every event emits the same few keys, so their pairs
collapse into one radix bucket and the comparison sort re-derives an
order the arena already has. termlanes.go collects their ids in per-key
ascending lanes and the run build merges them at their byte-order
positions; rows are byte-identical (differential-tested). Measured on
sac6000: -1.4ms/ledger in the events phase.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013zyTXU8wkocJgN6mBafnou
tamirms pushed a commit that referenced this pull request Aug 26, 2026
…s marshal, multithreaded zstd

The second latency pass: the commit phase's remaining floors — ~6,000 random txhash memtable keys, ~30k events write-path allocations, and the serial ledger encode — each go.

--- txhash: packed-row hot engine — window + sealed runs, point lookups
The hot tx-hash tier's storage engine, standalone (nothing wired): one
sorted 32B-hash row per ledger in the window, sealed every 256 ledgers
into hash-sorted 36B-record run files (TXHRUN01, CRC64, fsync-dir
ladder), routed by per-run blooms and a 16B-prefix page ladder — one
aligned pread per hit, full-hash verify, boundary-tie two-page rule.
Manifest-anchored recovery with drain-verified opens; the engine is
born DISARMED (sealing rights follow validation, the events-tier rule)
and enforces a dense per-ledger chain. No merge tier, no overlay, no
retired handles: a point lookup needs none of them.

--- txhash: one packed row per ledger — flip ingest, freeze, and reads
The atomic flip onto the packed-row engine. Ingest writes ONE
seq-keyed row per ledger (hashes sorted in PhaseExtract, fee-bump
outer+inner per #862) instead of ~6,000 hash-keyed Puts — the measured
~10.9ms of per-ledger memtable-insert CPU those random 32B keys cost
(commit p50 20.5 -> 7.9ms in the 1k solo cell). Warmup replays the
dense row chain, trips loudly on the old format, and arms sealing only
after validation; the post-commit hook applies txhash before events
(independent, by convention). Read-only opens disable txhash queries
structurally — the freeze never queries. The freeze merges the sealed
runs with each tail row as its own sorted source (RAM flat at any tail
size) and emits duplicates verbatim: cold .bin bytes stay identical to
the walk path, gated by the composition test and freeze-at-crash-point
identity tests. Freeze txhash arm: 1m25 -> 17.9s at full chunk scale.

--- events: flat-pair term accumulation, bucketed sort, derivation arenas
The events queue step drops its per-ledger map: the marshal loop
appends (term, id) pairs into writer-owned arenas (AppendTerms writes
term keys straight into a caller arena; topics walk via Count()+Raw()
with All()'s reject-on-truncation semantics preserved and golden-swept;
the per-event data key is one hoisted scratch), then a 256-bucket MSD
scatter + big-endian word-comparator sort over KEYS ONLY (sorting 20B
pairs measured 2-2.5x slower; stability = index tiebreak, property-
tested as the exact permutation of a stable bytes.Compare) and one
linear pass emits the exact-sized packed row. ApplyLedger takes the
sorted term-runs directly; promotion and overlay semantics are pinned
equivalent to the retired map path, which survives as test-only
reference code for the byte-identity differential gates. Phase p50
8.0 -> 5.9ms and ~30k write-path allocations/ledger -> ~1.

--- zstd: multithreaded ledger encode, config-selected (default 2 workers)
After the txhash/events/extract cuts, the ~21ms single-threaded zstd
encode became the binding floor of the pre-commit section (its join
wait absorbed further wins). WithWorkers enables libzstd's internal
multithreading — still ONE standard deterministic frame, so
FrameHeaderValid and the cold-inherits-hot verbatim-copy contract are
untouched, at a measured ~0.1% ratio cost. The workers count is a real
configuration field, not an env experiment: measurement settled on
default 2 (equal to 3 within noise — total p50 33.4 -> 29.2ms, join
7.35 -> 1.9ms — while claiming one fewer core). It is FORMAT-AFFECTING
for the stored ledger frames, so ONE resolved value
(storage.zstd_encode_workers in the daemon TOML, --zstd-workers on the
bench hot cell; 0 = explicit single-threaded, validated >= 0) feeds
BOTH encoders — hotchunk.Tuning.ZstdEncodeWorkers for hot ingest and
ingest.Config.ZstdEncodeWorkers for the walk/backfill cold writer —
because the freeze copies hot frames into the cold pack verbatim while
the walk re-encodes the same ledgers, and a chunk's pack must stay
byte-identical whichever materializer built it (the freeze-vs-walk
gates arbitrate, now exercising the MT default on both sides).
NewCompressor panics loudly on a non-MT libzstd rather than silently
degrading the format.

Route the two constant-key event terms (event type, topic count) around
the pair sort: every event emits the same few keys, so their pairs
collapse into one radix bucket and the comparison sort re-derives an
order the arena already has. termlanes.go collects their ids in per-key
ascending lanes and the run build merges them at their byte-order
positions; rows are byte-identical (differential-tested). Measured on
sac6000: -1.4ms/ledger in the events phase.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013zyTXU8wkocJgN6mBafnou
tamirms pushed a commit that referenced this pull request Aug 27, 2026
…s marshal, multithreaded zstd

The second latency pass: the commit phase's remaining floors — ~6,000 random txhash memtable keys, ~30k events write-path allocations, and the serial ledger encode — each go.

--- txhash: packed-row hot engine — window + sealed runs, point lookups
The hot tx-hash tier's storage engine, standalone (nothing wired): one
sorted 32B-hash row per ledger in the window, sealed every 256 ledgers
into hash-sorted 36B-record run files (TXHRUN01, CRC64, fsync-dir
ladder), routed by per-run blooms and a 16B-prefix page ladder — one
aligned pread per hit, full-hash verify, boundary-tie two-page rule.
Manifest-anchored recovery with drain-verified opens; the engine is
born DISARMED (sealing rights follow validation, the events-tier rule)
and enforces a dense per-ledger chain. No merge tier, no overlay, no
retired handles: a point lookup needs none of them.

--- txhash: one packed row per ledger — flip ingest, freeze, and reads
The atomic flip onto the packed-row engine. Ingest writes ONE
seq-keyed row per ledger (hashes sorted in PhaseExtract, fee-bump
outer+inner per #862) instead of ~6,000 hash-keyed Puts — the measured
~10.9ms of per-ledger memtable-insert CPU those random 32B keys cost
(commit p50 20.5 -> 7.9ms in the 1k solo cell). Warmup replays the
dense row chain, trips loudly on the old format, and arms sealing only
after validation; the post-commit hook applies txhash before events
(independent, by convention). Read-only opens disable txhash queries
structurally — the freeze never queries. The freeze merges the sealed
runs with each tail row as its own sorted source (RAM flat at any tail
size) and emits duplicates verbatim: cold .bin bytes stay identical to
the walk path, gated by the composition test and freeze-at-crash-point
identity tests. Freeze txhash arm: 1m25 -> 17.9s at full chunk scale.

--- events: flat-pair term accumulation, bucketed sort, derivation arenas
The events queue step drops its per-ledger map: the marshal loop
appends (term, id) pairs into writer-owned arenas (AppendTerms writes
term keys straight into a caller arena; topics walk via Count()+Raw()
with All()'s reject-on-truncation semantics preserved and golden-swept;
the per-event data key is one hoisted scratch), then a 256-bucket MSD
scatter + big-endian word-comparator sort over KEYS ONLY (sorting 20B
pairs measured 2-2.5x slower; stability = index tiebreak, property-
tested as the exact permutation of a stable bytes.Compare) and one
linear pass emits the exact-sized packed row. ApplyLedger takes the
sorted term-runs directly; promotion and overlay semantics are pinned
equivalent to the retired map path, which survives as test-only
reference code for the byte-identity differential gates. Phase p50
8.0 -> 5.9ms and ~30k write-path allocations/ledger -> ~1.

--- zstd: multithreaded ledger encode, config-selected (default 2 workers)
After the txhash/events/extract cuts, the ~21ms single-threaded zstd
encode became the binding floor of the pre-commit section (its join
wait absorbed further wins). WithWorkers enables libzstd's internal
multithreading — still ONE standard deterministic frame, so
FrameHeaderValid and the cold-inherits-hot verbatim-copy contract are
untouched, at a measured ~0.1% ratio cost. The workers count is a real
configuration field, not an env experiment: measurement settled on
default 2 (equal to 3 within noise — total p50 33.4 -> 29.2ms, join
7.35 -> 1.9ms — while claiming one fewer core). It is FORMAT-AFFECTING
for the stored ledger frames, so ONE resolved value
(storage.zstd_encode_workers in the daemon TOML, --zstd-workers on the
bench hot cell; 0 = explicit single-threaded, validated >= 0) feeds
BOTH encoders — hotchunk.Tuning.ZstdEncodeWorkers for hot ingest and
ingest.Config.ZstdEncodeWorkers for the walk/backfill cold writer —
because the freeze copies hot frames into the cold pack verbatim while
the walk re-encodes the same ledgers, and a chunk's pack must stay
byte-identical whichever materializer built it (the freeze-vs-walk
gates arbitrate, now exercising the MT default on both sides).
NewCompressor panics loudly on a non-MT libzstd rather than silently
degrading the format.

Route the two constant-key event terms (event type, topic count) around
the pair sort: every event emits the same few keys, so their pairs
collapse into one radix bucket and the comparison sort re-derives an
order the arena already has. termlanes.go collects their ids in per-key
ascending lanes and the run build merges them at their byte-order
positions; rows are byte-identical (differential-tested). Measured on
sac6000: -1.4ms/ledger in the events phase.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013zyTXU8wkocJgN6mBafnou
tamirms pushed a commit that referenced this pull request Aug 30, 2026
…s marshal, multithreaded zstd

The second latency pass: the commit phase's remaining floors — ~6,000 random txhash memtable keys, ~30k events write-path allocations, and the serial ledger encode — each go.

--- txhash: packed-row hot engine — window + sealed runs, point lookups
The hot tx-hash tier's storage engine, standalone (nothing wired): one
sorted 32B-hash row per ledger in the window, sealed every 256 ledgers
into hash-sorted 36B-record run files (TXHRUN01, CRC64, fsync-dir
ladder), routed by per-run blooms and a 16B-prefix page ladder — one
aligned pread per hit, full-hash verify, boundary-tie two-page rule.
Manifest-anchored recovery with drain-verified opens; the engine is
born DISARMED (sealing rights follow validation, the events-tier rule)
and enforces a dense per-ledger chain. No merge tier, no overlay, no
retired handles: a point lookup needs none of them.

--- txhash: one packed row per ledger — flip ingest, freeze, and reads
The atomic flip onto the packed-row engine. Ingest writes ONE
seq-keyed row per ledger (hashes sorted in PhaseExtract, fee-bump
outer+inner per #862) instead of ~6,000 hash-keyed Puts — the measured
~10.9ms of per-ledger memtable-insert CPU those random 32B keys cost
(commit p50 20.5 -> 7.9ms in the 1k solo cell). Warmup replays the
dense row chain, trips loudly on the old format, and arms sealing only
after validation; the post-commit hook applies txhash before events
(independent, by convention). Read-only opens disable txhash queries
structurally — the freeze never queries. The freeze merges the sealed
runs with each tail row as its own sorted source (RAM flat at any tail
size) and emits duplicates verbatim: cold .bin bytes stay identical to
the walk path, gated by the composition test and freeze-at-crash-point
identity tests. Freeze txhash arm: 1m25 -> 17.9s at full chunk scale.

--- events: flat-pair term accumulation, bucketed sort, derivation arenas
The events queue step drops its per-ledger map: the marshal loop
appends (term, id) pairs into writer-owned arenas (AppendTerms writes
term keys straight into a caller arena; topics walk via Count()+Raw()
with All()'s reject-on-truncation semantics preserved and golden-swept;
the per-event data key is one hoisted scratch), then a 256-bucket MSD
scatter + big-endian word-comparator sort over KEYS ONLY (sorting 20B
pairs measured 2-2.5x slower; stability = index tiebreak, property-
tested as the exact permutation of a stable bytes.Compare) and one
linear pass emits the exact-sized packed row. ApplyLedger takes the
sorted term-runs directly; promotion and overlay semantics are pinned
equivalent to the retired map path, which survives as test-only
reference code for the byte-identity differential gates. Phase p50
8.0 -> 5.9ms and ~30k write-path allocations/ledger -> ~1.

--- zstd: multithreaded ledger encode, config-selected (default 2 workers)
After the txhash/events/extract cuts, the ~21ms single-threaded zstd
encode became the binding floor of the pre-commit section (its join
wait absorbed further wins). WithWorkers enables libzstd's internal
multithreading — still ONE standard deterministic frame, so
FrameHeaderValid and the cold-inherits-hot verbatim-copy contract are
untouched, at a measured ~0.1% ratio cost. The workers count is a real
configuration field, not an env experiment: measurement settled on
default 2 (equal to 3 within noise — total p50 33.4 -> 29.2ms, join
7.35 -> 1.9ms — while claiming one fewer core). It is FORMAT-AFFECTING
for the stored ledger frames, so ONE resolved value
(storage.zstd_encode_workers in the daemon TOML, --zstd-workers on the
bench hot cell; 0 = explicit single-threaded, validated >= 0) feeds
BOTH encoders — hotchunk.Tuning.ZstdEncodeWorkers for hot ingest and
ingest.Config.ZstdEncodeWorkers for the walk/backfill cold writer —
because the freeze copies hot frames into the cold pack verbatim while
the walk re-encodes the same ledgers, and a chunk's pack must stay
byte-identical whichever materializer built it (the freeze-vs-walk
gates arbitrate, now exercising the MT default on both sides).
NewCompressor panics loudly on a non-MT libzstd rather than silently
degrading the format.

Route the two constant-key event terms (event type, topic count) around
the pair sort: every event emits the same few keys, so their pairs
collapse into one radix bucket and the comparison sort re-derives an
order the arena already has. termlanes.go collects their ids in per-key
ascending lanes and the run build merges them at their byte-order
positions; rows are byte-identical (differential-tested). Measured on
sac6000: -1.4ms/ledger in the events phase.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013zyTXU8wkocJgN6mBafnou
tamirms pushed a commit that referenced this pull request Sep 23, 2026
…s marshal, multithreaded zstd

The second latency pass: the commit phase's remaining floors — ~6,000 random txhash memtable keys, ~30k events write-path allocations, and the serial ledger encode — each go.

--- txhash: packed-row hot engine — window + sealed runs, point lookups
The hot tx-hash tier's storage engine, standalone (nothing wired): one
sorted 32B-hash row per ledger in the window, sealed every 256 ledgers
into hash-sorted 36B-record run files (TXHRUN01, CRC64, fsync-dir
ladder), routed by per-run blooms and a 16B-prefix page ladder — one
aligned pread per hit, full-hash verify, boundary-tie two-page rule.
Manifest-anchored recovery with drain-verified opens; the engine is
born DISARMED (sealing rights follow validation, the events-tier rule)
and enforces a dense per-ledger chain. No merge tier, no overlay, no
retired handles: a point lookup needs none of them.

--- txhash: one packed row per ledger — flip ingest, freeze, and reads
The atomic flip onto the packed-row engine. Ingest writes ONE
seq-keyed row per ledger (hashes sorted in PhaseExtract, fee-bump
outer+inner per #862) instead of ~6,000 hash-keyed Puts — the measured
~10.9ms of per-ledger memtable-insert CPU those random 32B keys cost
(commit p50 20.5 -> 7.9ms in the 1k solo cell). Warmup replays the
dense row chain, trips loudly on the old format, and arms sealing only
after validation; the post-commit hook applies txhash before events
(independent, by convention). Read-only opens disable txhash queries
structurally — the freeze never queries. The freeze merges the sealed
runs with each tail row as its own sorted source (RAM flat at any tail
size) and emits duplicates verbatim: cold .bin bytes stay identical to
the walk path, gated by the composition test and freeze-at-crash-point
identity tests. Freeze txhash arm: 1m25 -> 17.9s at full chunk scale.

--- events: flat-pair term accumulation, bucketed sort, derivation arenas
The events queue step drops its per-ledger map: the marshal loop
appends (term, id) pairs into writer-owned arenas (AppendTerms writes
term keys straight into a caller arena; topics walk via Count()+Raw()
with All()'s reject-on-truncation semantics preserved and golden-swept;
the per-event data key is one hoisted scratch), then a 256-bucket MSD
scatter + big-endian word-comparator sort over KEYS ONLY (sorting 20B
pairs measured 2-2.5x slower; stability = index tiebreak, property-
tested as the exact permutation of a stable bytes.Compare) and one
linear pass emits the exact-sized packed row. ApplyLedger takes the
sorted term-runs directly; promotion and overlay semantics are pinned
equivalent to the retired map path, which survives as test-only
reference code for the byte-identity differential gates. Phase p50
8.0 -> 5.9ms and ~30k write-path allocations/ledger -> ~1.

--- zstd: multithreaded ledger encode, config-selected (default 2 workers)
After the txhash/events/extract cuts, the ~21ms single-threaded zstd
encode became the binding floor of the pre-commit section (its join
wait absorbed further wins). WithWorkers enables libzstd's internal
multithreading — still ONE standard deterministic frame, so
FrameHeaderValid and the cold-inherits-hot verbatim-copy contract are
untouched, at a measured ~0.1% ratio cost. The workers count is a real
configuration field, not an env experiment: measurement settled on
default 2 (equal to 3 within noise — total p50 33.4 -> 29.2ms, join
7.35 -> 1.9ms — while claiming one fewer core). It is FORMAT-AFFECTING
for the stored ledger frames, so ONE resolved value
(storage.zstd_encode_workers in the daemon TOML, --zstd-workers on the
bench hot cell; 0 = explicit single-threaded, validated >= 0) feeds
BOTH encoders — hotchunk.Tuning.ZstdEncodeWorkers for hot ingest and
ingest.Config.ZstdEncodeWorkers for the walk/backfill cold writer —
because the freeze copies hot frames into the cold pack verbatim while
the walk re-encodes the same ledgers, and a chunk's pack must stay
byte-identical whichever materializer built it (the freeze-vs-walk
gates arbitrate, now exercising the MT default on both sides).
NewCompressor panics loudly on a non-MT libzstd rather than silently
degrading the format.

Route the two constant-key event terms (event type, topic count) around
the pair sort: every event emits the same few keys, so their pairs
collapse into one radix bucket and the comparison sort re-derives an
order the arena already has. termlanes.go collects their ids in per-key
ascending lanes and the run build merges them at their byte-order
positions; rows are byte-identical (differential-tested). Measured on
sac6000: -1.4ms/ledger in the events phase.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013zyTXU8wkocJgN6mBafnou
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants