Summary
With --failFast=false, seg integrity logs every problem it finds, returns success, and declares the snapshots publishable. The flag is meant to control when a check stops, but it also decides whether the run fails at all.
Observed on a sepolia archive node, full scan:
INFO [integrity] ReceiptsNoDups workers=11 chunkSize=100 blockRange=11800947 sampleRatio=1.000
EROR CheckReceiptsNoDups: non-monotonic cumGasUsed at txnum: 800000000, block: 9500808(799999819-800000035), cumGasUsed=33022, prevCumGasUsed=29084615
EROR CheckReceiptsNoDups: non-monotonic logIndex at txnum: 800000000, block: 9500808(799999819-800000035), logIdxAfterTx=1, prevLogIdxAfterTx=256
INFO [integrity] ReceiptsNoDups: done err=nil
INFO [integrity] snapshots are publishable
Two errors, err=nil, publishable, exit 0.
Mechanism
Every reporting site has this shape:
// db/integrity/receipts_no_duplicates.go
if failFast {
return err
}
log.Error(err.Error()) // logged, then dropped
so the check returns nil, and the caller treats that as success:
// cmd/utils/app/snapshots_cmd.go
if err := doIntegrity(ctx, cliCtx); err != nil {
log.Error("[integrity]", "err", err)
return err
}
log.Info("[integrity] snapshots are publishable")
This is not one check: no_gaps_in_canonical_headers.go, rcache_no_duplicates.go, rcache_receipt_root.go, receipts_no_duplicates.go and snap_blocks_read.go all log and continue, and none of them accumulates a verdict.
Why it matters
--failFast=false is the natural flag for "show me everything that is wrong", which is exactly what an operator reaches for before publishing. It is also the setting under which the command is least able to tell them the answer. A release pipeline that used it would publish defective snapshots on a green exit code.
The default failFast=true does fail correctly, so this is a trap rather than a total loss — but the safer-looking flag is the unsafe one.
Related: --integrity.budget has the same shape. A check that exhausts its slice logs budget exhausted at INFO and the command exits 0, so a budgeted run can report validation it never performed. Both let the gate say "publishable" without having established it.
Suggested fix
Track whether any problem was reported and return a non-nil error at the end of the run when one was, independently of failFast. failFast then means what its usage string says — "stop after first problem or print and continue" — without also deciding the exit code.
A cheap first step, if the aggregation is invasive: have doIntegrity refuse to print snapshots are publishable when any check logged a problem.
Summary
With
--failFast=false,seg integritylogs every problem it finds, returns success, and declares the snapshots publishable. The flag is meant to control when a check stops, but it also decides whether the run fails at all.Observed on a sepolia archive node, full scan:
Two errors,
err=nil,publishable, exit 0.Mechanism
Every reporting site has this shape:
so the check returns nil, and the caller treats that as success:
This is not one check:
no_gaps_in_canonical_headers.go,rcache_no_duplicates.go,rcache_receipt_root.go,receipts_no_duplicates.goandsnap_blocks_read.goall log and continue, and none of them accumulates a verdict.Why it matters
--failFast=falseis the natural flag for "show me everything that is wrong", which is exactly what an operator reaches for before publishing. It is also the setting under which the command is least able to tell them the answer. A release pipeline that used it would publish defective snapshots on a green exit code.The default
failFast=truedoes fail correctly, so this is a trap rather than a total loss — but the safer-looking flag is the unsafe one.Related:
--integrity.budgethas the same shape. A check that exhausts its slice logsbudget exhaustedat INFO and the command exits 0, so a budgeted run can report validation it never performed. Both let the gate say "publishable" without having established it.Suggested fix
Track whether any problem was reported and return a non-nil error at the end of the run when one was, independently of
failFast.failFastthen means what its usage string says — "stop after first problem or print and continue" — without also deciding the exit code.A cheap first step, if the aggregation is invasive: have
doIntegrityrefuse to printsnapshots are publishablewhen any check logged a problem.