You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
host_build_graph has no benchmark path at all. The whole benchmark story — the tools/benchmark_rounds.sh harness, the benchmark skill, and the example corpus —
is tensormap_and_ringbuffer-only. Add HBG support across all three, keeping T&R the
default and making HBG selectable.
Three separate gaps, in increasing order of difficulty:
1. The harness refuses the runtime.benchmark_rounds.sh dispatches on case "$RUNTIME" and the only arm is tensormap_and_ringbuffer; anything else exits
with "ERROR: unknown runtime … Use tensormap_and_ringbuffer"
(tools/benchmark_rounds.sh:165-172). The workload list is a single T&R-specific pair
of arrays, TMR_EXAMPLE_CASES / TMR_EXAMPLE_ORDER (:37-57).
2. There are almost no HBG benchmark cases. In the merged tree, not one of the six
T&R benchmark workloads has an HBG counterpart under tests/st/a2a3/host_build_graph/ or examples/a2a3/host_build_graph/. Two open PRs close part of this — see the PR section below:
The nearest available HBG scenes are bgemm (case default), matmul (default) and paged_attention (Case1, Case2) — related workloads, but not the same ones, so they
cannot be read as a cross-runtime comparison.
3. Four of the five reported metrics do not exist on HBG. The harness reports Host /
Device / Effective / Orch / Sched, and all but Host are derived from AicpuPhase
markers rendered by simpler_setup.tools.strace_timing --rounds-table. The HBG tree
emits zero AicpuPhase markers — grep -c AicpuPhase src/a2a3/runtime/host_build_graph/
returns nothing, against OrchWindow / SchedWindow / Preamble / ConfigValidate / ArenaWire / SmReset in the T&R tree. Only Host works, because the host-side STRACE
span lives in src/common/.
This one is not a wiring problem, it is a semantic one: HBG's orchestrator runs on the
host, so a device_wall.orch span is not merely missing, it is meaningless. The phases
worth timing in HBG are different — host graph build, pointer relocation, the single
H2D, boot classification, dispatch — plus the split between the S threads and the P
resolution thread, and record-vs-replay for Graph Execution.
Motivation / Use Case
A merged performance PR has no harness behind it.#1659 ("HBG: up to 99% host-side
overhead reduction") reports paged_attention host bind going 402.7 ms → ~7.8 ms. That
was measured ad hoc. Nobody can re-run it, confirm it, or notice when it regresses, and #1716 is about porting that same work to a5 — which will need the same measurement to
show it landed.
HBG's central claim is unverified. The runtime's thesis in #1706 is that host
orchestration is affordable because the build cost is amortized — by Graph replay
(#1444) and by prepared-successor pipelining. Both are performance arguments with no
performance evidence. Bucket 6 of #1706 exists precisely for this, and every other
performance sub-issue (#1713 on the recording path still paying full ring cost, #1717 on
overlapping device launch with orchestration) is a claim this harness would settle.
T&R-vs-HBG has never been measured on the same graph. That comparison is the input
to the "where should HBG land" question in #1706 — whether host orchestration plus Graph
replay beats device orchestration for repetitive, statically-shaped graphs.
Proposed API / Behavior
Harness. Add HBG_EXAMPLE_CASES / HBG_EXAMPLE_ORDER alongside the T&R arrays and a
second arm in the runtime dispatch. tensormap_and_ringbuffer stays the default, so ./tools/benchmark_rounds.sh with no -r behaves exactly as today; -r host_build_graph
selects the HBG corpus.
--serial-orch-sched must be rejected for HBG rather than silently accepted: it exists to
serialize a device-side orchestrator against the schedulers, and HBG has none.
Metrics. Decide the HBG column set before writing the cases, because it determines what
instrumentation is needed:
Option A — give HBG its own AicpuPhase markers for the phases it actually has, and
report an HBG-shaped table (host build / relocate / H2D / boot classify / dispatch,
with S-vs-P attribution). Most useful, most work.
Option B — report only what exists today (Host, plus device wall if it can be derived
without per-phase markers) and state plainly in the output that the Orch/Sched columns
are T&R-only. Cheap, honest, and unblocks regression tracking on the metric [Performance] HBG: up to 99% host-side overhead reduction #1659 was
about (host bind).
Either way the table must not print empty or zero columns that look like measurements.
Cases. Merge #1666 and #1667 first (see below) — they already supply qwen3_14b_decode, paged_attention and paged_attention_manual_scope in one-to-one
positions, which covers the workload #1659's claim was measured on and the
repetitive-decode workload Graph Execution targets. Everything still missing is authoring
work that overlaps the coverage gap in #1706 bucket 5, and each new copy must follow the
rule #1666 and #1667 already follow: same tree, same directory name, same case names, so
the HBG and T&R numbers stay directly comparable.
Graph Execution deserves a benchmark axis of its own: same graph with and without replay
is the direct measurement of the amortization claim.
Skill..claude/skills/benchmark/SKILL.md §"Runtime Selection" currently lists only tensormap_and_ringbuffer (default). Add host_build_graph, keep the default, and
document what changes about the reported columns when it is selected — a table whose
column meanings silently change per runtime is worse than one that says so. The same
applies to perf-runtime-device and perf-example-device, which accept a runtime
argument and would otherwise point users at a harness that rejects HBG.
Part of the corpus already exists in open PRs — merge them, do not re-author
Two open PRs already add HBG copies of T&R examples, and both place them one-to-one
with their T&R counterparts — same tree, same directory name, same file layout, only
the runtime directory differs. No separate case-authoring PR is needed for what they
cover; merging them is the work.
The Case3 omission is justified in #1667's own description and is not an HBG gap: it
is manual: True (so it never runs in CI), it uses head_dim: 256, and it fails on both runtimes — max_diff=nan on HBG and max_diff=0.1255 on the unmodified T&R
example, with a normal ~16 ms device_wall and no allocator fatal. That is a
pre-existing kernel numerics defect above head_dim > 128, so copying it would import a
known-failing case. #1667 reports 12/12 passing across both examples on a2a3 with golden
checking.
What this leaves for the benchmark corpus, against the current TMR_EXAMPLE_CASES:
qwen3_14b_decode is the most valuable single entry, since it is the repetitive-decode
shape Graph Execution is designed for. #1667's two examples are not in TMR_EXAMPLE_CASES today, but they mirror T&R examples exactly, so HBG_EXAMPLE_CASES
can include them and the numbers stay directly comparable — the HBG list does not have to
be a name-for-name copy of the T&R list, it has to consist of cases that exist one-to-one
on both sides.
One conflict to resolve before or with #1667.tests/st/a2a3/host_build_graph/paged_attention/
already exists in the tree, and T&R has no tests/st/.../paged_attention at all — its paged_attention lives only under examples/. So merging #1667 leaves HBG with two
different paged_attention scenes, only one of which corresponds to anything on the T&R
side. They are not the same content: the existing tests/st copy is 156 lines with cases Case1 / Case2 / small1 / small2, against #1667's 10 cases using T&R's CaseSmall1 / CaseSmall2 spelling. The tests/st copy should be removed in favour of #1667's, or its continued existence justified — otherwise the one-to-one placement rule is
broken the moment #1667 lands.
Alternatives Considered
Keep measuring ad hoc (status quo). This is what #1659 did. It produced a number
nobody can reproduce and no barrier against regressing it — and the port in #1716 has no
way to demonstrate parity.
Run the T&R harness unmodified against HBG. It exits at the runtime dispatch. Even if
that check were simply relaxed, the run would report Orch/Sched/Device columns with no
data behind them, which is worse than refusing: a table of zeros reads as a measurement.
Additional Context
Parent: #1706 (bucket 6, Benchmark). Related: #1659 (the unreproducible claim), #1716 (a5
port needing the same measurement), #1713 and #1717 (HBG performance claims this harness
would verify), #1712 (Graph step-2 work whose payoff needs measuring).
Incidental finding while surveying the harness — --serial-orch-sched has never
worked.benchmark_rounds.sh:249 runs the second pass with PTO2_SERIAL_ORCH_SCHED=1,
but the runtime reads SIMPLER_TMR_SERIAL_ORCH_SCHED_ENABLE
(src/{a2a3,a5}/runtime/tensormap_and_ringbuffer/host/runtime_maker.cpp:629). Nothing
consumes the variable the script sets, so both passes run in the default overlapped mode
and the "Serial vs Parallel Delta" table at :413-416 compares a configuration against
itself. The skill's own text (SKILL.md:120) names the correct variable, so the script is
the side that is wrong. This is a separate defect from the HBG work here and should get
its own issue.
Summary
host_build_graphhas no benchmark path at all. The whole benchmark story — thetools/benchmark_rounds.shharness, thebenchmarkskill, and the example corpus —is
tensormap_and_ringbuffer-only. Add HBG support across all three, keeping T&R thedefault and making HBG selectable.
Three separate gaps, in increasing order of difficulty:
1. The harness refuses the runtime.
benchmark_rounds.shdispatches oncase "$RUNTIME"and the only arm istensormap_and_ringbuffer; anything else exitswith "ERROR: unknown runtime … Use tensormap_and_ringbuffer"
(
tools/benchmark_rounds.sh:165-172). The workload list is a single T&R-specific pairof arrays,
TMR_EXAMPLE_CASES/TMR_EXAMPLE_ORDER(:37-57).2. There are almost no HBG benchmark cases. In the merged tree, not one of the six
T&R benchmark workloads has an HBG counterpart under
tests/st/a2a3/host_build_graph/orexamples/a2a3/host_build_graph/. Two open PRs close part of this — seethe PR section below:
alternating_matmul_add(Case1)benchmark_bgemm(Case0)paged_attention_unroll(Case1,Case2)paged_attention_unroll_manual_scope(Case1,Case2)batch_paged_attention(Case1)qwen3_14b_decode(StressBatch16Seq3500)The nearest available HBG scenes are
bgemm(casedefault),matmul(default) andpaged_attention(Case1,Case2) — related workloads, but not the same ones, so theycannot be read as a cross-runtime comparison.
3. Four of the five reported metrics do not exist on HBG. The harness reports Host /
Device / Effective / Orch / Sched, and all but Host are derived from
AicpuPhasemarkers rendered by
simpler_setup.tools.strace_timing --rounds-table. The HBG treeemits zero
AicpuPhasemarkers —grep -c AicpuPhase src/a2a3/runtime/host_build_graph/returns nothing, against
OrchWindow/SchedWindow/Preamble/ConfigValidate/ArenaWire/SmResetin the T&R tree. Only Host works, because the host-side STRACEspan lives in
src/common/.This one is not a wiring problem, it is a semantic one: HBG's orchestrator runs on the
host, so a
device_wall.orchspan is not merely missing, it is meaningless. The phasesworth timing in HBG are different — host graph build, pointer relocation, the single
H2D, boot classification, dispatch — plus the split between the S threads and the P
resolution thread, and record-vs-replay for Graph Execution.
Motivation / Use Case
A merged performance PR has no harness behind it. #1659 ("HBG: up to 99% host-side
overhead reduction") reports paged_attention host
bindgoing 402.7 ms → ~7.8 ms. Thatwas measured ad hoc. Nobody can re-run it, confirm it, or notice when it regresses, and
#1716 is about porting that same work to a5 — which will need the same measurement to
show it landed.
HBG's central claim is unverified. The runtime's thesis in #1706 is that host
orchestration is affordable because the build cost is amortized — by Graph replay
(#1444) and by prepared-successor pipelining. Both are performance arguments with no
performance evidence. Bucket 6 of #1706 exists precisely for this, and every other
performance sub-issue (#1713 on the recording path still paying full ring cost, #1717 on
overlapping device launch with orchestration) is a claim this harness would settle.
T&R-vs-HBG has never been measured on the same graph. That comparison is the input
to the "where should HBG land" question in #1706 — whether host orchestration plus Graph
replay beats device orchestration for repetitive, statically-shaped graphs.
Proposed API / Behavior
Harness. Add
HBG_EXAMPLE_CASES/HBG_EXAMPLE_ORDERalongside the T&R arrays and asecond arm in the runtime dispatch.
tensormap_and_ringbufferstays the default, so./tools/benchmark_rounds.shwith no-rbehaves exactly as today;-r host_build_graphselects the HBG corpus.
--serial-orch-schedmust be rejected for HBG rather than silently accepted: it exists toserialize a device-side orchestrator against the schedulers, and HBG has none.
Metrics. Decide the HBG column set before writing the cases, because it determines what
instrumentation is needed:
AicpuPhasemarkers for the phases it actually has, andreport an HBG-shaped table (host build / relocate / H2D / boot classify / dispatch,
with S-vs-P attribution). Most useful, most work.
without per-phase markers) and state plainly in the output that the Orch/Sched columns
are T&R-only. Cheap, honest, and unblocks regression tracking on the metric [Performance] HBG: up to 99% host-side overhead reduction #1659 was
about (host
bind).Either way the table must not print empty or zero columns that look like measurements.
Cases. Merge #1666 and #1667 first (see below) — they already supply
qwen3_14b_decode,paged_attentionandpaged_attention_manual_scopein one-to-onepositions, which covers the workload #1659's claim was measured on and the
repetitive-decode workload Graph Execution targets. Everything still missing is authoring
work that overlaps the coverage gap in #1706 bucket 5, and each new copy must follow the
rule #1666 and #1667 already follow: same tree, same directory name, same case names, so
the HBG and T&R numbers stay directly comparable.
Graph Execution deserves a benchmark axis of its own: same graph with and without replay
is the direct measurement of the amortization claim.
Skill.
.claude/skills/benchmark/SKILL.md§"Runtime Selection" currently lists onlytensormap_and_ringbuffer (default). Addhost_build_graph, keep the default, anddocument what changes about the reported columns when it is selected — a table whose
column meanings silently change per runtime is worse than one that says so. The same
applies to
perf-runtime-deviceandperf-example-device, which accept a runtimeargument and would otherwise point users at a harness that rejects HBG.
Part of the corpus already exists in open PRs — merge them, do not re-author
Two open PRs already add HBG copies of T&R examples, and both place them one-to-one
with their T&R counterparts — same tree, same directory name, same file layout, only
the runtime directory differs. No separate case-authoring PR is needed for what they
cover; merging them is the work.
examples/a2a3/host_build_graph/qwen3_14b_decode/examples/a2a3/tensormap_and_ringbuffer/qwen3_14b_decode/StressBatch16Seq3500, the exact caseTMR_EXAMPLE_CASESbenchmarksexamples/a2a3/host_build_graph/paged_attention/examples/a2a3/tensormap_and_ringbuffer/paged_attention/Case3deliberately omittedexamples/a2a3/host_build_graph/paged_attention_manual_scope/examples/a2a3/tensormap_and_ringbuffer/paged_attention_manual_scope/The
Case3omission is justified in #1667's own description and is not an HBG gap: itis
manual: True(so it never runs in CI), it useshead_dim: 256, and it fails onboth runtimes —
max_diff=nanon HBG andmax_diff=0.1255on the unmodified T&Rexample, with a normal ~16 ms
device_walland no allocator fatal. That is apre-existing kernel numerics defect above
head_dim > 128, so copying it would import aknown-failing case. #1667 reports 12/12 passing across both examples on a2a3 with golden
checking.
What this leaves for the benchmark corpus, against the current
TMR_EXAMPLE_CASES:qwen3_14b_decode(StressBatch16Seq3500)examples/alternating_matmul_add(Case1)tests/st/benchmark_bgemm(Case0)examples/paged_attention_unroll(Case1,Case2)tests/st/paged_attention_unroll_manual_scope(Case1,Case2)examples/batch_paged_attention(Case1)tests/st/qwen3_14b_decodeis the most valuable single entry, since it is the repetitive-decodeshape Graph Execution is designed for. #1667's two examples are not in
TMR_EXAMPLE_CASEStoday, but they mirror T&R examples exactly, soHBG_EXAMPLE_CASEScan include them and the numbers stay directly comparable — the HBG list does not have to
be a name-for-name copy of the T&R list, it has to consist of cases that exist one-to-one
on both sides.
One conflict to resolve before or with #1667.
tests/st/a2a3/host_build_graph/paged_attention/already exists in the tree, and T&R has no
tests/st/.../paged_attentionat all — itspaged_attentionlives only underexamples/. So merging #1667 leaves HBG with twodifferent
paged_attentionscenes, only one of which corresponds to anything on the T&Rside. They are not the same content: the existing
tests/stcopy is 156 lines with casesCase1/Case2/small1/small2, against #1667's 10 cases using T&R'sCaseSmall1/CaseSmall2spelling. Thetests/stcopy should be removed in favour of#1667's, or its continued existence justified — otherwise the one-to-one placement rule is
broken the moment #1667 lands.
Alternatives Considered
Keep measuring ad hoc (status quo). This is what #1659 did. It produced a number
nobody can reproduce and no barrier against regressing it — and the port in #1716 has no
way to demonstrate parity.
Run the T&R harness unmodified against HBG. It exits at the runtime dispatch. Even if
that check were simply relaxed, the run would report Orch/Sched/Device columns with no
data behind them, which is worse than refusing: a table of zeros reads as a measurement.
Additional Context
Parent: #1706 (bucket 6, Benchmark). Related: #1659 (the unreproducible claim), #1716 (a5
port needing the same measurement), #1713 and #1717 (HBG performance claims this harness
would verify), #1712 (Graph step-2 work whose payoff needs measuring).
Incidental finding while surveying the harness —
--serial-orch-schedhas neverworked.
benchmark_rounds.sh:249runs the second pass withPTO2_SERIAL_ORCH_SCHED=1,but the runtime reads
SIMPLER_TMR_SERIAL_ORCH_SCHED_ENABLE(
src/{a2a3,a5}/runtime/tensormap_and_ringbuffer/host/runtime_maker.cpp:629). Nothingconsumes the variable the script sets, so both passes run in the default overlapped mode
and the "Serial vs Parallel Delta" table at
:413-416compares a configuration againstitself. The skill's own text (
SKILL.md:120) names the correct variable, so the script isthe side that is wrong. This is a separate defect from the HBG work here and should get
its own issue.