Skip to content

[Feature] hbg: Add benchmark cases and harness/skill support (default stays tmr) #1727

Description

@ChaoZheng109

Summary

host_build_graph has no benchmark path at all. The whole benchmark story — the
tools/benchmark_rounds.sh harness, the benchmark skill, and the example corpus —
is tensormap_and_ringbuffer-only. Add HBG support across all three, keeping T&R the
default and making HBG selectable.

Three separate gaps, in increasing order of difficulty:

1. The harness refuses the runtime. benchmark_rounds.sh dispatches on
case "$RUNTIME" and the only arm is tensormap_and_ringbuffer; anything else exits
with "ERROR: unknown runtime … Use tensormap_and_ringbuffer"
(tools/benchmark_rounds.sh:165-172). The workload list is a single T&R-specific pair
of arrays, TMR_EXAMPLE_CASES / TMR_EXAMPLE_ORDER (:37-57).

2. There are almost no HBG benchmark cases. In the merged tree, not one of the six
T&R benchmark workloads has an HBG counterpart under tests/st/a2a3/host_build_graph/ or
examples/a2a3/host_build_graph/. Two open PRs close part of this — see
the PR section below:

T&R benchmark case HBG counterpart
alternating_matmul_add (Case1) —
benchmark_bgemm (Case0) —
paged_attention_unroll (Case1,Case2) —
paged_attention_unroll_manual_scope (Case1,Case2) —
batch_paged_attention (Case1) —
qwen3_14b_decode (StressBatch16Seq3500) open PR #1666

The nearest available HBG scenes are bgemm (case default), matmul (default) and
paged_attention (Case1, Case2) — related workloads, but not the same ones, so they
cannot be read as a cross-runtime comparison.

3. Four of the five reported metrics do not exist on HBG. The harness reports Host /
Device / Effective / Orch / Sched, and all but Host are derived from AicpuPhase
markers rendered by simpler_setup.tools.strace_timing --rounds-table. The HBG tree
emits zero AicpuPhase markers
— grep -c AicpuPhase src/a2a3/runtime/host_build_graph/
returns nothing, against OrchWindow / SchedWindow / Preamble / ConfigValidate /
ArenaWire / SmReset in the T&R tree. Only Host works, because the host-side STRACE
span lives in src/common/.

This one is not a wiring problem, it is a semantic one: HBG's orchestrator runs on the
host
, so a device_wall.orch span is not merely missing, it is meaningless. The phases
worth timing in HBG are different — host graph build, pointer relocation, the single
H2D, boot classification, dispatch — plus the split between the S threads and the P
resolution thread, and record-vs-replay for Graph Execution.

Motivation / Use Case

A merged performance PR has no harness behind it. #1659 ("HBG: up to 99% host-side
overhead reduction"
) reports paged_attention host bind going 402.7 ms → ~7.8 ms. That
was measured ad hoc. Nobody can re-run it, confirm it, or notice when it regresses, and
#1716 is about porting that same work to a5 — which will need the same measurement to
show it landed.

HBG's central claim is unverified. The runtime's thesis in #1706 is that host
orchestration is affordable because the build cost is amortized — by Graph replay
(#1444) and by prepared-successor pipelining. Both are performance arguments with no
performance evidence. Bucket 6 of #1706 exists precisely for this, and every other
performance sub-issue (#1713 on the recording path still paying full ring cost, #1717 on
overlapping device launch with orchestration) is a claim this harness would settle.

T&R-vs-HBG has never been measured on the same graph. That comparison is the input
to the "where should HBG land" question in #1706 — whether host orchestration plus Graph
replay beats device orchestration for repetitive, statically-shaped graphs.

Proposed API / Behavior

Harness. Add HBG_EXAMPLE_CASES / HBG_EXAMPLE_ORDER alongside the T&R arrays and a
second arm in the runtime dispatch. tensormap_and_ringbuffer stays the default, so
./tools/benchmark_rounds.sh with no -r behaves exactly as today; -r host_build_graph
selects the HBG corpus.

--serial-orch-sched must be rejected for HBG rather than silently accepted: it exists to
serialize a device-side orchestrator against the schedulers, and HBG has none.

Metrics. Decide the HBG column set before writing the cases, because it determines what
instrumentation is needed:

  • Option A — give HBG its own AicpuPhase markers for the phases it actually has, and
    report an HBG-shaped table (host build / relocate / H2D / boot classify / dispatch,
    with S-vs-P attribution). Most useful, most work.
  • Option B — report only what exists today (Host, plus device wall if it can be derived
    without per-phase markers) and state plainly in the output that the Orch/Sched columns
    are T&R-only. Cheap, honest, and unblocks regression tracking on the metric [Performance] HBG: up to 99% host-side overhead reduction #1659 was
    about (host bind).

Either way the table must not print empty or zero columns that look like measurements.

Cases. Merge #1666 and #1667 first (see below) — they already supply
qwen3_14b_decode, paged_attention and paged_attention_manual_scope in one-to-one
positions, which covers the workload #1659's claim was measured on and the
repetitive-decode workload Graph Execution targets. Everything still missing is authoring
work that overlaps the coverage gap in #1706 bucket 5, and each new copy must follow the
rule #1666 and #1667 already follow: same tree, same directory name, same case names, so
the HBG and T&R numbers stay directly comparable.

Graph Execution deserves a benchmark axis of its own: same graph with and without replay
is the direct measurement of the amortization claim.

Skill. .claude/skills/benchmark/SKILL.md §"Runtime Selection" currently lists only
tensormap_and_ringbuffer (default). Add host_build_graph, keep the default, and
document what changes about the reported columns when it is selected — a table whose
column meanings silently change per runtime is worse than one that says so. The same
applies to perf-runtime-device and perf-example-device, which accept a runtime
argument and would otherwise point users at a harness that rejects HBG.

Part of the corpus already exists in open PRs — merge them, do not re-author

Two open PRs already add HBG copies of T&R examples, and both place them one-to-one
with their T&R counterparts
— same tree, same directory name, same file layout, only
the runtime directory differs. No separate case-authoring PR is needed for what they
cover; merging them is the work.

PR Adds T&R counterpart Placement Cases
#1666 examples/a2a3/host_build_graph/qwen3_14b_decode/ examples/a2a3/tensormap_and_ringbuffer/qwen3_14b_decode/ one-to-one identical — all 38, including StressBatch16Seq3500, the exact case TMR_EXAMPLE_CASES benchmarks
#1667 examples/a2a3/host_build_graph/paged_attention/ examples/a2a3/tensormap_and_ringbuffer/paged_attention/ one-to-one 10 of 11 — Case3 deliberately omitted
#1667 examples/a2a3/host_build_graph/paged_attention_manual_scope/ examples/a2a3/tensormap_and_ringbuffer/paged_attention_manual_scope/ one-to-one same 10 of 11

The Case3 omission is justified in #1667's own description and is not an HBG gap: it
is manual: True (so it never runs in CI), it uses head_dim: 256, and it fails on
both runtimes — max_diff=nan on HBG and max_diff=0.1255 on the unmodified T&R
example, with a normal ~16 ms device_wall and no allocator fatal. That is a
pre-existing kernel numerics defect above head_dim > 128, so copying it would import a
known-failing case. #1667 reports 12/12 passing across both examples on a2a3 with golden
checking.

What this leaves for the benchmark corpus, against the current TMR_EXAMPLE_CASES:

T&R benchmark case Tree Covered by
qwen3_14b_decode (StressBatch16Seq3500) examples/ #1666
alternating_matmul_add (Case1) tests/st/ still to author
benchmark_bgemm (Case0) examples/ still to author
paged_attention_unroll (Case1,Case2) tests/st/ still to author
paged_attention_unroll_manual_scope (Case1,Case2) examples/ still to author
batch_paged_attention (Case1) tests/st/ still to author

qwen3_14b_decode is the most valuable single entry, since it is the repetitive-decode
shape Graph Execution is designed for. #1667's two examples are not in
TMR_EXAMPLE_CASES today, but they mirror T&R examples exactly, so HBG_EXAMPLE_CASES
can include them and the numbers stay directly comparable — the HBG list does not have to
be a name-for-name copy of the T&R list, it has to consist of cases that exist one-to-one
on both sides.

One conflict to resolve before or with #1667. tests/st/a2a3/host_build_graph/paged_attention/
already exists in the tree, and T&R has no tests/st/.../paged_attention at all — its
paged_attention lives only under examples/. So merging #1667 leaves HBG with two
different paged_attention scenes, only one of which corresponds to anything on the T&R
side. They are not the same content: the existing tests/st copy is 156 lines with cases
Case1 / Case2 / small1 / small2, against #1667's 10 cases using T&R's
CaseSmall1 / CaseSmall2 spelling. The tests/st copy should be removed in favour of
#1667's, or its continued existence justified — otherwise the one-to-one placement rule is
broken the moment #1667 lands.

Alternatives Considered

Keep measuring ad hoc (status quo). This is what #1659 did. It produced a number
nobody can reproduce and no barrier against regressing it — and the port in #1716 has no
way to demonstrate parity.

Run the T&R harness unmodified against HBG. It exits at the runtime dispatch. Even if
that check were simply relaxed, the run would report Orch/Sched/Device columns with no
data behind them, which is worse than refusing: a table of zeros reads as a measurement.

Additional Context

Parent: #1706 (bucket 6, Benchmark). Related: #1659 (the unreproducible claim), #1716 (a5
port needing the same measurement), #1713 and #1717 (HBG performance claims this harness
would verify), #1712 (Graph step-2 work whose payoff needs measuring).

Incidental finding while surveying the harness — --serial-orch-sched has never
worked.
benchmark_rounds.sh:249 runs the second pass with PTO2_SERIAL_ORCH_SCHED=1,
but the runtime reads SIMPLER_TMR_SERIAL_ORCH_SCHED_ENABLE
(src/{a2a3,a5}/runtime/tensormap_and_ringbuffer/host/runtime_maker.cpp:629). Nothing
consumes the variable the script sets, so both passes run in the default overlapped mode
and the "Serial vs Parallel Delta" table at :413-416 compares a configuration against
itself. The skill's own text (SKILL.md:120) names the correct variable, so the script is
the side that is wrong. This is a separate defect from the HBG work here and should get
its own issue.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

enhancementNew feature or request

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions