This issue tracks all host_build_graph (HBG) work, on a2a3 and a5 — what the runtime
is for, where it stands today, and the open work. Individual items are attached as
sub-issues below (the panel auto-updates as items are added or closed).
Scope: a2a3 and a5. HBG is one runtime with two arch trees, kept in sync. A sub-issue is
about HBG, not about an arch — file one, not two. See
Per-arch divergence for when that stops being true.
What HBG is
HBG separates graph construction from device execution
(src/{arch}/runtime/host_build_graph/docs/RUNTIME_LOGIC.md §1):
host: dlopen the orchestration SO, run it to completion → whole task graph
host: relocate the prebuilt graph image, copy it to device memory
device: attach, classify every task once, dispatch with the AICPU schedulers
host: collect outputs, destroy/reset per-run state
There is no device-side orchestrator thread. The AICPU threads split 3S+1P: the last
thread is a core-less resolution thread (P), the rest own AICore partitions and schedule
them (S). At least two threads are required (#1544). The host constructs the whole graph
before any device task can complete — §1 calls that ordering "the defining constraint of
the runtime", and everything below is a consequence of it.
Structurally HBG is the host-orchestration variant of tensormap_and_ringbuffer, not a
separate graph model: since #1185 it uses the same PTO2TaskDescriptor ring storage and the
same TensorMap-derived dependencies, with one ring and no device orchestrator. (The
"fixed Task[] array, up to 131,072 tasks, traverses dependencies" description still in
src/a2a3/docs/runtimes.md and in T&R's RUNTIME_LOGIC.md §1.1 is stale.)
Intent and vision
1. Take orchestration off the AICPU. No AICPU thread builds the graph, so the whole
device-side thread budget goes to dispatch and completion resolution. Graph construction
runs in ordinary host C++ — full STL, a host debugger, host-side DFX (dep_gen is captured
on the host orchestrator, #1492) — and no device-log budget is spent on build-time
diagnostics.
2. Pair the polling readiness model with the 3S+1P split. Dependency state is a per-slot
completion flag, a monotonic watermark, intrusive wake lists and inline integer fanin —
no dependency pool, no fanout lock, and no periodic sweep (#1435). That collapses
readiness resolution into one compact sequential job, which is exactly what can be handed
to a single thread: P owns all of it, so it is the sole producer of the ready queues, while
each S thread stays purely core-local — poll its own cores, dispatch them, hand finished
tasks to P over an SPSC queue (#1544). Whole-graph residency is what allows the split: all
threads classify disjoint slices of the finished graph once at boot (#1553), after which no
new task ever arrives. The gain is graph-shaped — a wide parallel graph wins, a serial chain
pays for the extra hop — so quantifying it belongs to the benchmark bucket.
3. Amortize the host build cost — this is the crux. Paying for a full host graph build
on every run is only acceptable if repeated work is not rebuilt. Two mechanisms exist:
4. Where this should land. README.md still positions HBG as the
"development, debugging" runtime against T&R's "production workloads". The vision is that
for repetitive, statically-shaped graphs, host orchestration plus Graph replay becomes a
first-class production path rather than the simple-but-slow option — while T&R remains the
runtime for streaming, dynamic, and flow-controlled workloads. Closing that gap is what the
sub-issues here are collectively for, and the README table is one of the last things to
change, not the first.
Work buckets
Sub-issues attach under these six. The list is the shape of the work, not a schedule.
1. New features — Graph and Pipeline.
Graph Execution step 2: every entry in GRAPH_EXECUTION.md "Current unsupported cases" is a
candidate sub-issue — dynamic boundary scalars, nested Graphs, dispatch predicates,
runtime-allocated boundary outputs, variable boundary metadata, cross-boundary explicit
dependencies, and the 16-Definition / 1024-node / 32-tensor ceilings. Pipelining: how deep
the slot pipeline should go beyond two, and what remains un-overlapped between runs.
2. Performance.
Host build time is now on the critical path. Open work: where a run's wall-clock actually
goes (build vs. relocation vs. H2D vs. device), whether successor pipelining genuinely
hides the build, Graph-replay speedup on a repeated layer, and the whole-graph capacity
model — sizing the task window / heap / fanin / TensorMap pool, and whether anything can
become reusable within a run rather than only at runtime destruction.
3. Redundant-code removal.
SchedulerContext::fanin_satisfied() has no callers anywhere in the HBG tree. Early
producer propagation and the shared scheduler's early-staging code are retained for T&R
parity but inactive here. Each such path gets a decision — wire it in or delete it — not
indefinite retention.
4. Naming and documentation conventions.
The tree is still pto_*.h / PTO2* throughout; retire it opportunistically per
.claude/rules/codestyle.md rules 9–10 — one identifier at a time, inside changes already
touching that code, no sweep PR. Documentation and comments that no longer match the
code, in both trees: RUNTIME_LOGIC.md §1 and the aicpu_executor.cpp dispatch comment
still say every AICPU thread schedules its own cores, which 3S+1P superseded; dep_pool /
DepListPool comments survive in runtime/pto_ring_buffer.h, pto_shared_memory.h,
pto_runtime2.h and scheduler/pto_scheduler.h describing structures #1435 deleted. Also
the stale HBG description in src/a2a3/docs/runtimes.md, the broken example path in
docs/getting-started.md, and the README positioning once the vision above is realized.
5. Test-case coverage.
Scene dirs today: 14 on a2a3 and 4 on a5, against 34 and 33 for T&R. Neither arch has an
examples/{arch}/host_build_graph/ tree at all. Which T&R scenes have no HBG counterpart,
which a2a3 scenes have no a5 counterpart, and which HBG-specific behaviors (Graph replay,
pipeline slots, capacity exhaustion, the latched error codes) have no scene at all.
6. Benchmark.
A repeatable HBG benchmark story: which workloads, which numbers, and against what
baseline — HBG vs. T&R on the same graph, and HBG with vs. without Graph replay. This
depends on DFX parity, so "which of the five DFX features (#995) work on HBG and which are
T&R-only" belongs here, along with the Graph Execution swimlane lanes described in
GRAPH_EXECUTION.md §DFX.
Per-arch divergence
Expect deliberate divergence later, where the silicon differs: AICPU thread counts and
topology, AICore geometry, the completion-mailbox ABI, available async engines. One
implementation for both stays the default; a divergence needs a stated hardware reason.
Open work
See the sub-issues panel above — kept current automatically as items are attached or
closed.
This issue tracks all
host_build_graph(HBG) work, on a2a3 and a5 — what the runtimeis for, where it stands today, and the open work. Individual items are attached as
sub-issues below (the panel auto-updates as items are added or closed).
Scope: a2a3 and a5. HBG is one runtime with two arch trees, kept in sync. A sub-issue is
about HBG, not about an arch — file one, not two. See
Per-arch divergence for when that stops being true.
What HBG is
HBG separates graph construction from device execution
(
src/{arch}/runtime/host_build_graph/docs/RUNTIME_LOGIC.md§1):There is no device-side orchestrator thread. The AICPU threads split 3S+1P: the last
thread is a core-less resolution thread (P), the rest own AICore partitions and schedule
them (S). At least two threads are required (#1544). The host constructs the whole graph
before any device task can complete — §1 calls that ordering "the defining constraint of
the runtime", and everything below is a consequence of it.
Structurally HBG is the host-orchestration variant of
tensormap_and_ringbuffer, not aseparate graph model: since #1185 it uses the same
PTO2TaskDescriptorring storage and thesame TensorMap-derived dependencies, with one ring and no device orchestrator. (The
"fixed
Task[]array, up to 131,072 tasks, traverses dependencies" description still insrc/a2a3/docs/runtimes.mdand in T&R'sRUNTIME_LOGIC.md§1.1 is stale.)Intent and vision
1. Take orchestration off the AICPU. No AICPU thread builds the graph, so the whole
device-side thread budget goes to dispatch and completion resolution. Graph construction
runs in ordinary host C++ — full STL, a host debugger, host-side DFX (
dep_genis capturedon the host orchestrator, #1492) — and no device-log budget is spent on build-time
diagnostics.
2. Pair the polling readiness model with the 3S+1P split. Dependency state is a per-slot
completion flag, a monotonic watermark, intrusive wake lists and inline integer fanin —
no dependency pool, no fanout lock, and no periodic sweep (#1435). That collapses
readiness resolution into one compact sequential job, which is exactly what can be handed
to a single thread: P owns all of it, so it is the sole producer of the ready queues, while
each S thread stays purely core-local — poll its own cores, dispatch them, hand finished
tasks to P over an SPSC queue (#1544). Whole-graph residency is what allows the split: all
threads classify disjoint slices of the finished graph once at boot (#1553), after which no
new task ever arrives. The gain is graph-shaped — a wide parallel graph wins, a serial chain
pays for the extra hop — so quantifying it belongs to the benchmark bucket.
3. Amortize the host build cost — this is the crux. Paying for a full host graph build
on every run is only acceptable if repeated work is not rebuilt. Two mechanisms exist:
docs/GRAPH_EXECUTION.md, Add Graph Execution to host_build_graph #1444): a Graph is a composite incoretask. The first invocation records the DAG; later invocations submit one
GRAPHtaskwhose internal nodes the device Scheduler expands from a saved, validated, pointer-free
POD Definition. Internal nodes consume no ring slots and are never re-submitted by the
host. This is what makes a repetitive workload (an LLM decoder layer submitted per layer,
per token) affordable under host orchestration.
with device execution of run N. HBG does not overlap orchestration and scheduling
within one run — this is the compensating mechanism.
4. Where this should land.
README.mdstill positions HBG as the"development, debugging" runtime against T&R's "production workloads". The vision is that
for repetitive, statically-shaped graphs, host orchestration plus Graph replay becomes a
first-class production path rather than the simple-but-slow option — while T&R remains the
runtime for streaming, dynamic, and flow-controlled workloads. Closing that gap is what the
sub-issues here are collectively for, and the README table is one of the last things to
change, not the first.
Work buckets
Sub-issues attach under these six. The list is the shape of the work, not a schedule.
1. New features — Graph and Pipeline.
Graph Execution step 2: every entry in
GRAPH_EXECUTION.md"Current unsupported cases" is acandidate sub-issue — dynamic boundary scalars, nested Graphs, dispatch predicates,
runtime-allocated boundary outputs, variable boundary metadata, cross-boundary explicit
dependencies, and the 16-Definition / 1024-node / 32-tensor ceilings. Pipelining: how deep
the slot pipeline should go beyond two, and what remains un-overlapped between runs.
2. Performance.
Host build time is now on the critical path. Open work: where a run's wall-clock actually
goes (build vs. relocation vs. H2D vs. device), whether successor pipelining genuinely
hides the build, Graph-replay speedup on a repeated layer, and the whole-graph capacity
model — sizing the task window / heap / fanin / TensorMap pool, and whether anything can
become reusable within a run rather than only at runtime destruction.
3. Redundant-code removal.
SchedulerContext::fanin_satisfied()has no callers anywhere in the HBG tree. Earlyproducer propagation and the shared scheduler's early-staging code are retained for T&R
parity but inactive here. Each such path gets a decision — wire it in or delete it — not
indefinite retention.
4. Naming and documentation conventions.
The tree is still
pto_*.h/PTO2*throughout; retire it opportunistically per.claude/rules/codestyle.mdrules 9–10 — one identifier at a time, inside changes alreadytouching that code, no sweep PR. Documentation and comments that no longer match the
code, in both trees:
RUNTIME_LOGIC.md§1 and theaicpu_executor.cppdispatch commentstill say every AICPU thread schedules its own cores, which 3S+1P superseded;
dep_pool/DepListPoolcomments survive inruntime/pto_ring_buffer.h,pto_shared_memory.h,pto_runtime2.handscheduler/pto_scheduler.hdescribing structures #1435 deleted. Alsothe stale HBG description in
src/a2a3/docs/runtimes.md, the broken example path indocs/getting-started.md, and the README positioning once the vision above is realized.5. Test-case coverage.
Scene dirs today: 14 on a2a3 and 4 on a5, against 34 and 33 for T&R. Neither arch has an
examples/{arch}/host_build_graph/tree at all. Which T&R scenes have no HBG counterpart,which a2a3 scenes have no a5 counterpart, and which HBG-specific behaviors (Graph replay,
pipeline slots, capacity exhaustion, the latched error codes) have no scene at all.
6. Benchmark.
A repeatable HBG benchmark story: which workloads, which numbers, and against what
baseline — HBG vs. T&R on the same graph, and HBG with vs. without Graph replay. This
depends on DFX parity, so "which of the five DFX features (#995) work on HBG and which are
T&R-only" belongs here, along with the Graph Execution swimlane lanes described in
GRAPH_EXECUTION.md§DFX.Per-arch divergence
Expect deliberate divergence later, where the silicon differs: AICPU thread counts and
topology, AICore geometry, the completion-mailbox ABI, available async engines. One
implementation for both stays the default; a divergence needs a stated hardware reason.
Open work
See the sub-issues panel above — kept current automatically as items are attached or
closed.