Skip to content

[Tracking] host_build_graph — intent, vision & change tracking #1706

Description

@ChaoZheng109

This issue tracks all host_build_graph (HBG) work, on a2a3 and a5 — what the runtime
is for, where it stands today, and the open work. Individual items are attached as
sub-issues below (the panel auto-updates as items are added or closed).

Scope: a2a3 and a5. HBG is one runtime with two arch trees, kept in sync. A sub-issue is
about HBG, not about an arch — file one, not two. See
Per-arch divergence for when that stops being true.

What HBG is

HBG separates graph construction from device execution
(src/{arch}/runtime/host_build_graph/docs/RUNTIME_LOGIC.md §1):

host: dlopen the orchestration SO, run it to completion → whole task graph
host: relocate the prebuilt graph image, copy it to device memory
device: attach, classify every task once, dispatch with the AICPU schedulers
host: collect outputs, destroy/reset per-run state

There is no device-side orchestrator thread. The AICPU threads split 3S+1P: the last
thread is a core-less resolution thread (P), the rest own AICore partitions and schedule
them (S). At least two threads are required (#1544). The host constructs the whole graph
before any device task can complete — §1 calls that ordering "the defining constraint of
the runtime", and everything below is a consequence of it.

Structurally HBG is the host-orchestration variant of tensormap_and_ringbuffer, not a
separate graph model: since #1185 it uses the same PTO2TaskDescriptor ring storage and the
same TensorMap-derived dependencies, with one ring and no device orchestrator. (The
"fixed Task[] array, up to 131,072 tasks, traverses dependencies" description still in
src/a2a3/docs/runtimes.md and in T&R's RUNTIME_LOGIC.md §1.1 is stale.)

Intent and vision

1. Take orchestration off the AICPU. No AICPU thread builds the graph, so the whole
device-side thread budget goes to dispatch and completion resolution. Graph construction
runs in ordinary host C++ — full STL, a host debugger, host-side DFX (dep_gen is captured
on the host orchestrator, #1492) — and no device-log budget is spent on build-time
diagnostics.

2. Pair the polling readiness model with the 3S+1P split. Dependency state is a per-slot
completion flag, a monotonic watermark, intrusive wake lists and inline integer fanin —
no dependency pool, no fanout lock, and no periodic sweep (#1435). That collapses
readiness resolution into one compact sequential job, which is exactly what can be handed
to a single thread: P owns all of it, so it is the sole producer of the ready queues, while
each S thread stays purely core-local — poll its own cores, dispatch them, hand finished
tasks to P over an SPSC queue (#1544). Whole-graph residency is what allows the split: all
threads classify disjoint slices of the finished graph once at boot (#1553), after which no
new task ever arrives. The gain is graph-shaped — a wide parallel graph wins, a serial chain
pays for the extra hop — so quantifying it belongs to the benchmark bucket.

3. Amortize the host build cost — this is the crux. Paying for a full host graph build
on every run is only acceptable if repeated work is not rebuilt. Two mechanisms exist:

4. Where this should land. README.md still positions HBG as the
"development, debugging" runtime against T&R's "production workloads". The vision is that
for repetitive, statically-shaped graphs, host orchestration plus Graph replay becomes a
first-class production path rather than the simple-but-slow option — while T&R remains the
runtime for streaming, dynamic, and flow-controlled workloads. Closing that gap is what the
sub-issues here are collectively for, and the README table is one of the last things to
change, not the first.

Work buckets

Sub-issues attach under these six. The list is the shape of the work, not a schedule.

1. New features — Graph and Pipeline.
Graph Execution step 2: every entry in GRAPH_EXECUTION.md "Current unsupported cases" is a
candidate sub-issue — dynamic boundary scalars, nested Graphs, dispatch predicates,
runtime-allocated boundary outputs, variable boundary metadata, cross-boundary explicit
dependencies, and the 16-Definition / 1024-node / 32-tensor ceilings. Pipelining: how deep
the slot pipeline should go beyond two, and what remains un-overlapped between runs.

2. Performance.
Host build time is now on the critical path. Open work: where a run's wall-clock actually
goes (build vs. relocation vs. H2D vs. device), whether successor pipelining genuinely
hides the build, Graph-replay speedup on a repeated layer, and the whole-graph capacity
model — sizing the task window / heap / fanin / TensorMap pool, and whether anything can
become reusable within a run rather than only at runtime destruction.

3. Redundant-code removal.
SchedulerContext::fanin_satisfied() has no callers anywhere in the HBG tree. Early
producer propagation and the shared scheduler's early-staging code are retained for T&R
parity but inactive here. Each such path gets a decision — wire it in or delete it — not
indefinite retention.

4. Naming and documentation conventions.
The tree is still pto_*.h / PTO2* throughout; retire it opportunistically per
.claude/rules/codestyle.md rules 9–10 — one identifier at a time, inside changes already
touching that code, no sweep PR. Documentation and comments that no longer match the
code, in both trees: RUNTIME_LOGIC.md §1 and the aicpu_executor.cpp dispatch comment
still say every AICPU thread schedules its own cores, which 3S+1P superseded; dep_pool /
DepListPool comments survive in runtime/pto_ring_buffer.h, pto_shared_memory.h,
pto_runtime2.h and scheduler/pto_scheduler.h describing structures #1435 deleted. Also
the stale HBG description in src/a2a3/docs/runtimes.md, the broken example path in
docs/getting-started.md, and the README positioning once the vision above is realized.

5. Test-case coverage.
Scene dirs today: 14 on a2a3 and 4 on a5, against 34 and 33 for T&R. Neither arch has an
examples/{arch}/host_build_graph/ tree at all. Which T&R scenes have no HBG counterpart,
which a2a3 scenes have no a5 counterpart, and which HBG-specific behaviors (Graph replay,
pipeline slots, capacity exhaustion, the latched error codes) have no scene at all.

6. Benchmark.
A repeatable HBG benchmark story: which workloads, which numbers, and against what
baseline — HBG vs. T&R on the same graph, and HBG with vs. without Graph replay. This
depends on DFX parity, so "which of the five DFX features (#995) work on HBG and which are
T&R-only" belongs here, along with the Graph Execution swimlane lanes described in
GRAPH_EXECUTION.md §DFX.

Per-arch divergence

Expect deliberate divergence later, where the silicon differs: AICPU thread counts and
topology, AICore geometry, the completion-mailbox ABI, available async engines. One
implementation for both stays the default; a divergence needs a stated hardware reason.

Open work

See the sub-issues panel above — kept current automatically as items are attached or
closed.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

No labels
No labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions