Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
25 changes: 17 additions & 8 deletions src/a5/runtime/host_build_graph/docs/GRAPH_EXECUTION.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,10 +4,18 @@ Graph Execution is available only in the `host_build_graph` runtime. A Graph is
a composite incore task: it is submitted and completed once like an AIC, AIV,
MIX, or SPMD task, but contains a recorded task DAG.

The first invocation executes normally and records the DAG. A later invocation
places one `GRAPH` task in the host task window. The device Scheduler expands
the saved topology and dispatches its internal nodes; the Host Orchestrator does
not submit those nodes again.
Every invocation places exactly one `GRAPH` task in the host task window. The
first invocation records the DAG off the ring — its internal submissions build
host-only node metadata and reserve scratch output buffers instead of consuming
task-window slots — then emits the outer `GRAPH` task from the freshly built
Definition. Later invocations reuse the cached Definition and emit the same one
`GRAPH` task directly. In both cases the device Scheduler expands the saved
topology and dispatches the internal nodes; the Host Orchestrator never submits
those nodes as ring tasks.

A recording that hits an unsupported construct is discarded and the body re-runs
on the ordinary task-submit path so its work is still submitted; the internal
nodes then occupy the ring only for that one fallback invocation.

## API

Expand Down Expand Up @@ -175,10 +183,11 @@ void decode_three_layers(
}
```

The first layer records ordinary task submissions. Layers two and three submit
one Graph task each when their ChipTensor metadata and boundary scalar count
match. Each replay patches the current layer's `token_position`; its value is
not part of the Graph key.
All three layers submit one Graph task each: the first records the sub-DAG off
the ring and emits its Graph task, layers two and three replay the cached
Definition when their ChipTensor metadata and boundary scalar count match. Each
invocation patches the current layer's `token_position`; it is a dynamic
boundary scalar refreshed on every submission and is not part of the Graph key.

## Definition

Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -94,7 +94,7 @@ typedef struct PTO2RuntimeOps {
int32_t (*available_cluster_count)(PTO2Runtime *rt);
int32_t (*available_aiv_count)(PTO2Runtime *rt);
GraphScopeResult (*graph_begin)(PTO2Runtime *rt, uint64_t graph_key, const CoreTaskArgs &args);
void (*graph_end)(PTO2Runtime *rt);
bool (*graph_end)(PTO2Runtime *rt);
void (*graph_commit)(PTO2Runtime *rt);

// Stash the call-site of the next PTO2ScopeGuard so the [ScopeStats]
Expand Down Expand Up @@ -221,12 +221,17 @@ static inline GraphScopeResult rt_graph_begin(uint64_t graph_key, const CoreTask
return rt->ops->graph_begin(rt, graph_key, args);
}

static inline void rt_graph_end() {
// Finish the recording pass. Returns true when the recorded sub-DAG was emitted
// as a single outer GRAPH task (the body must not run again); false when the
// recording was unsupported and the caller must re-run the body on the ordinary
// path. A fatal runtime is terminal, so it reports true to suppress a pointless
// re-run.
static inline bool rt_graph_end() {
PTO2Runtime *rt = current_runtime();
if (rt->ops->is_fatal(rt) || rt->ops->graph_end == nullptr) {
return;
return true;
}
rt->ops->graph_end(rt);
return rt->ops->graph_end(rt);
}

static inline void rt_graph_commit() {
Expand Down Expand Up @@ -408,10 +413,17 @@ static inline GraphSubmitResult rt_submit_graph_impl(uint64_t graph_key, const C
return GraphSubmitResult{};
}
GraphScopeResult result = rt_graph_begin(graph_key, args);
if (result.execute_block) invoke();
if (result.recording) {
rt_graph_end();
// First invocation: record the sub-DAG off the ring, then emit one outer
// GRAPH task. If the recording hit an unsupported construct, fall back and
// run the body on the ordinary path so its work is still submitted.
invoke();
if (!rt_graph_end()) invoke();
} else if (result.execute_block) {
// Un-cacheable at begin, or the Definition cache is full: ordinary path.
invoke();
}
// Cache hit: execute_block and recording are both false; the body is skipped.
rt_graph_commit();
return result;
}
Expand Down
Loading
Loading