Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
68 changes: 46 additions & 22 deletions src/a2a3/runtime/host_build_graph/docs/GRAPH_EXECUTION.md
Original file line number Diff line number Diff line change
Expand Up @@ -9,7 +9,7 @@ places one `GRAPH` task in the host task window. The device Scheduler expands
the saved topology and dispatches its internal nodes; the Host Orchestrator does
not submit those nodes again.

## Step-1 API
## API

A Graph uses `CoreTaskArgs`, the existing incore argument type:

Expand All @@ -27,7 +27,7 @@ void graph_function(const CoreTaskArgs &args, int variant) {
CoreTaskArgs matmul_args;
matmul_args.add_input(input, weight);
matmul_args.add_output(intermediate);
matmul_args.add_scalar(uint32_t{16}); // fixed Definition data
matmul_args.copy_scalars_from(args, 0, 1); // current invocation's value
TaskOutputTensors matmul = rt_submit_aic_task(
variant == 0 ? FUNC_MATMUL : FUNC_MATMUL_TRANSPOSED,
matmul_args
Expand Down Expand Up @@ -69,19 +69,39 @@ the same key for different functions can select the wrong recorded topology.
There are no public `GraphArgs`, `GraphBindings`, `Patch`, or `ScalarRef`
types. The boundary is represented by `CoreTaskArgs`.

## Supported dynamic and static data
Boundary scalars are pass-through bindings. Forward them directly with
`node_args.add_scalar(args.scalar(i))` or `copy_scalars_from(args, i, count)`
so recording can retain their source indices.

Ordinary C++ value transformations do not retain boundary provenance. Both
`node_args.add_scalar(args.scalar(i) + 1)` and copying `args.scalar(i)` into a
local arithmetic variable before calling `add_scalar` produce an ordinary
static node scalar. That value is stored in the Definition, and later cache
hits reuse the first invocation's value without a warning. The runtime cannot
distinguish such a derived value from an intentional static literal after the
C++ expression has produced a plain arithmetic value. Compute the derived value
before constructing the Graph boundary and pass it as another boundary scalar,
perform the transformation in a kernel, or use a construction parameter when
the value changes the Graph structure.

Access through a non-const `scalar()` invalidates inherited boundary provenance
conservatively, because returning a mutable reference cannot distinguish a read
from a later write. A Graph containing such an invalidated binding is not
cached, which prevents replay from silently replacing the transformed value
with the unmodified boundary value.

Step 1 deliberately supports a narrow, safe contract:
## Supported dynamic and static data

- Boundary ChipTensor addresses may change for every invocation.
- Boundary scalar values may change for every invocation. Their count is fixed
by the recorded boundary contract. Unused boundary scalars are allowed and do
not create internal scalar patches.
- A Graph boundary contains at least one ChipTensor.
- Construction parameters are part of Graph identity and may control the
function's task count, kernel selection, or other structural choices.
- Boundary ChipTensor shape, stride, dtype, size, direction, contiguity, and alias
partition must match the first invocation.
- Scalars inside internal task args are fixed Definition data.
- Scalars in the boundary `CoreTaskArgs` are not cacheable yet. Such a call uses
the ordinary task-submit path.
- Internal task scalars with no boundary source are fixed Definition data.
- Boundary storage is caller-owned. `INPUT`, `INOUT`, `OUTPUT_EXISTING`, and
`NO_DEP` are supported. A boundary `TensorCreateInfo` tagged `OUTPUT` is not.
- Early-resolve hints apply while recording the first invocation. Replayed
Expand Down Expand Up @@ -116,7 +136,7 @@ void qwen_decoder_layer(const CoreTaskArgs &args) {
CoreTaskArgs attention_args;
attention_args.add_input(hidden, attention_weight);
attention_args.add_output(attention_out);
attention_args.add_scalar(uint32_t{16}); // fixed model configuration
attention_args.copy_scalars_from(args, 0, 1); // dynamic token position
TaskOutputTensors attention =
rt_submit_aic_task(FUNC_ATTENTION, attention_args);

Expand All @@ -138,7 +158,8 @@ void decode_three_layers(
const std::array<ChipTensor, 3> &hidden,
const std::array<ChipTensor, 3> &attention_weight,
const std::array<ChipTensor, 3> &mlp_weight,
const std::array<ChipTensor, 3> &output
const std::array<ChipTensor, 3> &output,
const std::array<uint32_t, 3> &token_position
) {
for (std::size_t layer = 0; layer < hidden.size(); ++layer) {
CoreTaskArgs args;
Expand All @@ -148,15 +169,16 @@ void decode_three_layers(
mlp_weight[layer]
);
args.add_output(output[layer]);
args.add_scalar(token_position[layer]);
submit_qwen_decoder_layer(args);
}
}
```

The first layer records ordinary task submissions. Layers two and three submit
one Graph task each when their ChipTensor metadata matches. A per-layer or
per-token scalar is not dynamic in step 1; use ordinary submission or a
different fixed Graph function/key until dynamic scalar support is added.
one Graph task each when their ChipTensor metadata and boundary scalar count
match. Each replay patches the current layer's `token_position`; its value is
not part of the Graph key.

## Definition

Expand All @@ -179,7 +201,7 @@ Definition. It contains:
- one packed-heap offset per node;
- each node's ChipTensor source:
`BOUNDARY_EXACT`, `BOUNDARY_VIEW`, `INTERNAL`, or `OWN_OUTPUT`;
- fixed scalar values;
- fixed scalar values plus boundary-scalar source indices;
- fixed boundary signatures and alias representatives.

The header also carries a content hash of the complete Definition image. The
Expand Down Expand Up @@ -211,8 +233,8 @@ For a cache hit, the Host Orchestrator:
5. emits one outer `GRAPH` task;
6. stages the exact-size POD submission image for upload after orchestration;
7. asks the host runtime for an aligned execution block sized from the recorded
node count, tensor-address patch count, and Definition bytes, then writes
that device address into the submission wire image.
node count, Tensor-address and scalar patch capacities, and Definition
bytes, then writes that device address into the submission wire image.

Internal nodes consume no ring task-window slots. Their descriptor, payload,
and slot state are built in host-owned GM. The runtime retains one grow-only
Expand All @@ -226,12 +248,14 @@ The `GraphSubmission` wire POD carries the aligned device address and usable
byte capacity explicitly. The Scheduler validates both before placement-
constructing `GraphExecution`; it never allocates execution storage from the
AICPU process heap. A block whose prior Definition key and content hash match
retains the local Definition, static node fields, and the address patch table
generated during its first materialization. That graph-affine replay skips
retains the local Definition, static node fields, and the Tensor-address and
scalar patch tables generated during its first materialization. That
graph-affine replay skips
topology binding, per-node count/offset validation, tensor-source
classification, tensor wire validation, static field stores, and scalar
classification, tensor wire validation, static field stores, and static scalar
copies. It refreshes only task IDs, packed-buffer bases, boundary/internal
tensor addresses, scheduling state, dispatch atomics, and wake registrations.
tensor addresses, boundary scalar bindings, scheduling state, dispatch
atomics, and wake registrations.

The retained blocks are addressed directly by `(pipeline slot, Graph key,
occurrence index)`. Occurrence numbering restarts deterministically for every
Expand Down Expand Up @@ -274,8 +298,8 @@ Internal dependency readiness borrows the completion-state polling idea, but
dependency wiring remains an Orchestrator responsibility:

- recording constructs both fanin and fanout CSR in the immutable Definition;
- first materialization builds static runnable node state plus a compact
boundary/internal address patch table; affine replay applies that table and
- first materialization builds static runnable node state plus compact Tensor
address and scalar patch tables; affine replay applies those tables and
resets only dynamic runnable state;
- materialization registers each non-root on one producer selected from its
saved fanin CSR;
Expand Down Expand Up @@ -315,7 +339,6 @@ error instead of leaving an already-submitted outer Graph unable to complete.
These cases assert in debug builds and execute through the ordinary path in a
release build:

- dynamic boundary scalars;
- an empty Graph boundary;
- variable ChipTensor shape or metadata;
- changed boundary aliasing;
Expand All @@ -325,6 +348,7 @@ release build:
- cross-boundary explicit dependencies that are not represented by a boundary
ChipTensor's creator;
- an unclassifiable internal ChipTensor source;
- a boundary-derived scalar accessed through mutable `scalar()`;
- more than 16 Definitions, 1024 internal nodes, or 32 boundary Tensors;
- insufficient task-window or heap capacity detected before outer submission.

Expand Down
4 changes: 2 additions & 2 deletions src/a2a3/runtime/host_build_graph/host/runtime_maker.cpp
Original file line number Diff line number Diff line change
Expand Up @@ -417,8 +417,8 @@ bool upload_graph_submissions(Runtime *runtime, const HostApi *api, GraphHostSta
if (definition == nullptr || definition->full_key != submission->graph_key || definition->task_count == 0 ||
definition->task_count > GRAPH_MAX_NODES ||
!graph_execution_storage_bytes(
static_cast<int32_t>(definition->task_count), definition->tensor_arg_count, definition->total_bytes,
&execution_bytes
static_cast<int32_t>(definition->task_count), definition->tensor_arg_count,
definition->scalar_arg_count, definition->total_bytes, &execution_bytes
)) {
LOG_ERROR("host-orch: invalid Graph execution storage request");
return false;
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -391,10 +391,8 @@ template <typename Invoke>
static inline GraphSubmitResult rt_submit_graph_impl(uint64_t graph_key, const CoreTaskArgs &args, Invoke invoke) {
debug_assert(!args.has_error && "Graph boundary CoreTaskArgs construction failed");
debug_assert(
args.tensor_count() <= static_cast<int32_t>(GRAPH_MAX_TENSOR_ARGS) &&
"Graph boundary exceeds the step-1 tensor limit"
args.tensor_count() <= static_cast<int32_t>(GRAPH_MAX_TENSOR_ARGS) && "Graph boundary exceeds the tensor limit"
);
debug_assert(args.scalar_count() == 0 && "Dynamic Graph boundary scalars are not supported in step 1");
debug_assert(
args.explicit_dep_count() == 0 && "Explicit dependencies crossing the Graph boundary are not supported"
);
Expand Down
4 changes: 1 addition & 3 deletions src/a2a3/runtime/host_build_graph/runtime/graph_cache.h
Original file line number Diff line number Diff line change
Expand Up @@ -46,10 +46,8 @@ constexpr uint64_t graph_const_hash_impl(const char *s, uint64_t h) {
constexpr uint64_t GRAPH_KEY(const char *s) { return graph_const_hash_impl(s, 1469598103934665603ULL); }

inline bool rt_graph_args_cacheable(const CoreTaskArgs &args) {
// Step 1 supports dynamic tensor addresses only. Kernel scalars are
// literals inside the Graph function and become immutable Definition data.
if (args.has_error || args.tensor_count() <= 0 ||
args.tensor_count() > static_cast<int32_t>(GRAPH_MAX_TENSOR_ARGS) || args.scalar_count() != 0) {
args.tensor_count() > static_cast<int32_t>(GRAPH_MAX_TENSOR_ARGS)) {
return false;
}
for (int32_t i = 0; i < args.tensor_count(); ++i) {
Expand Down
75 changes: 65 additions & 10 deletions src/a2a3/runtime/host_build_graph/runtime/graph_execution.h
Original file line number Diff line number Diff line change
Expand Up @@ -62,6 +62,17 @@ struct GraphTensorSourceRef {
uint64_t packed_offset;
};

enum class GraphScalarSource : uint8_t {
STATIC_VALUE = 0,
BOUNDARY = 1,
};

struct GraphScalarSourceRef {
uint16_t source_index;
uint8_t source;
uint8_t reserved;
};

struct GraphNodeDefinition {
int32_t kernel_id[PTO2_SUBTASK_SLOT_COUNT];
uint8_t active_mask;
Expand Down Expand Up @@ -98,6 +109,7 @@ struct GraphDefinition {
uint32_t edge_count;
uint32_t root_count;
uint32_t boundary_count;
uint32_t boundary_scalar_count;
uint32_t tensor_arg_count;
uint32_t scalar_arg_count;
uint32_t off_fanout_offsets;
Expand All @@ -110,6 +122,7 @@ struct GraphDefinition {
uint32_t off_tensors;
uint32_t off_tensor_sources;
uint32_t off_scalars;
uint32_t off_scalar_sources;
uint32_t off_boundary_signatures;
};

Expand All @@ -123,13 +136,16 @@ struct GraphSubmission {
uint32_t definition_offset;
uint32_t tensors_offset;
uint32_t tensor_count;
uint32_t reserved;
uint32_t scalars_offset;
uint32_t scalar_count;
};

static_assert(std::is_trivially_copyable_v<GraphTensorSourceRef>);
static_assert(std::is_standard_layout_v<GraphTensorSourceRef>);
static_assert(std::is_trivially_copyable_v<GraphTensor>);
static_assert(std::is_standard_layout_v<GraphTensor>);
static_assert(std::is_trivially_copyable_v<GraphScalarSourceRef>);
static_assert(std::is_standard_layout_v<GraphScalarSourceRef>);
static_assert(std::is_trivially_copyable_v<GraphNodeDefinition>);
static_assert(std::is_standard_layout_v<GraphNodeDefinition>);
static_assert(std::is_trivially_copyable_v<GraphBoundarySignature>);
Expand Down Expand Up @@ -243,6 +259,18 @@ inline const GraphTensor *graph_submission_tensors(const GraphSubmission &submis
);
}

inline const uint64_t *graph_submission_scalars(const GraphSubmission &submission) {
if (submission.scalar_count == 0) return nullptr;
if (submission.scalars_offset == 0 || submission.scalars_offset % alignof(uint64_t) != 0 ||
submission.scalars_offset > submission.total_bytes ||
submission.scalar_count > (submission.total_bytes - submission.scalars_offset) / sizeof(uint64_t)) {
return nullptr;
}
return reinterpret_cast<const uint64_t *>(
reinterpret_cast<const uint8_t *>(&submission) + submission.scalars_offset
);
}

enum class GraphExecutionState : uint8_t {
SUBMITTED = 0,
MATERIALIZING = 1,
Expand Down Expand Up @@ -277,6 +305,18 @@ static_assert(std::is_trivially_copyable_v<GraphTensorAddressPatch>);
static_assert(std::is_standard_layout_v<GraphTensorAddressPatch>);
static_assert(sizeof(GraphTensorAddressPatch) == 16);

struct GraphScalarPatch {
uint16_t node_index;
uint8_t node_scalar_index;
uint8_t boundary_scalar_index;
};

static_assert(std::is_trivially_copyable_v<GraphScalarPatch>);
static_assert(std::is_standard_layout_v<GraphScalarPatch>);
static_assert(sizeof(GraphScalarPatch) == 4);
static_assert(GRAPH_MAX_NODES <= UINT16_MAX);
static_assert(MAX_SCALAR_ARGS <= UINT8_MAX);

struct alignas(64) GraphNodeStorage {
PTO2TaskDescriptor task;
PTO2TaskPayload payload;
Expand All @@ -300,6 +340,9 @@ struct GraphExecution {
uint32_t tensor_patch_capacity{0};
uint32_t materialized_tensor_patches{0};
uint32_t materialized_tensor_patch_count{0};
uint32_t scalar_patch_capacity{0};
uint32_t materialized_scalar_patches{0};
uint32_t materialized_scalar_patch_count{0};
size_t allocation_bytes{0};
size_t definition_capacity{0};
uint64_t graph_key{0};
Expand All @@ -312,25 +355,30 @@ struct GraphExecution {
GraphNodeStorage *nodes{nullptr};
GraphNodeStorage *node_storage{nullptr};
GraphTensorAddressPatch *tensor_patches{nullptr};
GraphScalarPatch *scalar_patches{nullptr};
void *definition_storage{nullptr};
const GraphDefinition *definition{nullptr};
const uint32_t *fanin_offsets{nullptr};
const uint16_t *fanin_indices{nullptr};
const GraphTensor *boundary_tensors{nullptr};
uint32_t boundary_tensor_count{0};
const uint64_t *boundary_scalars{nullptr};
uint32_t boundary_scalar_count{0};
};

static_assert(std::is_trivially_destructible_v<GraphNodeStorage>);
static_assert(std::is_trivially_destructible_v<GraphExecution>);

inline bool graph_execution_storage_layout(
int32_t node_capacity, uint32_t tensor_patch_capacity, size_t definition_capacity, size_t *nodes_offset,
size_t *tensor_patches_offset, size_t *definition_offset, size_t *storage_bytes
int32_t node_capacity, uint32_t tensor_patch_capacity, uint32_t scalar_patch_capacity, size_t definition_capacity,
size_t *nodes_offset, size_t *tensor_patches_offset, size_t *scalar_patches_offset, size_t *definition_offset,
size_t *storage_bytes
) {
if (nodes_offset == nullptr || tensor_patches_offset == nullptr || definition_offset == nullptr ||
storage_bytes == nullptr || node_capacity <= 0 ||
if (nodes_offset == nullptr || tensor_patches_offset == nullptr || scalar_patches_offset == nullptr ||
definition_offset == nullptr || storage_bytes == nullptr || node_capacity <= 0 ||
static_cast<size_t>(node_capacity) > SIZE_MAX / sizeof(GraphNodeStorage) ||
tensor_patch_capacity > GRAPH_MAX_NODES * MAX_TENSOR_ARGS) {
tensor_patch_capacity > GRAPH_MAX_NODES * MAX_TENSOR_ARGS ||
scalar_patch_capacity > GRAPH_MAX_NODES * MAX_SCALAR_ARGS) {
return false;
}
auto checked_align_up = [](size_t value, size_t alignment, size_t *result) {
Expand All @@ -340,26 +388,33 @@ inline bool graph_execution_storage_layout(
};
const size_t nodes_bytes = static_cast<size_t>(node_capacity) * sizeof(GraphNodeStorage);
const size_t tensor_patches_bytes = static_cast<size_t>(tensor_patch_capacity) * sizeof(GraphTensorAddressPatch);
const size_t scalar_patches_bytes = static_cast<size_t>(scalar_patch_capacity) * sizeof(GraphScalarPatch);
if (!checked_align_up(sizeof(GraphExecution), alignof(GraphNodeStorage), nodes_offset) ||
*nodes_offset > SIZE_MAX - nodes_bytes ||
!checked_align_up(*nodes_offset + nodes_bytes, alignof(GraphTensorAddressPatch), tensor_patches_offset) ||
*tensor_patches_offset > SIZE_MAX - tensor_patches_bytes ||
!checked_align_up(*tensor_patches_offset + tensor_patches_bytes, alignof(GraphDefinition), definition_offset) ||
!checked_align_up(
*tensor_patches_offset + tensor_patches_bytes, alignof(GraphScalarPatch), scalar_patches_offset
) ||
*scalar_patches_offset > SIZE_MAX - scalar_patches_bytes ||
!checked_align_up(*scalar_patches_offset + scalar_patches_bytes, alignof(GraphDefinition), definition_offset) ||
*definition_offset > SIZE_MAX - definition_capacity) {
return false;
}
return checked_align_up(*definition_offset + definition_capacity, alignof(GraphNodeStorage), storage_bytes);
}

inline bool graph_execution_storage_bytes(
int32_t node_capacity, uint32_t tensor_patch_capacity, size_t definition_capacity, size_t *storage_bytes
int32_t node_capacity, uint32_t tensor_patch_capacity, uint32_t scalar_patch_capacity, size_t definition_capacity,
size_t *storage_bytes
) {
size_t nodes_offset = 0;
size_t tensor_patches_offset = 0;
size_t scalar_patches_offset = 0;
size_t definition_offset = 0;
return graph_execution_storage_layout(
node_capacity, tensor_patch_capacity, definition_capacity, &nodes_offset, &tensor_patches_offset,
&definition_offset, storage_bytes
node_capacity, tensor_patch_capacity, scalar_patch_capacity, definition_capacity, &nodes_offset,
&tensor_patches_offset, &scalar_patches_offset, &definition_offset, storage_bytes
);
}

Expand Down
Loading
Loading