Skip to content
2 changes: 1 addition & 1 deletion docs/dfx/dep_gen.md
Original file line number Diff line number Diff line change
Expand Up @@ -361,7 +361,7 @@ list; only the dep_gen replay graph loses the tail.
| Shared-mem layout | `src/{a2a3,a5}/platform/include/common/dep_gen.h` | `DepGenRecord` (2624 B base, cache-line aligned, ≤64 inline explicit_deps) + `DepGenOverflowRecord` chain view (≤326 deps per slot) + SPSC ring + per-thread ready queue. Byte-identical layout across platforms. |
| AICPU writer | `src/{a2a3,a5}/platform/{include,shared}/aicpu/dep_gen_collector_aicpu.{h,cpp}` | Single-instance write path; weak-fallback exported to host build. a5 reuses the a2a3 source verbatim — the writer accesses its own device-side view of shared memory, independent of how host↔device transport is implemented. |
| Host collector | `src/{a2a3,a5}/platform/{include/host,shared/host}/dep_gen_collector.{h,cpp}` | `ProfilerBase<DepGenCollector, DepGenModule>` — drains ring → `records_` vector. On a5 (no SVM) it uses the base `alloc_paired_buffer`, which malloc's a host shadow + `copy_to_device`'s it and registers it via `add_malloc_shadow` so teardown can free it; `reconcile_counters` explicitly `copy_from_device`'s the BufferState before reading, and `finalize` lets `BufferPoolManager::clear_mappings()` release all shadows as the single source of truth. |
| Capture call site | `src/{a2a3,a5}/runtime/tensormap_and_ringbuffer/runtime/pto_orchestrator.cpp` `submit_task_common` | One conditional block that snapshots inputs into the ring when `is_dep_gen_enabled()`; fires for both `submit_task` and `submit_dummy_task`. Dep-only tasks land in the record stream with valid tensor/dep info but no kernel_id field (the schema does not carry kernel_id), so replay treats them as ordinary dep nodes — viewers do not currently distinguish dummy from real tasks. |
| Capture call site | `src/{a2a3,a5}/runtime/tensormap_and_ringbuffer/runtime/pto_orchestrator.cpp` `submit_task_common` | One conditional block that snapshots inputs into the ring when `is_dep_gen_enabled()`; fires for both `submit_task` and `submit_dummy_task`. **a2a3 only:** the schema additionally carries `kernel_id[3] = {aic, aiv0, aiv1}` so the swimlane post-processor can resolve `task_id → kernel` from `deps.json` at level=1 where the AICore record is the sole device-side identity source. Inactive subslots stay at `INVALID_KERNEL_ID = -1`. |
| Replay | `src/{a2a3,a5}/runtime/tensormap_and_ringbuffer/host/dep_gen_replay.{h,cpp}` | Pure CPU; runs dual-pass differential replay — `compute_task_fanin` (oracle) + inlined STEP A/B mirror (annotated) against two `PTO2TensorMap` instances. Emits `deps.json` when both passes agree per record. Platform-agnostic — a5 reuses the a2a3 source verbatim. |
| Device-runner hookup | `src/{a2a3,a5}/platform/{onboard,sim}/host/device_runner.cpp` | post-`reconcile_counters` calls `dep_gen_replay_emit_deps_json(records.data(), records.size(), deps_path)` |
| Viewer | `simpler_setup/tools/deps_to_graph.py` | `deps.json` → pan/zoom HTML |
Expand Down
87 changes: 86 additions & 1 deletion simpler_setup/tools/swimlane_converter.py
Original file line number Diff line number Diff line change
Expand Up @@ -162,6 +162,74 @@ def load_deps_json(deps_path):
return dict(by_pred)


def load_deps_kernel_map(deps_path):
"""Build a ``task_id → kernel_ids[3]`` map from deps.json's ``tasks[]``.

a2a3 dep_gen captures per-task ``kernel_ids = [aic, aiv0, aiv1]`` so the
swimlane post-processor can resolve ``func_id`` at AICORE_TIMING (level=1)
where the AICore record alone is on disk and carries ``func_id == -1``.
The trace generator uses the per-record ``core_type`` to pick the right
subslot: ``aic → kernel_ids[0]``, ``aiv → kernel_ids[1]`` (falling back
to ``[2]`` if AIV0 is inactive). Same pattern fanout edges already use
(deps.json is the offline-joined identity source).

Returns:
dict[int, list[int]] mapping ``task_id_raw → [aic, aiv0, aiv1]``,
or ``None`` if the file is missing / unreadable / lacks the field.
Entries without ``kernel_ids`` (pre-schema deps.json from older
runs) are silently skipped — the caller treats a missing map as
"no override available" and emits the ``func_-1_(...)`` fallback.
"""
deps_path = Path(deps_path)
if not deps_path.exists():
return None
try:
with deps_path.open() as f:
data = json.load(f)
except (OSError, ValueError):
return None
tasks = data.get("tasks")
if not isinstance(tasks, list):
return None
kmap: dict[int, list[int]] = {}
for task in tasks:
if not isinstance(task, dict):
continue
tid = normalize_pto2_task_id_int(task.get("task_id"))
kids = task.get("kernel_ids")
if tid is None or not isinstance(kids, list) or len(kids) != 3:
continue
kmap[tid] = [int(k) for k in kids]
return kmap if kmap else None


def resolve_func_id_from_kernel_map(task_id, core_type, kernel_map):
"""Look up the active ``func_id`` for an AICORE_TIMING record via dep_gen.

Picks the kernel_ids[3] subslot by record ``core_type``. Returns the
resolved func_id (>= 0) on a hit, or -1 if no usable subslot was found
(caller keeps the original -1 and emits the ``func_-1_(...)`` fallback
name). The choice for ``aiv`` prefers AIV0 ([1]) and falls back to AIV1
([2]) — works for pure-AIV and MIX-with-single-AIV records; for MIX
records that span both AIVs the host swimlane record only tells us the
lane is "aiv", so the resolver may name an AIV1 lane after AIV0's
kernel. Acceptable trade-off until the host emits a lane-disambiguated
core_type ("aiv0" / "aiv1").
"""
if kernel_map is None or task_id is None:
return -1
kids = kernel_map.get(int(task_id))
if not kids:
return -1
if core_type == "aic":
return kids[0] if kids[0] >= 0 else -1
# "aiv": prefer AIV0, fall back to AIV1.
for idx in (1, 2):
if kids[idx] >= 0:
return kids[idx]
return -1


def load_kernel_config(config_path):
"""Load kernel configuration from kernel_config.py file.

Expand Down Expand Up @@ -381,6 +449,7 @@ def generate_chrome_trace_json( # noqa: PLR0912, PLR0915
core_to_thread=None,
orchestrator_name=None,
deps_edges=None,
deps_kernel_map=None,
):
"""Generate Chrome Trace Event Format JSON from task data.

Expand Down Expand Up @@ -468,8 +537,16 @@ def generate_chrome_trace_json( # noqa: PLR0912, PLR0915
ts = task["start_time_us"]
dur = task["duration_us"]

# Get function name if available
# Get function name if available. At AICORE_TIMING (level=1) the
# host emits func_id=-1; recover the real func_id from dep_gen's
# per-task kernel_ids[3] using the record's core_type to pick the
# active subslot. See resolve_func_id_from_kernel_map() for the
# AIV0-vs-AIV1 tie-break and the host-side contract.
func_id = task["func_id"]
if int(func_id) < 0 and deps_kernel_map is not None:
resolved = resolve_func_id_from_kernel_map(task["task_id"], task.get("core_type"), deps_kernel_map)
if resolved >= 0:
func_id = resolved
tdisp = format_task_display(task["task_id"])
if func_id_to_name and str(func_id) in func_id_to_name:
func_name = func_id_to_name[str(func_id)]
Expand Down Expand Up @@ -1247,9 +1324,16 @@ def main():

deps_path = Path(args.deps_json) if args.deps_json else Path(input_path).parent / "deps.json"
deps_edges = load_deps_json(deps_path)
# Load the per-task kernel_ids map separately so the trace generator
# can resolve func_id=-1 records (AICORE_TIMING / level=1) back to
# the real kernel name. Optional — pre-schema deps.json without
# kernel_ids and AICPU_TIMING+ runs both leave this at None.
deps_kernel_map = load_deps_kernel_map(deps_path)
if deps_edges is not None:
if args.verbose:
print(f" Using deps.json edges ({sum(len(v) for v in deps_edges.values())} total) from {deps_path}")
if deps_kernel_map is not None:
print(f" Using deps.json kernel_ids for {len(deps_kernel_map)} tasks (level=1 name recovery)")
else:
print(
f"Warning: no usable deps.json at {deps_path}; Perfetto trace will have no dependency arrows. "
Expand All @@ -1267,6 +1351,7 @@ def main():
orchestrator_phases=data.get("aicpu_orchestrator_phases"),
core_to_thread=data.get("core_to_thread"),
deps_edges=deps_edges,
deps_kernel_map=deps_kernel_map,
)

print("\n✓ Conversion complete")
Expand Down
62 changes: 43 additions & 19 deletions src/a2a3/platform/include/aicore/l2_swimlane_collector_aicore.h
Original file line number Diff line number Diff line change
Expand Up @@ -65,27 +65,45 @@ struct L2SwimlaneAicoreLocalState {
* `docs/dfx/l2-swimlane-profiling.md`.
*
* Race avoidance: AICPU rotates strictly before `write_reg(DATA_MAIN_BASE)`
* for the first task of a new BUFFER_SIZE batch. The runtime's
* completion-before-dispatch invariant guarantees all prior tasks have FIN'd,
* so AICore has already finished writing their records before AICPU enqueues
* the old buffer to the ready queue.
* for the first task of a new BUFFER_SIZE batch — driven by AICPU's own
* per-core dispatch count (no AICore-side signal). The runtime's
* completion-before-dispatch invariant (AICore per core is single-threaded
* and AICPU does not dispatch task K+1 until K FIN'd) guarantees all prior
* tasks have FIN'd at rotation time, so AICore has already finished writing
* their records and dcci'd them out before AICPU enqueues the old buffer to
* the ready queue.
*
* @param head Per-core L2SwimlaneActiveHead channel — lazy-resolved on
* the executor's first-task branch via
* get_l2_swimlane_aicore_head(), which deref's the slot the
* kernel entry stashed from
* KernelArgs::l2_swimlane_aicore_rotation_table[block_idx].
* (Kernel entry can't deref directly — AICPU init runs
* concurrently with kernel entry, so the slot may not yet
* hold a valid address at that point.)
* @param local Per-core AICore-local state (caller-owned static)
* @param task_id Register dispatch id (DATA_MAIN_BASE), low 32 bits
* @param start_time Start timestamp (get_sys_cnt)
* @param end_time End timestamp
* @param head Per-core L2SwimlaneActiveHead channel — lazy-resolved on
* the executor's first-task branch via
* get_l2_swimlane_aicore_head(), which deref's the slot
* the kernel entry stashed from
* KernelArgs::l2_swimlane_aicore_rotation_table[block_idx].
* (Kernel entry can't deref directly — AICPU init runs
* concurrently with kernel entry, so the slot may not yet
* hold a valid address at that point.)
* @param local Per-core AICore-local state (caller-owned static)
* @param task_token_raw Full task identity (PTO2 encoding for tensormap_and_ringbuffer
* runtime: `(ring_id << 32) | local_id`; plain task index
* zero-extended for host_build_graph). The caller in the
* ringbuffer runtime reads this from
* `exec_payload->local_context.async_ctx.task_token.raw`
* which is already in AICore cache (it was just dcci'd for
* the kernel call), so no extra GM load.
* @param reg_task_id Per-core dispatch token (low 32 bits of the per-core
* monotonic dispatch_seq). Per-dispatch unique within
* a core; serves as the host-side join key against the
* AICPU record stream. Required because SPMD with
* `block_num > num_cores` (and MIX cluster spread)
* dispatch the same `task_token_raw` multiple times to
* the same core — each dispatch needs its own AICore
* record matched to its own AICPU record, which
* task_token_raw alone cannot disambiguate.
* @param start_time Start timestamp (get_sys_cnt)
* @param end_time End timestamp
*/
__aicore__ __attribute__((always_inline)) static inline void l2_swimlane_aicore_record_task(
__gm__ L2SwimlaneActiveHead *head, L2SwimlaneAicoreLocalState *local, uint32_t task_id, uint64_t start_time,
uint64_t end_time
__gm__ L2SwimlaneActiveHead *head, L2SwimlaneAicoreLocalState *local, uint64_t task_token_raw, uint32_t reg_task_id,
uint64_t start_time, uint64_t end_time
) {
// Re-fetch head channel each task; cheap relative to the
// baseline `dcci(payload, ENTIRE_DATA_CACHE)` we already pay per task.
Expand Down Expand Up @@ -113,10 +131,16 @@ __aicore__ __attribute__((always_inline)) static inline void l2_swimlane_aicore_
__gm__ L2SwimlaneAicoreTaskRecord *record = &local->cached_buf->records[slot];
record->start_time = start_time;
record->end_time = end_time;
record->task_id = task_id;
record->task_token_raw = task_token_raw;
record->reg_task_id = reg_task_id;
local->slot_within_buf = slot + 1;

// Flush record to GM so host can read it after the buffer is enqueued.
// No buffer-full signal is needed: AICPU drives rotation from its own
// per-core dispatch count (it knows how many DATA_MAIN_BASE writes it has
// sent to this core, and rotates before crossing a BUFFER_SIZE boundary).
// The completion-before-dispatch invariant guarantees this dcci has hit
// GM before AICPU enqueues the buffer.
dcci(record, SINGLE_CACHE_LINE, CACHELINE_OUT);
dsb((mem_dsb_t)0);
}
Expand Down
7 changes: 6 additions & 1 deletion src/a2a3/platform/include/aicpu/dep_gen_collector_aicpu.h
Original file line number Diff line number Diff line change
Expand Up @@ -97,10 +97,15 @@ void dep_gen_aicpu_init();
* @param explicit_dep_count Number of explicit_deps — no static cap; truncated only when the
* chain would not fit in a single DepGenBuffer
* @param explicit_deps_raw Per-dep PTO2TaskId::raw (length = explicit_dep_count)
* @param kernel_ids Per-subslot kernel id triple {AIC, AIV0, AIV1};
* inactive subslots use INVALID_KERNEL_ID (-1).
* Captured here so the host swimlane post-processor
* can resolve (task_id → kernel) without the AICore
* hot path writing identity fields itself.
*/
void dep_gen_aicpu_record_submit(
uint64_t task_id_raw, bool in_manual_scope, int tensor_count, const void *const *tensor_ptrs,
const uint8_t *arg_types, int explicit_dep_count, const uint64_t *explicit_deps_raw
const uint8_t *arg_types, int explicit_dep_count, const uint64_t *explicit_deps_raw, const int32_t kernel_ids[3]
);

/**
Expand Down
67 changes: 35 additions & 32 deletions src/a2a3/platform/include/aicpu/l2_swimlane_collector_aicpu.h
Original file line number Diff line number Diff line change
Expand Up @@ -77,55 +77,58 @@ L2SwimlaneLevel get_l2_swimlane_level();
void l2_swimlane_aicpu_init(int worker_count);

/**
* Rotate the AICore buffer for a given core, if needed.
* Pre-dispatch hook for AICore buffer rotation and per-pool stats.
*
* Called from the dispatch path (scheduler_dispatch in tensormap_and_ringbuffer,
* aicpu_executor in host_build_graph) immediately before write_reg(DATA_MAIN_BASE)
* for each task. Increments the per-core dispatch counter and, when it crosses
* a PLATFORM_AICORE_BUFFER_SIZE boundary, enqueues the current AICore buffer
* to the ready queue (kind=2) and pops the next one from free_queue.
*
* Race safety: rotation happens BEFORE the dispatch register write, so by the
* runtime's completion-before-dispatch invariant all prior tasks have FIN'd
* (and AICore has finished writing their records into the old buffer) before
* the old buffer enters the ready queue.
*
* Called regardless of l2_swimlane_level — internally gates on AICORE_TIMING.
*
* @param core_id Core index
* @param thread_idx Owning AICPU thread (target ready-queue)
* aicpu_executor in host_build_graph) immediately BEFORE `write_reg(DATA_MAIN_BASE)`
* for each AICore task. Two responsibilities:
*
* 1. Maintain the per-core AICPU-side dispatch count.
* 2. Rotate the AICore buffer when the count is about to cross a
* PLATFORM_AICORE_BUFFER_SIZE boundary — enqueue the just-filled buffer
* to the ready queue and pop the next one from free_queue.
* 3. Bump the AICore pool's `total_record_count` so host reconcile
* (total == collected + dropped) stays accurate at all levels —
* including AICORE_TIMING (level=1), where `complete_task` is bypassed.
*
* Race safety: rotation runs BEFORE the dispatch register write. The runtime's
* completion-before-dispatch invariant (AICore per core is single-threaded
* and AICPU does not dispatch task K+1 until K FIN'd) guarantees AICore has
* FIN'd — and dcci'd out — every record in the old buffer by then. No
* AICore-side signal is needed; AICPU has full dispatch visibility itself.
*
* No-op if l2_swimlane is disabled or `core_id` is out of range.
*
* @param core_id Core index this dispatch targets
* @param thread_idx Owning AICPU thread (target ready-queue for rotation)
*/
void l2_swimlane_aicpu_maybe_rotate_aicore(int core_id, int thread_idx);
void l2_swimlane_aicpu_on_aicore_dispatch(int core_id, int thread_idx);

/**
* Complete a L2SwimlaneAicpuTaskRecord with AICPU-side metadata after AICore task completion
* Commit an AICPU-side timing record for one completed task.
*
* AICore-as-producer: AICore writes start/end/task_id directly into the
* per-core L2SwimlaneAicoreTaskBuffer at `records[reg_task_id % SIZE]`. AICPU does
* NOT read that buffer on the hot path — it only writes AICPU-owned fields
* (task_id, reg_task_id, func_id, core_type, dispatch_time, finish_time)
* here, leaving start/end as zero. The host post-processor joins the AICore
* stream into the L2SwimlaneAicpuTaskRecord stream by `reg_task_id` at flush time.
* AICore-as-producer: identity (task_token_raw) and AICore-side timing
* (start/end) live in the per-core L2SwimlaneAicoreTaskRecord stream;
* core_type is published once by the host into the collector
* (L2SwimlaneCollector::set_core_types); func_id is resolved post-process
* from deps.json. This function therefore only needs to record the two
* AICPU-only timestamps plus the host-side join key.
*
* Per-core counter accounting:
* total_record_count++ — every commit attempt (success or failure)
* dropped_record_count++ — capacity-driven drop (no free buffer / queue
* full); actionable via
* dropped_record_count++ — capacity-driven drop (no free buffer /
* queue full); actionable via
* PLATFORM_PROF_BUFFERS_PER_CORE
*
* @param core_id Core index — used to resolve buffer state and update counters
* @param thread_idx Owning AICPU thread (used when rotating records buffer)
* @param expected_reg_task_id Register dispatch token (low 32 bits) — written
* into L2SwimlaneAicpuTaskRecord.reg_task_id as the join key
* @param task_id Task identifier to write (PTO2 encoding or plain id)
* @param func_id Kernel function identifier
* @param core_type Core type (AIC/AIV)
* @param reg_task_id Per-core dispatch token (low 32 bits) — host join
* key against the AICore record stream
* @param dispatch_time AICPU timestamp when task was dispatched
* @param finish_time AICPU timestamp when task completion was observed
*/
int l2_swimlane_aicpu_complete_task(
int core_id, int thread_idx, uint32_t expected_reg_task_id, uint64_t task_id, uint32_t func_id, CoreType core_type,
uint64_t dispatch_time, uint64_t finish_time
int core_id, int thread_idx, uint32_t reg_task_id, uint64_t dispatch_time, uint64_t finish_time
);

/**
Expand Down
Loading