Conversation
`src/a2a3/platform/onboard/aicpu/kernel.cpp` propagates dump_tensor /
l2_swimlane / pmu from `k_args->enable_profiling_flag` to AICPU globals
on every kernel entry, but the two equivalent calls for dep_gen were
missing — so `is_dep_gen_enabled()` stayed false on hardware, the
orchestrator's `dep_gen_aicpu_record_submit` early-returned for every
submit, the ring stayed empty, and the host replay wrote
`{"version":1,"edges":[]}` for every run. The fanout ⊆ deps gate
introduced in PR hw-native-sys#737 then ran against an empty deps set — vacuously
passing only when fanout was also empty, and silent-FAILing otherwise.
Sim was unaffected: its host device_runner already dlsym's
`set_platform_dep_gen_base` and `set_dep_gen_enabled` into the AICPU
.so. Onboard had no equivalent wiring.
Changes
- `aicpu/kernel.cpp`: include `dep_gen_collector_aicpu.h` and add the
two missing setters next to the existing dump_tensor / l2_swimlane /
pmu propagation.
- `dep_gen_capture/test_dep_gen_capture.py`: add `"a2a3"` to the case's
platforms list (was sim-only — the reason this never surfaced in CI)
and override `test_run` so the pytest path actually invokes
`_post_validate` when `--enable-dep-gen` is passed. Without that
override, pytest collected the test and ran orchestration but never
asserted anything about deps.json, so a future regression of the same
shape would still slip through.
- `.github/workflows/ci.yml`: add a dedicated dep_gen smoke step to the
st-onboard-a2a3 job that runs the capture test with
`--enable-dep-gen --enable-l2-swimlane`, so the fanout ⊆ deps gate
exercises hardware.
Verified on a2a3 onboard: `deps.json` now contains the six edges
documented in `example_orchestration.cpp` (t0->t1, t0->t2, t1->t3,
t2->t3, t0->t4, t3->t4), the smoke test passes with the gate flags,
and the broader scene-test sweep is unchanged.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
There was a problem hiding this comment.
Code Review
This pull request enables the dep_gen profiling feature in the AICPU kernel by integrating the collector header and initializing the base address and enablement flags from kernel arguments. Additionally, the test suite is updated to support the a2a3 platform and includes an override for test_run to ensure that dep_gen output is validated when enabled. Feedback suggests verifying that profiling is correctly handled when the number of execution rounds is greater than one, as per profiling guidelines.
Addresses gemini-code-assist on PR hw-native-sys#742: the override in TestDepGenCapture.test_run re-read --enable-dep-gen directly from request.config and ignored the framework's --rounds > 1 disable rule, so `pytest --enable-dep-gen --rounds 2` would run super().test_run() without dep_gen (no deps.json produced), then call _post_validate which asserts deps.json exists — turning a silent regression into a confusing AssertionError. Add SceneTestCase._effective_enable_dep_gen(request, *, warn=False) as the single source of truth for the resolved flag (CLI value AND the rounds-guard). The framework's test_run consumes it with warn=True (it owns the user-facing "disabled because rounds > 1" message); the subclass override consumes it with warn=False (super() already warned). The other three profiling flags (l2_swimlane / dump_tensor / pmu) keep their inline rounds-guard — no subclass needs them yet, so factoring those out would be premature. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
ChaoWao
approved these changes
May 12, 2026
ChaoWao
added a commit
to ChaoWao/simpler-fork
that referenced
this pull request
May 12, 2026
Follow-up to hw-native-sys#737 addressing post-merge review findings: 1. **Dedup replay fanin per-successor** ``dep_gen_replay_emit_deps_json`` now mirrors the runtime's ``PTO2FaninBuilder::append_fanin_or_fail`` semantics: an ``std::unordered_set<uint64_t>`` tracks predecessor task ids seen so far for the current successor, and both STEP 1 (``explicit_deps``) and STEP 3 (creator retention + tensormap lookup) push through a single ``emit_unique`` lambda. Previously an ``explicit_dep`` that the tensormap also surfaced (via ``owner_task_id`` or an overlap hit) emitted two edges, which double-counted ``deps.json`` and made ``swimlane_converter.py`` draw duplicate flow events. 2. **Document the OUTPUT-slot safety contract at the capture site** ``dep_gen_replay.cpp`` sets ``tref_buf[i].ptr`` for every captured tensor slot including OUTPUT — the on-disk blob for OUTPUT is zeroed by the AICPU writer. Added an inline comment pointing at ``pto_dep_compute.h``'s per-tag dispatch (which is what makes the never-dereferenced-on-OUTPUT contract hold) so the next reader of the arg_types width-fix area doesn't have to re-derive it. 3. **``make_deps_json_path`` helper** Both onboard + sim device_runner used to build ``deps.json`` with ``output_prefix_ + "/deps.json"`` inline — out of step with the ``make_<feature>_path()`` convention shared by PMU and (previously) submit_trace. Added ``make_deps_json_path`` in ``dep_gen_collector.h``; both call sites now go through it, and the helper also handles ``create_directories`` so the path is safe even when the output dir hasn't been touched by anything else yet. 4. **``_task_id(ring, local)`` helper in the validation test** The 6-edge expected-set in ``test_dep_gen_capture.py`` was open- coded ``1 << 32`` arithmetic at every call. One helper, layout stated once. 5. **``docs/dep_gen.md``** First user-facing doc for the feature. Covers motivation (links hw-native-sys#599), enable flags, ``deps.json`` format, ``deps_to_graph.py`` usage + node shape/color legend (AIC cube box, AIV vector ellipse, mix diamond, alloc dashed note), the ``fanout ⊆ deps`` validation gate, and the architecture touchpoints. 6. **CI smoke step for dep_gen** (``.github/workflows/ci.yml``) The default ``pytest tests/st`` invocation does not pass any DFX flag, so the dep_gen capture path never executed in CI. Added a second pytest step in ``st-sim-a2a3`` that re-runs only ``test_dep_gen_capture`` with ``--enable-dep-gen --enable-l2-swimlane``, which forces the full capture → replay → deps.json → fanout ⊆ deps gate path to execute. Same pattern is the model for future PMU / tensor_dump / swimlane smoke steps once those grow dedicated tests. 7. **UT roundtrip for ``enable_dep_gen``** (``tests/ut/py/test_chip_worker.py``) Adds the missing nanobind setter / getter / repr roundtrip for the ``enable_dep_gen`` ``CallConfig`` field, alongside the existing round-trips for ``enable_l2_swimlane`` / ``enable_dump_tensor`` / ``enable_pmu``. Catches binding-ABI regressions at the UT layer (e.g. the ``bool`` vs ``int32`` Python-wrapper bug that broke 9 st jobs earlier in the PR series — 10 s ut would have caught it). 8. **Pytest-friendly test hook via ``test_run`` override** instead of a framework-level ``_post_validate`` callback. ``test_dep_gen_capture`` overrides the inherited ``SceneTestCase.test_run`` to call ``super()`` then walk its cases and assert the dep_gen artifacts. Keeps the framework unchanged (no implicit ``hasattr`` check on every SceneTestCase), localises the test-specific behavior, and the ``--enable-dep-gen`` flag gate prevents the assertion from firing in default invocations. 9. **Unified DFX smoke directory + symmetric a2a3sim/a2a3 coverage** (``tests/st/a2a3/tensormap_and_ringbuffer/dfx/{dep_gen,l2_swimlane,pmu,tensor_dump}/``) Moved ``dep_gen_capture/`` → ``dfx/dep_gen/`` and added three sibling per-feature subdirs matching the repo's "one feature, one directory" convention (cf. ``spmd_sync_start/``, ``paged_attention_unroll/`` etc.). Each subdir holds one ``test_<feature>.py``: - ``dep_gen/test_dep_gen.py`` — asserts ``deps.json`` edges + fanout ⊆ deps gate - ``l2_swimlane/test_l2_swimlane.py`` — asserts ``l2_perf_records.json`` shape - ``pmu/test_pmu.py`` — asserts ``pmu.csv`` header + row count - ``tensor_dump/test_tensor_dump.py`` — asserts ``tensor_dump/`` manifest + bin Each test uses ``vector_example`` as a fixed 5-task workload, opens only its own flag, asserts only its own artifact. ``platforms`` field is ``["a2a3sim", "a2a3"]`` on all four — DFX features matter most on real hardware (the race window in dep_gen never fires on sim's deterministic timing, and post-hw-native-sys#742 we have empirical evidence the gate catches missed edges only on hardware), so both st-sim-a2a3 and st-onboard-a2a3 get four matching smoke steps. Each smoke is its own pytest invocation because L2PerfCollector::initialize trips its dup-init guard (``shm_host_ != nullptr``) when the L2 worker pool reuses the DeviceRunner across tests in the same session — fresh process per smoke side-steps that state leak; making the runtime init re-init-safe is a separate concern. 10. **Sibling helper consistency** (``make_pmu_csv_path``) Converted ``make_pmu_csv_path`` in ``pmu_collector.h`` to use ``std::filesystem::path`` operator/ instead of bare string concat, matching the new ``make_deps_json_path`` convention. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
ChaoWao
added a commit
to ChaoWao/simpler-fork
that referenced
this pull request
May 12, 2026
Follow-up to hw-native-sys#737 addressing post-merge review findings: 1. **Dedup replay fanin per-successor** ``dep_gen_replay_emit_deps_json`` now mirrors the runtime's ``PTO2FaninBuilder::append_fanin_or_fail`` semantics: an ``std::unordered_set<uint64_t>`` tracks predecessor task ids seen so far for the current successor, and both STEP 1 (``explicit_deps``) and STEP 3 (creator retention + tensormap lookup) push through a single ``emit_unique`` lambda. Previously an ``explicit_dep`` that the tensormap also surfaced (via ``owner_task_id`` or an overlap hit) emitted two edges, which double-counted ``deps.json`` and made ``swimlane_converter.py`` draw duplicate flow events. 2. **Document the OUTPUT-slot safety contract at the capture site** ``dep_gen_replay.cpp`` sets ``tref_buf[i].ptr`` for every captured tensor slot including OUTPUT — the on-disk blob for OUTPUT is zeroed by the AICPU writer. Added an inline comment pointing at ``pto_dep_compute.h``'s per-tag dispatch (which is what makes the never-dereferenced-on-OUTPUT contract hold) so the next reader of the arg_types width-fix area doesn't have to re-derive it. 3. **``make_deps_json_path`` helper** Both onboard + sim device_runner used to build ``deps.json`` with ``output_prefix_ + "/deps.json"`` inline — out of step with the ``make_<feature>_path()`` convention shared by PMU and (previously) submit_trace. Added ``make_deps_json_path`` in ``dep_gen_collector.h``; both call sites now go through it, and the helper also handles ``create_directories`` so the path is safe even when the output dir hasn't been touched by anything else yet. 4. **``_task_id(ring, local)`` helper in the validation test** The 6-edge expected-set in ``test_dep_gen_capture.py`` was open- coded ``1 << 32`` arithmetic at every call. One helper, layout stated once. 5. **``docs/dep_gen.md``** First user-facing doc for the feature. Covers motivation (links hw-native-sys#599), enable flags, ``deps.json`` format, ``deps_to_graph.py`` usage + node shape/color legend (AIC cube box, AIV vector ellipse, mix diamond, alloc dashed note), the ``fanout ⊆ deps`` validation gate, and the architecture touchpoints. 6. **CI smoke step for dep_gen** (``.github/workflows/ci.yml``) The default ``pytest tests/st`` invocation does not pass any DFX flag, so the dep_gen capture path never executed in CI. Added a second pytest step in ``st-sim-a2a3`` that re-runs only ``test_dep_gen_capture`` with ``--enable-dep-gen --enable-l2-swimlane``, which forces the full capture → replay → deps.json → fanout ⊆ deps gate path to execute. Same pattern is the model for future PMU / tensor_dump / swimlane smoke steps once those grow dedicated tests. 7. **UT roundtrip for ``enable_dep_gen``** (``tests/ut/py/test_chip_worker.py``) Adds the missing nanobind setter / getter / repr roundtrip for the ``enable_dep_gen`` ``CallConfig`` field, alongside the existing round-trips for ``enable_l2_swimlane`` / ``enable_dump_tensor`` / ``enable_pmu``. Catches binding-ABI regressions at the UT layer (e.g. the ``bool`` vs ``int32`` Python-wrapper bug that broke 9 st jobs earlier in the PR series — 10 s ut would have caught it). 8. **Pytest-friendly test hook via ``test_run`` override** instead of a framework-level ``_post_validate`` callback. ``test_dep_gen_capture`` overrides the inherited ``SceneTestCase.test_run`` to call ``super()`` then walk its cases and assert the dep_gen artifacts. Keeps the framework unchanged (no implicit ``hasattr`` check on every SceneTestCase), localises the test-specific behavior, and the ``--enable-dep-gen`` flag gate prevents the assertion from firing in default invocations. 9. **Unified DFX smoke directory + symmetric a2a3sim/a2a3 coverage** (``tests/st/a2a3/tensormap_and_ringbuffer/dfx/{dep_gen,l2_swimlane,pmu,tensor_dump}/``) Moved ``dep_gen_capture/`` → ``dfx/dep_gen/`` and added three sibling per-feature subdirs matching the repo's "one feature, one directory" convention (cf. ``spmd_sync_start/``, ``paged_attention_unroll/`` etc.). Each subdir holds one ``test_<feature>.py``: - ``dep_gen/test_dep_gen.py`` — asserts ``deps.json`` edges + fanout ⊆ deps gate - ``l2_swimlane/test_l2_swimlane.py`` — asserts ``l2_perf_records.json`` shape - ``pmu/test_pmu.py`` — asserts ``pmu.csv`` header + row count - ``tensor_dump/test_tensor_dump.py`` — asserts ``tensor_dump/`` manifest + bin Each test uses ``vector_example`` as a fixed 5-task workload, opens only its own flag, asserts only its own artifact. ``platforms`` field is ``["a2a3sim", "a2a3"]`` on all four — DFX features matter most on real hardware (the race window in dep_gen never fires on sim's deterministic timing, and post-hw-native-sys#742 we have empirical evidence the gate catches missed edges only on hardware), so both st-sim-a2a3 and st-onboard-a2a3 get four matching smoke steps. Each smoke is its own pytest invocation because L2PerfCollector::initialize trips its dup-init guard (``shm_host_ != nullptr``) when the L2 worker pool reuses the DeviceRunner across tests in the same session — fresh process per smoke side-steps that state leak; making the runtime init re-init-safe is a separate concern. 10. **Sibling helper consistency** (``make_pmu_csv_path``) Converted ``make_pmu_csv_path`` in ``pmu_collector.h`` to use ``std::filesystem::path`` operator/ instead of bare string concat, matching the new ``make_deps_json_path`` convention. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
ChaoWao
added a commit
that referenced
this pull request
May 12, 2026
Follow-up to #737 addressing post-merge review findings: 1. **Dedup replay fanin per-successor** ``dep_gen_replay_emit_deps_json`` now mirrors the runtime's ``PTO2FaninBuilder::append_fanin_or_fail`` semantics: an ``std::unordered_set<uint64_t>`` tracks predecessor task ids seen so far for the current successor, and both STEP 1 (``explicit_deps``) and STEP 3 (creator retention + tensormap lookup) push through a single ``emit_unique`` lambda. Previously an ``explicit_dep`` that the tensormap also surfaced (via ``owner_task_id`` or an overlap hit) emitted two edges, which double-counted ``deps.json`` and made ``swimlane_converter.py`` draw duplicate flow events. 2. **Document the OUTPUT-slot safety contract at the capture site** ``dep_gen_replay.cpp`` sets ``tref_buf[i].ptr`` for every captured tensor slot including OUTPUT — the on-disk blob for OUTPUT is zeroed by the AICPU writer. Added an inline comment pointing at ``pto_dep_compute.h``'s per-tag dispatch (which is what makes the never-dereferenced-on-OUTPUT contract hold) so the next reader of the arg_types width-fix area doesn't have to re-derive it. 3. **``make_deps_json_path`` helper** Both onboard + sim device_runner used to build ``deps.json`` with ``output_prefix_ + "/deps.json"`` inline — out of step with the ``make_<feature>_path()`` convention shared by PMU and (previously) submit_trace. Added ``make_deps_json_path`` in ``dep_gen_collector.h``; both call sites now go through it, and the helper also handles ``create_directories`` so the path is safe even when the output dir hasn't been touched by anything else yet. 4. **``_task_id(ring, local)`` helper in the validation test** The 6-edge expected-set in ``test_dep_gen_capture.py`` was open- coded ``1 << 32`` arithmetic at every call. One helper, layout stated once. 5. **``docs/dep_gen.md``** First user-facing doc for the feature. Covers motivation (links #599), enable flags, ``deps.json`` format, ``deps_to_graph.py`` usage + node shape/color legend (AIC cube box, AIV vector ellipse, mix diamond, alloc dashed note), the ``fanout ⊆ deps`` validation gate, and the architecture touchpoints. 6. **CI smoke step for dep_gen** (``.github/workflows/ci.yml``) The default ``pytest tests/st`` invocation does not pass any DFX flag, so the dep_gen capture path never executed in CI. Added a second pytest step in ``st-sim-a2a3`` that re-runs only ``test_dep_gen_capture`` with ``--enable-dep-gen --enable-l2-swimlane``, which forces the full capture → replay → deps.json → fanout ⊆ deps gate path to execute. Same pattern is the model for future PMU / tensor_dump / swimlane smoke steps once those grow dedicated tests. 7. **UT roundtrip for ``enable_dep_gen``** (``tests/ut/py/test_chip_worker.py``) Adds the missing nanobind setter / getter / repr roundtrip for the ``enable_dep_gen`` ``CallConfig`` field, alongside the existing round-trips for ``enable_l2_swimlane`` / ``enable_dump_tensor`` / ``enable_pmu``. Catches binding-ABI regressions at the UT layer (e.g. the ``bool`` vs ``int32`` Python-wrapper bug that broke 9 st jobs earlier in the PR series — 10 s ut would have caught it). 8. **Pytest-friendly test hook via ``test_run`` override** instead of a framework-level ``_post_validate`` callback. ``test_dep_gen_capture`` overrides the inherited ``SceneTestCase.test_run`` to call ``super()`` then walk its cases and assert the dep_gen artifacts. Keeps the framework unchanged (no implicit ``hasattr`` check on every SceneTestCase), localises the test-specific behavior, and the ``--enable-dep-gen`` flag gate prevents the assertion from firing in default invocations. 9. **Unified DFX smoke directory + symmetric a2a3sim/a2a3 coverage** (``tests/st/a2a3/tensormap_and_ringbuffer/dfx/{dep_gen,l2_swimlane,pmu,tensor_dump}/``) Moved ``dep_gen_capture/`` → ``dfx/dep_gen/`` and added three sibling per-feature subdirs matching the repo's "one feature, one directory" convention (cf. ``spmd_sync_start/``, ``paged_attention_unroll/`` etc.). Each subdir holds one ``test_<feature>.py``: - ``dep_gen/test_dep_gen.py`` — asserts ``deps.json`` edges + fanout ⊆ deps gate - ``l2_swimlane/test_l2_swimlane.py`` — asserts ``l2_perf_records.json`` shape - ``pmu/test_pmu.py`` — asserts ``pmu.csv`` header + row count - ``tensor_dump/test_tensor_dump.py`` — asserts ``tensor_dump/`` manifest + bin Each test uses ``vector_example`` as a fixed 5-task workload, opens only its own flag, asserts only its own artifact. ``platforms`` field is ``["a2a3sim", "a2a3"]`` on all four — DFX features matter most on real hardware (the race window in dep_gen never fires on sim's deterministic timing, and post-#742 we have empirical evidence the gate catches missed edges only on hardware), so both st-sim-a2a3 and st-onboard-a2a3 get four matching smoke steps. Each smoke is its own pytest invocation because L2PerfCollector::initialize trips its dup-init guard (``shm_host_ != nullptr``) when the L2 worker pool reuses the DeviceRunner across tests in the same session — fresh process per smoke side-steps that state leak; making the runtime init re-init-safe is a separate concern. 10. **Sibling helper consistency** (``make_pmu_csv_path``) Converted ``make_pmu_csv_path`` in ``pmu_collector.h`` to use ``std::filesystem::path`` operator/ instead of bare string concat, matching the new ``make_deps_json_path`` convention.
indigo1973
added a commit
to indigo1973/simpler
that referenced
this pull request
Jun 1, 2026
Port the dep_gen (SubmitTrace) feature from a2a3 to a5 so the
tensormap_and_ringbuffer runtime on a5 can produce deps.json and feed
flow events into swimlane_converter.py. Without this, --enable-dep-gen
was a no-op on a5 and merged_swimlane_*.json had no dependency arrows.
Reused from a2a3 (code-identical; the .cpp files are byte-for-byte
copies, the headers differ only in include-guard names and
L2Swimlane->L2Perf / a2a3->a5 comment refs):
- Shared-memory ABI: common/dep_gen.h (DepGenRecord 2624 B, overflow
chain, SPSC free_queue, per-thread ready_queue) — layout byte-identical
- AICPU writer: aicpu/dep_gen_collector_aicpu.{h,cpp} (.cpp identical)
- Runtime replay: runtime/tensormap_and_ringbuffer/host/dep_gen_replay
(.cpp identical)
- Orchestrator capture point + aicpu_executor lifecycle hooks
- 5 platform_config constants + PROFILING_FLAG_DEP_GEN bit
Specialized for a5 (no SVM, see profiling_common diff vs a2a3):
- dep_gen_collector.cpp uses alloc_single_buffer (malloc shadow +
profiling_copy_to_device) instead of identity-mapping when
register_cb is null — matches a5's PMU/L2Perf/Dump collectors.
- Two-phase set_memory_context: callbacks first, then shm pointers
once the region is committed, so start(tf) gates correctly.
- reconcile_counters explicitly copy_from_device's the BufferState +
current_buf before reading (mgmt thread is stopped by then).
- finalize lets BufferPoolManager::clear_mappings() be the single
source of truth for host-shadow lifetime — no per-collector dedup.
Sim path: dlsym set_platform_dep_gen_base / set_dep_gen_enabled out of
the AICPU .so and forward kernel_args.dep_gen_data_base + enable flag
at boot, mirroring the existing pmu / dump / l2_perf setters.
Onboard kernel.cpp adds two lines to forward dep_gen_data_base +
PROFILING_FLAG_DEP_GEN into the AICPU writer's globals, mirroring the
existing PMU / L2 / Dump setters.
c_api: a5's DeviceRunner now overrides set_dep_gen_enabled() (was the
base no-op), so run_prepared's already-present enable_dep_gen call is
honored on a5 instead of being silently dropped. The shared
c_api_shared.cpp / device_runner_base.h changes are comment-only —
refreshing stale "a2a3-only" notes to reflect both arches now wire it.
Tests:
- tests/st/a5/.../dfx/dep_gen/test_dep_gen.py: 6-edge validation
against vector_example orchestration (byte-identical to a2a3 — same
expected edge set).
- tests/st/a5/.../dfx/dep_gen/test_dep_gen_chain.py
(+ kernels/orchestration/chain_barrier_orch.cpp): overflow chain
regression for >64 explicit deps.
- tests/st/a2a3/.../dfx/dep_gen/test_dep_gen.py: harden the existing
a2a3 gate — a missing deps.json under --enable-dep-gen now fails
loudly (assert) instead of silently skipping, so a
capture->reconcile->replay regression can no longer pass green
(hw-native-sys#742).
Docs:
- docs/dfx/dep_gen.md: §8 Architecture Touchpoints now lists both
platforms; "Currently a2a3 only" line removed.
- src/a5/runtime/.../docs/profiling_levels.md: Code Locations point
at src/a5/ (was stale src/a2a3/ refs from PR hw-native-sys#777 cleanup) and
add a dep_gen entry.
indigo1973
added a commit
to indigo1973/simpler
that referenced
this pull request
Jun 1, 2026
Port the dep_gen (SubmitTrace) feature from a2a3 to a5 so the
tensormap_and_ringbuffer runtime on a5 can produce deps.json and feed
flow events into swimlane_converter.py. Without this, --enable-dep-gen
was a no-op on a5 and merged_swimlane_*.json had no dependency arrows.
Reused from a2a3 (code-identical; the .cpp files are byte-for-byte
copies, the headers differ only in include-guard names and
L2Swimlane->L2Perf / a2a3->a5 comment refs):
- Shared-memory ABI: common/dep_gen.h (DepGenRecord 2624 B, overflow
chain, SPSC free_queue, per-thread ready_queue) — layout byte-identical
- AICPU writer: aicpu/dep_gen_collector_aicpu.{h,cpp} (.cpp identical)
- Runtime replay: runtime/tensormap_and_ringbuffer/host/dep_gen_replay
(.cpp identical)
- Orchestrator capture point + aicpu_executor lifecycle hooks
- 5 platform_config constants + PROFILING_FLAG_DEP_GEN bit
Specialized for a5 (no SVM, see profiling_common diff vs a2a3):
- dep_gen_collector.cpp uses alloc_single_buffer (malloc shadow +
profiling_copy_to_device) instead of identity-mapping when
register_cb is null — matches a5's PMU/L2Perf/Dump collectors.
- Two-phase set_memory_context: callbacks first, then shm pointers
once the region is committed, so start(tf) gates correctly.
- reconcile_counters explicitly copy_from_device's the BufferState +
current_buf before reading (mgmt thread is stopped by then).
- finalize lets BufferPoolManager::clear_mappings() be the single
source of truth for host-shadow lifetime — no per-collector dedup.
Sim path: dlsym set_platform_dep_gen_base / set_dep_gen_enabled out of
the AICPU .so and forward kernel_args.dep_gen_data_base + enable flag
at boot, mirroring the existing pmu / dump / l2_perf setters.
Onboard kernel.cpp adds two lines to forward dep_gen_data_base +
PROFILING_FLAG_DEP_GEN into the AICPU writer's globals, mirroring the
existing PMU / L2 / Dump setters.
c_api: a5's DeviceRunner now overrides set_dep_gen_enabled() (was the
base no-op), so run_prepared's already-present enable_dep_gen call is
honored on a5 instead of being silently dropped. The shared
c_api_shared.cpp / device_runner_base.h changes are comment-only —
refreshing stale "a2a3-only" notes to reflect both arches now wire it.
Tests:
- tests/st/a5/.../dfx/dep_gen/test_dep_gen.py: 6-edge validation
against vector_example orchestration (byte-identical to a2a3 — same
expected edge set).
- tests/st/a5/.../dfx/dep_gen/test_dep_gen_chain.py
(+ kernels/orchestration/chain_barrier_orch.cpp): overflow chain
regression for >64 explicit deps.
- tests/st/a2a3/.../dfx/dep_gen/test_dep_gen.py: harden the existing
a2a3 gate — a missing deps.json under --enable-dep-gen now fails
loudly (assert) instead of silently skipping, so a
capture->reconcile->replay regression can no longer pass green
(hw-native-sys#742).
Docs:
- docs/dfx/dep_gen.md: §8 Architecture Touchpoints now lists both
platforms; "Currently a2a3 only" line removed.
- src/a5/runtime/.../docs/profiling_levels.md: Code Locations point
at src/a5/ (was stale src/a2a3/ refs from PR hw-native-sys#777 cleanup) and
add a dep_gen entry.
indigo1973
added a commit
to indigo1973/simpler
that referenced
this pull request
Jun 1, 2026
Port the dep_gen (SubmitTrace) feature from a2a3 to a5 so the
tensormap_and_ringbuffer runtime on a5 can produce deps.json and feed
flow events into swimlane_converter.py. Without this, --enable-dep-gen
was a no-op on a5 and merged_swimlane_*.json had no dependency arrows.
Reused from a2a3 (code-identical; the .cpp files are byte-for-byte
copies, the headers differ only in include-guard names and
L2Swimlane->L2Perf / a2a3->a5 comment refs):
- Shared-memory ABI: common/dep_gen.h (DepGenRecord 2624 B, overflow
chain, SPSC free_queue, per-thread ready_queue) — layout byte-identical
- AICPU writer: aicpu/dep_gen_collector_aicpu.{h,cpp} (.cpp identical)
- Runtime replay: runtime/tensormap_and_ringbuffer/host/dep_gen_replay
(.cpp identical)
- Orchestrator capture point + aicpu_executor lifecycle hooks
- 5 platform_config constants + PROFILING_FLAG_DEP_GEN bit
Specialized for a5 (no SVM, see profiling_common diff vs a2a3):
- dep_gen_collector.cpp uses alloc_single_buffer (malloc shadow +
profiling_copy_to_device) instead of identity-mapping when
register_cb is null — matches a5's PMU/L2Perf/Dump collectors.
- Two-phase set_memory_context: callbacks first, then shm pointers
once the region is committed, so start(tf) gates correctly.
- reconcile_counters explicitly copy_from_device's the BufferState +
current_buf before reading (mgmt thread is stopped by then).
- finalize lets BufferPoolManager::clear_mappings() be the single
source of truth for host-shadow lifetime — no per-collector dedup.
Sim path: dlsym set_platform_dep_gen_base / set_dep_gen_enabled out of
the AICPU .so and forward kernel_args.dep_gen_data_base + enable flag
at boot, mirroring the existing pmu / dump / l2_perf setters.
Onboard kernel.cpp adds two lines to forward dep_gen_data_base +
PROFILING_FLAG_DEP_GEN into the AICPU writer's globals, mirroring the
existing PMU / L2 / Dump setters.
c_api: a5's DeviceRunner now overrides set_dep_gen_enabled() (was the
base no-op), so run_prepared's already-present enable_dep_gen call is
honored on a5 instead of being silently dropped. The shared
c_api_shared.cpp / device_runner_base.h changes are comment-only —
refreshing stale "a2a3-only" notes to reflect both arches now wire it.
Tests:
- tests/st/a5/.../dfx/dep_gen/test_dep_gen.py: 6-edge validation
against vector_example orchestration (byte-identical to a2a3 — same
expected edge set).
- tests/st/a5/.../dfx/dep_gen/test_dep_gen_chain.py
(+ kernels/orchestration/chain_barrier_orch.cpp): overflow chain
regression for >64 explicit deps.
- tests/st/a2a3/.../dfx/dep_gen/test_dep_gen.py: harden the existing
a2a3 gate — a missing deps.json under --enable-dep-gen now fails
loudly (assert) instead of silently skipping, so a
capture->reconcile->replay regression can no longer pass green
([hw-native-sys#742](hw-native-sys#742)).
Docs:
- docs/dfx/dep_gen.md: §8 Architecture Touchpoints now lists both
platforms; "Currently a2a3 only" line removed.
- src/a5/runtime/.../docs/profiling_levels.md: Code Locations point
at src/a5/ (was stale src/a2a3/ refs from PR [hw-native-sys#777](hw-native-sys#777) cleanup) and
add a dep_gen entry.
indigo1973
added a commit
to indigo1973/simpler
that referenced
this pull request
Jun 1, 2026
Port the dep_gen (SubmitTrace) feature from a2a3 to a5 so the
tensormap_and_ringbuffer runtime on a5 can produce deps.json and feed
flow events into swimlane_converter.py. Without this, --enable-dep-gen
was a no-op on a5 and merged_swimlane_*.json had no dependency arrows.
Reused from a2a3 (code-identical; the .cpp files are byte-for-byte
copies, the headers differ only in include-guard names and
L2Swimlane->L2Perf / a2a3->a5 comment refs):
- Shared-memory ABI: common/dep_gen.h (DepGenRecord 2624 B, overflow
chain, SPSC free_queue, per-thread ready_queue) — layout byte-identical
- AICPU writer: aicpu/dep_gen_collector_aicpu.{h,cpp} (.cpp identical)
- Runtime replay: runtime/tensormap_and_ringbuffer/host/dep_gen_replay
(.cpp identical)
- Orchestrator capture point + aicpu_executor lifecycle hooks
- 5 platform_config constants + PROFILING_FLAG_DEP_GEN bit
Specialized for a5 (no SVM, see profiling_common diff vs a2a3):
- dep_gen_collector.cpp uses alloc_single_buffer (malloc shadow +
profiling_copy_to_device) instead of identity-mapping when
register_cb is null — matches a5's PMU/L2Perf/Dump collectors.
- Two-phase set_memory_context: callbacks first, then shm pointers
once the region is committed, so start(tf) gates correctly.
- reconcile_counters explicitly copy_from_device's the BufferState +
current_buf before reading (mgmt thread is stopped by then).
- finalize lets BufferPoolManager::clear_mappings() be the single
source of truth for host-shadow lifetime — no per-collector dedup.
Sim path: dlsym set_platform_dep_gen_base / set_dep_gen_enabled out of
the AICPU .so and forward kernel_args.dep_gen_data_base + enable flag
at boot, mirroring the existing pmu / dump / l2_perf setters.
Onboard kernel.cpp adds two lines to forward dep_gen_data_base +
PROFILING_FLAG_DEP_GEN into the AICPU writer's globals, mirroring the
existing PMU / L2 / Dump setters.
c_api: a5's DeviceRunner now overrides set_dep_gen_enabled() (was the
base no-op), so run_prepared's already-present enable_dep_gen call is
honored on a5 instead of being silently dropped. The shared
c_api_shared.cpp / device_runner_base.h changes are comment-only —
refreshing stale "a2a3-only" notes to reflect both arches now wire it.
Tests:
- tests/st/a5/.../dfx/dep_gen/test_dep_gen.py: 6-edge validation
against vector_example orchestration (byte-identical to a2a3 — same
expected edge set).
- tests/st/a5/.../dfx/dep_gen/test_dep_gen_chain.py
(+ kernels/orchestration/chain_barrier_orch.cpp): overflow chain
regression for >64 explicit deps.
- tests/st/a2a3/.../dfx/dep_gen/test_dep_gen.py: harden the existing
a2a3 gate — a missing deps.json under --enable-dep-gen now fails
loudly (assert) instead of silently skipping, so a
capture->reconcile->replay regression can no longer pass green
([[hw-native-sys#742](https://github.com/hw-native-sys/simpler/issues/742)](https://github.com/hw-native-sys/simpler/issues/742)).
Docs:
- docs/dfx/dep_gen.md: §8 Architecture Touchpoints now lists both
platforms; "Currently a2a3 only" line removed.
- src/a5/runtime/.../docs/profiling_levels.md: Code Locations point
ChaoZheng109
pushed a commit
that referenced
this pull request
Jun 2, 2026
Port the dep_gen (SubmitTrace) feature from a2a3 to a5 so the
tensormap_and_ringbuffer runtime on a5 can produce deps.json and feed
flow events into swimlane_converter.py. Without this, --enable-dep-gen
was a no-op on a5 and merged_swimlane_*.json had no dependency arrows.
Reused from a2a3 (code-identical; the .cpp files are byte-for-byte
copies, the headers differ only in include-guard names and
L2Swimlane->L2Perf / a2a3->a5 comment refs):
- Shared-memory ABI: common/dep_gen.h (DepGenRecord 2624 B, overflow
chain, SPSC free_queue, per-thread ready_queue) — layout byte-identical
- AICPU writer: aicpu/dep_gen_collector_aicpu.{h,cpp} (.cpp identical)
- Runtime replay: runtime/tensormap_and_ringbuffer/host/dep_gen_replay
(.cpp identical)
- Orchestrator capture point + aicpu_executor lifecycle hooks
- 5 platform_config constants + PROFILING_FLAG_DEP_GEN bit
Specialized for a5 (no SVM, see profiling_common diff vs a2a3):
- dep_gen_collector.cpp uses alloc_single_buffer (malloc shadow +
profiling_copy_to_device) instead of identity-mapping when
register_cb is null — matches a5's PMU/L2Perf/Dump collectors.
- Two-phase set_memory_context: callbacks first, then shm pointers
once the region is committed, so start(tf) gates correctly.
- reconcile_counters explicitly copy_from_device's the BufferState +
current_buf before reading (mgmt thread is stopped by then).
- finalize lets BufferPoolManager::clear_mappings() be the single
source of truth for host-shadow lifetime — no per-collector dedup.
Sim path: dlsym set_platform_dep_gen_base / set_dep_gen_enabled out of
the AICPU .so and forward kernel_args.dep_gen_data_base + enable flag
at boot, mirroring the existing pmu / dump / l2_perf setters.
Onboard kernel.cpp adds two lines to forward dep_gen_data_base +
PROFILING_FLAG_DEP_GEN into the AICPU writer's globals, mirroring the
existing PMU / L2 / Dump setters.
c_api: a5's DeviceRunner now overrides set_dep_gen_enabled() (was the
base no-op), so run_prepared's already-present enable_dep_gen call is
honored on a5 instead of being silently dropped. The shared
c_api_shared.cpp / device_runner_base.h changes are comment-only —
refreshing stale "a2a3-only" notes to reflect both arches now wire it.
Tests:
- tests/st/a5/.../dfx/dep_gen/test_dep_gen.py: 6-edge validation
against vector_example orchestration (byte-identical to a2a3 — same
expected edge set).
- tests/st/a5/.../dfx/dep_gen/test_dep_gen_chain.py
(+ kernels/orchestration/chain_barrier_orch.cpp): overflow chain
regression for >64 explicit deps.
- tests/st/a2a3/.../dfx/dep_gen/test_dep_gen.py: harden the existing
a2a3 gate — a missing deps.json under --enable-dep-gen now fails
loudly (assert) instead of silently skipping, so a
capture->reconcile->replay regression can no longer pass green
([[#742](https://github.com/hw-native-sys/simpler/issues/742)](https://github.com/hw-native-sys/simpler/issues/742)).
Docs:
- docs/dfx/dep_gen.md: §8 Architecture Touchpoints now lists both
platforms; "Currently a2a3 only" line removed.
- src/a5/runtime/.../docs/profiling_levels.md: Code Locations point
nalinaly
pushed a commit
to nalinaly/simpler
that referenced
this pull request
Jul 31, 2026
) `src/a2a3/platform/onboard/aicpu/kernel.cpp` propagates dump_tensor / l2_swimlane / pmu from `k_args->enable_profiling_flag` to AICPU globals on every kernel entry, but the two equivalent calls for dep_gen were missing — so `is_dep_gen_enabled()` stayed false on hardware, the orchestrator's `dep_gen_aicpu_record_submit` early-returned for every submit, the ring stayed empty, and the host replay wrote `{"version":1,"edges":[]}` for every run. The fanout ⊆ deps gate introduced in PR hw-native-sys#737 then ran against an empty deps set — vacuously passing only when fanout was also empty, and silent-FAILing otherwise. Sim was unaffected: its host device_runner already dlsym's `set_platform_dep_gen_base` and `set_dep_gen_enabled` into the AICPU .so. Onboard had no equivalent wiring. Changes - `aicpu/kernel.cpp`: include `dep_gen_collector_aicpu.h` and add the two missing setters next to the existing dump_tensor / l2_swimlane / pmu propagation. - `scene_test.py`: add `SceneTestCase._effective_enable_dep_gen(request, *, warn=False)` as the single source of truth for the CLI flag after the `--rounds > 1` disable rule. The framework's `test_run` consumes it with `warn=True` (it owns the user-facing "disabled because rounds > 1" message); subclass overrides consume it with `warn=False` so the gating logic can't drift and the warning fires once. - `dep_gen_capture/test_dep_gen_capture.py`: add `"a2a3"` to the case's platforms list (was sim-only — the reason this never surfaced in CI) and override `test_run` so the pytest path actually invokes `_post_validate` when `--enable-dep-gen` is passed, via the new helper. Without that override, pytest collected the test and ran orchestration but never asserted anything about deps.json, so a future regression of the same shape would still slip through. - `.github/workflows/ci.yml`: add a dedicated dep_gen smoke step to the st-onboard-a2a3 job that runs the capture test with `--enable-dep-gen --enable-l2-swimlane`, so the fanout ⊆ deps gate exercises hardware on every PR touching a2a3. Verified on a2a3 onboard: `deps.json` now contains the six edges documented in `example_orchestration.cpp` (t0->t1, t0->t2, t1->t3, t2->t3, t0->t4, t3->t4), the smoke test passes with the gate flags, and the broader scene-test sweep is unchanged. Co-authored-by: wcwxy <26245345+ChaoWao@users.noreply.github.com>
nalinaly
pushed a commit
to nalinaly/simpler
that referenced
this pull request
Jul 31, 2026
…ve-sys#740) Follow-up to hw-native-sys#737 addressing post-merge review findings: 1. **Dedup replay fanin per-successor** ``dep_gen_replay_emit_deps_json`` now mirrors the runtime's ``PTO2FaninBuilder::append_fanin_or_fail`` semantics: an ``std::unordered_set<uint64_t>`` tracks predecessor task ids seen so far for the current successor, and both STEP 1 (``explicit_deps``) and STEP 3 (creator retention + tensormap lookup) push through a single ``emit_unique`` lambda. Previously an ``explicit_dep`` that the tensormap also surfaced (via ``owner_task_id`` or an overlap hit) emitted two edges, which double-counted ``deps.json`` and made ``swimlane_converter.py`` draw duplicate flow events. 2. **Document the OUTPUT-slot safety contract at the capture site** ``dep_gen_replay.cpp`` sets ``tref_buf[i].ptr`` for every captured tensor slot including OUTPUT — the on-disk blob for OUTPUT is zeroed by the AICPU writer. Added an inline comment pointing at ``pto_dep_compute.h``'s per-tag dispatch (which is what makes the never-dereferenced-on-OUTPUT contract hold) so the next reader of the arg_types width-fix area doesn't have to re-derive it. 3. **``make_deps_json_path`` helper** Both onboard + sim device_runner used to build ``deps.json`` with ``output_prefix_ + "/deps.json"`` inline — out of step with the ``make_<feature>_path()`` convention shared by PMU and (previously) submit_trace. Added ``make_deps_json_path`` in ``dep_gen_collector.h``; both call sites now go through it, and the helper also handles ``create_directories`` so the path is safe even when the output dir hasn't been touched by anything else yet. 4. **``_task_id(ring, local)`` helper in the validation test** The 6-edge expected-set in ``test_dep_gen_capture.py`` was open- coded ``1 << 32`` arithmetic at every call. One helper, layout stated once. 5. **``docs/dep_gen.md``** First user-facing doc for the feature. Covers motivation (links hw-native-sys#599), enable flags, ``deps.json`` format, ``deps_to_graph.py`` usage + node shape/color legend (AIC cube box, AIV vector ellipse, mix diamond, alloc dashed note), the ``fanout ⊆ deps`` validation gate, and the architecture touchpoints. 6. **CI smoke step for dep_gen** (``.github/workflows/ci.yml``) The default ``pytest tests/st`` invocation does not pass any DFX flag, so the dep_gen capture path never executed in CI. Added a second pytest step in ``st-sim-a2a3`` that re-runs only ``test_dep_gen_capture`` with ``--enable-dep-gen --enable-l2-swimlane``, which forces the full capture → replay → deps.json → fanout ⊆ deps gate path to execute. Same pattern is the model for future PMU / tensor_dump / swimlane smoke steps once those grow dedicated tests. 7. **UT roundtrip for ``enable_dep_gen``** (``tests/ut/py/test_chip_worker.py``) Adds the missing nanobind setter / getter / repr roundtrip for the ``enable_dep_gen`` ``CallConfig`` field, alongside the existing round-trips for ``enable_l2_swimlane`` / ``enable_dump_tensor`` / ``enable_pmu``. Catches binding-ABI regressions at the UT layer (e.g. the ``bool`` vs ``int32`` Python-wrapper bug that broke 9 st jobs earlier in the PR series — 10 s ut would have caught it). 8. **Pytest-friendly test hook via ``test_run`` override** instead of a framework-level ``_post_validate`` callback. ``test_dep_gen_capture`` overrides the inherited ``SceneTestCase.test_run`` to call ``super()`` then walk its cases and assert the dep_gen artifacts. Keeps the framework unchanged (no implicit ``hasattr`` check on every SceneTestCase), localises the test-specific behavior, and the ``--enable-dep-gen`` flag gate prevents the assertion from firing in default invocations. 9. **Unified DFX smoke directory + symmetric a2a3sim/a2a3 coverage** (``tests/st/a2a3/tensormap_and_ringbuffer/dfx/{dep_gen,l2_swimlane,pmu,tensor_dump}/``) Moved ``dep_gen_capture/`` → ``dfx/dep_gen/`` and added three sibling per-feature subdirs matching the repo's "one feature, one directory" convention (cf. ``spmd_sync_start/``, ``paged_attention_unroll/`` etc.). Each subdir holds one ``test_<feature>.py``: - ``dep_gen/test_dep_gen.py`` — asserts ``deps.json`` edges + fanout ⊆ deps gate - ``l2_swimlane/test_l2_swimlane.py`` — asserts ``l2_perf_records.json`` shape - ``pmu/test_pmu.py`` — asserts ``pmu.csv`` header + row count - ``tensor_dump/test_tensor_dump.py`` — asserts ``tensor_dump/`` manifest + bin Each test uses ``vector_example`` as a fixed 5-task workload, opens only its own flag, asserts only its own artifact. ``platforms`` field is ``["a2a3sim", "a2a3"]`` on all four — DFX features matter most on real hardware (the race window in dep_gen never fires on sim's deterministic timing, and post-hw-native-sys#742 we have empirical evidence the gate catches missed edges only on hardware), so both st-sim-a2a3 and st-onboard-a2a3 get four matching smoke steps. Each smoke is its own pytest invocation because L2PerfCollector::initialize trips its dup-init guard (``shm_host_ != nullptr``) when the L2 worker pool reuses the DeviceRunner across tests in the same session — fresh process per smoke side-steps that state leak; making the runtime init re-init-safe is a separate concern. 10. **Sibling helper consistency** (``make_pmu_csv_path``) Converted ``make_pmu_csv_path`` in ``pmu_collector.h`` to use ``std::filesystem::path`` operator/ instead of bare string concat, matching the new ``make_deps_json_path`` convention.
nalinaly
pushed a commit
to nalinaly/simpler
that referenced
this pull request
Jul 31, 2026
Port the dep_gen (SubmitTrace) feature from a2a3 to a5 so the
tensormap_and_ringbuffer runtime on a5 can produce deps.json and feed
flow events into swimlane_converter.py. Without this, --enable-dep-gen
was a no-op on a5 and merged_swimlane_*.json had no dependency arrows.
Reused from a2a3 (code-identical; the .cpp files are byte-for-byte
copies, the headers differ only in include-guard names and
L2Swimlane->L2Perf / a2a3->a5 comment refs):
- Shared-memory ABI: common/dep_gen.h (DepGenRecord 2624 B, overflow
chain, SPSC free_queue, per-thread ready_queue) — layout byte-identical
- AICPU writer: aicpu/dep_gen_collector_aicpu.{h,cpp} (.cpp identical)
- Runtime replay: runtime/tensormap_and_ringbuffer/host/dep_gen_replay
(.cpp identical)
- Orchestrator capture point + aicpu_executor lifecycle hooks
- 5 platform_config constants + PROFILING_FLAG_DEP_GEN bit
Specialized for a5 (no SVM, see profiling_common diff vs a2a3):
- dep_gen_collector.cpp uses alloc_single_buffer (malloc shadow +
profiling_copy_to_device) instead of identity-mapping when
register_cb is null — matches a5's PMU/L2Perf/Dump collectors.
- Two-phase set_memory_context: callbacks first, then shm pointers
once the region is committed, so start(tf) gates correctly.
- reconcile_counters explicitly copy_from_device's the BufferState +
current_buf before reading (mgmt thread is stopped by then).
- finalize lets BufferPoolManager::clear_mappings() be the single
source of truth for host-shadow lifetime — no per-collector dedup.
Sim path: dlsym set_platform_dep_gen_base / set_dep_gen_enabled out of
the AICPU .so and forward kernel_args.dep_gen_data_base + enable flag
at boot, mirroring the existing pmu / dump / l2_perf setters.
Onboard kernel.cpp adds two lines to forward dep_gen_data_base +
PROFILING_FLAG_DEP_GEN into the AICPU writer's globals, mirroring the
existing PMU / L2 / Dump setters.
c_api: a5's DeviceRunner now overrides set_dep_gen_enabled() (was the
base no-op), so run_prepared's already-present enable_dep_gen call is
honored on a5 instead of being silently dropped. The shared
c_api_shared.cpp / device_runner_base.h changes are comment-only —
refreshing stale "a2a3-only" notes to reflect both arches now wire it.
Tests:
- tests/st/a5/.../dfx/dep_gen/test_dep_gen.py: 6-edge validation
against vector_example orchestration (byte-identical to a2a3 — same
expected edge set).
- tests/st/a5/.../dfx/dep_gen/test_dep_gen_chain.py
(+ kernels/orchestration/chain_barrier_orch.cpp): overflow chain
regression for >64 explicit deps.
- tests/st/a2a3/.../dfx/dep_gen/test_dep_gen.py: harden the existing
a2a3 gate — a missing deps.json under --enable-dep-gen now fails
loudly (assert) instead of silently skipping, so a
capture->reconcile->replay regression can no longer pass green
([[hw-native-sys#742](https://github.com/hw-native-sys/simpler/issues/742)](https://github.com/hw-native-sys/simpler/issues/742)).
Docs:
- docs/dfx/dep_gen.md: §8 Architecture Touchpoints now lists both
platforms; "Currently a2a3 only" line removed.
- src/a5/runtime/.../docs/profiling_levels.md: Code Locations point
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
src/a2a3/platform/onboard/aicpu/kernel.cpppropagatesdump_tensor/l2_swimlane/pmufromk_args->enable_profiling_flagon every kernel entry, but the two equivalent setters for dep_gen were missing — so on hardwareis_dep_gen_enabled()stayed false,dep_gen_aicpu_record_submitearly-returned for every submit, and the host replay wrote{\"version\":1,\"edges\":[]}for every run. Sim works because its host device_runner already dlsym'sset_platform_dep_gen_base+set_dep_gen_enabledinto the AICPU .so; onboard had no equivalent wiring.CASES[0].platforms == [\"a2a3sim\"]), so this never surfaced in CI.Changes
aicpu/kernel.cpp: includedep_gen_collector_aicpu.hand addset_platform_dep_gen_base(k_args->dep_gen_data_base)+set_dep_gen_enabled(GET_PROFILING_FLAG(k_args->enable_profiling_flag, PROFILING_FLAG_DEP_GEN))next to the existing dump_tensor / l2_swimlane / pmu propagation.tests/st/a2a3/tensormap_and_ringbuffer/dep_gen_capture/test_dep_gen_capture.py: add\"a2a3\"to the case's platforms list and overridetest_runso the pytest path invokes_post_validatewhen--enable-dep-genis passed. Without the override the pytest collection ran orchestration but never asserted deps.json — a future regression of the same shape would still slip through..github/workflows/ci.yml: add a dedicated dep_gen smoke step to thest-onboard-a2a3job that runs the capture test with--enable-dep-gen --enable-l2-swimlane, so the fanout ⊆ deps gate exercises hardware on every PR touching a2a3.Test plan
build_runtimes.py --platform a2a3).python -m pytest tests/st/a2a3/tensormap_and_ringbuffer/dep_gen_capture --platform a2a3 --enable-dep-gen --enable-l2-swimlanepasses on a2a3 hardware.outputs/TestDepGenCapture_default_*/deps.jsoncontains the six expected edges (t0→t1, t0→t2, t1→t3, t2→t3, t0→t4, t3→t4) — matchesexample_orchestration.cppexactly.st-onboard-a2a3job's new "Run dep_gen smoke" step passes once this PR runs through CI.st-sim-a2a3continues to pass — only the test case's platforms list changes for sim (no behavior change on sim).Context
Found while running a sweep to look for hardware cases where
|deps - fanout| > 0(the kind of race-window evidence dep_gen was designed to surface). Every onboarddeps.jsoncame back empty, which traced to the missing wiring inaicpu/kernel.cpp. The sweep needs to be re-run after this lands.🤖 Generated with Claude Code