Skip to content

Fix: capture dep_gen on the host orchestrator for host_build_graph - #1492

Merged
ChaoWao merged 1 commit into
hw-native-sys:mainfrom
ChaoWao:fix/issue-1487-hbg-dep-gen-init
Jul 26, 2026
Merged

ChaoWao merged 1 commit into
hw-native-sys:mainfrom
ChaoWao:fix/issue-1487-hbg-dep-gen-init

Conversation

@ChaoWao

@ChaoWao ChaoWao commented Jul 26, 2026

Copy link
Copy Markdown
Collaborator

Fixes #1487

What

host_build_graph called dep_gen_aicpu_init() from its scheduler cold path but never called dep_gen_aicpu_set_orch_thread_idx() or dep_gen_aicpu_flush() — the subsystem was initialised and then left half-wired.

It cannot be finished the tensormap_and_ringbuffer way: hbg orchestrates on the host (run_host_orchestration dlsym's and runs the whole orchestration before any scheduler thread starts; the AICPU boot thread only attaches the prebuilt arena), so there is no device orchestrator thread to own a ready queue and nothing device-side to flush.

So this PR gives hbg dep_gen for real, in the shape its architecture allows: capture the graph as the host orchestrator builds it.

  • compute_task_fanin takes an Annotate hook that fires alongside its existing emit, so creator-retention and tensormap-lookup producers are recorded as the runtime resolves them; submit_task_common opens the task entry and records declared dependencies. The graph is the runtime's own answer, not a replay's reconstruction, so it cannot drift from compute_task_fanin semantics.
  • No ring, no collector, no reconcile — orchestration completes before scheduling starts, so nothing can be dropped under back-pressure. The runner picks the shape via dep_gen_host_graph_active() and skips device collector init/start/reconcile (and the device DFX flag) for the host-orch runtime.
  • deps.json keeps the exact schema the device-orch replay emits, so deps_viewer and the swimlane join read both runtimes identically.
  • hbg's copy of dep_gen_replay.{h,cpp} is removed — which is what the device runners' weak stubs already claimed ("host_build_graph has no replay implementation today").

Separately, the shared AICPU writer now guards the negative orchestrator index: a record arriving before dep_gen_aicpu_set_orch_thread_idx() has no ready queue to reach, so it is charged to dropped_record_count instead of filling a buffer nothing can publish and surfacing later as a misleading "ready_queue full" error. Note the queues[-1] write the issue predicted was already impossible — DeviceProfilerEngine::enqueue_ready → wait_for_ready_queue_space rejects a negative index before write_ready_entry.

Equivalence check

The new st case runs the same vector_example orchestration the tensormap_and_ringbuffer dep_gen test uses, and asserts the same 6 edges. Cross-checking the two artifacts directly (topology by submit order, since hbg keeps the inner manual scope on ring 0 where trb moves it to ring 1):

tmr (device-orch replay): [(0,1,creator), (0,2,creator), (0,4,creator), (1,3,creator), (2,3,creator), (3,4,creator)]
hbg (host-orch direct)  : [(0,1,creator), (0,2,creator), (0,4,creator), (1,3,creator), (2,3,creator), (3,4,creator)]
identical: True

Testing

Suite Result
cpput (-LE requires_hardware) 61/61 pass, incl. new test_dep_gen_collector_aicpu (written failing first)
pyut (tests/ut/py) 824 passed, 2 skipped
sim — tests/st/a2a3/host_build_graph 13 passed
sim — hbg dep_gen case with --enable-dep-gen passed
sim — tmr dep_gen cases with --enable-dep-gen 2 passed
onboard (a2a3, via task-submit) — hbg dep_gen + vector_example with --enable-dep-gen 2 passed
onboard (a2a3) — tmr dep_gen with --enable-dep-gen 2 passed
onboard (a2a3) — full tests/st/a2a3/host_build_graph 13 passed, 1 skipped

The full --runtime tensormap_and_ringbuffer sim sweep shows 7 failures / 25 errors on this box (a5 cases collected under a2a3sim, multi-device L3 cases against a 1-device sim pool). A baseline worktree at upstream/main without this change reproduces the identical set (7 failed, 66 passed, 2 skipped, 25 errors), so they are pre-existing environment limits, not regressions.

🤖 Generated with Claude Code

@coderabbitai

coderabbitai Bot commented Jul 26, 2026 •

Copy link
Copy Markdown

Warning

Review limit reached

@ChaoWao, you've reached your PR review limit, so we couldn't start this review.

Next review available in: 34 minutes

Enable usage-based reviews in Billing to review now. Otherwise, wait until the next included review is available.
You're only billed for reviews past your plan's rate limits ($0.25/file).

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro Plus

Run ID: 29b283b4-f686-43c6-8d41-fd7521165f41

📥 Commits

Reviewing files that changed from the base of the PR and between b28bd58 and 7b75109.

📒 Files selected for processing (22)
  • .github/workflows/ci.yml
  • docs/dfx/dep_gen.md
  • src/a2a3/platform/onboard/host/device_runner.cpp
  • src/a2a3/platform/onboard/host/device_runner.h
  • src/a2a3/platform/sim/host/device_runner.cpp
  • src/a2a3/platform/sim/host/device_runner.h
  • src/a2a3/runtime/host_build_graph/aicpu/aicpu_executor.cpp
  • src/a2a3/runtime/host_build_graph/docs/RUNTIME_LOGIC.md
  • src/a2a3/runtime/host_build_graph/host/dep_gen_host_graph.cpp
  • src/a2a3/runtime/host_build_graph/host/dep_gen_replay.cpp
  • src/a2a3/runtime/host_build_graph/host/dep_gen_replay.h
  • src/a2a3/runtime/host_build_graph/host/runtime_maker.cpp
  • src/a2a3/runtime/host_build_graph/runtime/dep_gen_host_graph.h
  • src/a2a3/runtime/host_build_graph/runtime/orchestrator_core/pto_orchestrator.cpp
  • src/a2a3/runtime/host_build_graph/runtime/pto_dep_compute.h
  • src/a2a3/runtime/host_build_graph/runtime/scheduler/scheduler_cold_path.cpp
  • src/common/platform/onboard/host/c_api_shared.cpp
  • src/common/platform/shared/aicpu/dep_gen_collector_aicpu.cpp
  • src/common/platform/sim/host/c_api_shared.cpp
  • tests/st/a2a3/host_build_graph/dfx/dep_gen/test_dep_gen.py
  • tests/ut/cpp/CMakeLists.txt
  • tests/ut/cpp/common/test_dep_gen_collector_aicpu.cpp

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@ChaoWao

ChaoWao commented Jul 26, 2026

Copy link
Copy Markdown
Collaborator Author

CI note — the first attempt of both no-hardware ut jobs failed on
tests/ut/py/test_worker/test_l4_recursive.py::TestL4WithOwnSubs::test_l4_sub_and_l3_dispatch
with assert 1 == 2 (the shared counter saw only one of the two sub-callable
increments). They pass on retry, and everything is green now.

Not this PR: the same assertion fails on an unrelated branch —
run 30186923896
(fix-depgen-init-race, macos ut) — and this PR touches no Python worker code
(git diff upstream/main...HEAD -- python/ simpler_setup/ is empty); the test is
a pure multi-process fork/shared-memory case with no dep_gen involvement. It also
passed 8/8 locally.

Flagging rather than waving it through: it looks like a real race in the L4→L3
dispatch teardown (one path not having completed when w4.close() returns) that
happens to reproduce more readily on the hosted runners. Happy to open a separate
issue for it if maintainers want it tracked.

@ChaoWao
ChaoWao force-pushed the fix/issue-1487-hbg-dep-gen-init branch from 1ba1035 to b40dac3 Compare July 26, 2026 11:15
@ChaoWao

ChaoWao commented Jul 26, 2026

Copy link
Copy Markdown
Collaborator Author

Self-review pass (with codex + gemini cross-check) turned up findings; all fixed and re-pushed.

Must fix — the new capture had zero CI coverage. All four dep_gen smoke steps hardcoded the tensormap_and_ringbuffer path, and the default pytest examples tests/st passes no --enable-dep-gen, so the new st case was collected, passed, and asserted nothing. Exactly the failure mode the smoke steps were added for in #742. Both a2a3 smoke steps (sim + onboard) now run the host_build_graph dep_gen test alongside the tmr one.

Should fix — process-global capture state. HostGraphState was a function-local static, but both c_api_shared.cpp files key the runner off a pthread_key_t; two runners on two threads would have wiped and raced each other's graph. Not reachable from Python today (ChipWorker.run holds the GIL, sub-workers are forked), but the device-orch counterpart is per-runner (dep_gen_collector_ is a member) and there was no reason for this half to be weaker. Now thread_local, with the isolation contract documented in the header. Independently flagged by gemini.

Should fix — silent empty graph. emit() wrote whatever was in the tables without checking capture was ever armed, so a mis-wired arm produced an empty-but-valid deps.json that looks like success (I hit exactly this during development). It now returns non-zero and says so.

Should fix — st covered only one of three edge sources. The shared vector_example orchestration produces only creator edges, leaving the tensormap (Step B) and explicit (STEP 1) hooks unasserted. Added TestDepGenHostBuildGraphEdgeSources, which runs predicated_dispatch and asserts {(0,2,explicit), (1,2,tensormap), (2,3,tensormap)} plus the producer-side slice annotation — the latter being what would catch the entry read drifting after the INOUT+COVERED remove_entry().

Should fix — doc drift I introduced. docs/dfx/dep_gen.md §1 still defined dep_gen as "capture inputs + replay offline", true of only one of the two shapes after §2 split them.

Also folded in: block_num now guards before the widening cast (a negative int16_t would have become ~4e9 instead of clamping to 1); reset() releases the graph's memory instead of clear()ing it (a captured graph is one heap-allocated arg vector per task — bgemm already hits 128 tasks); an explicit end_task() closes the task entry so a stray edge can't attach to the previous task; and the dedup-key asymmetry (runtime keys fanin on (ring, slot), capture on producer task id — equal only because hbg never reuses a slot at build time) is now stated in the header.

Re-verified after the fixes: hbg sim suite 14 passed; both dep_gen st cases pass with and without --enable-dep-gen; tmr dfx suite 7 passed with --enable-dep-gen; cpput 62/62; onboard a2a3 run of the exact new CI command (both runtimes' dep_gen tests) passes. A sweep of every hbg example with capture on exercises all three edge sources (bgemm: 128 tasks / 112 edges, creator 64 + tensormap 48; predicated_dispatch: explicit 1 + tensormap 2) with no edge referencing a tensor_id absent from tensors[].

One note on the earlier ut flake: still believed pre-existing (same assertion fails on fix-depgen-init-race), unrelated to this diff.

@ChaoWao
ChaoWao force-pushed the fix/issue-1487-hbg-dep-gen-init branch from b40dac3 to 7b75109 Compare July 26, 2026 11:38
Fixes hw-native-sys#1487

host_build_graph called dep_gen_aicpu_init() from its scheduler cold path
but never called dep_gen_aicpu_set_orch_thread_idx() or _flush(), leaving
the subsystem initialised and half-wired. It cannot be finished the
tensormap_and_ringbuffer way: host_build_graph orchestrates on the host, so
there is no device orchestrator thread to own a ready queue, nothing on the
device to flush, and the AICPU never submits a task.

Capture the graph where this runtime actually builds it. compute_task_fanin
takes an Annotate hook that fires alongside its existing emit, so creator
retention and tensormap lookups are recorded as the runtime resolves them,
and submit_task_common opens the task entry and records declared
dependencies. What lands in deps.json is therefore the runtime's own
dependency resolution rather than a replay's reconstruction of it, and it
cannot drift from compute_task_fanin semantics. No ring, no collector, no
reconcile: the orchestration completes before any scheduler thread starts,
so nothing can be dropped under back-pressure.

The capture state is thread-local, matching the per-thread runner both
c_api_shared.cpp files already key off a pthread_key_t — two runners on two
threads build two independent graphs, the same per-runner isolation the
device-orch shape gets from DeviceRunner::dep_gen_collector_ being a
member. emit() refuses to write when no capture ran on its thread, so a
mis-wired arm reports instead of producing an empty-but-valid deps.json.

deps.json keeps the schema the device-orch replay emits, so deps_viewer and
the swimlane join read both runtimes identically. Two st cases cover it:
one runs the same vector_example orchestration as the
tensormap_and_ringbuffer dep_gen test and asserts the same 6 edges (which
is what would catch the two shapes diverging), the other runs
predicated_dispatch for the tensormap and explicit edge sources and their
producer-side slice annotation. Both are wired into the a2a3 dep_gen smoke
steps, sim and onboard — the default st sweep passes no --enable-dep-gen,
so a capture regression is only visible there.

The runner picks the shape via dep_gen_host_graph_active() and skips device
collector init/start/reconcile for the host-orch one; host_build_graph's
copy of dep_gen_replay.{h,cpp} is removed, matching what the device
runners' weak stubs already claimed.

Also guard the shared AICPU writer: a record arriving before
dep_gen_aicpu_set_orch_thread_idx() has no ready queue to reach, so it is
now charged to dropped_record_count instead of filling a buffer nothing can
publish and surfacing later as a misleading "ready_queue full" error.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@ChaoWao
ChaoWao merged commit 4d39cbc into hw-native-sys:main Jul 26, 2026
16 checks passed
@ChaoWao
ChaoWao deleted the fix/issue-1487-hbg-dep-gen-init branch July 26, 2026 12:12
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug] host_build_graph inits dep_gen but never sets its orch thread index, leaving queues[-1]

1 participant