Skip to content

Refactor: move per-run simpler_run state to its correct lifecycle - #1242

Merged
ChaoWao merged 1 commit into
hw-native-sys:mainfrom
ChaoWao:worktree-serene-napping-unicorn
Jul 2, 2026
Merged

ChaoWao merged 1 commit into
hw-native-sys:mainfrom
ChaoWao:worktree-serene-napping-unicorn

Conversation

@ChaoWao

@ChaoWao ChaoWao commented Jul 1, 2026

Copy link
Copy Markdown
Collaborator

Summary

Four related cleanups that stop simpler_run from doing per-run work that belongs to longer-lived scopes, plus a shrink of the host_build_graph Runtime upload. Each is independent; grouped here as one reviewable change.

  • HostApi built once at load time. The 12-pointer HostApi table was reassembled on every simpler_run (onboard + sim). Its wrappers are context-free — each recovers its runner from a thread-local — so a single static const table is valid for every runner and run. Completes the Refactor: move HostApi out of Runtime into a shared explicit parameter #1227 direction (which moved HostApi off Runtime but left it a per-run local).

  • hbg Runtime upload shrunk to the populated task prefix. host_build_graph uploaded the whole ~92 MB Runtime image every run, dominated by the fixed tasks[131072] array. Members are reordered so all device-read fields form a contiguous prefix ending in tasks[], with a host-only tail after it. runtime_device_copy_size() now returns offsetof(tasks) + get_task_count()*sizeof(Task), uploading only populated slots; the AICPU cache-invalidate matches. offsetof static_asserts guard the layout. Safe because the device never reads tasks[i] for i >= next_task_id (get_task() bounds-checks). The trb dev sub-struct pattern is inapplicable here because hbg embeds std::atomic<int> fanin in the uploaded image, so this is an explicit prefix boundary rather than a nested descriptor.

  • Two-step callable bind merged into one facade. bind_callable_to_runtime returned a BindCallableResult carrying a raw pointer into CallableState's cached signature, which the c_api then handed to bind_callable_to_runtime_impl. Both halves are folded into one runner facade so that pointer no longer crosses the c_api boundary; BindCallableResult is removed.

  • CallConfig threaded through simpler_run. The C ABI unpacked CallConfig into 14 positional params, and the c_api re-scattered them across six runner->set_*_enabled() calls. simpler_run now takes const CallConfig*; run() takes const CallConfig& and latches the diagnostic enables via apply_call_config() at entry. The enable_*_ members are kept (≈50 downstream readers in the collector paths).

Testing

  • Simulation tests pass — host_build_graph on a2a3sim (10 passed) and a5sim (7 passed)
  • Hardware tests pass — a2a3 onboard host_build_graph (10 passed, 1 skipped) and tensormap_and_ringbuffer (33 passed, 1 skipped)

Note: onboard was validated per-refactor before a rebase onto #1234; after rebasing, both sim suites were re-run green. tests/ut/cpp has a pre-existing gtest link-ABI failure in this environment (affects untouched binaries too) and was not run.

@coderabbitai

coderabbitai Bot commented Jul 1, 2026 •

Copy link
Copy Markdown

Review Change Stack

Important

Review skipped

Auto incremental reviews are disabled on this repository.

Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: 5410d5a4-8092-423a-b9b2-7677500d5c34

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review
📝 Walkthrough

Walkthrough

Introduces a CallConfig object that replaces individual launch parameters (block_dim, aicpu_thread_num, feature flags, ring pointers, output_prefix) across DeviceRunner::run, bind_callable_to_runtime, and the public simpler_run C API. Also restructures the Runtime struct to define a device-read prefix ending at tasks[], enabling offsetof-based truncated host-to-device copy sizing instead of copying the full struct. Documentation and unit tests updated accordingly.

Changes

CallConfig-Driven Run/Bind Refactor and Runtime Layout

Layer / File(s) Summary
Base interface changes
src/common/platform/onboard/host/device_runner_base.{h,cpp}, src/common/platform/sim/host/device_runner_base.{h,cpp}, src/common/task_interface/prepare_callable_common.h
run() now takes const CallConfig&; bind_callable_to_runtime returns int and accepts HostApi*, orch_args, ring pointers; apply_call_config added; BindCallableResult struct removed.
C API and shared HostApi table
src/common/platform/onboard/host/c_api_shared.cpp, src/common/platform/sim/host/c_api_shared.cpp, src/common/worker/pto_runtime_c_api.h, src/common/worker/chip_worker.{h,cpp}, src/common/platform/include/common/host_api.h, tests/ut/cpp/stubs/test_stubs.cpp
simpler_run now takes a single const CallConfig*; a shared static g_host_api table replaces per-call construction; SimplerRunFn typedef simplified; test stub added for linking.
Platform runner implementations
src/a2a3/platform/{onboard,sim}/host/device_runner.{h,cpp}, src/a5/platform/{onboard,sim}/host/device_runner.{h,cpp}, tests/ut/cpp/common/test_l3_l2_orch_comm_sim_runner.cpp, tests/ut/cpp/hardware/test_l3_l2_orch_comm_onboard_runner.cpp
Concrete run() overrides updated to accept CallConfig, call apply_call_config, and derive block_dim/launch_aicpu_num from config; test doubles updated to match.
Runtime device-read prefix layout
src/a2a3/runtime/host_build_graph/runtime/runtime.{h,cpp}, src/a5/runtime/host_build_graph/runtime/runtime.{h,cpp}, src/a2a3/runtime/host_build_graph/aicpu/aicpu_executor.cpp, src/a5/runtime/host_build_graph/aicpu/aicpu_executor.cpp, */host/runtime_maker.cpp
Runtime members reordered so tasks[] is the last device-read member; offsetof-based static_asserts added; runtime_device_copy_size and cache invalidation truncated to populated prefix instead of sizeof(Runtime).
Documentation updates
docs/chip-level-arch.md, docs/dynamic-linking.md
API examples and lifecycle diagrams updated to reflect config-based run/simpler_run calls.

Estimated code review effort: 4 (Complex) | ~60 minutes

Possibly related PRs

  • hw-native-sys/simpler#928: Extracts shared onboard C-API glue and DeviceRunnerBase interface directly reused by this refactor.
  • hw-native-sys/simpler#932: Introduces the sim shared host glue (SimDeviceRunnerBase, c_api_shared.cpp) that this PR converts to the CallConfig-driven flow.
  • hw-native-sys/simpler#1042: Extends CallConfig with runtime_env ring sizing threaded through the same bind/run call path introduced here.

Poem

A rabbit hops through configs new,
One bundle where param lists once grew.
Tasks trimmed tight, no bytes to spare,
offsetof whispers "just what's there."
Thump thump — the runtime's lean and neat! 🐇✨

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Title check ✅ Passed The title clearly captures the main refactor around moving per-run simpler_run state into longer-lived lifecycle scopes.
Description check ✅ Passed The description is directly related and accurately summarizes the refactor and runtime upload changes.
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@gemini-code-assist

Copy link
Copy Markdown

Warning

Gemini encountered an error creating the review. You can try again by commenting /gemini review.

@ChaoWao
ChaoWao force-pushed the worktree-serene-napping-unicorn branch from 007a452 to 2828962 Compare July 2, 2026 00:50

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (2)
src/a5/platform/onboard/host/device_runner.h (1)

86-108: 📐 Maintainability & Code Quality | 🟡 Minor | ⚡ Quick win

Stale Doxygen params on run().

The @param block_dim / @param launch_aicpu_num entries no longer match the const CallConfig &config signature, and config is undocumented.

📝 Proposed doc fix
-     * `@param` runtime             Runtime to execute (will be modified to
-     * initialize workers)
-     * `@param` block_dim            Number of blocks (1 block = 1 AIC + 2 AIV)
-     * `@param` launch_aicpu_num     Number of AICPU instances (default: 1)
+     * `@param` runtime  Runtime to execute (will be modified to initialize workers)
+     * `@param` config   Per-run CallConfig: block_dim (0 = auto), aicpu_thread_num,
+     *                 diagnostic enables + output_prefix (latched via
+     *                 apply_call_config), and ring sizing overrides.
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@src/a5/platform/onboard/host/device_runner.h` around lines 86 - 108, Update
the Doxygen for DeviceRunner::run so it matches the current signature: remove
the stale `@param` entries for block_dim and launch_aicpu_num, and document the
actual const CallConfig &config parameter instead. Keep the existing run()
behavior description intact and make sure the parameter list reflects the
symbols in DeviceRunner::run and CallConfig.
src/a2a3/platform/onboard/host/device_runner.h (1)

93-115: 📐 Maintainability & Code Quality | 🟡 Minor | ⚡ Quick win

Stale Doxygen params on run().

The @param block_dim / @param launch_aicpu_num entries no longer match the new const CallConfig &config signature, and config is undocumented.

📝 Proposed doc fix
-     * `@param` runtime             Runtime to execute (will be modified to
-     * initialize workers)
-     * `@param` block_dim            Number of blocks (1 block = 1 AIC + 2 AIV)
-     * `@param` launch_aicpu_num     Number of AICPU instances (default: 1)
+     * `@param` runtime  Runtime to execute (will be modified to initialize workers)
+     * `@param` config   Per-run CallConfig: block_dim (0 = auto), aicpu_thread_num,
+     *                 diagnostic enables + output_prefix (latched via
+     *                 apply_call_config), and ring sizing overrides.
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@src/a2a3/platform/onboard/host/device_runner.h` around lines 93 - 115, The
Doxygen for DeviceRunner::run is stale because it still documents block_dim and
launch_aicpu_num even though the signature now takes const CallConfig &config.
Update the comment to describe config instead, document its fields or intent,
and remove the obsolete parameter entries so the docs match run(Runtime
&runtime, const CallConfig &config) and the implementation context in
DeviceRunner.
🧹 Nitpick comments (2)
src/a5/runtime/host_build_graph/aicpu/aicpu_executor.cpp (1)

1152-1157: 📐 Maintainability & Code Quality | 🔵 Trivial | ⚡ Quick win

Reuse runtime_device_copy_size() for the invalidate range.

This path should stay byte-for-byte aligned with the upload/allocation sizing contract; duplicating the offsetof + task_count * sizeof(Task) formula makes future layout changes easy to miss.

Proposed refactor
-#pragma GCC diagnostic push
-#pragma GCC diagnostic ignored "-Winvalid-offsetof"
-    const size_t runtime_prefix_bytes =
-        offsetof(Runtime, tasks) + static_cast<size_t>(runtime->get_task_count()) * sizeof(Task);
-#pragma GCC diagnostic pop
+    const size_t runtime_prefix_bytes = runtime_device_copy_size(*runtime);
     cache_invalidate_range(runtime, runtime_prefix_bytes);
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@src/a5/runtime/host_build_graph/aicpu/aicpu_executor.cpp` around lines 1152 -
1157, The invalidate-range calculation is duplicating the runtime
upload/allocation sizing contract, so update the AICPU executor path to reuse
runtime_device_copy_size() instead of recomputing offsetof(Runtime, tasks) plus
task_count * sizeof(Task). Make the cache_invalidate_range call consume that
shared size value directly so the invalidation stays byte-for-byte aligned with
the existing contract and future Runtime layout changes only need to be fixed in
one place.
src/a2a3/runtime/host_build_graph/aicpu/aicpu_executor.cpp (1)

1157-1162: 📐 Maintainability & Code Quality | 🔵 Trivial | ⚡ Quick win

Reuse runtime_device_copy_size() for the invalidate range.

This path should stay byte-for-byte aligned with the upload/allocation sizing contract; duplicating the offsetof + task_count * sizeof(Task) formula makes future layout changes easy to miss.

Proposed refactor
-#pragma GCC diagnostic push
-#pragma GCC diagnostic ignored "-Winvalid-offsetof"
-    const size_t runtime_prefix_bytes =
-        offsetof(Runtime, tasks) + static_cast<size_t>(runtime->get_task_count()) * sizeof(Task);
-#pragma GCC diagnostic pop
+    const size_t runtime_prefix_bytes = runtime_device_copy_size(*runtime);
     cache_invalidate_range(runtime, runtime_prefix_bytes);
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@src/a2a3/runtime/host_build_graph/aicpu/aicpu_executor.cpp` around lines 1157
- 1162, The invalidate-range size in the runtime upload path is duplicating the
allocation sizing formula, so update the `aicpu_executor.cpp` logic to call
`runtime_device_copy_size()` instead of recomputing `offsetof(Runtime, tasks) +
task_count * sizeof(Task)`. Keep `cache_invalidate_range(runtime, ...)` using
that shared helper so the `Runtime`/`Task` layout contract stays centralized and
in sync.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Outside diff comments:
In `@src/a2a3/platform/onboard/host/device_runner.h`:
- Around line 93-115: The Doxygen for DeviceRunner::run is stale because it
still documents block_dim and launch_aicpu_num even though the signature now
takes const CallConfig &config. Update the comment to describe config instead,
document its fields or intent, and remove the obsolete parameter entries so the
docs match run(Runtime &runtime, const CallConfig &config) and the
implementation context in DeviceRunner.

In `@src/a5/platform/onboard/host/device_runner.h`:
- Around line 86-108: Update the Doxygen for DeviceRunner::run so it matches the
current signature: remove the stale `@param` entries for block_dim and
launch_aicpu_num, and document the actual const CallConfig &config parameter
instead. Keep the existing run() behavior description intact and make sure the
parameter list reflects the symbols in DeviceRunner::run and CallConfig.

---

Nitpick comments:
In `@src/a2a3/runtime/host_build_graph/aicpu/aicpu_executor.cpp`:
- Around line 1157-1162: The invalidate-range size in the runtime upload path is
duplicating the allocation sizing formula, so update the `aicpu_executor.cpp`
logic to call `runtime_device_copy_size()` instead of recomputing
`offsetof(Runtime, tasks) + task_count * sizeof(Task)`. Keep
`cache_invalidate_range(runtime, ...)` using that shared helper so the
`Runtime`/`Task` layout contract stays centralized and in sync.

In `@src/a5/runtime/host_build_graph/aicpu/aicpu_executor.cpp`:
- Around line 1152-1157: The invalidate-range calculation is duplicating the
runtime upload/allocation sizing contract, so update the AICPU executor path to
reuse runtime_device_copy_size() instead of recomputing offsetof(Runtime, tasks)
plus task_count * sizeof(Task). Make the cache_invalidate_range call consume
that shared size value directly so the invalidation stays byte-for-byte aligned
with the existing contract and future Runtime layout changes only need to be
fixed in one place.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: a9cca73d-16d6-4c5b-8d54-d5563116def3

📥 Commits

Reviewing files that changed from the base of the PR and between 91c46a6 and 2828962.

📒 Files selected for processing (32)
  • docs/chip-level-arch.md
  • docs/dynamic-linking.md
  • src/a2a3/platform/onboard/host/device_runner.cpp
  • src/a2a3/platform/onboard/host/device_runner.h
  • src/a2a3/platform/sim/host/device_runner.cpp
  • src/a2a3/platform/sim/host/device_runner.h
  • src/a2a3/runtime/host_build_graph/aicpu/aicpu_executor.cpp
  • src/a2a3/runtime/host_build_graph/host/runtime_maker.cpp
  • src/a2a3/runtime/host_build_graph/runtime/runtime.cpp
  • src/a2a3/runtime/host_build_graph/runtime/runtime.h
  • src/a5/platform/onboard/host/device_runner.cpp
  • src/a5/platform/onboard/host/device_runner.h
  • src/a5/platform/sim/host/device_runner.cpp
  • src/a5/platform/sim/host/device_runner.h
  • src/a5/runtime/host_build_graph/aicpu/aicpu_executor.cpp
  • src/a5/runtime/host_build_graph/host/runtime_maker.cpp
  • src/a5/runtime/host_build_graph/runtime/runtime.cpp
  • src/a5/runtime/host_build_graph/runtime/runtime.h
  • src/common/platform/include/common/host_api.h
  • src/common/platform/onboard/host/c_api_shared.cpp
  • src/common/platform/onboard/host/device_runner_base.cpp
  • src/common/platform/onboard/host/device_runner_base.h
  • src/common/platform/sim/host/c_api_shared.cpp
  • src/common/platform/sim/host/device_runner_base.cpp
  • src/common/platform/sim/host/device_runner_base.h
  • src/common/task_interface/prepare_callable_common.h
  • src/common/worker/chip_worker.cpp
  • src/common/worker/chip_worker.h
  • src/common/worker/pto_runtime_c_api.h
  • tests/ut/cpp/common/test_l3_l2_orch_comm_sim_runner.cpp
  • tests/ut/cpp/hardware/test_l3_l2_orch_comm_onboard_runner.cpp
  • tests/ut/cpp/stubs/test_stubs.cpp
💤 Files with no reviewable changes (1)
  • src/common/task_interface/prepare_callable_common.h

Four related cleanups that stop simpler_run from doing per-run work that
belongs to longer-lived scopes, and shrink the hbg Runtime upload:

- HostApi is built once at load time as a file-scope `static const` table
  instead of reassembling its 12 context-free function pointers on every
  simpler_run (each wrapper recovers its runner from a thread-local, so one
  table is valid for every runner and run). Completes hw-native-sys#1227.

- host_build_graph Runtime is reordered so all device-read fields form a
  contiguous prefix ending in tasks[], with a host-only tail after it.
  runtime_device_copy_size() now uploads only offsetof(tasks) +
  get_task_count()*sizeof(Task), truncating the unpopulated slots of the
  fixed 131072-entry tasks[] array; the AICPU cache-invalidate matches.
  offsetof static_asserts guard the layout. Safe because the device never
  reads tasks[i] for i >= next_task_id.

- The two-step callable bind (bind_callable_to_runtime returning a
  BindCallableResult, then bind_callable_to_runtime_impl) is merged into one
  runner facade so the CallableState-owned signature pointer no longer
  crosses the c_api boundary; BindCallableResult is removed.

- CallConfig is threaded through simpler_run as a single pointer instead of
  14 unpacked C-ABI params plus six runner->set_*_enabled() calls. run()
  takes const CallConfig& and latches the diagnostic enables via
  apply_call_config() at entry.

Verified on a2a3 silicon (host_build_graph + tensormap_and_ringbuffer) and
in simulation on a2a3 + a5.
@ChaoWao
ChaoWao force-pushed the worktree-serene-napping-unicorn branch from 2828962 to 9052868 Compare July 2, 2026 01:58
@ChaoWao
ChaoWao merged commit 04d3dca into hw-native-sys:main Jul 2, 2026
16 checks passed
@ChaoWao
ChaoWao deleted the worktree-serene-napping-unicorn branch July 2, 2026 02:40
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant