Platform: a2a3 (Ascend 910B/C hardware)
Runtime Variant: tensormap_and_ringbuffer
Host: Linux (aarch64)
Commit: a94d514 (runtime submodule HEAD, detached)
Summary
validate_runtime_impl (post-run() host finalize) is taking ~50ms on average and up to ~91ms per chip invocation, which is ~2x the cost of the entire chip execution (runner_run avg 26.6ms incl. AICPU+AICore stream sync). Over a single 16-token decode of Qwen3-314B (1-layer config), it consumed 2.43s of host time across 48 calls — i.e. ~26% of the entire decode budget is spent here, after the chip is already idle.
Because run_prepared() serializes runner->run() -> validate_runtime_impl() on the host thread, this overhead is on the critical path of every token: it cannot be hidden behind chip execution.
Reproduction
Driven via pypto-lib running Qwen3-314B 1-layer decode on a2a3 hardware:
cd pypto-lib
python models/qwen3/314b/qwen3_314b_decode.py -p a2a3 -d 0 \
--num-tokens 16 --num-layers 1
To get the timing breakdown, build runtime/ with local instrumentation in src/a2a3/platform/onboard/host/pto_runtime_c_api.cpp and device_runner.cpp (steady_clock timing around validate_runtime_impl, kernel launches, and stream syncs, gated by SIMPLER_CHIP_TIMING=1). Then:
SIMPLER_CHIP_TIMING=1 python ... -d 0 2> timing.log
Expected Performance
validate_runtime_impl is pure host finalize after the chip has synced. For a 1-layer decode kernel with a small number of output tensors, expected cost is <5ms (a few aclrtMemcpy D2H calls + a few device frees).
Actual Performance
Measured across 48 consecutive run_prepared() calls in the same process (no rebuild / no driver re-init between calls):
validate_runtime_impl: n=48 total=2426.8ms avg=50.56ms min=14.49ms max=90.62ms
runner_run (chip+sync): n=48 total=1274.5ms avg=26.55ms min=17.89ms max=64.30ms
A clear 3-phase periodic pattern is visible across the trace, one phase per kernel invoked per decode step (decode_layer, final_rms, lm_head):
decode_layer validate ~ 14-20 ms (small output tensors)
final_rms validate ~ 82-91 ms (largest output)
lm_head validate ~ 48-56 ms (vocab-sized output)
This is consistent with the dominant cost being per-tensor D2H copy + per-tensor device_free inside the host loop at src/a2a3/runtime/tensormap_and_ringbuffer/host/runtime_maker.cpp:367-410:
runtime->host_api.copy_from_device(...) is called per tensor pair
runtime->host_api.device_free(...) is called per tensor pair
- both are presumably synchronous ACL calls — no batching, no async, no parallelism, and no skip path for tensors whose host_ptr was already up-to-date.
Profiling Data
Per-run breakdown extract ([c_api_timing] lines, omitted bind_callable which is always 0.00ms):
run= 0 prepare_ctx=0.83 runner_run=64.30 validate=76.21 total=492.60
run= 1 prepare_ctx=0.70 runner_run=33.69 validate=14.62 total=51.62
run= 2 prepare_ctx=0.74 runner_run=35.67 validate=83.08 total=214.28
run= 3 prepare_ctx=0.69 runner_run=34.40 validate=51.88 total=117.21
...
run=44 prepare_ctx=0.58 runner_run=19.72 validate=90.62 total=150.38
run=45 prepare_ctx=0.60 runner_run=34.32 validate=48.28 total=109.97
run=46 prepare_ctx=0.61 runner_run=18.22 validate=15.97 total=36.60
run=47 prepare_ctx=0.60 runner_run=20.21 validate=82.17 total=142.39
Run-level downstream impact in this trace:
[api] run_decode total : 8.37s (15 steps, avg 558.0 ms/step, 1.79 step/s)
[phase] throughput (e2e) : 0.89 tok/s
So removing the validate host overhead (~150ms per token = 14+82+50) could in principle bring per-step from ~558ms toward ~408ms — about a 27% improvement on a single chip / single layer.
Additional Context
- This is host-side overhead AFTER
aclrtSynchronizeStreamWithTimeout on both streams. The chip is already idle when validate runs.
- The
total field (total=492.60 on run 0, ~150 on later runs) is much larger than prepare_ctx + runner_run + validate. The remainder is on the Python side (worker.py store_to_host flush + mailbox transitions). That overhead is separate and not the subject of this issue.
- Possible angles to investigate:
- Are all tensor pairs that
validate_runtime_impl iterates actually outputs? If the loop is also copying back inputs / weights / KV-cache entries with stale host_ptr, that is the bug.
- Can
copy_from_device be issued asynchronously and joined once at the end, rather than serially?
- Can
device_free be batched, or deferred to the next init_runtime_impl (free-on-reuse)?
- Tensors whose
host_ptr == dev_ptr (unified addr) or whose output is already known to live in the packed graph_output buffer should skip the per-tensor copy.
Happy to provide the full timing log on request.
Platform: a2a3 (Ascend 910B/C hardware)
Runtime Variant: tensormap_and_ringbuffer
Host: Linux (aarch64)
Commit: a94d514 (runtime submodule HEAD, detached)
Summary
validate_runtime_impl(post-run()host finalize) is taking ~50ms on average and up to ~91ms per chip invocation, which is ~2x the cost of the entire chip execution (runner_run avg 26.6ms incl. AICPU+AICore stream sync). Over a single 16-token decode of Qwen3-314B (1-layer config), it consumed 2.43s of host time across 48 calls — i.e. ~26% of the entire decode budget is spent here, after the chip is already idle.Because
run_prepared()serializesrunner->run()->validate_runtime_impl()on the host thread, this overhead is on the critical path of every token: it cannot be hidden behind chip execution.Reproduction
Driven via pypto-lib running Qwen3-314B 1-layer decode on a2a3 hardware:
cd pypto-lib python models/qwen3/314b/qwen3_314b_decode.py -p a2a3 -d 0 \ --num-tokens 16 --num-layers 1To get the timing breakdown, build
runtime/with local instrumentation insrc/a2a3/platform/onboard/host/pto_runtime_c_api.cppanddevice_runner.cpp(steady_clock timing aroundvalidate_runtime_impl, kernel launches, and stream syncs, gated bySIMPLER_CHIP_TIMING=1). Then:SIMPLER_CHIP_TIMING=1 python ... -d 0 2> timing.logExpected Performance
validate_runtime_implis pure host finalize after the chip has synced. For a 1-layer decode kernel with a small number of output tensors, expected cost is <5ms (a fewaclrtMemcpyD2H calls + a few device frees).Actual Performance
Measured across 48 consecutive
run_prepared()calls in the same process (no rebuild / no driver re-init between calls):A clear 3-phase periodic pattern is visible across the trace, one phase per kernel invoked per decode step (
decode_layer,final_rms,lm_head):This is consistent with the dominant cost being per-tensor D2H copy + per-tensor
device_freeinside the host loop atsrc/a2a3/runtime/tensormap_and_ringbuffer/host/runtime_maker.cpp:367-410:runtime->host_api.copy_from_device(...)is called per tensor pairruntime->host_api.device_free(...)is called per tensor pairProfiling Data
Per-run breakdown extract (
[c_api_timing]lines, omittedbind_callablewhich is always 0.00ms):Run-level downstream impact in this trace:
So removing the validate host overhead (~150ms per token = 14+82+50) could in principle bring per-step from ~558ms toward ~408ms — about a 27% improvement on a single chip / single layer.
Additional Context
aclrtSynchronizeStreamWithTimeouton both streams. The chip is already idle when validate runs.totalfield (total=492.60on run 0,~150on later runs) is much larger thanprepare_ctx + runner_run + validate. The remainder is on the Python side (worker.pystore_to_host flush + mailbox transitions). That overhead is separate and not the subject of this issue.validate_runtime_impliterates actually outputs? If the loop is also copying back inputs / weights / KV-cache entries with stalehost_ptr, that is the bug.copy_from_devicebe issued asynchronously and joined once at the end, rather than serially?device_freebe batched, or deferred to the nextinit_runtime_impl(free-on-reuse)?host_ptr == dev_ptr(unified addr) or whose output is already known to live in the packed graph_output buffer should skip the per-tensor copy.Happy to provide the full timing log on request.