Skip to content

[Performance] validate_runtime_impl is ~2x the chip-side cost per run (avg 50ms, max 91ms) on Qwen3-314B decode #796

Description

@wangqin1723-max

Platform: a2a3 (Ascend 910B/C hardware)
Runtime Variant: tensormap_and_ringbuffer
Host: Linux (aarch64)
Commit: a94d514 (runtime submodule HEAD, detached)

Summary

validate_runtime_impl (post-run() host finalize) is taking ~50ms on average and up to ~91ms per chip invocation, which is ~2x the cost of the entire chip execution (runner_run avg 26.6ms incl. AICPU+AICore stream sync). Over a single 16-token decode of Qwen3-314B (1-layer config), it consumed 2.43s of host time across 48 calls — i.e. ~26% of the entire decode budget is spent here, after the chip is already idle.

Because run_prepared() serializes runner->run() -> validate_runtime_impl() on the host thread, this overhead is on the critical path of every token: it cannot be hidden behind chip execution.

Reproduction

Driven via pypto-lib running Qwen3-314B 1-layer decode on a2a3 hardware:

cd pypto-lib
python models/qwen3/314b/qwen3_314b_decode.py -p a2a3 -d 0 \
    --num-tokens 16 --num-layers 1

To get the timing breakdown, build runtime/ with local instrumentation in src/a2a3/platform/onboard/host/pto_runtime_c_api.cpp and device_runner.cpp (steady_clock timing around validate_runtime_impl, kernel launches, and stream syncs, gated by SIMPLER_CHIP_TIMING=1). Then:

SIMPLER_CHIP_TIMING=1 python ... -d 0 2> timing.log

Expected Performance

validate_runtime_impl is pure host finalize after the chip has synced. For a 1-layer decode kernel with a small number of output tensors, expected cost is <5ms (a few aclrtMemcpy D2H calls + a few device frees).

Actual Performance

Measured across 48 consecutive run_prepared() calls in the same process (no rebuild / no driver re-init between calls):

validate_runtime_impl: n=48  total=2426.8ms  avg=50.56ms  min=14.49ms  max=90.62ms
runner_run (chip+sync): n=48  total=1274.5ms  avg=26.55ms  min=17.89ms  max=64.30ms

A clear 3-phase periodic pattern is visible across the trace, one phase per kernel invoked per decode step (decode_layer, final_rms, lm_head):

decode_layer  validate ~ 14-20 ms  (small output tensors)
final_rms     validate ~ 82-91 ms  (largest output)
lm_head       validate ~ 48-56 ms  (vocab-sized output)

This is consistent with the dominant cost being per-tensor D2H copy + per-tensor device_free inside the host loop at src/a2a3/runtime/tensormap_and_ringbuffer/host/runtime_maker.cpp:367-410:

  • runtime->host_api.copy_from_device(...) is called per tensor pair
  • runtime->host_api.device_free(...) is called per tensor pair
  • both are presumably synchronous ACL calls — no batching, no async, no parallelism, and no skip path for tensors whose host_ptr was already up-to-date.

Profiling Data

Per-run breakdown extract ([c_api_timing] lines, omitted bind_callable which is always 0.00ms):

run= 0  prepare_ctx=0.83  runner_run=64.30   validate=76.21  total=492.60
run= 1  prepare_ctx=0.70  runner_run=33.69   validate=14.62  total=51.62
run= 2  prepare_ctx=0.74  runner_run=35.67   validate=83.08  total=214.28
run= 3  prepare_ctx=0.69  runner_run=34.40   validate=51.88  total=117.21
...
run=44  prepare_ctx=0.58  runner_run=19.72   validate=90.62  total=150.38
run=45  prepare_ctx=0.60  runner_run=34.32   validate=48.28  total=109.97
run=46  prepare_ctx=0.61  runner_run=18.22   validate=15.97  total=36.60
run=47  prepare_ctx=0.60  runner_run=20.21   validate=82.17  total=142.39

Run-level downstream impact in this trace:

[api]   run_decode total : 8.37s (15 steps, avg 558.0 ms/step, 1.79 step/s)
[phase] throughput (e2e) : 0.89 tok/s

So removing the validate host overhead (~150ms per token = 14+82+50) could in principle bring per-step from ~558ms toward ~408ms — about a 27% improvement on a single chip / single layer.

Additional Context

  • This is host-side overhead AFTER aclrtSynchronizeStreamWithTimeout on both streams. The chip is already idle when validate runs.
  • The total field (total=492.60 on run 0, ~150 on later runs) is much larger than prepare_ctx + runner_run + validate. The remainder is on the Python side (worker.py store_to_host flush + mailbox transitions). That overhead is separate and not the subject of this issue.
  • Possible angles to investigate:
    1. Are all tensor pairs that validate_runtime_impl iterates actually outputs? If the loop is also copying back inputs / weights / KV-cache entries with stale host_ptr, that is the bug.
    2. Can copy_from_device be issued asynchronously and joined once at the end, rather than serially?
    3. Can device_free be batched, or deferred to the next init_runtime_impl (free-on-reuse)?
    4. Tensors whose host_ptr == dev_ptr (unified addr) or whose output is already known to live in the packed graph_output buffer should skip the per-tensor copy.

Happy to provide the full timing log on request.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    performancePerformance regression or optimization

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions