Platform
a2a3 (Ascend 910B/C hardware) and a2a3sim (Ascend 910B/C simulation) — fails on both.
Runtime Variant
tensormap_and_ringbuffer
Description
tests/st/a2a3/tensormap_and_ringbuffer/spmd_paged_attention_highperf/test_spmd_paged_attention_highperf.py::TestSpmdPagedAttentionHighPerf::test_run fails on the only non-manual case, b1_h32_kv8_s128_bs128_fp16 (batch=1, num_heads=32, kv_heads=8, head_dim=128, kv_seq=128, block_size=128, fp16).
This is a regression on main: the st-sim-a2a3 and st-onboard-a2a3 jobs both passed at commit 17fa04aa (2026-06-11, run 27324277512) and fail at the current HEAD 19b2c0be. The failure reproduces deterministically on simulation, so it is not hardware flakiness.
The failure presents differently on the two backends:
- a2a3sim: scheduler stall — the workload never completes and the runtime times out.
- a2a3 (onboard): numerical golden mismatch on
out, and the resulting hang/fault poisons the device (cascading 507018 into the next test in the same worker).
Steps to Reproduce
# simulation (deterministic)
pytest tests/st/a2a3/tensormap_and_ringbuffer/spmd_paged_attention_highperf/test_spmd_paged_attention_highperf.py \
--platform a2a3sim --device 0-15 -v
# onboard
python -m pytest tests/st/a2a3/tensormap_and_ringbuffer/spmd_paged_attention_highperf/test_spmd_paged_attention_highperf.py \
--platform a2a3 --device <range> -v
Observed in CI run 27612538258:
Expected Behavior
test_run passes for b1_h32_kv8_s128_bs128_fp16 on both a2a3sim and a2a3, as it did at 17fa04aa.
Actual Behavior
a2a3sim:
E RuntimeError: run_prepared failed with code -100
[ERROR] handle_timeout_exit: [scheduler_cold_path.cpp:378] [STALL thread=1 idle_iterations=324] TIMEOUT_EXIT after_idle_iterations=324
[ERROR] aicpu_execute: PTO2 runtime failed with rc=-100
[ERROR] validate_runtime_impl: PTO2 runtime failed: orch_error_code=0 sched_error_code=100 runtime_status=-100
a2a3 (onboard):
E AssertionError: Golden mismatch on 'out': max_diff=0.39404296875, rtol=0.005, atol=0.02
... (an earlier run showed max_diff=3.859375)
[ERROR] sync_run_streams: aclrtSynchronizeStreamWithTimeout (AICPU) failed: 507018
[ERROR] recover_device_or_mark_unusable: Device unrecoverable after AICore error 507018 ... force-reset the card
Git Commit ID
19b2c0b
Host Platform
Linux (aarch64)
Additional Context
Platform
a2a3 (Ascend 910B/C hardware) and a2a3sim (Ascend 910B/C simulation) — fails on both.
Runtime Variant
tensormap_and_ringbuffer
Description
tests/st/a2a3/tensormap_and_ringbuffer/spmd_paged_attention_highperf/test_spmd_paged_attention_highperf.py::TestSpmdPagedAttentionHighPerf::test_runfails on the only non-manual case,b1_h32_kv8_s128_bs128_fp16(batch=1, num_heads=32, kv_heads=8, head_dim=128, kv_seq=128, block_size=128, fp16).This is a regression on
main: thest-sim-a2a3andst-onboard-a2a3jobs both passed at commit17fa04aa(2026-06-11, run 27324277512) and fail at the current HEAD19b2c0be. The failure reproduces deterministically on simulation, so it is not hardware flakiness.The failure presents differently on the two backends:
out, and the resulting hang/fault poisons the device (cascading 507018 into the next test in the same worker).Steps to Reproduce
Observed in CI run 27612538258:
Expected Behavior
test_runpasses forb1_h32_kv8_s128_bs128_fp16on both a2a3sim and a2a3, as it did at17fa04aa.Actual Behavior
a2a3sim:
a2a3 (onboard):
Git Commit ID
19b2c0b
Host Platform
Linux (aarch64)
Additional Context
17fa04aa(2026-06-11) → failing at19b2c0be. Commits landed in between include Add: per-task ring sizing via CallConfig.runtime_env #1042 (per-task ring sizing), fix(l2_swimlane): unify func_id resolve & gate kernel names behind --enable-dep-gen #1048, feat(dfx): dependency- & MIX-aware scheduler-overhead model #1039, Fix: direction-aware host<->device tensor staging (#1047) #1051 (direction-aware host<->device tensor staging), Update: raise CORE_MAX_TENSOR_ARGS to 32, lower scalars to 16 #1056 (CORE_MAX_TENSOR_ARGS 32 / scalars 16). Not yet bisected — listed only to bound the search.b1_h32_kv8_s128_bs128_fp16case"manual": True(CI runs--manual exclude, so it is skipped on both sim and onboard while still runnable via--manual only). The fix PR must remove that"manual": Trueand confirm the case passes on a2a3sim and a2a3.