Platform
a2a3 (Ascend 910B/C hardware)
Runtime Variant
tensormap_and_ringbuffer
Summary
#899 added a high-performance Paged Attention variant
(spmd_paged_attention_highperf). We need to benchmark it against the existing PA unroll
variant (paged_attention_unroll) on a2a3 and determine which is faster (and under which
shapes), so we can decide which to keep / recommend.
Git Commit ID
f257850
Host Platform
Linux (aarch64)
Reproduction
Run both on the same locked device(s) and compare. (Wrap onboard runs in task-submit; gate
the arch with onboard-arch-precheck.)
.claude/skills/onboard-arch-precheck/check.sh a2a3 || exit 1
# High-perf PA (#899)
task-submit --device auto --device-num 1 --run "\
python tests/st/a2a3/tensormap_and_ringbuffer/spmd_paged_attention_highperf/test_spmd_paged_attention_highperf.py \
-p a2a3 -d \$TASK_DEVICE --rounds 10 --skip-golden"
# PA unroll (baseline to compare)
task-submit --device auto --device-num 1 --run "\
python tests/st/a2a3/tensormap_and_ringbuffer/paged_attention_unroll/test_paged_attention_unroll.py \
-p a2a3 -d \$TASK_DEVICE --rounds 10 --skip-golden"
Expected Performance
Open question — this issue is to produce the numbers. Acceptance: a side-by-side latency
comparison of spmd_paged_attention_highperf vs paged_attention_unroll across the shared
shapes, with a clear verdict on which wins (overall and per-shape if it varies).
Actual Performance
Not yet measured.
Additional Context
Platform
a2a3 (Ascend 910B/C hardware)
Runtime Variant
tensormap_and_ringbuffer
Summary
#899 added a high-performance Paged Attention variant
(
spmd_paged_attention_highperf). We need to benchmark it against the existing PA unrollvariant (
paged_attention_unroll) on a2a3 and determine which is faster (and under whichshapes), so we can decide which to keep / recommend.
Git Commit ID
f257850
Host Platform
Linux (aarch64)
Reproduction
Run both on the same locked device(s) and compare. (Wrap onboard runs in
task-submit; gatethe arch with
onboard-arch-precheck.)Expected Performance
Open question — this issue is to produce the numbers. Acceptance: a side-by-side latency
comparison of
spmd_paged_attention_highperfvspaged_attention_unrollacross the sharedshapes, with a clear verdict on which wins (overall and per-shape if it varies).
Actual Performance
Not yet measured.
Additional Context
tests/st/a2a3/tensormap_and_ringbuffer/spmd_paged_attention_highperf/(from High performance Paged Attention A2A3 ST Test #899)tests/st/a2a3/tensormap_and_ringbuffer/paged_attention_unroll/(also
paged_attention_unroll_4dims)