Skip to content

[Performance] Benchmark high-perf Paged Attention (#899) vs PA unroll on a2a3 #998

Description

@ChaoZheng109

Platform

a2a3 (Ascend 910B/C hardware)

Runtime Variant

tensormap_and_ringbuffer

Summary

#899 added a high-performance Paged Attention variant
(spmd_paged_attention_highperf). We need to benchmark it against the existing PA unroll
variant (paged_attention_unroll) on a2a3 and determine which is faster (and under which
shapes), so we can decide which to keep / recommend.

Git Commit ID

f257850

Host Platform

Linux (aarch64)

Reproduction

Run both on the same locked device(s) and compare. (Wrap onboard runs in task-submit; gate
the arch with onboard-arch-precheck.)

.claude/skills/onboard-arch-precheck/check.sh a2a3 || exit 1

# High-perf PA (#899)
task-submit --device auto --device-num 1 --run "\
  python tests/st/a2a3/tensormap_and_ringbuffer/spmd_paged_attention_highperf/test_spmd_paged_attention_highperf.py \
    -p a2a3 -d \$TASK_DEVICE --rounds 10 --skip-golden"

# PA unroll (baseline to compare)
task-submit --device auto --device-num 1 --run "\
  python tests/st/a2a3/tensormap_and_ringbuffer/paged_attention_unroll/test_paged_attention_unroll.py \
    -p a2a3 -d \$TASK_DEVICE --rounds 10 --skip-golden"

Expected Performance

Open question — this issue is to produce the numbers. Acceptance: a side-by-side latency
comparison of spmd_paged_attention_highperf vs paged_attention_unroll across the shared
shapes, with a clear verdict on which wins (overall and per-shape if it varies).

Actual Performance

Not yet measured.

Additional Context

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    performancePerformance regression or optimization

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions