Skip to content

[Performance] Scene-test warm-up leaves large callables serial despite --compile-workers #1834

Description

@doraemonmj

Platform

a5 (Ascend 950 hardware)

Runtime Variant

All / Unknown

Summary

scene_test_compile --compile-workers N applies N at the SceneTestCase class level. A single large callable still compiles its orchestration and every incore sequentially, so it can become the batch critical path while most of the configured compile budget is idle.

PR #1833 exposes this with TestQwen314BDecodeHostBuildGraph: the class contains 37 incores plus one orchestration. Its cold build is one serial chain of 38 compiler invocations even though CI passes --compile-workers 8.

This is not a first-run performance regression from #1675: the serial incore loop predates that PR. #1675 correctly moved compilation outside the NPU lock and added class-level parallel warm-up, but the scheduling granularity does not cover a single heavyweight class.

Git Commit ID

87d9d0e69a1e9c93d24ec0835ebd70c58968f447

CANN Version

N/A — the issue is in host-side compilation scheduling and reproduces across runner/compiler versions.

Driver Version

N/A — compilation occurs before device acquisition.

Host Platform

Other (please specify): self-hosted Linux runners on both x86_64 and aarch64.

Reproduction

On commit 87d9d0e69a1e9c93d24ec0835ebd70c58968f447, start from a cold build/cache/kernels entry for the HBG Qwen scene test and run:

.venv/bin/python -m simpler_setup.tools.scene_test_compile \
  examples/a5/host_build_graph/qwen3_14b_decode \
  --platform a5 --exclude-level 4 --require-pto-isa \
  --compile-workers 8 -q

Observe compiler processes or add timing around KernelCompiler.compile_orchestration() and compile_incore():

  • compile_collected_scene_tests() submits one task per selected class to a ThreadPoolExecutor(max_workers=8).
  • _compile_chip_callable_from_spec() compiles orchestration first, then iterates through all incores with a serial for loop.
  • With only this Qwen class selected, at most one callable compilation advances at a time despite --compile-workers 8.

The existing unit test test_compile_collected_scene_tests_uses_configured_workers checks only that the executor receives the configured max_workers; it does not exercise multiple compilation units within one class or measure actual compiler concurrency.

Expected Performance

--compile-workers N should act as one global upper bound on independent compiler subprocesses across the whole warm-up, including orchestration/incore units belonging to the same class.

For one callable containing many independent incores, up to N units should be able to compile concurrently, while total compiler concurrency across multiple classes must remain bounded by N rather than multiplying through nested pools.

Actual Performance

PR #1833 CI run 31763338460, job 94654574945 spent 316 seconds in Compile scene-test kernels (a5) on the aarch64 runner a5ci8p-2.

Across the most recent 77 applicable A5 PR jobs inspected when filing this issue:

Runner Architecture Jobs Average compile step Range
a5-npu-1 x86_64 47 74.7 s 52–118 s
a5ci8p-2 aarch64 30 168.1 s 101–316 s

Runner speed changes the magnitude, but not the structural limit: one large class cannot consume more than one class worker.

Profiling Data (Optional)

Additional Context

Related: #1604, #1772

#1604 introduced the original persistent-cache/device-lock work and was largely addressed by #1675. #1772 tracks broader CI wall-time work. This issue is narrower: make the configured compilation budget effective for a single large callable.

A suitable implementation should:

  • flatten orchestration/incore work into a shared compilation scheduler or otherwise use a shared global permit pool;
  • avoid nested N classes × N incores oversubscription;
  • preserve deterministic CoreCallable child ordering even if binaries finish out of order;
  • retain per-callable cache locking and atomic publication behavior;
  • preserve the current failure attribution to the owning scene-test class;
  • add a test with one class containing multiple compilation units that proves observed concurrency is greater than one and never exceeds N.

Per-kernel persistent caching and cross-class binary reuse could further reduce the Qwen HBG cold build, but are intentionally not required by this issue.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    performancePerformance regression or optimization

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions