Platform
a5 (Ascend 950 hardware)
Runtime Variant
All / Unknown
Summary
scene_test_compile --compile-workers N applies N at the SceneTestCase class level. A single large callable still compiles its orchestration and every incore sequentially, so it can become the batch critical path while most of the configured compile budget is idle.
PR #1833 exposes this with TestQwen314BDecodeHostBuildGraph: the class contains 37 incores plus one orchestration. Its cold build is one serial chain of 38 compiler invocations even though CI passes --compile-workers 8.
This is not a first-run performance regression from #1675: the serial incore loop predates that PR. #1675 correctly moved compilation outside the NPU lock and added class-level parallel warm-up, but the scheduling granularity does not cover a single heavyweight class.
Git Commit ID
87d9d0e69a1e9c93d24ec0835ebd70c58968f447
CANN Version
N/A — the issue is in host-side compilation scheduling and reproduces across runner/compiler versions.
Driver Version
N/A — compilation occurs before device acquisition.
Host Platform
Other (please specify): self-hosted Linux runners on both x86_64 and aarch64.
Reproduction
On commit 87d9d0e69a1e9c93d24ec0835ebd70c58968f447, start from a cold build/cache/kernels entry for the HBG Qwen scene test and run:
.venv/bin/python -m simpler_setup.tools.scene_test_compile \
examples/a5/host_build_graph/qwen3_14b_decode \
--platform a5 --exclude-level 4 --require-pto-isa \
--compile-workers 8 -q
Observe compiler processes or add timing around KernelCompiler.compile_orchestration() and compile_incore():
compile_collected_scene_tests() submits one task per selected class to a ThreadPoolExecutor(max_workers=8).
_compile_chip_callable_from_spec() compiles orchestration first, then iterates through all incores with a serial for loop.
- With only this Qwen class selected, at most one callable compilation advances at a time despite
--compile-workers 8.
The existing unit test test_compile_collected_scene_tests_uses_configured_workers checks only that the executor receives the configured max_workers; it does not exercise multiple compilation units within one class or measure actual compiler concurrency.
Expected Performance
--compile-workers N should act as one global upper bound on independent compiler subprocesses across the whole warm-up, including orchestration/incore units belonging to the same class.
For one callable containing many independent incores, up to N units should be able to compile concurrently, while total compiler concurrency across multiple classes must remain bounded by N rather than multiplying through nested pools.
Actual Performance
PR #1833 CI run 31763338460, job 94654574945 spent 316 seconds in Compile scene-test kernels (a5) on the aarch64 runner a5ci8p-2.
Across the most recent 77 applicable A5 PR jobs inspected when filing this issue:
| Runner |
Architecture |
Jobs |
Average compile step |
Range |
a5-npu-1 |
x86_64 |
47 |
74.7 s |
52–118 s |
a5ci8p-2 |
aarch64 |
30 |
168.1 s |
101–316 s |
Runner speed changes the magnitude, but not the structural limit: one large class cannot consume more than one class worker.
Profiling Data (Optional)
Additional Context
Related: #1604, #1772
#1604 introduced the original persistent-cache/device-lock work and was largely addressed by #1675. #1772 tracks broader CI wall-time work. This issue is narrower: make the configured compilation budget effective for a single large callable.
A suitable implementation should:
- flatten orchestration/incore work into a shared compilation scheduler or otherwise use a shared global permit pool;
- avoid nested
N classes × N incores oversubscription;
- preserve deterministic
CoreCallable child ordering even if binaries finish out of order;
- retain per-callable cache locking and atomic publication behavior;
- preserve the current failure attribution to the owning scene-test class;
- add a test with one class containing multiple compilation units that proves observed concurrency is greater than one and never exceeds
N.
Per-kernel persistent caching and cross-class binary reuse could further reduce the Qwen HBG cold build, but are intentionally not required by this issue.
Platform
a5 (Ascend 950 hardware)
Runtime Variant
All / Unknown
Summary
scene_test_compile --compile-workers NappliesNat theSceneTestCaseclass level. A single large callable still compiles its orchestration and every incore sequentially, so it can become the batch critical path while most of the configured compile budget is idle.PR #1833 exposes this with
TestQwen314BDecodeHostBuildGraph: the class contains 37 incores plus one orchestration. Its cold build is one serial chain of 38 compiler invocations even though CI passes--compile-workers 8.This is not a first-run performance regression from #1675: the serial incore loop predates that PR. #1675 correctly moved compilation outside the NPU lock and added class-level parallel warm-up, but the scheduling granularity does not cover a single heavyweight class.
Git Commit ID
87d9d0e69a1e9c93d24ec0835ebd70c58968f447CANN Version
N/A — the issue is in host-side compilation scheduling and reproduces across runner/compiler versions.
Driver Version
N/A — compilation occurs before device acquisition.
Host Platform
Other (please specify): self-hosted Linux runners on both x86_64 and aarch64.
Reproduction
On commit
87d9d0e69a1e9c93d24ec0835ebd70c58968f447, start from a coldbuild/cache/kernelsentry for the HBG Qwen scene test and run:Observe compiler processes or add timing around
KernelCompiler.compile_orchestration()andcompile_incore():compile_collected_scene_tests()submits one task per selected class to aThreadPoolExecutor(max_workers=8)._compile_chip_callable_from_spec()compiles orchestration first, then iterates through all incores with a serialforloop.--compile-workers 8.The existing unit test
test_compile_collected_scene_tests_uses_configured_workerschecks only that the executor receives the configuredmax_workers; it does not exercise multiple compilation units within one class or measure actual compiler concurrency.Expected Performance
--compile-workers Nshould act as one global upper bound on independent compiler subprocesses across the whole warm-up, including orchestration/incore units belonging to the same class.For one callable containing many independent incores, up to
Nunits should be able to compile concurrently, while total compiler concurrency across multiple classes must remain bounded byNrather than multiplying through nested pools.Actual Performance
PR #1833 CI run 31763338460, job 94654574945 spent 316 seconds in
Compile scene-test kernels (a5)on the aarch64 runnera5ci8p-2.Across the most recent 77 applicable A5 PR jobs inspected when filing this issue:
a5-npu-1a5ci8p-2Runner speed changes the magnitude, but not the structural limit: one large class cannot consume more than one class worker.
Profiling Data (Optional)
examples/a5/host_build_graph/qwen3_14b_decode/test_qwen3_14b_decode.pyreuses a callable containing 37 incores.simpler_setup/tools/scene_test_compile.py: class-granularThreadPoolExecutor.simpler_setup/scene_test.py: serial orchestration followed by serial incore loop.Additional Context
Related: #1604, #1772
#1604 introduced the original persistent-cache/device-lock work and was largely addressed by #1675. #1772 tracks broader CI wall-time work. This issue is narrower: make the configured compilation budget effective for a single large callable.
A suitable implementation should:
N classes × N incoresoversubscription;CoreCallablechild ordering even if binaries finish out of order;N.Per-kernel persistent caching and cross-class binary reuse could further reduce the Qwen HBG cold build, but are intentionally not required by this issue.