Skip to content

Fix: size CoreTracker by the device's clusters, not by a uint64_t - #1477

Merged
ChaoWao merged 1 commit into
hw-native-sys:mainfrom
ChaoWao:fix-coretracker-cluster-capacity
Jul 26, 2026
Merged

ChaoWao merged 1 commit into
hw-native-sys:mainfrom
ChaoWao:fix-coretracker-cluster-capacity

Conversation

@ChaoWao

@ChaoWao ChaoWao commented Jul 25, 2026

Copy link
Copy Markdown
Collaborator

The bug

CoreTracker packs three state bits per cluster (AIC, AIV0, AIV1) into a single
uint64_t, so it holds floor(64/3) = 21 clusters. The two literals that
encoded that came from the width of the backing word, not from any device:

static inline int32_t MAX_CORE_PER_THREAD = 63;
static constexpr int32_t MAX_CLUSTERS = 63 / 3;
uint64_t states_{0};
int32_t core_id_map_[63];  // bit_position -> worker_id, max 21 clusters * 3

a2a3 has 24 clusters and a5 has 36. Neither fits.

assign_cores_to_threads() checks the ceiling (scheduler_cold_path.cpp:990).
assign_own_clusters() — the barrier-free path the decoupled scheduler actually
takes — does not. Give one scheduler thread more than 21 clusters and
set_cluster() writes past core_id_map_ into the next CoreTracker, which on
that path is the orchestrator's.

At 24 clusters the store lands exactly on core_trackers_[1].cluster_count_
(sizeof(CoreTracker)==320, offsetof(core_id_map_)==40, so
core_id_map_[70] → byte 320). The orchestrator's tracker is never assigned and
its shutdown() documents the assumption it breaks:

// Orchestrator threads have core_trackers_[thread_idx].core_num() == 0 -> no-op.

It instead reports 210 cores, walks a zeroed id map, and sends AICORE_EXIT to
core 0 two hundred times. If the scheduler still has a task dispatched there it
polls that core's COND for a FIN that can never arrive:

Thread 1: Orchestrator completed
Thread 1: Shutting down 210 cores          <-- must be 0
[STALL thread=0] CLUSTER cluster_id=0 aic=core0(busy ... ANOMALY cond_tok=2147483646 ...)

cond_tok=2147483646 = 0x7FFFFFFE = AICORE_EXIT_TASK_ID. Surfaces as
sched_error_code=100 SCHEDULER_TIMEOUT → host 507018. No deadlock or
capacity detector fires — counts of Task Allocator Deadlock, SPIN
Timeout (N cycles) and HandleTaskTimeout are all 0 — which is why this reads
as a mystery stall rather than an overflow.

Bisected onboard by clamping the resolved width:

clusters result
21 pass
22 pass
23 pass
24 fail

Exactly the byte-offset prediction: only cluster_idx == 23 reaches
core_id_map_[70]. 22 and 23 are already corrupt — the bitmask aliases once
past 63 bits (1ULL << 66 becomes LSL #2 on aarch64) — the write just lands
in tail padding and survives.

Nothing reaches this today because every scene test pins block_dim low enough
to stay under 21 clusters per thread. The bound is a property of the hardware,
not of the test suite, so any caller passing aicpu_thread_num: 2 on a
24-cluster device hits it.

The fix

  • MAX_CLUSTERS derives from PLATFORM_MAX_BLOCKDIM and
    MAX_CORE_PER_THREAD from it, so one scheduler thread can own a whole device
    on either arch. A static_assert ties the pair to the storage width.
    (MAX_CORE_PER_THREAD was also a mutable static inline used as a constant —
    now constexpr.)
  • BitStates is backed by unsigned __int128 (a2a3 needs 72 bits, a5 108).
    The hot path shifts the whole mask by 1 and 2 to test cluster co-residency
    (scheduler_types.h:260/262/427/432); letting the compiler generate those
    crossings is what keeps the widening honest. Only popcount and ctz split by
    hand. __uint128_t is already used on this target in
    src/common/platform/include/aicpu/device_time.h.
  • Callers can no longer build a mask with 1ULL << offset, undefined once
    offset reaches 64 — there were 13 such sites and the widening would have made
    every one of them UB. BitStates::bit() owns the shift.
  • assign_own_clusters() gets the guard its serial sibling already had
    (a2a3 + a5), and init() / set_cluster() assert the bound, so exceeding it
    fails by name instead of corrupting a neighbour.

Verification

Repro before the fix: dummy_task, predicated_dispatch (tmr and hbg),
dep_gen_chain at aicpu_thread_num: 2 with the width auto-resolved to 24 —
100% reproducible. After: 5/5 pass, and Shutting down 210 cores is gone from
the device log.

Suite Result
pytest examples tests/st --platform a2a3 (onboard) 60 passed, 0 failed
pytest examples tests/st --platform a2a3sim 51 passed, 0 failed
pytest examples tests/st --platform a5sim 39 passed, 0 failed
pytest tests/ut -m "not requires_hardware" 813 passed
ctest C++ unit tests 60/60
pre-commit (all hooks) passed

hbg reaches the same ceiling through the guarded path, so it fails cleanly today
(Can't assign more then 64 cores in per scheduler) rather than corrupting —
the widening lifts that limit too.

@coderabbitai

coderabbitai Bot commented Jul 25, 2026 •

Copy link
Copy Markdown

Review Change Stack

Important

Review skipped

Auto incremental reviews are disabled on this repository.

Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro Plus

Run ID: f45d96a5-adee-46f4-818c-c0dab8d6a387

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review
📝 Walkthrough

Walkthrough

CoreTracker capacity now derives from platform limits and uses 128-bit bitmasks across the a2a3 and a5 schedulers. Bit operations and MIX placement masks were updated, while cluster ownership now rejects counts exceeding tracker capacity.

Changes

CoreTracker capacity and scheduling safety

Layer / File(s) Summary
Capacity and BitStates foundation
src/a2a3/.../scheduler/scheduler_types.h, src/a2a3/.../scheduler/scheduler_types.h, src/a5/.../scheduler/scheduler_types.h
Platform-derived limits, 128-bit BitStates storage and operations, and resized core_id_map_ arrays are introduced.
CoreTracker and MIX mask integration
src/a2a3/.../scheduler/scheduler_types.h, src/a5/.../scheduler/scheduler_types.h
Idle-state, pending-state, cluster classification, and MIX mask construction use BitStates::bit() and widened-mask operations.
Cluster ownership capacity guard
src/a2a3/.../scheduler/scheduler_cold_path.cpp, src/a5/.../scheduler/scheduler_cold_path.cpp
assign_own_clusters rejects oversized ownership counts before tracker initialization, records handshake failure, and returns.

Estimated code review effort: 4 (Complex) | ~45 minutes

Poem

I’m a bunny with bits in a bright new array,
Hopping past sixty-four without losing my way.
Wide masks now gather each core in their fold,
While guards stop clusters from spilling uncontrolled.
MIX finds its pathways, neat, safe, and bold!

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Title check ✅ Passed The title accurately summarizes the main change: expanding CoreTracker capacity beyond a uint64_t-sized limit.
Description check ✅ Passed The description clearly matches the implemented fix and its motivation, including the wider bitmask and added bounds checks.
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In
`@src/a2a3/runtime/tensormap_and_ringbuffer/runtime/scheduler/scheduler_types.h`:
- Around line 199-213: The CoreTracker validation accepts negative cluster
counts and indices. Update CoreTracker::init and CoreTracker::set_cluster in
src/a2a3/runtime/tensormap_and_ringbuffer/runtime/scheduler/scheduler_types.h:199-213,
src/a2a3/runtime/host_build_graph/runtime/scheduler/scheduler_types.h:210-224,
and
src/a5/runtime/tensormap_and_ringbuffer/runtime/scheduler/scheduler_types.h:200-214
to require cluster_count and cluster_idx to be non-negative as well as below
MAX_CLUSTERS, preserving the existing assertions and initialization behavior for
valid values.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro Plus

Run ID: 579c2387-d4c6-4da1-a3d5-7107d5b2c4c0

📥 Commits

Reviewing files that changed from the base of the PR and between f8e2067 and b21b394.

📒 Files selected for processing (5)
  • src/a2a3/runtime/host_build_graph/runtime/scheduler/scheduler_types.h
  • src/a2a3/runtime/tensormap_and_ringbuffer/runtime/scheduler/scheduler_cold_path.cpp
  • src/a2a3/runtime/tensormap_and_ringbuffer/runtime/scheduler/scheduler_types.h
  • src/a5/runtime/tensormap_and_ringbuffer/runtime/scheduler/scheduler_cold_path.cpp
  • src/a5/runtime/tensormap_and_ringbuffer/runtime/scheduler/scheduler_types.h

Comment thread src/a2a3/runtime/tensormap_and_ringbuffer/runtime/scheduler/scheduler_types.h Outdated
CoreTracker packs three state bits per cluster (AIC, AIV0, AIV1) into a
single uint64_t, so it held floor(64/3) = 21 clusters — and the two
literals that encoded that, MAX_CORE_PER_THREAD = 63 and
core_id_map_[63], came from the width of the backing word rather than
from any device. a2a3 has 24 clusters and a5 has 36, so neither fits.

assign_cores_to_threads() checks the ceiling; assign_own_clusters(), the
barrier-free path the decoupled scheduler actually takes, does not. A
run that gives one scheduler thread more than 21 clusters therefore
writes past core_id_map_ into the next CoreTracker, which on that path
is the orchestrator's. For 24 clusters the store lands exactly on
core_trackers_[1].cluster_count_, so the orchestrator — whose tracker is
never assigned and whose shutdown() documents "core_num() == 0 -> no-op"
— reports 210 cores, walks a zeroed id map, and sends AICORE_EXIT to
core 0 two hundred times. If the scheduler still has a task dispatched
there it polls that core's COND for a FIN that can no longer arrive, and
the run dies on the scheduler's forward-progress timeout.

Nothing reaches this today because every scene test pins block_dim low
enough to stay under 21 clusters per thread, but the bound is a property
of the hardware, not of the test suite.

- MAX_CLUSTERS now derives from PLATFORM_MAX_BLOCKDIM and
  MAX_CORE_PER_THREAD from it, so one scheduler thread can own a whole
  device on either arch. A static_assert ties the pair to the storage.
- BitStates is backed by unsigned __int128 (a2a3 needs 72 bits, a5 108).
  The hot path shifts the whole mask by 1 and 2 to test cluster
  co-residency, and letting the compiler generate those crossings is
  what keeps the widening honest; only popcount and ctz split by hand.
- Callers can no longer build a mask with `1ULL << offset`, undefined
  once offset reaches 64. BitStates::bit() owns the shift.
- assign_own_clusters() gets the guard its serial sibling already had,
  and init()/set_cluster() assert the bound at both ends, so exceeding it
  fails by name instead of corrupting a neighbour. The lower bound
  matters as much as the upper one: a negative cluster_idx would index
  core_id_map_ backwards, and a negative cluster_count would leave every
  dispatch loop visiting no clusters at all.
@ChaoWao
ChaoWao merged commit ae2611d into hw-native-sys:main Jul 26, 2026
15 of 16 checks passed
@ChaoWao
ChaoWao deleted the fix-coretracker-cluster-capacity branch July 26, 2026 01:21
ChaoWao added a commit to ChaoWao/simpler-fork that referenced this pull request Jul 27, 2026
The guide covered capacity codes and stalls but had nothing for an AICore
addressing fault, which is the third way a run reaches the same generic
507018. hw-native-sys#1489 spent an investigation re-deriving the procedure from
scratch and still could not close, so record it.

Three steps, ordered so the cheap disqualifications come first:

- F1 separates "the kernel computed a bad address" from "the core was
  made to execute something that is not the kernel". `binSize` in
  GetBinAndKernelNameExceptionArgs matching the runtime's own
  aicore_kernel.o means the report is naming the polling-dispatch
  executor, so the fault is a dispatch-payload problem and no amount of
  reading the kernel's arithmetic will find it. hw-native-sys#1036 is the worked
  example.
- F2 rules the kernel in or out statically. Constant TASSIGN bases plus
  template tile extents cannot produce an out-of-range address, and
  runtime values that only shrink a tile move nothing — so summing the
  highest byte reached against the UB size settles it without
  instrumenting anything. This is what cleared the allreduce collectives
  in hw-native-sys#1489.
- F3 separates a real fault from a post-mortem register dump of a core
  reaped mid-spin, by counting which detector actually fired.

hw-native-sys#1489 has since been closed as very likely fixed by hw-native-sys#1477 — a CoreTracker
whose uint64_t state overran past 21 clusters while those cases ran at the
device's full 24. F2's verdict held: the kernels were never at fault, and
the answer was in F1's second case all along. The section says so, because
a worked example that names its own outcome is worth more than one that
trails off.

Also state where the device log has to be redirected and why the harness
will not do it: outputs/<case>_<ts>/ exists only when a DFX flag is on, so
a plain onboard run leaves its log in the shared ~/ascend/log/debug/ with
every other user's, and the evidence is unattributable once the run ends.
That is how hw-native-sys#1489's only real occurrence lost the one artifact that would
have decided F1.

Renumber the trailing section and fix the in-page link that pointed at
its old number.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
ChaoWao added a commit to ChaoWao/simpler-fork that referenced this pull request Jul 27, 2026
The guide covered capacity codes and stalls but had nothing for an AICore
addressing fault, which is the third way a run reaches the same generic
507018. hw-native-sys#1489 spent an investigation re-deriving the procedure from
scratch and still could not close, so record it.

Three steps, ordered so the cheap disqualifications come first:

- F1 separates "the kernel computed a bad address" from "the core was
  made to execute something that is not the kernel". `binSize` in
  GetBinAndKernelNameExceptionArgs matching the runtime's own
  aicore_kernel.o means the report is naming the polling-dispatch
  executor, so the fault is a dispatch-payload problem and no amount of
  reading the kernel's arithmetic will find it. hw-native-sys#1036 is the worked
  example.
- F2 rules the kernel in or out statically. Constant TASSIGN bases plus
  template tile extents do not make an out-of-range address impossible —
  the constants can overrun on their own — but they make it decidable on
  paper, because runtime values that only shrink a tile move no base and
  grow no extent. Summing the highest byte reached and comparing against
  the UB size settles it either way without touching the device. This is
  what cleared the allreduce collectives in hw-native-sys#1489.
- F3 separates a real fault from a post-mortem register dump of a core
  reaped mid-spin, by counting which detector actually fired.

hw-native-sys#1489 has since been closed as very likely fixed by hw-native-sys#1477 — a CoreTracker
whose uint64_t state overran past 21 clusters while those cases ran at the
device's full 24. F2's verdict held: the kernels were never at fault, and
the answer was in F1's second case all along. The section says so, because
a worked example that names its own outcome is worth more than one that
trails off.

Also state where the device log has to be redirected and why the harness
will not do it: outputs/<case>_<ts>/ exists only when a DFX flag is on, so
a plain onboard run leaves its log in the shared ~/ascend/log/debug/ with
every other user's, and the evidence is unattributable once the run ends.
That is how hw-native-sys#1489's only real occurrence lost the one artifact that would
have decided F1.

Renumber the trailing section and fix the in-page link that pointed at
its old number.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
ChaoWao added a commit that referenced this pull request Jul 27, 2026
The guide covered capacity codes and stalls but had nothing for an AICore
addressing fault, which is the third way a run reaches the same generic
507018. #1489 spent an investigation re-deriving the procedure from
scratch and still could not close, so record it.

Three steps, ordered so the cheap disqualifications come first:

- F1 separates "the kernel computed a bad address" from "the core was
  made to execute something that is not the kernel". `binSize` in
  GetBinAndKernelNameExceptionArgs matching the runtime's own
  aicore_kernel.o means the report is naming the polling-dispatch
  executor, so the fault is a dispatch-payload problem and no amount of
  reading the kernel's arithmetic will find it. #1036 is the worked
  example.
- F2 rules the kernel in or out statically. Constant TASSIGN bases plus
  template tile extents do not make an out-of-range address impossible —
  the constants can overrun on their own — but they make it decidable on
  paper, because runtime values that only shrink a tile move no base and
  grow no extent. Summing the highest byte reached and comparing against
  the UB size settles it either way without touching the device. This is
  what cleared the allreduce collectives in #1489.
- F3 separates a real fault from a post-mortem register dump of a core
  reaped mid-spin, by counting which detector actually fired.

#1489 has since been closed as very likely fixed by #1477 — a CoreTracker
whose uint64_t state overran past 21 clusters while those cases ran at the
device's full 24. F2's verdict held: the kernels were never at fault, and
the answer was in F1's second case all along. The section says so, because
a worked example that names its own outcome is worth more than one that
trails off.

Also state where the device log has to be redirected and why the harness
will not do it: outputs/<case>_<ts>/ exists only when a DFX flag is on, so
a plain onboard run leaves its log in the shared ~/ascend/log/debug/ with
every other user's, and the evidence is unattributable once the run ends.
That is how #1489's only real occurrence lost the one artifact that would
have decided F1.

Renumber the trailing section and fix the in-page link that pointed at
its old number.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant