Skip to content

Refactor: remove the user-facing block_dim knob - #1309

Merged
ChaoWao merged 1 commit into
hw-native-sys:mainfrom
doraemonmj:blockdim-auto
Jul 26, 2026
Merged

ChaoWao merged 1 commit into
hw-native-sys:mainfrom
doraemonmj:blockdim-auto

Conversation

@doraemonmj

@doraemonmj doraemonmj commented Jul 9, 2026 •

Copy link
Copy Markdown
Contributor

Stacked on #1478. Its commit is currently in this branch and drops out
once it merges. #1477 (CoreTracker capacity) has landed, which is what lets
the four cases keeping aicpu_thread_num: 2 run at the full 24-cluster
width.

What

A run now always takes the whole device. block_dim was a per-call knob that
every scene test pinned to the platform maximum anyway, and it forced host
orchestration to guess a width before the device had reported one.

The knob, end to end

  • CallConfig loses the field. The packed wire layout is one int32 shorter,
    so the remote-L3 PROTOCOL_VERSION goes 1 → 2 in both codecs and the fork
    mailbox format string drops an i. The compile-time layout guard in
    call_config.h is updated with it — it is what caught the drift.
  • DeviceRunner resolves the width unconditionally: every cluster the
    AICore stream reports onboard, SIM_AUTO_BLOCKDIM on sim.
  • The nanobind binding, scene_test's config plumbing, and the l0_swimlane
    replay lose their block_dim paths.

Two latent problems surfaced while doing it and are fixed here:

  • remote_l3_session.py hardcoded protocol_version=1 in the HELLO payload
    instead of referencing the constant — a version bump would have silently
    missed it.
  • spmd_basic's Case2_AutoBlockDim became an exact duplicate of Case1 once
    the knob was gone, so it is dropped.

Scene tests

178 pinned values across 75 files removed. Kernel-side params["block_dim"] — a
tiling argument, not a runtime knob — is untouched.

l0_swimlane can no longer infer an SPMD replay width from the test file, since
a cohort now sizes itself on device: --spmd-block-num is required to replay
one, and the help/doc say so.

Coverage note

Sim resolves to SIM_AUTO_BLOCKDIM (8) and onboard to the device's real width
(24 on a2a3), so sim no longer exercises wide-device behaviour. That is a
deliberate trade — sim runs one OS thread per AICore, and 24-36 clusters is
72-108 threads per case — but it is why #1477's bug was invisible until an
onboard run. Wide-device coverage rests on the onboard job.

Verification

Measured on top of #1478 (#1477 is now in main):

Suite Result
pytest examples tests/st --platform a2a3 (onboard) 60 passed, 0 failed
pytest examples tests/st --platform a2a3sim 51 passed, 0 failed
pytest examples tests/st --platform a5sim 39 passed, 0 failed
pytest tests/ut -m "not requires_hardware" 813 passed
pre-commit (all hooks) passed

The four cases that keep aicpu_thread_num: 2 (dummy_task,
predicated_dispatch on both runtimes, dep_gen_chain) pass at the full
24-cluster width with no thread-count workaround, which is the check that #1477
actually fixed the cause rather than moving it.

History

This PR started as "zero out the hardcoded block_dim". That premise stopped
holding when #1458 landed SIM_AUTO_BLOCKDIM = 8: on sim the auto sentinel no
longer resolves to PLATFORM_MAX_BLOCKDIM, so zeroing the pins is a behaviour
change rather than a no-op. Removing the knob outright makes "a run takes the
whole device" the only shape there is. @doraemonmj's cohort-width analysis lives
on in #1478.

Cost: the under-count check from #1472

#1472 added a Pinned case to the three available_aicore_counts tests that
fixed block_dim so the host held an expectation independent of what the device
reported — the only thing able to catch an under-reported count. That handle
is exactly the knob this PR removes, so the case goes with it.

What survives: the count is range-checked, aiv == 2 * clusters is checked, and
the cohort is actually spent, so an over-reported count still trips the
require_sync_start deadlock guard on device. What is lost: an under-reported
count is now self-consistent — fewer blocks launch, fewer slots are expected,
the tail is zero on both sides. The test docstrings say so rather than implying
coverage that is not there.

Restoring it needs a host-side handle on the device width that is not a second
source of truth for a device property. There is not one today.

@gemini-code-assist

Copy link
Copy Markdown

Warning

Gemini encountered an error creating the review. You can try again by commenting /gemini review.

@coderabbitai

coderabbitai Bot commented Jul 9, 2026 •

Copy link
Copy Markdown

Review Change Stack

Important

Review skipped

Auto incremental reviews are disabled on this repository.

Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro Plus

Run ID: 64a7bc4f-1430-40d5-a7f2-d42a146bcb95

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review
📝 Walkthrough

Walkthrough

This change updates block_dim values in test case configuration dictionaries across numerous example and system test files under tensormap_and_ringbuffer, host_build_graph, and spmd_* suites. All affected cases previously set block_dim to 24 or 36 and now set it to 0, with other configuration fields unchanged.

Changes

block_dim configuration reset

Layer / File(s) Summary
Example test suites
examples/a2a3/tensormap_and_ringbuffer/benchmark_bgemm/test_benchmark_bgemm.py, examples/a2a3/tensormap_and_ringbuffer/paged_attention*/test_paged_attention*.py, examples/a5/tensormap_and_ringbuffer/paged_attention_unroll_manual_scope/test_paged_attention_unroll.py
Multiple CASES entries across bgemm and paged-attention example tests change config.block_dim from 24 (or 36 in a5) to 0.
Host build graph tests
tests/st/a2a3/host_build_graph/bgemm/test_bgemm.py, tests/st/a2a3/host_build_graph/paged_attention/test_paged_attention.py
config.block_dim changes from 24 to 0 for host build graph bgemm and paged-attention cases.
ST tensormap_and_ringbuffer core tests
tests/st/a2a3/tensormap_and_ringbuffer/alternating_matmul_add/..., batch_paged_attention/..., fanin_lookup_perf/..., multi_round_paged_attention/..., paged_attention_unroll/..., paged_attention_unroll_4dims/..., tests/st/a5/tensormap_and_ringbuffer/paged_attention_unroll/test_paged_attention_unroll.py
CASES entries across these core ST test suites change block_dim from 24 (or 36 in a5) to 0.
ST spmd test suites
tests/st/a2a3/tensormap_and_ringbuffer/spmd_basic/..., spmd_batch_dispatch_oob/..., spmd_multiblock_aiv/..., spmd_multiblock_mix/..., spmd_paged_attention/..., spmd_paged_attention_highperf/..., spmd_starvation/..., spmd_sync_start*/...
CASES entries across SPMD test suites change config.block_dim from 24 to 0.

Estimated code review effort: 1 (Trivial) | ~5 minutes

Possibly related PRs

  • hw-native-sys/simpler#899: Related to the SPMD paged-attention highperf suite where block_dim is read by orchestration logic to set launch block counts.
  • hw-native-sys/simpler#1076: Both PRs edit the same TestSpmdPagedAttentionHighPerf.CASES configuration for overlapping highperf paged-attention test cases.
  • hw-native-sys/simpler#1265: Both PRs touch tests/st/a2a3/tensormap_and_ringbuffer/spmd_paged_attention/test_spmd_paged_attention.py, one adjusting pytestmark, the other config.block_dim.

Poem

A hop, a skip, through configs I dash,
Twenty-four to zero, no more block clash 🐇
From bgemm to spmd, cases align,
Every block_dim now reads a clean zero line.
Thump-thump goes my heart, tests pass with glee!

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Title check ✅ Passed The title matches the PR’s main theme: removing the hardcoded user-facing block_dim setting in favor of automatic sizing.
Description check ✅ Passed The description is clearly about the same block_dim auto-sizing change and its related runtime/test updates.

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🧹 Nitpick comments (1)
tests/st/a2a3/tensormap_and_ringbuffer/spmd_basic/test_spmd_basic.py (1)

74-86: 📐 Maintainability & Code Quality | 🔵 Trivial | 💤 Low value

Case1 and Case2_AutoBlockDim are now functionally equivalent.

Case1 explicitly sets block_dim: 0 and Case2_AutoBlockDim omits block_dim (defaulting to 0 via _build_config). Both resolve to the same CallConfig, making the two cases redundant in terms of runtime configuration. If the intent is to test both the explicit-0 and omitted paths as distinct regression cases, consider adding a comment to Case1 clarifying that distinction. Otherwise, Case2_AutoBlockDim could be removed.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@tests/st/a2a3/tensormap_and_ringbuffer/spmd_basic/test_spmd_basic.py` around
lines 74 - 86, Case1 and Case2_AutoBlockDim are redundant because both end up
using the same CallConfig through block_dim=0. Update the table entry in
test_spmd_basic to either remove Case2_AutoBlockDim or add a clear comment near
Case1 and Case2 explaining that one exercises the explicit block_dim=0 path and
the other exercises the omitted-field path via _build_config, so the distinction
is intentional.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Nitpick comments:
In `@tests/st/a2a3/tensormap_and_ringbuffer/spmd_basic/test_spmd_basic.py`:
- Around line 74-86: Case1 and Case2_AutoBlockDim are redundant because both end
up using the same CallConfig through block_dim=0. Update the table entry in
test_spmd_basic to either remove Case2_AutoBlockDim or add a clear comment near
Case1 and Case2 explaining that one exercises the explicit block_dim=0 path and
the other exercises the omitted-field path via _build_config, so the distinction
is intentional.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: 9a577d69-2578-4554-831e-3134c73b3635

📥 Commits

Reviewing files that changed from the base of the PR and between 29595b1 and 06417ee.

📒 Files selected for processing (26)
  • examples/a2a3/tensormap_and_ringbuffer/benchmark_bgemm/test_benchmark_bgemm.py
  • examples/a2a3/tensormap_and_ringbuffer/paged_attention/test_paged_attention.py
  • examples/a2a3/tensormap_and_ringbuffer/paged_attention_manual_scope/test_paged_attention.py
  • examples/a2a3/tensormap_and_ringbuffer/paged_attention_ringbuffer/test_paged_attention_ringbuffer.py
  • examples/a2a3/tensormap_and_ringbuffer/paged_attention_unroll_manual_scope/test_paged_attention_unroll.py
  • examples/a5/tensormap_and_ringbuffer/paged_attention_unroll_manual_scope/test_paged_attention_unroll.py
  • tests/st/a2a3/host_build_graph/bgemm/test_bgemm.py
  • tests/st/a2a3/host_build_graph/paged_attention/test_paged_attention.py
  • tests/st/a2a3/tensormap_and_ringbuffer/alternating_matmul_add/test_alternating_matmul_add.py
  • tests/st/a2a3/tensormap_and_ringbuffer/batch_paged_attention/test_batch_paged_attention.py
  • tests/st/a2a3/tensormap_and_ringbuffer/fanin_lookup_perf/test_fanin_lookup_perf.py
  • tests/st/a2a3/tensormap_and_ringbuffer/multi_round_paged_attention/test_multi_round_paged_attention.py
  • tests/st/a2a3/tensormap_and_ringbuffer/paged_attention_unroll/test_paged_attention_unroll.py
  • tests/st/a2a3/tensormap_and_ringbuffer/paged_attention_unroll_4dims/test_paged_attention_unroll_4dims.py
  • tests/st/a2a3/tensormap_and_ringbuffer/spmd_basic/test_spmd_basic.py
  • tests/st/a2a3/tensormap_and_ringbuffer/spmd_batch_dispatch_oob/test_spmd_batch_dispatch_oob.py
  • tests/st/a2a3/tensormap_and_ringbuffer/spmd_multiblock_aiv/test_spmd_multiblock_aiv.py
  • tests/st/a2a3/tensormap_and_ringbuffer/spmd_multiblock_mix/test_spmd_multiblock_mix.py
  • tests/st/a2a3/tensormap_and_ringbuffer/spmd_paged_attention/test_spmd_paged_attention.py
  • tests/st/a2a3/tensormap_and_ringbuffer/spmd_paged_attention_highperf/test_spmd_paged_attention_highperf.py
  • tests/st/a2a3/tensormap_and_ringbuffer/spmd_starvation/test_spmd_starvation.py
  • tests/st/a2a3/tensormap_and_ringbuffer/spmd_sync_start/test_spmd_sync_start.py
  • tests/st/a2a3/tensormap_and_ringbuffer/spmd_sync_start_aiv/test_spmd_sync_start_aiv.py
  • tests/st/a2a3/tensormap_and_ringbuffer/spmd_sync_start_edge/test_spmd_sync_start_edge.py
  • tests/st/a2a3/tensormap_and_ringbuffer/spmd_sync_start_stress/test_spmd_sync_start_stress.py
  • tests/st/a5/tensormap_and_ringbuffer/paged_attention_unroll/test_paged_attention_unroll.py

@ChaoWao

ChaoWao commented Jul 25, 2026

Copy link
Copy Markdown
Collaborator

Heads-up: #1458 landed and it invalidates this PR's core premise on sim.

This PR's rationale is that block_dim=0 already resolves to the literal being
removed:

Sim: uses PLATFORM_MAX_BLOCKDIM directly (a2a3=24, a5=36).

That is no longer true. #1458 added SIM_AUTO_BLOCKDIM = 8
(src/common/platform/sim/host/device_runner_base.h): sim now resolves the
auto sentinel to 8, not 24/36, because sim runs one OS thread per AICore and
taking the whole modelled chip costs 72-108 threads per case. An explicit
block_dim is still honoured up to PLATFORM_MAX_BLOCKDIM, so this PR's
zeroing is the exact operation that flips behaviour.

So on sim this stops being a no-op cleanup, and at least two cases break
outright rather than merely running narrower:

Case Submits After zeroing (sim)
spmd_sync_start (pins 24) block_num=12, require_sync_start 12 > limit 8 → REQUIRE_SYNC_START_INVALID
spmd_sync_start_edge (pins 24) block_num=23, require_sync_start 23 > limit 8 → same

The fix is the rework this PR's follow-up was always going to need: SPMD
cohorts must size themselves from rt_available_cluster_count() (added in
#1458) instead of hardcoding a width, and cases whose block count is load-bearing
need their incore tiling adjusted to match. I'm preparing that change and will
link it here.

Until then, please don't merge this as-is on top of current main — the two
together are a red sim suite. Either hold it for the cohort rework, or update
the body so the sim behaviour change is stated rather than described as a no-op.

@ChaoWao

ChaoWao commented Jul 25, 2026

Copy link
Copy Markdown
Collaborator

Heads-up: I've force-pushed onto this branch, replacing 775d2800. Your work is
not discarded — the commit carries Co-authored-by: majin0824, and your semantic
analysis of the cohort widths (stress at total/6, /2, /3, /6; early_dispatch in
waves of at most cluster_count) is what the new orchestrations implement. Sorry
for landing on top of yours rather than alongside; happy to restructure if you'd
prefer your commit kept separate underneath.

What changed relative to your version

Two things, one additive and one a design swap.

Additive: the PR now also removes the block_dim field itself, not just the
pinned values — CallConfig, both remote-L3 wire codecs (PROTOCOL_VERSION
1→2, the payload is an int32 shorter), the fork-mailbox format string, the
nanobind binding, scene_test, and l0_swimlane. Zeroing the values without
removing the knob leaves the sim behaviour change (24/36 → 8) implicit; removing
it makes "a run takes the whole device" the only shape there is.

Design swap: simpler_setup/available_block_dim.py is gone. Instead each
orchestration writes the geometry it actually used into a layout output
tensor, and the test rebuilds its golden from that (via the compare_outputs
hook from #1470). The reason is that the helper made the host a second source of
truth for a device property:

  • _SIM_AUTO_BLOCKDIM = 8 and _PLATFORM_MAX_BLOCKDIM = {a2a3: 24, a5: 36} are
    hand-mirrored from device_runner_base.h / platform_config.h, with nothing
    enforcing they stay in sync.
  • The onboard path opens a second ACL client inside the pytest process
    (acl.init() → set_device → create_stream) defaulting to device_id=0,
    but the resource scheduler hands each case whatever devices are free. On a
    shared box that both asks the wrong device and touches one the run does not
    hold.
  • Every failure path falls back to ceiling, which is exactly the wrong answer
    on a device that reports fewer cores than the arch maximum.

With the device reporting its own width there is nothing to keep in sync and no
second ACL context.

State: in progress — 4 of 6 sync_start cases converted (spmd_sync_start,
_aiv, _edge, _mix_spill). _stress and _early_dispatch are still on the
old literals, so CI is red until those land. Pushing now for visibility rather
than sitting on it.

@ChaoWao ChaoWao changed the title test: use block_dim=0 (auto) instead of hardcoded platform max Update: size SPMD sync_start cohorts from the device, not literals Jul 25, 2026
@ChaoWao ChaoWao changed the title Update: size SPMD sync_start cohorts from the device, not literals Refactor: remove the user-facing block_dim knob Jul 25, 2026
@ChaoWao
ChaoWao force-pushed the blockdim-auto branch 3 times, most recently from 8514da0 to 8a54fed Compare July 26, 2026 02:02
@ChaoWao
ChaoWao force-pushed the blockdim-auto branch 2 times, most recently from 573c75b to 6ebdec1 Compare July 26, 2026 04:20
block_dim let a caller pin how many AICore blocks a task occupied, but
the value is a property of the device, not of the task: every in-repo
caller either left it at the "auto" sentinel or pinned the width of the
one device it was written for. Pinning it below the device width simply
wasted cores, and pinning it above was rejected at run time, so the knob
only ever expressed what the runtime already knows.

- Drop CallConfig::block_dim and its plumbing through the task
  interface, the packed mailbox wire layout, the remote-L3 protocol, and
  both platform device runners; the runner now always resolves the
  block count from the stream's capacity.
- Strip the pinned "block_dim" entries from the scene tests and
  examples that carried them, and update the fixtures that constructed a
  CallConfig with one. The wire contract was asserted in three places —
  the header's static_assert, a C++ unit test that restated the same
  sizeof expression, and a remote-wire round-trip that set the field —
  and all three move with the layout.
- The two spmd_sync_start_mix_spill docstrings stop quoting core counts.
  They named a5's 72 AIV / 24 clusters and a2a3's 48 / 24, which were
  already inconsistent with the block_dim those cases pinned and are
  meaningless once the width is the device's own. (The a5 one also
  carried a paragraph twice.)
- available_aicore_counts loses its Pinned case: it fixed block_dim so
  the host held an expectation independent of the reported count, which
  was the only way to catch an under-report. Nothing replaces it, and
  the docstrings say so rather than implying coverage that is gone.
- Sweep the docs and skills that documented the knob.

Co-authored-by: majin0824 <majin15@huawei.com>
@ChaoWao
ChaoWao merged commit e17b29a into hw-native-sys:main Jul 26, 2026
16 checks passed
ChaoWao added a commit to ChaoWao/simpler-fork that referenced this pull request Jul 27, 2026
…tree

An editable install pins the compiled extension at install time
(`editable.rebuild = false`) while `python/simpler/*.py` is read live. Switching
branches or rebasing therefore moves the Python out from under a fixed binary,
nothing rebuilds, and a changed struct layout makes attributes read as 0 with no
error anywhere.

That is not hypothetical. A worktree installed before hw-native-sys#1309 (which dropped
`block_dim` from `CallConfig`) and then branched onto a main containing it read
`aicpu_thread_num` as 0, and the whole onboard L2 suite failed with
`launch_aicpu_num (0) must be in range [1, 4]` — a plausible-looking runtime
rejection that reads as a product bug. It cost a full bisect, and an A/B against
main "reproduced" it because both arms shared the same stale binary. The same
skew hit again a few hours later as `AttributeError:
run_stream_set_create_count` after a rebase across hw-native-sys#1464.

`_task_interface` now records the commit it was built from, and
`simpler.task_interface` compares it against the working tree at import, raising
with the one command that fixes it. Loud beats silent here: the alternative is
not an error, it is wrong values. A warning would also have been the wrong
choice — pytest relegates import-time warnings to its end-of-run summary, and in
the failure above the same warning would have appeared in *both* arms of the A/B
and been dismissed as noise.

A *missing* stamp raises too, rather than being treated as "cannot tell". The
attribute is absent only on an extension compiled before it existed, which in a
checkout new enough to run the check is by definition a different revision —
and is the state of every already-installed worktree the day this lands.

Keyed on git HEAD, matching `RuntimeBuilder._build_cache_stamp` — mtimes do not
survive a branch switch, which is why that stamp is git-based too. Deliberately
the whole HEAD rather than a subset: excluding paths means maintaining a second
"what cannot affect the build" list beside the one CI already keeps, and the
guard exists precisely so correctness does not rest on such judgments. It costs
little — 120 of the last 200 commits touch the ABI surface anyway, so most of
the reinstalls it forces were owed regardless and merely became visible; the
rest cost one ~20 s reinstall.

Inert outside a source tree: a wheel has no `.git` to compare against, and a
build made without git carries an empty stamp.

The rebuild table gained the trigger it was missing and no longer contradicts
itself. It was framed as "what did you change", so it had no row for the case
where you changed nothing and the tree moved — the one that bites, and the one
where verifying that `import simpler` resolves into your worktree gives false
confidence. Its "no rebuild needed" rows now say what is actually true: nothing
to recompile, but the guard keys on HEAD, so a commit still needs a reinstall
before the next import.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
ChaoWao added a commit to ChaoWao/simpler-fork that referenced this pull request Jul 27, 2026
…tree

An editable install pins the compiled extension at install time
(`editable.rebuild = false`) while `python/simpler/*.py` is read live. Switching
branches or rebasing therefore moves the Python out from under a fixed binary,
nothing rebuilds, and a changed struct layout makes attributes read as 0 with no
error anywhere.

That is not hypothetical. A worktree installed before hw-native-sys#1309 (which dropped
`block_dim` from `CallConfig`) and then branched onto a main containing it read
`aicpu_thread_num` as 0, and the whole onboard L2 suite failed with
`launch_aicpu_num (0) must be in range [1, 4]` — a plausible-looking runtime
rejection that reads as a product bug. It cost a full bisect, and an A/B against
main "reproduced" it because both arms shared the same stale binary. The same
skew hit again a few hours later as `AttributeError:
run_stream_set_create_count` after a rebase across hw-native-sys#1464.

`_task_interface` now records the commit it was built from, and
`simpler.task_interface` compares it against the working tree at import, raising
with the one command that fixes it. Loud beats silent here: the alternative is
not an error, it is wrong values. A warning would also have been the wrong
choice — pytest relegates import-time warnings to its end-of-run summary, and in
the failure above the same warning would have appeared in *both* arms of the A/B
and been dismissed as noise.

A *missing* stamp raises too, rather than being treated as "cannot tell". The
attribute is absent only on an extension compiled before it existed, which in a
checkout new enough to run the check is by definition a different revision —
and is the state of every already-installed worktree the day this lands.

Keyed on git HEAD, matching `RuntimeBuilder._build_cache_stamp` — mtimes do not
survive a branch switch, which is why that stamp is git-based too. Deliberately
the whole HEAD rather than a subset: excluding paths means maintaining a second
"what cannot affect the build" list beside the one CI already keeps, and the
guard exists precisely so correctness does not rest on such judgments. It costs
little — 120 of the last 200 commits touch the ABI surface anyway, so most of
the reinstalls it forces were owed regardless and merely became visible; the
rest cost one ~20 s reinstall.

Inert outside a source tree: a wheel has no `.git` to compare against, and a
build made without git carries an empty stamp.

The rebuild table gained the trigger it was missing and no longer contradicts
itself. It was framed as "what did you change", so it had no row for the case
where you changed nothing and the tree moved — the one that bites, and the one
where verifying that `import simpler` resolves into your worktree gives false
confidence. Its "no rebuild needed" rows now say what is actually true: nothing
to recompile, but the guard keys on HEAD, so a commit still needs a reinstall
before the next import.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
ChaoWao added a commit that referenced this pull request Jul 27, 2026
…tree (#1523)

An editable install pins the compiled extension at install time
(`editable.rebuild = false`) while `python/simpler/*.py` is read live. Switching
branches or rebasing therefore moves the Python out from under a fixed binary,
nothing rebuilds, and a changed struct layout makes attributes read as 0 with no
error anywhere.

That is not hypothetical. A worktree installed before #1309 (which dropped
`block_dim` from `CallConfig`) and then branched onto a main containing it read
`aicpu_thread_num` as 0, and the whole onboard L2 suite failed with
`launch_aicpu_num (0) must be in range [1, 4]` — a plausible-looking runtime
rejection that reads as a product bug. It cost a full bisect, and an A/B against
main "reproduced" it because both arms shared the same stale binary. The same
skew hit again a few hours later as `AttributeError:
run_stream_set_create_count` after a rebase across #1464.

`_task_interface` now records the commit it was built from, and
`simpler.task_interface` compares it against the working tree at import, raising
with the one command that fixes it. Loud beats silent here: the alternative is
not an error, it is wrong values. A warning would also have been the wrong
choice — pytest relegates import-time warnings to its end-of-run summary, and in
the failure above the same warning would have appeared in *both* arms of the A/B
and been dismissed as noise.

A *missing* stamp raises too, rather than being treated as "cannot tell". The
attribute is absent only on an extension compiled before it existed, which in a
checkout new enough to run the check is by definition a different revision —
and is the state of every already-installed worktree the day this lands.

Keyed on git HEAD, matching `RuntimeBuilder._build_cache_stamp` — mtimes do not
survive a branch switch, which is why that stamp is git-based too. Deliberately
the whole HEAD rather than a subset: excluding paths means maintaining a second
"what cannot affect the build" list beside the one CI already keeps, and the
guard exists precisely so correctness does not rest on such judgments. It costs
little — 120 of the last 200 commits touch the ABI surface anyway, so most of
the reinstalls it forces were owed regardless and merely became visible; the
rest cost one ~20 s reinstall.

Inert outside a source tree: a wheel has no `.git` to compare against, and a
build made without git carries an empty stamp.

The rebuild table gained the trigger it was missing and no longer contradicts
itself. It was framed as "what did you change", so it had no row for the case
where you changed nothing and the tree moved — the one that bites, and the one
where verifying that `import simpler` resolves into your worktree gives false
confidence. Its "no rebuild needed" rows now say what is actually true: nothing
to recompile, but the guard keys on HEAD, so a commit still needs a reinstall
before the next import.
@doraemonmj
doraemonmj deleted the blockdim-auto branch August 7, 2026 01:37
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants