Add: host_build_graph regression test and an under-count check - #1472
Conversation
The scene test derived its expected block slots from the count the runtime reported, so a count that was too small was self-consistent and passed. It also ran on tensormap_and_ringbuffer only, leaving the host_build_graph read path — the one that was returning zero — uncovered. - A Pinned case fixes block_dim, giving the host an expectation independent of what the runtime reported. It is the only case that can catch an under-count; the Auto case, which takes whatever the platform resolves, cannot. The pinned width differs from every platform's auto width, so a count sourced from the ceiling rather than from this run fails it too. - The same case runs on host_build_graph. Against the pre-fix ordering it fails at `block_num >= 1` in submit_task, because host orchestration read a cluster count of zero and asked for a cohort that wide. - The shape tensor no longer takes a dummy_task producer. host_build_graph runs its orchestrator to completion on the host before the device executes anything, so a producer's task_state can never reach COMPLETED and set_tensor_data's wait_for_tensor_ready spins out to PTO2_TENSOR_DATA_TIMEOUT_CYCLES. The MIX cohort already keeps the graph non-empty, and shape has neither producer nor consumer.
📝 WalkthroughWalkthroughAdds a host-build orchestration fixture that reports available AICore counts, validates pinned and automatic execution modes, and removes dummy shape producers from existing A2A3 and A5 fixtures. ChangesAvailable AICore Count Validation
Estimated code review effort: 3 (Moderate) | ~25 minutes Sequence Diagram(s)sequenceDiagram
participant HostTest
participant Orchestration
participant Runtime
participant MixedKernels
participant OutputValidation
HostTest->>Orchestration: submit blocks and shape tensors
Orchestration->>Runtime: query available cluster and AIV counts
Orchestration->>MixedKernels: launch configured mixed kernels
Orchestration->>HostTest: write counts into shape
HostTest->>OutputValidation: validate counts and block slots
Possibly related PRs
Poem
🚥 Pre-merge checks | ✅ 4 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (4 passed)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
There was a problem hiding this comment.
🧹 Nitpick comments (1)
tests/st/a2a3/host_build_graph/available_aicore_counts/test_available_aicore_counts.py (1)
61-90: 📐 Maintainability & Code Quality | 🔵 Trivial | ⚡ Quick winAnnotate mutable test metadata as
ClassVar.Ruff reports RUF012 for
CALLABLEandCASES. AddClassVarannotations, or confirm this rule is not enforced for scene-test metadata.Also applies to: 92-105
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@tests/st/a2a3/host_build_graph/available_aicore_counts/test_available_aicore_counts.py` around lines 61 - 90, Annotate the class-level mutable metadata dictionaries CALLABLE and CASES with typing.ClassVar, preserving their existing structures and values. Import ClassVar from typing if needed, and apply the annotation to both symbols to satisfy Ruff RUF012.Source: Linters/SAST tools
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Nitpick comments:
In
`@tests/st/a2a3/host_build_graph/available_aicore_counts/test_available_aicore_counts.py`:
- Around line 61-90: Annotate the class-level mutable metadata dictionaries
CALLABLE and CASES with typing.ClassVar, preserving their existing structures
and values. Import ClassVar from typing if needed, and apply the annotation to
both symbols to satisfy Ruff RUF012.
ℹ️ Review info
⚙️ Run configuration
Configuration used: Organization UI
Review profile: CHILL
Plan: Pro Plus
Run ID: 3bc762ea-f977-4d8a-9902-ef016268cd3a
📒 Files selected for processing (6)
tests/st/a2a3/host_build_graph/available_aicore_counts/kernels/orchestration/available_aicore_counts_orch.cpptests/st/a2a3/host_build_graph/available_aicore_counts/test_available_aicore_counts.pytests/st/a2a3/tensormap_and_ringbuffer/available_aicore_counts/kernels/orchestration/available_aicore_counts_orch.cpptests/st/a2a3/tensormap_and_ringbuffer/available_aicore_counts/test_available_aicore_counts.pytests/st/a5/tensormap_and_ringbuffer/available_aicore_counts/kernels/orchestration/available_aicore_counts_orch.cpptests/st/a5/tensormap_and_ringbuffer/available_aicore_counts/test_available_aicore_counts.py
Follow-up to #1458, addressing two review findings that landed after it merged.
The under-count check did not check under-counts
The scene test built its expected block slots from the count the runtime
reported:
A count that is too small is self-consistent under that comparison: the
orchestration launches fewer blocks, the host expects fewer blocks, the tail
stays zero on both sides, green. The docstring and #1458's body claimed the
tail zeros caught it. They did not.
Fix: a
Pinnedcase fixesblock_dim, so the host has an expectation thatdoes not come from the runtime and can assert the exact value. The pinned width
(4) differs from every platform's auto width — 8 on sim, the driver's answer
onboard — so a count sourced from the ceiling rather than from this run fails it
too. The
Autocase keeps the range / ratio / spend checks and is nowdocumented as unable to detect an under-count.
Over-counting was and remains covered in both: the cohort asks for
require_sync_start, so it needs every block co-resident and the deadlock guardfires on device.
The host_build_graph fix had no test
#1458's central bugfix was that host orchestration read a cluster count of zero,
and it shipped with coverage on
tensormap_and_ringbufferonly. The same casenow runs on
host_build_graph.Verified it is a real regression test, not just an extra green tick — against
the pre-fix ordering (shape publication moved back after the bind, host reading
worker_count / 3) it fails:because host orchestration read 0 and asked for a cohort that wide.
Dropping the dummy_task producer
shapeno longer gets adummy_taskproducer.host_build_graphruns itsorchestrator to completion on the host, before the device executes anything,
so a producer's
task_statecan never reachCOMPLETEDandset_tensor_data'swait_for_tensor_readyspins out toPTO2_TENSOR_DATA_TIMEOUT_CYCLES(orch_error_code=8 TENSOR_WAIT_TIMEOUT).The MIX cohort already keeps the graph non-empty and
shapehas neitherproducer nor consumer, so the write goes straight through on both runtimes.
Verification
Tests only — no
src/changes.pytest examples tests/st --platform a2a3simpytest examples tests/st --platform a5simpytest examples tests/st --platform a2a3(onboard, task-submit)All three new/changed cases confirmed executed, not silently deselected.
Interaction with #1309
Commented there: #1458's
SIM_AUTO_BLOCKDIM = 8invalidates #1309's premisethat
block_dim=0resolves toPLATFORM_MAX_BLOCKDIMon sim. Zeroing thepinned values now breaks
spmd_sync_start(block_num=12) andspmd_sync_start_edge(block_num=23) against a limit of 8. The SPMD cohortrework that fixes it is next.