Skip to content

feat(a5): add URMA deferred completion demo for tensormap_and_ringbuffer - #1421

Merged
jvjhfhg merged 1 commit into
hw-native-sys:mainfrom
wxwnnzdyd:a5-urma-deferred-completion-demo
Jul 22, 2026
Merged

jvjhfhg merged 1 commit into
hw-native-sys:mainfrom
wxwnnzdyd:a5-urma-deferred-completion-demo

Conversation

@wxwnnzdyd

Copy link
Copy Markdown
Contributor

Summary

  • Add a minimal A5 URMA deferred-completion demo under examples/a5/tensormap_and_ringbuffer/urma_deferred_completion_demo
  • Exercise the canonical two-rank path: URMA TGET producer -> deferred completion -> dependent consumer
  • Keep the demo isolated from CI until the A5 URMA workspace environment is enabled for [Code Health] Require upgraded a5 environment supporting sdma #1315

Details

This demo mirrors the existing SDMA deferred-completion smoke test shape, but uses the URMA backend:

  • each rank stages an input tensor in an allocated communication domain
  • the producer issues one deferred UrmaTget from the peer rank's input window into local out
  • the consumer depends on out and computes result = out + 1
  • host-side validation checks both out and result against the peer input

The pytest case skips unless SIMPLER_ENABLE_PTO_URMA_WORKSPACE is enabled, so the demo can land before CI enables the URMA workspace overlay by default.

Testing

  • python3 -m py_compile examples/a5/tensormap_and_ringbuffer/urma_deferred_completion_demo/test_urma_deferred_completion_demo.py
  • SIMPLER_ENABLE_PTO_URMA_WORKSPACE=OFF .venv/bin/python -m pytest -q examples/a5/tensormap_and_ringbuffer/urma_deferred_completion_demo/test_urma_deferred_completion_demo.py -> 1 skipped
  • Local A5 hardware: SIMPLER_ENABLE_PTO_URMA_WORKSPACE=ON PTO_ISA_ROOT=/home/hdy/simpler/build/pto-isa python examples/a5/tensormap_and_ringbuffer/urma_deferred_completion_demo/test_urma_deferred_completion_demo.py -p a5 -d 0,1 -> max_out=0.000e+00 max_result=0.000e+00 on both ranks

@coderabbitai

coderabbitai Bot commented Jul 21, 2026 •

Copy link
Copy Markdown

Review Change Stack

📝 Walkthrough

Walkthrough

Adds a two-rank URMA deferred-completion demo with asynchronous producer and tiled consumer AICore kernels, orchestration that chains both tasks, and an A5 Python smoke test with pytest and CLI execution paths.

Changes

URMA deferred completion workflow

Layer / File(s) Summary
URMA producer kernel
examples/a5/tensormap_and_ringbuffer/urma_deferred_completion_demo/kernels/aiv/kernel_urma_tget_async.cpp
Validates the two-rank communication setup, computes local and peer tensor addresses, and submits an asynchronous URMA TGET request.
Orchestration and consumer pipeline
examples/a5/tensormap_and_ringbuffer/urma_deferred_completion_demo/kernels/orchestration/urma_deferred_completion_orch.cpp, examples/a5/tensormap_and_ringbuffer/urma_deferred_completion_demo/kernels/aiv/kernel_consumer.cpp
Exports orchestration configuration, submits producer and consumer tasks sequentially, and adds 1.0f to transferred tensor tiles before storing results.
Build, execution, and validation harness
examples/a5/tensormap_and_ringbuffer/urma_deferred_completion_demo/test_urma_deferred_completion_demo.py
Builds the callable, allocates two-rank URMA resources, runs the workflow, validates peer-derived outputs, and provides pytest and CLI entry points.

Estimated code review effort: 4 (Complex) | ~45 minutes

Sequence Diagram(s)

sequenceDiagram
  participant Orchestration
  participant ProducerKernel
  participant URMA
  participant ConsumerKernel
  Orchestration->>ProducerKernel: submit input and CommContext
  ProducerKernel->>URMA: submit asynchronous peer TGET
  URMA-->>ProducerKernel: transfer data to intermediate tensor
  Orchestration->>ConsumerKernel: submit intermediate tensor
  ConsumerKernel->>ConsumerKernel: add 1.0 to each tile
  ConsumerKernel-->>Orchestration: write final result
Loading

Possibly related PRs

Poem

I’m a bunny hopping through tensors bright,
Fetching peer data in a URMA flight.
Two kernels dance, then add one more,
While tests check every rank and score.
/ \_/\\
( o.o )

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Title check ✅ Passed The title clearly matches the main change: adding an A5 URMA deferred-completion demo for tensormap_and_ringbuffer.
Description check ✅ Passed The description is directly aligned with the changeset and accurately summarizes the demo, workflow, and test gating.

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces a new URMA deferred completion demo and smoke test for the a5 platform, consisting of a producer kernel initiating asynchronous TGET operations, a consumer kernel processing the data, orchestration logic, and a Python validation test. Feedback is provided regarding the orchestration argument validation, which should explicitly verify the counts of tensors and scalars separately to prevent potential out-of-bounds access.

Important

The consumer version of Gemini Code Assist on GitHub is being sunset. Starting June 18, 2026, new organization installations will be blocked, and all code review activity will officially cease on July 17, 2026.
For more details on the timeline and next steps, please review the Help Documentation.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🧹 Nitpick comments (1)
examples/a5/tensormap_and_ringbuffer/urma_deferred_completion_demo/test_urma_deferred_completion_demo.py (1)

67-67: 📐 Maintainability & Code Quality | 🔵 Trivial | 💤 Low value

Prefer iterable unpacking over list concatenation.

Static analysis (Ruff RUF005) flags this concatenation pattern.

🧹 Proposed fix
-    extra_includes = list(include_dirs) + [str(kc.project_root / "src" / "common")]
+    extra_includes = [*include_dirs, str(kc.project_root / "src" / "common")]
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In
`@examples/a5/tensormap_and_ringbuffer/urma_deferred_completion_demo/test_urma_deferred_completion_demo.py`
at line 67, Update the extra_includes assignment to use iterable unpacking
instead of list concatenation, preserving the existing include_dirs entries
followed by the common source directory path and satisfying Ruff RUF005.

Source: Linters/SAST tools

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Nitpick comments:
In
`@examples/a5/tensormap_and_ringbuffer/urma_deferred_completion_demo/test_urma_deferred_completion_demo.py`:
- Line 67: Update the extra_includes assignment to use iterable unpacking
instead of list concatenation, preserving the existing include_dirs entries
followed by the common source directory path and satisfying Ruff RUF005.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: 14602917-7590-4feb-b6b9-227af6c7909e

📥 Commits

Reviewing files that changed from the base of the PR and between 0ffb88f and d26c1c8.

📒 Files selected for processing (4)
  • examples/a5/tensormap_and_ringbuffer/urma_deferred_completion_demo/kernels/aiv/kernel_consumer.cpp
  • examples/a5/tensormap_and_ringbuffer/urma_deferred_completion_demo/kernels/aiv/kernel_urma_tget_async.cpp
  • examples/a5/tensormap_and_ringbuffer/urma_deferred_completion_demo/kernels/orchestration/urma_deferred_completion_orch.cpp
  • examples/a5/tensormap_and_ringbuffer/urma_deferred_completion_demo/test_urma_deferred_completion_demo.py

@wxwnnzdyd
wxwnnzdyd force-pushed the a5-urma-deferred-completion-demo branch from d26c1c8 to 9900a03 Compare July 21, 2026 09:27
- Add a minimal two-rank URMA TGET deferred-completion example
- Reuse the SDMA demo structure with one producer and one consumer
- Keep the pytest path skipped unless SIMPLER_ENABLE_PTO_URMA_WORKSPACE is enabled
@wxwnnzdyd
wxwnnzdyd force-pushed the a5-urma-deferred-completion-demo branch from 9900a03 to 29c3579 Compare July 22, 2026 02:18
@jvjhfhg jvjhfhg changed the title Add A5 URMA deferred completion demo feat(a5): add URMA deferred completion demo for tensormap_and_ringbuffer Jul 22, 2026
@jvjhfhg
jvjhfhg merged commit a7a29ee into hw-native-sys:main Jul 22, 2026
28 of 29 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants