Skip to content

[Feature] swimlane_converter: render SPMD deps at block granularity, not a per-instance crossbar #1049

Description

@lyfne123

Summary

The runtime resolves SPMD inter-block dependencies at per-SPMD-block (scope) granularity, but swimlane_converter expands each block-level edge into a dense per-instance crossbar of flow arrows. Represent the dependency at block granularity (or clearly annotate it) so the trace doesn't imply per-instance dependency resolution.

Motivation / Use Case

  • For a whole kernel (DSv4 decode_sparse_attn), deps.json captures the dependency graph as 9 task nodes / 23 tensor-mediated edges (source: tensormap, region overlap: covered/other) — i.e. one node per SPMD block.
  • The same run's merged_swimlane_*.json contains 613,820 flow events (cat: flow, name: dependency) — every block-level edge is blown up into a per-instance crossbar (e.g. 128 producer instances × 64 consumer instances).
  • Empirically the runtime does a per-scope barrier: dependent stages run sequentially with only a small dispatch gap and do not overlap at the instance level, while a dependency-free input-only pre-pass correctly floats ahead and overlaps. So the per-instance arrows do not correspond to real runtime edges.
  • This misleads performance analysis — a reader looking at the swimlane concludes the runtime resolves a full per-instance dependency between consecutive SPMD stages, and reasons about dependency-resolution cost / overlap potential incorrectly. It also bloats the merged trace from ~2 MB to ~110 MB.

Proposed API / Behavior

Any subset:

  1. Collapse the per-instance flow arrows to one representative edge per (producer-block, consumer-block) pair, annotated with the fan degree (e.g. 128→64).
  2. Add a block-level dependency view / toggle alongside the per-instance one.
  3. At minimum, document (tool help + output) that the flow arrows are a conservative visualization expansion of block-level tensormap edges, not runtime per-instance edges.

Reducing the arrow count would also shrink the merged trace dramatically.

Additional Context

  • Repro: run a kernel with --enable-l2-swimlane --enable-dep-gen (e.g. pypto-lib models/deepseek/v4/decode_sparse_attn.py -p a2a3), then compare deps.json (9 nodes / 23 edges) against merged_swimlane_*.json (613,820 flow events) in the same dfx_outputs/.
  • runtime (simpler) commit: 48980572.
  • Related: simpler#1001 (text-based dep_gen rendering) — adjacent tooling, but a different concern (output format of the dep_gen graph).

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

enhancementNew feature or request

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions