Summary
The runtime resolves SPMD inter-block dependencies at per-SPMD-block (scope) granularity, but swimlane_converter expands each block-level edge into a dense per-instance crossbar of flow arrows. Represent the dependency at block granularity (or clearly annotate it) so the trace doesn't imply per-instance dependency resolution.
Motivation / Use Case
- For a whole kernel (DSv4
decode_sparse_attn), deps.json captures the dependency graph as 9 task nodes / 23 tensor-mediated edges (source: tensormap, region overlap: covered/other) — i.e. one node per SPMD block.
- The same run's
merged_swimlane_*.json contains 613,820 flow events (cat: flow, name: dependency) — every block-level edge is blown up into a per-instance crossbar (e.g. 128 producer instances × 64 consumer instances).
- Empirically the runtime does a per-scope barrier: dependent stages run sequentially with only a small dispatch gap and do not overlap at the instance level, while a dependency-free input-only pre-pass correctly floats ahead and overlaps. So the per-instance arrows do not correspond to real runtime edges.
- This misleads performance analysis — a reader looking at the swimlane concludes the runtime resolves a full per-instance dependency between consecutive SPMD stages, and reasons about dependency-resolution cost / overlap potential incorrectly. It also bloats the merged trace from ~2 MB to ~110 MB.
Proposed API / Behavior
Any subset:
- Collapse the per-instance flow arrows to one representative edge per (producer-block, consumer-block) pair, annotated with the fan degree (e.g.
128→64).
- Add a block-level dependency view / toggle alongside the per-instance one.
- At minimum, document (tool help + output) that the flow arrows are a conservative visualization expansion of block-level
tensormap edges, not runtime per-instance edges.
Reducing the arrow count would also shrink the merged trace dramatically.
Additional Context
- Repro: run a kernel with
--enable-l2-swimlane --enable-dep-gen (e.g. pypto-lib models/deepseek/v4/decode_sparse_attn.py -p a2a3), then compare deps.json (9 nodes / 23 edges) against merged_swimlane_*.json (613,820 flow events) in the same dfx_outputs/.
- runtime (simpler) commit:
48980572.
- Related: simpler#1001 (text-based dep_gen rendering) — adjacent tooling, but a different concern (output format of the dep_gen graph).
Summary
The runtime resolves SPMD inter-block dependencies at per-SPMD-block (scope) granularity, but
swimlane_converterexpands each block-level edge into a dense per-instance crossbar of flow arrows. Represent the dependency at block granularity (or clearly annotate it) so the trace doesn't imply per-instance dependency resolution.Motivation / Use Case
decode_sparse_attn),deps.jsoncaptures the dependency graph as 9 task nodes / 23 tensor-mediated edges (source: tensormap, regionoverlap: covered/other) — i.e. one node per SPMD block.merged_swimlane_*.jsoncontains 613,820 flow events (cat: flow, name: dependency) — every block-level edge is blown up into a per-instance crossbar (e.g. 128 producer instances × 64 consumer instances).Proposed API / Behavior
Any subset:
128→64).tensormapedges, not runtime per-instance edges.Reducing the arrow count would also shrink the merged trace dramatically.
Additional Context
--enable-l2-swimlane --enable-dep-gen(e.g. pypto-libmodels/deepseek/v4/decode_sparse_attn.py -p a2a3), then comparedeps.json(9 nodes / 23 edges) againstmerged_swimlane_*.json(613,820 flow events) in the samedfx_outputs/.48980572.