Summary
In tensormap_and_ringbuffer, one physical task dependency currently carries two independent semantics:
- WAIT / readiness: the consumer cannot become ready until the producer completes.
- RETAIN / lifetime: the producer task slot and its packed output buffer must remain alive until the consumer releases the reference.
Coupling these semantics prevents safe transitive reduction.
For example:
If C reads an output tensor allocated by A, the direct A -> C WAIT is redundant because A -> B -> C already provides the ordering. However, the RETAIN relation is not redundant: removing the whole edge can allow A and its packed output to be reclaimed before C finishes reading it.
Introduce orthogonal WAIT and RETAIN edge semantics so that transitive reduction can turn A -> C from WAIT | RETAIN into RETAIN instead of either keeping the redundant readiness edge or dropping the required lifetime reference.
Motivation / Use Case
Today, creator dependencies derived from Tensor::owner_task_id and modifier dependencies derived from TensorMap lookup are ultimately stored in the same fanin/fanout representation. Every retained creator edge therefore also participates in readiness accounting and producer fanout traversal.
This has two consequences:
- Pure ordering edges cannot be transitively reduced without first separating lifetime ownership.
- Dense graphs pay unnecessary fanin/fanout storage, dep-pool usage, wiring work, atomics, and completion-time consumer traversal.
The intended optimization must preserve this invariant: ordering can be transitive, but tensor lifetime cannot. In particular, A -> B -> C proves that A completes before C starts; it does not prove that A's output remains allocated until C completes.
Related: #1202, #545
Proposed API / Behavior
Represent the two properties independently, for example:
enum class DepFlags : uint8_t {
NONE = 0,
WAIT = 1 << 0,
RETAIN = 1 << 1,
};
The exact storage may use flags, separate lists, or another compact representation, but it must support all three meaningful states:
WAIT only: ordering without retaining the producer until consumer completion.
RETAIN only: lifetime reference without affecting readiness.
WAIT | RETAIN: both ordering and lifetime.
Dependency construction
- Creator /
owner_task_id dependency: initially WAIT | RETAIN.
- TensorMap modifier dependency:
WAIT.
- Explicit dependency: allow the caller/code generator to select flags; preserve the existing API as a conservative
WAIT | RETAIN default.
- If the same
(producer, consumer) pair is discovered for multiple reasons, accumulate flags with OR before deduplication. Do not rely on first-wins ordering.
- In manual scopes, remain conservative unless creator provenance is known; an explicit dependency that protects a runtime-created output must not be reduced to WAIT-only.
Transitive reduction
- Run reduction over the WAIT graph only.
- If another WAIT path already connects the producer to the consumer, clear WAIT on the direct edge.
- Keep the edge as RETAIN-only when lifetime is still required; remove it only when no flags remain.
Runtime accounting
- Readiness
fanin_count and producer fanout notifications must include WAIT edges only.
- Consumer completion must release RETAIN edges only.
- A RETAIN-only edge must not create a readiness fanout node.
- WAIT-only edges may use a temporary submit-to-wire pin to protect producer slot generation, as explored by the reference branch, but that pin must be released after wiring and in the all-producers-completed fast path.
- Implement against the current orchestrator-side wiring path rather than mechanically moving the old scheduler-side implementation.
- Inline and spill representations must preserve flags through payload initialization and slot reuse.
Acceptance criteria
- With pure ordering
A -> B -> C plus A -> C, the direct WAIT can be removed.
- If
C reads A's runtime-created output, the direct relation becomes RETAIN-only and A's output remains alive until C releases it, including after scope end.
- Creator + modifier and creator + explicit-WAIT discovery for the same producer produce
WAIT | RETAIN.
- A WAIT-only producer can be consumed after completion and wiring without waiting for the consumer to complete.
- Inline and spill edges round-trip all supported flag combinations.
- Tests cover the complete submit -> payload init -> wire -> complete -> release path, not only hand-constructed payloads.
- a2a3 and a5
tensormap_and_ringbuffer implementations remain semantically aligned.
- Dependency DFX/replay preserves edge flags or provenance instead of comparing only the producer-ID set.
Alternatives Considered
The reference branch introduces two kinds:
RESOURCE = WAIT | RETAIN
EXECUTION = WAIT
This is useful for allowing modifier producers to release their lifetime pin after wiring, but it cannot express RETAIN-only. Therefore it does not enable the target reduction: A -> C must either remain a full RESOURCE readiness edge or lose lifetime protection.
Deleting the complete transitive edge is unsafe, while retaining all current edges is correct but preserves the scheduling and storage overhead.
Additional Context
Reference prototype:
Useful ideas from the prototype include classifying creator dependencies separately from TensorMap modifier dependencies, compact inline/spill kind storage, making the stronger lifetime requirement win, and releasing the temporary pin for ordering-only dependencies after wiring.
The prototype must not be ported as-is: it writes fanin_inline_dep_kind_mask before payload.init(), while payload.init() resets that mask to zero. With its encoding, the first 64 RESOURCE edges are consequently decoded as EXECUTION and can release creator lifetime too early. A full-path regression test is required for this case.
The branch also predates the current mainline orchestrator-side wiring flow, so the design should be reimplemented on current main rather than resolved as a mechanical cherry-pick.
Summary
In
tensormap_and_ringbuffer, one physical task dependency currently carries two independent semantics:Coupling these semantics prevents safe transitive reduction.
For example:
If
Creads an output tensor allocated byA, the directA -> CWAIT is redundant becauseA -> B -> Calready provides the ordering. However, the RETAIN relation is not redundant: removing the whole edge can allowAand its packed output to be reclaimed beforeCfinishes reading it.Introduce orthogonal WAIT and RETAIN edge semantics so that transitive reduction can turn
A -> CfromWAIT | RETAINintoRETAINinstead of either keeping the redundant readiness edge or dropping the required lifetime reference.Motivation / Use Case
Today, creator dependencies derived from
Tensor::owner_task_idand modifier dependencies derived from TensorMap lookup are ultimately stored in the same fanin/fanout representation. Every retained creator edge therefore also participates in readiness accounting and producer fanout traversal.This has two consequences:
The intended optimization must preserve this invariant: ordering can be transitive, but tensor lifetime cannot. In particular,
A -> B -> Cproves thatAcompletes beforeCstarts; it does not prove thatA's output remains allocated untilCcompletes.Related: #1202, #545
Proposed API / Behavior
Represent the two properties independently, for example:
The exact storage may use flags, separate lists, or another compact representation, but it must support all three meaningful states:
WAITonly: ordering without retaining the producer until consumer completion.RETAINonly: lifetime reference without affecting readiness.WAIT | RETAIN: both ordering and lifetime.Dependency construction
owner_task_iddependency: initiallyWAIT | RETAIN.WAIT.WAIT | RETAINdefault.(producer, consumer)pair is discovered for multiple reasons, accumulate flags with OR before deduplication. Do not rely on first-wins ordering.Transitive reduction
Runtime accounting
fanin_countand producer fanout notifications must include WAIT edges only.Acceptance criteria
A -> B -> CplusA -> C, the direct WAIT can be removed.CreadsA's runtime-created output, the direct relation becomes RETAIN-only andA's output remains alive untilCreleases it, including after scope end.WAIT | RETAIN.tensormap_and_ringbufferimplementations remain semantically aligned.Alternatives Considered
The reference branch introduces two kinds:
RESOURCE = WAIT | RETAINEXECUTION = WAITThis is useful for allowing modifier producers to release their lifetime pin after wiring, but it cannot express RETAIN-only. Therefore it does not enable the target reduction:
A -> Cmust either remain a full RESOURCE readiness edge or lose lifetime protection.Deleting the complete transitive edge is unsafe, while retaining all current edges is correct but preserves the scheduling and storage overhead.
Additional Context
Reference prototype:
Useful ideas from the prototype include classifying creator dependencies separately from TensorMap modifier dependencies, compact inline/spill kind storage, making the stronger lifetime requirement win, and releasing the temporary pin for ordering-only dependencies after wiring.
The prototype must not be ported as-is: it writes
fanin_inline_dep_kind_maskbeforepayload.init(), whilepayload.init()resets that mask to zero. With its encoding, the first 64 RESOURCE edges are consequently decoded as EXECUTION and can release creator lifetime too early. A full-path regression test is required for this case.The branch also predates the current mainline orchestrator-side wiring flow, so the design should be reimplemented on current
mainrather than resolved as a mechanical cherry-pick.