Category
Observability / DFX accuracy (tech debt)
Component
dep_gen — DepGenRecord capture + dep_gen_replay (tensormap_and_ringbuffer)
Description
PR #1806 (part of #1375) added orthogonal DepFlags { WAIT, RETAIN } per fanin edge and
threaded them into deps.json and the dep_gen differential gate. Creator and tensormap
(modifier) edge flags are re-derivable at replay time (arg_types + owner_task_id +
tensormap), so those are recorded accurately.
Explicit-dep flags are not. DepGenRecord records explicit_deps[] as bare
PTO2TaskId::raw values with no per-dep DepFlags. So the replay hard-codes every explicit
edge to DEP_WAIT | DEP_RETAIN. Consequences:
- A dep added via the ordering-only
CoreTaskArgsWithDeps::add_dep_wait() API (a genuine
["wait"] edge at runtime) is recorded in deps.json as ["wait","retain"].
- The dual-pass differential gate cannot catch an explicit-kind divergence, because both the
oracle and the annotated pass read the same hard-coded constant.
This is functionally harmless (runtime behaviour is correct; only the DFX artifact is
imprecise) and is currently documented as a known replay limitation in docs/dfx/dep-gen.md,
but it means deps.json cannot be trusted for the WAIT/RETAIN classification of explicit
edges.
Related, and better handled inside the follow-up reduction PR rather than here: dep_gen
reconstructs edges by re-running compute_task_fanin (construction only), so it does not model
the planned device-side transitive reduction. Once reduction lands, deps.json flags will be
the pre-reduction values unless the replay also models the reduction. (Tracking that with the
reduction work, noted here for context.)
Location
src/common/platform/include/common/dep_gen.h — struct DepGenRecord (has
explicit_deps[], no kinds; static_assert(sizeof(DepGenRecord) == 4672) locks the wire
layout)
src/{a2a3,a5}/runtime/tensormap_and_ringbuffer/host/dep_gen_replay.cpp — hard-codes
explicit edges to DEP_WAIT | DEP_RETAIN
src/common/platform/shared/host/dep_gen_collector.cpp — capture path
docs/dfx/dep-gen.md — documents the limitation
Proposed Fix
Add explicit_dep_kinds[DEP_GEN_MAX_EXPLICIT_DEPS] (one uint8_t DepFlags per explicit dep)
to DepGenRecord (and the overflow-record chain), populate it at capture, read it in
dep_gen_replay instead of the constant, and make the differential gate explicit-kind-aware.
Update the sizeof(DepGenRecord) static_assert + docs/dfx/dep-gen.md, and drop the
"explicit edges always recorded as wait+retain" caveat. Priority is low — only the DFX
artifact accuracy for add_dep_wait edges is affected, not runtime correctness.
Category
Observability / DFX accuracy (tech debt)
Component
dep_gen—DepGenRecordcapture +dep_gen_replay(tensormap_and_ringbuffer)Description
PR #1806 (part of #1375) added orthogonal
DepFlags { WAIT, RETAIN }per fanin edge andthreaded them into
deps.jsonand the dep_gen differential gate. Creator and tensormap(modifier) edge flags are re-derivable at replay time (
arg_types+owner_task_id+tensormap), so those are recorded accurately.
Explicit-dep flags are not.
DepGenRecordrecordsexplicit_deps[]as barePTO2TaskId::rawvalues with no per-depDepFlags. So the replay hard-codes every explicitedge to
DEP_WAIT | DEP_RETAIN. Consequences:CoreTaskArgsWithDeps::add_dep_wait()API (a genuine["wait"]edge at runtime) is recorded indeps.jsonas["wait","retain"].oracle and the annotated pass read the same hard-coded constant.
This is functionally harmless (runtime behaviour is correct; only the DFX artifact is
imprecise) and is currently documented as a known replay limitation in
docs/dfx/dep-gen.md,but it means
deps.jsoncannot be trusted for the WAIT/RETAIN classification of explicitedges.
Related, and better handled inside the follow-up reduction PR rather than here:
dep_genreconstructs edges by re-running
compute_task_fanin(construction only), so it does not modelthe planned device-side transitive reduction. Once reduction lands,
deps.jsonflags will bethe pre-reduction values unless the replay also models the reduction. (Tracking that with the
reduction work, noted here for context.)
Location
src/common/platform/include/common/dep_gen.h—struct DepGenRecord(hasexplicit_deps[], no kinds;static_assert(sizeof(DepGenRecord) == 4672)locks the wirelayout)
src/{a2a3,a5}/runtime/tensormap_and_ringbuffer/host/dep_gen_replay.cpp— hard-codesexplicit edges to
DEP_WAIT | DEP_RETAINsrc/common/platform/shared/host/dep_gen_collector.cpp— capture pathdocs/dfx/dep-gen.md— documents the limitationProposed Fix
Add
explicit_dep_kinds[DEP_GEN_MAX_EXPLICIT_DEPS](oneuint8_tDepFlagsper explicit dep)to
DepGenRecord(and the overflow-record chain), populate it at capture, read it indep_gen_replayinstead of the constant, and make the differential gate explicit-kind-aware.Update the
sizeof(DepGenRecord)static_assert +docs/dfx/dep-gen.md, and drop the"explicit edges always recorded as wait+retain" caveat. Priority is low — only the DFX
artifact accuracy for
add_dep_waitedges is affected, not runtime correctness.