Platform
All / Unknown (the copy logic is mirrored on both a5 and a2a3)
Runtime Variant
tensormap_and_ringbuffer
Summary
At the orchestration entry point, host↔device tensor transfer does not look at
the per-argument direction (ArgDirection / TensorArgType: IN / OUT /
INOUT). Every non-child_memory tensor is copied host→device at start and
device→host at end, regardless of its direction. This wastes:
- H2D bandwidth uploading pure
OUTPUT buffers — the host-side content is
uninitialized/garbage, so copying it up is meaningless.
- D2H bandwidth copying back pure
INPUT buffers — read-only inputs are
never modified by the kernel, so copying them back is wasted work.
Only INOUT genuinely needs both directions; child_memory tensors need
neither (already on device).
The signature (ArgDirection[]) is already plumbed end-to-end to the host
runtime, but is dropped at the consumption site — the comment in
c_api_shared.cpp states it is "plumbed end-to-end for per-tensor direction
decisions in runtime_maker but is currently unconsumed on both runtimes".
Git Commit ID
7c2f35f
Host Platform
Linux (aarch64)
Reproduction
The waste is observable directly in the host runtime's INFO logs — every tensor
is logged once on upload (H2D) and once on copy-back (D2H), regardless of
direction:
```bash
Run any example that has input-only tensors alongside output tensors.
python tests/st/a5/tensormap_and_ringbuffer//test_*.py -p a5 -d 1
In the host log:
bind_callable_to_runtime_impl -> "Tensor i: bytes at 0x..." (H2D upload, printed for EVERY tensor)
validate_runtime_impl -> "Tensor i: bytes copied to host" (D2H copy-back, printed for EVERY tensor)
=> pure-INPUT tensors are copied back, and pure-OUTPUT tensors are copied up.
```
Expected Performance
Copies should be gated on ArgDirection:
IN -> H2D copy-in only; no copy-back.
OUT -> device buffer only (skip the H2D copy-in of garbage); D2H copy-back only.
INOUT -> both directions (today's behavior).
child_memory -> neither direction.
For a typical kernel with mostly read-only inputs (weights, KV-cache) plus a
small output, this removes the dominant share of both the H2D and D2H traffic.
Actual Performance
All non-child_memory tensors are copied both directions unconditionally:
- H2D (start) — `bind_callable_to_runtime_impl` loops over every tensor and
does `device_malloc` + `copy_to_device`, branching only on `is_child_memory()`,
never on direction. Pure outputs are uploaded too.
- D2H (end) — `validate_runtime_impl` loops over every recorded
`tensor_pairs_` entry and does `copy_from_device`, gated only on null
host/dev pointers, never on direction. Pure inputs are copied back too.
Profiling Data (Optional)
Related closed issue #796 ("validate_runtime_impl is ~2x the chip-side cost per
run") measured ~50ms of copy-back overhead and explicitly hypothesized angle #1:
"Are all tensor pairs that validate_runtime_impl iterates actually outputs? If
the loop is also copying back inputs/weights/KV-cache entries ... that is the
bug." This issue is the root-cause counterpart: the direction info needed to
avoid those copies is available but unconsumed.
Additional Context
Code locations (mirrored on both runtimes):
- a5: `src/a5/runtime/tensormap_and_ringbuffer/host/runtime_maker.cpp`
- `bind_callable_to_runtime_impl` (H2D loop) — signature param is `const ArgDirection * /signature/` (unused)
- `validate_runtime_impl` (D2H loop)
- a2a3: `src/a2a3/runtime/tensormap_and_ringbuffer/host/runtime_maker.cpp` (same structure)
- Plumbing: `src/common/platform/onboard/host/c_api_shared.cpp` passes `bind_result.signature` into `bind_callable_to_runtime_impl`.
Suggested fix: consume the already-plumbed `signature`/`sig_count` in
`bind_callable_to_runtime_impl` to skip H2D for `OUT`, and tag each
`tensor_pairs_` entry with its direction so `validate_runtime_impl` skips D2H for
`IN`. Note the existing `graph_output_ptr` first-output-tensor heuristic in
`validate_runtime_impl` should be revisited together, since it currently assumes
the first iterated pair is an output.
Related: #796
Platform
All / Unknown (the copy logic is mirrored on both a5 and a2a3)
Runtime Variant
tensormap_and_ringbuffer
Summary
At the orchestration entry point, host↔device tensor transfer does not look at
the per-argument direction (
ArgDirection/TensorArgType:IN/OUT/INOUT). Every non-child_memorytensor is copied host→device at start anddevice→host at end, regardless of its direction. This wastes:
OUTPUTbuffers — the host-side content isuninitialized/garbage, so copying it up is meaningless.
INPUTbuffers — read-only inputs arenever modified by the kernel, so copying them back is wasted work.
Only
INOUTgenuinely needs both directions;child_memorytensors needneither (already on device).
The
signature(ArgDirection[]) is already plumbed end-to-end to the hostruntime, but is dropped at the consumption site — the comment in
c_api_shared.cppstates it is "plumbed end-to-end for per-tensor directiondecisions in runtime_maker but is currently unconsumed on both runtimes".
Git Commit ID
7c2f35f
Host Platform
Linux (aarch64)
Reproduction
The waste is observable directly in the host runtime's INFO logs — every tensor
is logged once on upload (H2D) and once on copy-back (D2H), regardless of
direction:
```bash
Run any example that has input-only tensors alongside output tensors.
python tests/st/a5/tensormap_and_ringbuffer//test_*.py -p a5 -d 1
In the host log:
bind_callable_to_runtime_impl -> "Tensor i: bytes at 0x..." (H2D upload, printed for EVERY tensor)
validate_runtime_impl -> "Tensor i: bytes copied to host" (D2H copy-back, printed for EVERY tensor)
=> pure-INPUT tensors are copied back, and pure-OUTPUT tensors are copied up.
```
Expected Performance
Copies should be gated on
ArgDirection:IN-> H2D copy-in only; no copy-back.OUT-> device buffer only (skip the H2D copy-in of garbage); D2H copy-back only.INOUT-> both directions (today's behavior).child_memory-> neither direction.For a typical kernel with mostly read-only inputs (weights, KV-cache) plus a
small output, this removes the dominant share of both the H2D and D2H traffic.
Actual Performance
All non-
child_memorytensors are copied both directions unconditionally:does `device_malloc` + `copy_to_device`, branching only on `is_child_memory()`,
never on direction. Pure outputs are uploaded too.
`tensor_pairs_` entry and does `copy_from_device`, gated only on null
host/dev pointers, never on direction. Pure inputs are copied back too.
Profiling Data (Optional)
Related closed issue #796 ("validate_runtime_impl is ~2x the chip-side cost per
run") measured ~50ms of copy-back overhead and explicitly hypothesized angle #1:
"Are all tensor pairs that validate_runtime_impl iterates actually outputs? If
the loop is also copying back inputs/weights/KV-cache entries ... that is the
bug." This issue is the root-cause counterpart: the direction info needed to
avoid those copies is available but unconsumed.
Additional Context
Code locations (mirrored on both runtimes):
Suggested fix: consume the already-plumbed `signature`/`sig_count` in
`bind_callable_to_runtime_impl` to skip H2D for `OUT`, and tag each
`tensor_pairs_` entry with its direction so `validate_runtime_impl` skips D2H for
`IN`. Note the existing `graph_output_ptr` first-output-tensor heuristic in
`validate_runtime_impl` should be revisited together, since it currently assumes
the first iterated pair is an output.
Related: #796