Skip to content

[Performance] Orchestration entry copies every tensor host↔device both directions, ignoring ArgDirection #1047

Description

@ChaoZheng109

Platform

All / Unknown (the copy logic is mirrored on both a5 and a2a3)

Runtime Variant

tensormap_and_ringbuffer

Summary

At the orchestration entry point, host↔device tensor transfer does not look at
the per-argument direction
(ArgDirection / TensorArgType: IN / OUT /
INOUT). Every non-child_memory tensor is copied host→device at start and
device→host at end, regardless of its direction. This wastes:

  • H2D bandwidth uploading pure OUTPUT buffers — the host-side content is
    uninitialized/garbage, so copying it up is meaningless.
  • D2H bandwidth copying back pure INPUT buffers — read-only inputs are
    never modified by the kernel, so copying them back is wasted work.

Only INOUT genuinely needs both directions; child_memory tensors need
neither (already on device).

The signature (ArgDirection[]) is already plumbed end-to-end to the host
runtime, but is dropped at the consumption site — the comment in
c_api_shared.cpp states it is "plumbed end-to-end for per-tensor direction
decisions in runtime_maker but is currently unconsumed on both runtimes".

Git Commit ID

7c2f35f

Host Platform

Linux (aarch64)

Reproduction

The waste is observable directly in the host runtime's INFO logs — every tensor
is logged once on upload (H2D) and once on copy-back (D2H), regardless of
direction:

```bash

Run any example that has input-only tensors alongside output tensors.

python tests/st/a5/tensormap_and_ringbuffer//test_*.py -p a5 -d 1

In the host log:

bind_callable_to_runtime_impl -> "Tensor i: bytes at 0x..." (H2D upload, printed for EVERY tensor)

validate_runtime_impl -> "Tensor i: bytes copied to host" (D2H copy-back, printed for EVERY tensor)

=> pure-INPUT tensors are copied back, and pure-OUTPUT tensors are copied up.

```

Expected Performance

Copies should be gated on ArgDirection:

  • IN -> H2D copy-in only; no copy-back.
  • OUT -> device buffer only (skip the H2D copy-in of garbage); D2H copy-back only.
  • INOUT -> both directions (today's behavior).
  • child_memory -> neither direction.

For a typical kernel with mostly read-only inputs (weights, KV-cache) plus a
small output, this removes the dominant share of both the H2D and D2H traffic.

Actual Performance

All non-child_memory tensors are copied both directions unconditionally:

  • H2D (start) — `bind_callable_to_runtime_impl` loops over every tensor and
    does `device_malloc` + `copy_to_device`, branching only on `is_child_memory()`,
    never on direction. Pure outputs are uploaded too.
  • D2H (end) — `validate_runtime_impl` loops over every recorded
    `tensor_pairs_` entry and does `copy_from_device`, gated only on null
    host/dev pointers, never on direction. Pure inputs are copied back too.

Profiling Data (Optional)

Related closed issue #796 ("validate_runtime_impl is ~2x the chip-side cost per
run") measured ~50ms of copy-back overhead and explicitly hypothesized angle #1:
"Are all tensor pairs that validate_runtime_impl iterates actually outputs? If
the loop is also copying back inputs/weights/KV-cache entries ... that is the
bug." This issue is the root-cause counterpart: the direction info needed to
avoid those copies is available but unconsumed.

Additional Context

Code locations (mirrored on both runtimes):

  • a5: `src/a5/runtime/tensormap_and_ringbuffer/host/runtime_maker.cpp`
    • `bind_callable_to_runtime_impl` (H2D loop) — signature param is `const ArgDirection * /signature/` (unused)
    • `validate_runtime_impl` (D2H loop)
  • a2a3: `src/a2a3/runtime/tensormap_and_ringbuffer/host/runtime_maker.cpp` (same structure)
  • Plumbing: `src/common/platform/onboard/host/c_api_shared.cpp` passes `bind_result.signature` into `bind_callable_to_runtime_impl`.

Suggested fix: consume the already-plumbed `signature`/`sig_count` in
`bind_callable_to_runtime_impl` to skip H2D for `OUT`, and tag each
`tensor_pairs_` entry with its direction so `validate_runtime_impl` skips D2H for
`IN`. Note the existing `graph_output_ptr` first-output-tensor heuristic in
`validate_runtime_impl` should be revisited together, since it currently assumes
the first iterated pair is an output.

Related: #796

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

performancePerformance regression or optimization

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions