Skip to content

[Bug] 0.51 lowers cmo.cacheinvalid on a partition_tensor_view to an uncompilable DCCI call #995

Description

@luohuan19

Component

EmitC / Codegen (lib/PTO/Transforms/PTOToEmitC.cpp)

Description

Since 0.51, pto.cmo.cacheinvalid <partition_tensor_view> single_cache_line is
lowered to a call that PTOAS's own emitted helper cannot compile.

The helper PTOAS emits is pointer-only:

template <typename Ptr>
static AICORE inline void PTOAS__DCCI_SINGLE_CACHE_LINE(Ptr ptr) {
  dcci((__gm__ void*)ptr, cache_line_t::SINGLE_CACHE_LINE);
}

but for a partition_tensor_view operand the argument passed in is the
materialised GlobalTensor<...> object, not its address. GlobalTensor has no
conversion to __gm__ void*, so every kernel carrying this op fails to compile.

I checked pto-isa at both the pinned commit 83d01313 and current main
(439faf48): GlobalTensor exposes data() and SetAddr() and has no
conversion operator to __gm__ void* in either, so this is not a pto-isa version
skew — no pto-isa revision makes the emitted cast valid.

0.50 also accepted this IR but emitted no call at all for it (only the helper
template definition), so the op was silently a no-op. Either behaviour is a
problem, but the 0.51 one is a hard build break.

Impact on the PyPTO side: this takes down 36 dist-system-tests (all L3
collectives) and the DeepSeek-V4-Flash moe model example when moving the
toolchain pin from 0.50 to 0.51. PyPTO's InsertCommFence pass emits this op to
implement the publish side of the data-before-signal contract; we have had to
suspend emitting it entirely to get off 0.50.

Reproduction (minimal)

repro.pto:

module attributes {pto.target_arch = "a2a3"} {
  func.func @k(%arg0: !pto.ptr<f32>) attributes {pto.kernel_kind = #pto.kernel_kind<vector>} {
  %c16_index = arith.constant 16 : index
  %c1_index = arith.constant 1 : index
  %c0_index = arith.constant 0 : index
  %view = pto.make_tensor_view %arg0, shape = [%c16_index, %c16_index], strides = [%c16_index, %c1_index] {layout = #pto.layout<nd>}: !pto.tensor_view<?x?xf32>
  %pview = pto.partition_view %view, offsets = [%c0_index, %c0_index], sizes = [%c16_index, %c16_index] : !pto.tensor_view<?x?xf32> -> !pto.partition_tensor_view<16x16xf32>
  pto.cmo.cacheinvalid %pview single_cache_line : !pto.partition_tensor_view<16x16xf32>
  return
  }
}
ptoas repro.pto -o repro.cpp --pto-level=level3

Then compile repro.cpp for the target (or inspect the emitted call directly —
the type error is visible without compiling).

Expected behavior

pto.cmo.cacheinvalid on a partition_tensor_view lowers to a dcci on the
view's base address — e.g. passing v7.data(), or giving the helper a
GlobalTensor overload that extracts the address.

If tensor-view operands are not intended to be supported at all, PTOAS should
reject the op with a diagnostic rather than emit uncompilable C++ (0.51) or
silently drop it (0.50).

Actual behavior / error logs

0.51 emits:

GlobalTensor<float, pto::Shape<1, 1, 1, 16, 16>, pto::Stride<256, 256, 256, 16, 1>, pto::Layout::ND> v7 = ...;
PTOAS__DCCI_SINGLE_CACHE_LINE(v7);

which fails to compile:

dispatch_meta.cpp:54:8: error: cannot cast from type
  'pto::GlobalTensor<int, pto::Shape<1, 1, 1, 1, 32>, pto::Stride<32, 32, 32, 32, 1>, pto::Layout::ND>'
  to pointer type '__gm__ void *'
  dcci((__gm__ void*)ptr, cache_line_t::SINGLE_CACHE_LINE);
       ^~~~~~~~~~~~~~~~~
dispatch_meta.cpp:144:5: note: in instantiation of function template specialization
  'PTOAS__DCCI_SINGLE_CACHE_LINE<pto::GlobalTensor<...>>' requested here
    PTOAS__DCCI_SINGLE_CACHE_LINE(v45);
    ^

0.50, same input and flags, emits the helper template but no call site — diffing
the two outputs shows exactly one added line:

   TSTORE(v13, v6);
+  PTOAS__DCCI_SINGLE_CACHE_LINE(v13);

Related: the pointer form is not a workaround

Feeding a raw pointer instead is rejected by both versions:

  • pto.cmo.cacheinvalid %ptr single_cache_line (no type annotation, which is what
    PyPTO currently emits for the pointer form) — parse error expected ':' on 0.50
    and 0.51 alike.
  • with an explicit : !pto.ptr<f32> it parses, then fails in a pass on both:
    error: addptr must feed make_tensor_view, initialize_l2g2l_pipe(gm_addr) or load/store_scalar for lowering.

So there is currently no form of pto.cmo.cacheinvalid on a specific region that
produces working code. The whole-GM form
(pto.cmo.cacheinvalid all #pto.address_space<gm>) is fine — both versions lower
it identically to dcci((__gm__ void*)0, cache_line_t::ENTIRE_DATA_CACHE).

Git commit

Release binaries: ptoas 0.51 (ptoas-bin-aarch64.tar.gz,
sha256 39f3fb6da1fa1dc9f206d77f11f5be5b6412e1ac1a71b86e1555907e75bf3a32),
compared against ptoas 0.50
(sha256 acf7b316bedccf0689971d2dc92f9f80621ab5eba131d805317e01c766c1dc2c).

Target Ascend arch

a2a3

PTOAS build level

level3

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't working

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions