Component
EmitC / Codegen (lib/PTO/Transforms/PTOToEmitC.cpp)
Description
Since 0.51, pto.cmo.cacheinvalid <partition_tensor_view> single_cache_line is
lowered to a call that PTOAS's own emitted helper cannot compile.
The helper PTOAS emits is pointer-only:
template <typename Ptr>
static AICORE inline void PTOAS__DCCI_SINGLE_CACHE_LINE(Ptr ptr) {
dcci((__gm__ void*)ptr, cache_line_t::SINGLE_CACHE_LINE);
}
but for a partition_tensor_view operand the argument passed in is the
materialised GlobalTensor<...> object, not its address. GlobalTensor has no
conversion to __gm__ void*, so every kernel carrying this op fails to compile.
I checked pto-isa at both the pinned commit 83d01313 and current main
(439faf48): GlobalTensor exposes data() and SetAddr() and has no
conversion operator to __gm__ void* in either, so this is not a pto-isa version
skew — no pto-isa revision makes the emitted cast valid.
0.50 also accepted this IR but emitted no call at all for it (only the helper
template definition), so the op was silently a no-op. Either behaviour is a
problem, but the 0.51 one is a hard build break.
Impact on the PyPTO side: this takes down 36 dist-system-tests (all L3
collectives) and the DeepSeek-V4-Flash moe model example when moving the
toolchain pin from 0.50 to 0.51. PyPTO's InsertCommFence pass emits this op to
implement the publish side of the data-before-signal contract; we have had to
suspend emitting it entirely to get off 0.50.
Reproduction (minimal)
repro.pto:
module attributes {pto.target_arch = "a2a3"} {
func.func @k(%arg0: !pto.ptr<f32>) attributes {pto.kernel_kind = #pto.kernel_kind<vector>} {
%c16_index = arith.constant 16 : index
%c1_index = arith.constant 1 : index
%c0_index = arith.constant 0 : index
%view = pto.make_tensor_view %arg0, shape = [%c16_index, %c16_index], strides = [%c16_index, %c1_index] {layout = #pto.layout<nd>}: !pto.tensor_view<?x?xf32>
%pview = pto.partition_view %view, offsets = [%c0_index, %c0_index], sizes = [%c16_index, %c16_index] : !pto.tensor_view<?x?xf32> -> !pto.partition_tensor_view<16x16xf32>
pto.cmo.cacheinvalid %pview single_cache_line : !pto.partition_tensor_view<16x16xf32>
return
}
}
ptoas repro.pto -o repro.cpp --pto-level=level3
Then compile repro.cpp for the target (or inspect the emitted call directly —
the type error is visible without compiling).
Expected behavior
pto.cmo.cacheinvalid on a partition_tensor_view lowers to a dcci on the
view's base address — e.g. passing v7.data(), or giving the helper a
GlobalTensor overload that extracts the address.
If tensor-view operands are not intended to be supported at all, PTOAS should
reject the op with a diagnostic rather than emit uncompilable C++ (0.51) or
silently drop it (0.50).
Actual behavior / error logs
0.51 emits:
GlobalTensor<float, pto::Shape<1, 1, 1, 16, 16>, pto::Stride<256, 256, 256, 16, 1>, pto::Layout::ND> v7 = ...;
PTOAS__DCCI_SINGLE_CACHE_LINE(v7);
which fails to compile:
dispatch_meta.cpp:54:8: error: cannot cast from type
'pto::GlobalTensor<int, pto::Shape<1, 1, 1, 1, 32>, pto::Stride<32, 32, 32, 32, 1>, pto::Layout::ND>'
to pointer type '__gm__ void *'
dcci((__gm__ void*)ptr, cache_line_t::SINGLE_CACHE_LINE);
^~~~~~~~~~~~~~~~~
dispatch_meta.cpp:144:5: note: in instantiation of function template specialization
'PTOAS__DCCI_SINGLE_CACHE_LINE<pto::GlobalTensor<...>>' requested here
PTOAS__DCCI_SINGLE_CACHE_LINE(v45);
^
0.50, same input and flags, emits the helper template but no call site — diffing
the two outputs shows exactly one added line:
TSTORE(v13, v6);
+ PTOAS__DCCI_SINGLE_CACHE_LINE(v13);
Related: the pointer form is not a workaround
Feeding a raw pointer instead is rejected by both versions:
pto.cmo.cacheinvalid %ptr single_cache_line (no type annotation, which is what
PyPTO currently emits for the pointer form) — parse error expected ':' on 0.50
and 0.51 alike.
- with an explicit
: !pto.ptr<f32> it parses, then fails in a pass on both:
error: addptr must feed make_tensor_view, initialize_l2g2l_pipe(gm_addr) or load/store_scalar for lowering.
So there is currently no form of pto.cmo.cacheinvalid on a specific region that
produces working code. The whole-GM form
(pto.cmo.cacheinvalid all #pto.address_space<gm>) is fine — both versions lower
it identically to dcci((__gm__ void*)0, cache_line_t::ENTIRE_DATA_CACHE).
Git commit
Release binaries: ptoas 0.51 (ptoas-bin-aarch64.tar.gz,
sha256 39f3fb6da1fa1dc9f206d77f11f5be5b6412e1ac1a71b86e1555907e75bf3a32),
compared against ptoas 0.50
(sha256 acf7b316bedccf0689971d2dc92f9f80621ab5eba131d805317e01c766c1dc2c).
Target Ascend arch
a2a3
PTOAS build level
level3
Component
EmitC / Codegen (lib/PTO/Transforms/PTOToEmitC.cpp)
Description
Since 0.51,
pto.cmo.cacheinvalid <partition_tensor_view> single_cache_lineislowered to a call that PTOAS's own emitted helper cannot compile.
The helper PTOAS emits is pointer-only:
but for a
partition_tensor_viewoperand the argument passed in is thematerialised
GlobalTensor<...>object, not its address.GlobalTensorhas noconversion to
__gm__ void*, so every kernel carrying this op fails to compile.I checked pto-isa at both the pinned commit
83d01313and currentmain(
439faf48):GlobalTensorexposesdata()andSetAddr()and has noconversion operator to
__gm__ void*in either, so this is not a pto-isa versionskew — no pto-isa revision makes the emitted cast valid.
0.50 also accepted this IR but emitted no call at all for it (only the helper
template definition), so the op was silently a no-op. Either behaviour is a
problem, but the 0.51 one is a hard build break.
Impact on the PyPTO side: this takes down 36
dist-system-tests(all L3collectives) and the DeepSeek-V4-Flash
moemodel example when moving thetoolchain pin from 0.50 to 0.51. PyPTO's
InsertCommFencepass emits this op toimplement the publish side of the data-before-signal contract; we have had to
suspend emitting it entirely to get off 0.50.
Reproduction (minimal)
repro.pto:Then compile
repro.cppfor the target (or inspect the emitted call directly —the type error is visible without compiling).
Expected behavior
pto.cmo.cacheinvalidon apartition_tensor_viewlowers to adccion theview's base address — e.g. passing
v7.data(), or giving the helper aGlobalTensoroverload that extracts the address.If tensor-view operands are not intended to be supported at all, PTOAS should
reject the op with a diagnostic rather than emit uncompilable C++ (0.51) or
silently drop it (0.50).
Actual behavior / error logs
0.51 emits:
which fails to compile:
0.50, same input and flags, emits the helper template but no call site — diffing
the two outputs shows exactly one added line:
TSTORE(v13, v6); + PTOAS__DCCI_SINGLE_CACHE_LINE(v13);Related: the pointer form is not a workaround
Feeding a raw pointer instead is rejected by both versions:
pto.cmo.cacheinvalid %ptr single_cache_line(no type annotation, which is whatPyPTO currently emits for the pointer form) — parse error
expected ':'on 0.50and 0.51 alike.
: !pto.ptr<f32>it parses, then fails in a pass on both:error: addptr must feed make_tensor_view, initialize_l2g2l_pipe(gm_addr) or load/store_scalar for lowering.So there is currently no form of
pto.cmo.cacheinvalidon a specific region thatproduces working code. The whole-GM form
(
pto.cmo.cacheinvalid all #pto.address_space<gm>) is fine — both versions lowerit identically to
dcci((__gm__ void*)0, cache_line_t::ENTIRE_DATA_CACHE).Git commit
Release binaries:
ptoas 0.51(ptoas-bin-aarch64.tar.gz,sha256
39f3fb6da1fa1dc9f206d77f11f5be5b6412e1ac1a71b86e1555907e75bf3a32),compared against
ptoas 0.50(sha256
acf7b316bedccf0689971d2dc92f9f80621ab5eba131d805317e01c766c1dc2c).Target Ascend arch
a2a3
PTOAS build level
level3