Repository navigation
Suboptimal ASM emmited for Vector256<T>.Zero and Vector128<T>.Zero #76067
Description
Activity
- ghost addedarea-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMICLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI
on Sep 23, 2022 - ghost addeduntriagedNew issue has not been triaged by the area ownerNew issue has not been triaged by the area owner
on Sep 23, 2022 This is a known issue that we might eventually address with constant materialization during LSRA.
We don't currently do CSE forVector<>.Zerobecause 1) it's cheap to materialize 2) we plan to rely on it in optimizations in lower*.cpp- removeduntriagedNew issue has not been triaged by the area ownerNew issue has not been triaged by the area owner
on Sep 23, 2022 Thanks for the info @EgorBo
Discussed with @tannergooding for a bit regarding this, as an immediate thing we could do, is a peephole optimization to eliminate unnecessary
vxorps ymm1, ymm1, ymm1instructions by looking back to see if the same instruction occurred, as long as we know thatymm1was not written to in any other way.Another manifistation of this problem:
void Foo(ref byte b) { Unsafe.As<byte, Vector128<byte>>(ref Unsafe.Add(ref b, 0)) = default; Unsafe.As<byte, Vector128<byte>>(ref Unsafe.Add(ref b, 16)) = default; Unsafe.As<byte, Vector128<byte>>(ref Unsafe.Add(ref b, 32)) = default; Unsafe.As<byte, Vector128<byte>>(ref Unsafe.Add(ref b, 48)) = default; }
Emits:
movi v16.4s, #0 str q16, [x1] movi v16.4s, #0 str q16, [x1, #0x10] movi v16.4s, #0 str q16, [x1, #0x20] movi v16.4s, #0 str q16, [x1, #0x30]
Expected:
movi v16.4s, #0 str q16, q16, [x1] str q16, q16, [x1, #0x20]
eliminate unnecessary vxorps ymm1, ymm1, ymm1 instructions by looking back to see if the same instruction occurred, as long as we know that ymm1 was not written to in any other way.
TLDR: that didn't work due to ABI issues, although, it won't be a real fix anyway because it wouldn't help with loop hoisting.
When using Vector256.Zero, I would expect it to be kept in a fixed register and reused. Instead, what I see is a
vxorpsoperation emitted every time.AVX2 has 16 YMM registers.
Bellow is one example. I can get the desired behavior by forcing a zero vector variable instead of using Vector256.Zero;
Assigning
Vector256<byte>.Zeroto a variable alone does not do the trick. Only the extra xor operation ensures it stays in a fixed register.category:cq
theme:cse
skill-level:intermediate
cost:medium
impact:small