Skip to content

JIT: LSRA never reuses an enregistered 256/512-bit vector constant on windows-x64, even in call-free code #134540

Description

@AndyAyersMS

LinearScan::isMatchingConstant (src/coreclr/jit/lsra.cpp) rejects every GT_CNS_VEC whose register type needs a partial callee save:

case GT_CNS_VEC:
{
    return
#if FEATURE_PARTIAL_SIMD_CALLEE_SAVE
        !Compiler::varTypeNeedsPartialCalleeSave(physRegRecord->assignedInterval->registerType) &&
#endif
        GenTreeVecCon::Equals(refPosition->treeNode->AsVecCon(), otherTreeNode->AsVecCon());
}

Compiler::varTypeNeedsPartialCalleeSave returns true for every TYP_SIMD32/TYP_SIMD64 on x64 and for TYP_SIMD16/TYP_SIMD12 on ARM64, so this is a pure type test — it never asks whether a clobber actually occurred. The guard was added by #74110 for a genuine hazard: the Windows x64 ABI preserves only the low 128 bits of xmm6–xmm15, and freeKilledRegs clears the constant bit only for the registers in a call's kill set (RBM_FLT_CALLEE_TRASH = xmm0–xmm5), so a wide constant parked in ymm6–ymm15 would survive a call in LSRA's bookkeeping while its upper half was destroyed. But the rejection also fires when no call — and therefore no clobber — can have happened.

Two consequences:

  • The identical Vector128 source shape already reuses the register today, so the JIT emits strictly more code for the wider type.
  • FEATURE_PARTIAL_SIMD_CALLEE_SAVE is 0 under UNIX_AMD64_ABI (src/coreclr/jit/targetamd64.h), so the whole guard is compiled out on linux-x64 and the reuse already happens there. Windows-x64 emits more instructions than linux-x64 for the same source.

Minimal repro

using System;
using System.Runtime.CompilerServices;
using System.Runtime.Intrinsics;
using System.Runtime.Intrinsics.X86;

public static class VecRepro
{
    [MethodImpl(MethodImplOptions.NoInlining)]
    public static bool TwoEmpty256(Vector256<int> a, Vector256<int> b)
        => Avx.TestZ(a, Vector256<int>.AllBitsSet)
         & Avx.TestZ(b, Vector256<int>.AllBitsSet);

    // Control: the identical 128-bit shape already reuses the register today.
    [MethodImpl(MethodImplOptions.NoInlining)]
    public static bool TwoEmpty128(Vector128<int> a, Vector128<int> b)
        => Sse41.TestZ(a, Vector128<int>.AllBitsSet)
         & Sse41.TestZ(b, Vector128<int>.AllBitsSet);

    [MethodImpl(MethodImplOptions.NoInlining)]
    public static int CountPairs256(Vector256<int>[] a, Vector256<int>[] b)
    {
        int n = 0;
        for (int i = 0; i < a.Length; i++)
        {
            if (Avx.TestZ(a[i], Vector256<int>.AllBitsSet)
              & Avx.TestZ(b[i], Vector256<int>.AllBitsSet))
            {
                n++;
            }
        }
        return n;
    }

    [MethodImpl(MethodImplOptions.NoInlining)]
    private static int Opaque(int v) => v + 1;

    // Negative case: the constant must NOT be reused across a call, because a call
    // destroys the upper half of the partially callee-saved float registers.
    [MethodImpl(MethodImplOptions.NoInlining)]
    public static bool AcrossCall256(Vector256<int> a, Vector256<int> b, out int sink)
    {
        bool first = Avx.TestZ(a, Vector256<int>.AllBitsSet);
        sink = Opaque(first ? 1 : 0);
        return first & Avx.TestZ(b, Vector256<int>.AllBitsSet);
    }

    public static int Main()
    {
        if (!Avx2.IsSupported)
        {
            Console.WriteLine("AVX2 required.");
            return 100;
        }

        var a256 = new Vector256<int>[8];
        var b256 = new Vector256<int>[8];
        var a128 = new Vector128<int>[8];
        a256[3] = Vector256.Create(7);
        a128[3] = Vector128.Create(7);

        Console.WriteLine($"TwoEmpty256   = {TwoEmpty256(a256[0], a256[3])}");
        Console.WriteLine($"TwoEmpty128   = {TwoEmpty128(a128[0], a128[3])}");
        Console.WriteLine($"CountPairs256 = {CountPairs256(a256, b256)}");
        Console.WriteLine($"AcrossCall256 = {AcrossCall256(a256[0], a256[3], out int s)} {s}");
        return 100;
    }
}

Run with DOTNET_TieredCompilation=0, DOTNET_ReadyToRun=0 and DOTNET_JitDisasm=TwoEmpty256 TwoEmpty128 CountPairs256 on windows-x64. Reproduced both with DOTNET_EnableAVX512=0 (x64 + VEX) and with AVX-512 enabled (x64 + VEX + EVEX) — the listings are byte-identical, so this is not an AVX-512 artifact.

Current codegen

TwoEmpty256 — 36 bytes, PerfScore 22.75, 11 instructions:

       vpcmpeqd ymm0, ymm0, ymm0        ; constant #1
       vptest   ymm0, ymmword ptr [rcx]
       sete     al
       movzx    rax, al
       vpcmpeqd ymm0, ymm0, ymm0        ; redundant: ymm0 already holds all-ones
       vptest   ymm0, ymmword ptr [rdx]

The redundancy also lands inside loop bodies — the CountPairs256 loop block is 41 bytes against 37 for the 128-bit control, at bbWeight 3.96.

Expected codegen

What the JIT already produces for the 128-bit control TwoEmpty128 (29 bytes, PerfScore 15.25, 9 instructions), and what a prototype relaxation produces for TwoEmpty256 (32 bytes, PerfScore 22.25, 10 instructions):

       vpcmpeqd ymm0, ymm0, ymm0
       vptest   ymm0, ymmword ptr [rcx]
       sete     al
       movzx    rax, al
       vptest   ymm0, ymmword ptr [rdx]  ; register reused

Impact

Measured on 4856f0c16f89d8cc625a62ee7bab7ed4afe14f9d, windows-x64, with a prototype that replaces the type test with a clobber test (patch below). All numbers are static — code size, instruction count, and PerfScore.

Repro methods (base → prototype):

Method Bytes PerfScore Instructions
TwoEmpty256 36 → 32 (−11.1%) 22.75 → 22.25 11 → 10
FourEmpty256 70 → 58 (−17.1%) 43.75 → 42.25 21 → 18
CountPairs256 169 → 161 (−4.7%) 102.23 → 100.23 50 → 48
TwoEmpty128 / FourEmpty128 / CountPairs128 (controls) unchanged unchanged unchanged
AcrossCall256 (negative case) 74 → 74 33.25 → 33.25 26 → 26

SuperPMI asm diffs, checked base JIT vs. checked prototype JIT, 12 default collections for JIT-EE fbbaf45f-5b0e-4767-b962-5084c8caae77.windows.x64:

Collection Contexts with diffs Size impr / regr Bytes
benchmarks.run 9 9 / 0 −190
benchmarks.run_pgo 1 1 / 0 −4
benchmarks.run_pgo_optrepeat 9 9 / 0 −190
coreclr_tests.run 27 27 / 0 −282
libraries.crossgen2 9 9 / 0 −40
libraries.pmi 14 14 / 0 −104
libraries_tests.run 75 75 / 0 −498
libraries_tests_no_tiered_compilation.run 230 230 / 0 −1,210
realworld.run 45 45 / 0 −1,400
aspire.nativeaot / aspnet2.run / smoke_tests.nativeaot 0 – –
Total 419 419 / 0 −3,918

Across those 419 contexts: 411 PerfScore improvements, 0 PerfScore regressions, 8 unchanged. Largest individual wins:

Base → diff PerfScore Method
1325 → 1197 (−128) 510.71 → 491.04 TensorPrimitives:<Aggregate>g__Vectorized256|72_2
937 → 833 (−104) 135.58 → 126.92 BepuPhysics TwoBodyConstraintBenchmarks:Contact4Nonconvex() (26 redundant vxorps ymm0, ymm0, ymm0 removed, 146 → 120 instructions)
900 → 816 (−84) 286.92 → 279.92 BepuPhysics OneBodyConstraintBenchmarks:Contact4NonconvexOneBody()

Known regressions and measurement limitations:

  • Code size: no size or PerfScore regression appeared in any of the 12 collections in this run. A separate Release-JIT diff pair, run against a baseline clrjit.dll binary built in a different environment, reported 2 regressing contexts (+253 bytes) on libraries_tests.run with asymmetric missing-data counts; the same signature appeared for unrelated patches in the same batch, so it is attributed to baseline build provenance rather than to this change. It has not been independently reproduced against a same-environment baseline.
  • JIT throughput: tpdiff (PIN, and against that same externally built baseline) shows MinOpts +0.40% to +0.94% on all 12 collections, with overall −0.05% to +0.18% and FullOpts −0.10% to +0.12%. The prototype does add MinOpts work for no MinOpts benefit — processKills and updateAssignedInterval are shared with allocateRegistersMinimal, while the reuse itself is gated on opts.OptimizationEnabled() — so this should not be dismissed as baseline noise without a same-environment re-run. A real fix should probably gate the new bookkeeping on OptimizationEnabled(); that variant was not measured.
  • No runtime measurement. No BenchmarkDotNet or other timed benchmark was run. vpcmpeqd r,r,r and vxorps r,r,r are dependency-breaking idioms, so the wall-clock effect is expected to be small even where a hot loop loses an instruction per iteration. Nothing here supports a speedup claim.
  • windows-x64 only. ARM64 and linux-x64 were not measured.
  • No JitStress / JitStressRegs / GC-stress runs; the assert-enabled checked SuperPMI replay over all 12 collections is the only failure gate that was exercised.

Notes

Prototype patch

Experimental only — it has not been through stress testing, it is windows-x64-measured only, and the MinOpts throughput question above is unresolved.

Experimental patch
diff --git a/src/coreclr/jit/lsra.cpp b/src/coreclr/jit/lsra.cpp
index faf1b4b10c8..37889cf9403 100644
--- a/src/coreclr/jit/lsra.cpp
+++ b/src/coreclr/jit/lsra.cpp
@@ -2757,11 +2757,18 @@ bool LinearScan::isMatchingConstant(RegRecord* physRegRecord, RefPosition* refPo
 #if defined(FEATURE_SIMD)
         case GT_CNS_VEC:
         {
-            return
 #if FEATURE_PARTIAL_SIMD_CALLEE_SAVE
-                !Compiler::varTypeNeedsPartialCalleeSave(physRegRecord->assignedInterval->registerType) &&
+            // Only part of a wide vector register is preserved across a call, so such a register
+            // may only be reused for an identical constant if no call-like kill has occurred
+            // since the constant was materialized.
+            if ((Compiler::varTypeNeedsPartialCalleeSave(physRegRecord->assignedInterval->registerType) ||
+                 Compiler::varTypeNeedsPartialCalleeSave(interval->registerType)) &&
+                !isWideConstantRegValid(physRegRecord->regNum, physRegRecord->assignedInterval->registerType))
+            {
+                return false;
+            }
 #endif
-                GenTreeVecCon::Equals(refPosition->treeNode->AsVecCon(), otherTreeNode->AsVecCon());
+            return GenTreeVecCon::Equals(refPosition->treeNode->AsVecCon(), otherTreeNode->AsVecCon());
         }
 #endif // FEATURE_SIMD
 
@@ -3831,6 +3838,17 @@ void LinearScan::processKills(RefPosition* killRefPosition)
 #endif
 
     regsBusyUntilKill &= ~killRefPosition->getKilledRegisters();
+
+#if FEATURE_PARTIAL_SIMD_CALLEE_SAVE
+    // A call trashes the non-preserved part of every float register, including the ones that are
+    // not in its kill set (only the low part of the callee-saved float registers is preserved).
+    // Any kill that trashes float registers therefore invalidates every wide constant register.
+    if (!(killedRegs & RBM_FLT_CALLEE_TRASH).IsEmpty())
+    {
+        m_WideConstantsValid = RBM_NONE;
+    }
+#endif // FEATURE_PARTIAL_SIMD_CALLEE_SAVE
+
     INDEBUG(dumpLsraAllocationEvent(LSRA_EVENT_KILL_REGS, nullptr, REG_NA, nullptr, NONE,
                                     killRefPosition->getKilledRegisters()));
 }
@@ -6850,6 +6868,16 @@ void LinearScan::updateAssignedInterval(RegRecord* reg, Interval* interval ARM_A
     if (interval->isConstant)
     {
         setConstantReg(reg->regNum, interval->registerType);
+#if FEATURE_PARTIAL_SIMD_CALLEE_SAVE
+        if (Compiler::varTypeNeedsPartialCalleeSave(interval->registerType))
+        {
+            setWideConstantReg(reg->regNum, interval->registerType);
+        }
+        else
+        {
+            clearWideConstantReg(reg->regNum, interval->registerType);
+        }
+#endif
     }
     else
     {
diff --git a/src/coreclr/jit/lsra.h b/src/coreclr/jit/lsra.h
index 3541e1a2309..03ee7b2aae0 100644
--- a/src/coreclr/jit/lsra.h
+++ b/src/coreclr/jit/lsra.h
@@ -1763,6 +1763,9 @@ private:
     {
         m_AvailableRegs          = allAvailableRegs;
         m_RegistersWithConstants = RBM_NONE;
+#if FEATURE_PARTIAL_SIMD_CALLEE_SAVE
+        m_WideConstantsValid = RBM_NONE;
+#endif
     }
 
     bool isRegAvailable(regNumber reg, var_types regType)
@@ -1802,9 +1805,30 @@ private:
                                                  DEBUG_ARG(regNumber assignedReg));
 
     regMaskTP m_RegistersWithConstants;
-    void      clearConstantReg(regNumber reg, var_types regType)
+#if FEATURE_PARTIAL_SIMD_CALLEE_SAVE
+    // The subset of `m_RegistersWithConstants` that holds a constant whose register is only
+    // partially preserved across a call. Such a register may only be reused for an identical
+    // constant while no call-like kill has occurred since the constant was materialized.
+    regMaskTP m_WideConstantsValid;
+    void      setWideConstantReg(regNumber reg, var_types regType)
+    {
+        m_WideConstantsValid.AddRegNum(reg, regType);
+    }
+    void clearWideConstantReg(regNumber reg, var_types regType)
+    {
+        m_WideConstantsValid.RemoveRegNum(reg, regType);
+    }
+    bool isWideConstantRegValid(regNumber reg, var_types regType)
+    {
+        return m_WideConstantsValid.IsRegNumPresent(reg, regType);
+    }
+#endif // FEATURE_PARTIAL_SIMD_CALLEE_SAVE
+    void clearConstantReg(regNumber reg, var_types regType)
     {
         m_RegistersWithConstants.RemoveRegNum(reg, regType);
+#if FEATURE_PARTIAL_SIMD_CALLEE_SAVE
+        m_WideConstantsValid.RemoveRegNum(reg, regType);
+#endif
     }
     void setConstantReg(regNumber reg, var_types regType)
     {

Note

This issue was generated with GitHub Copilot.

Activity

  1. dotnet-policy-service commented on Sep 23, 2026

    @dotnet-policy-service
    Contributor

    Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
    See info in area-owners.md if you want to be subscribed.

  2. added this to the 12.0.0 milestone on Sep 25, 2026
  3. removed
    untriagedNew issue has not been triaged by the area owner
    on Sep 25, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMIperformance

    Type

    No type

    Projects

    No projects

      Milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions