Repository navigation
Unnecessary sign-extension for LeadingZeroCount(UInt64) #119699
Description
Activity
- addeduntriagedNew issue has not been triaged by the area ownerNew issue has not been triaged by the area owner
on Sep 14, 2025 - addedneeds-area-labelAn area label is needed to ensure this gets routed to the appropriate area ownersAn area label is needed to ensure this gets routed to the appropriate area owners
on Sep 14, 2025 - addedarea-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMICLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMIand removedneeds-area-labelAn area label is needed to ensure this gets routed to the appropriate area ownersAn area label is needed to ensure this gets routed to the appropriate area owners
on Sep 14, 2025 dotnet-policy-service commented
on Sep 14, 2025 ContributorMore actionsTagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.- addedhelp wanted[up-for-grabs] Good issue for external contributors[up-for-grabs] Good issue for external contributorsand removeduntriagedNew issue has not been triaged by the area ownerNew issue has not been triaged by the area owner
on Sep 15, 2025 Presumably should be trivial to fix, thanks for filing
Confusingly the intrinsic is now named
NI_AVX2_X64_LeadingZeroCount, despiteLZCNTnot being related to AVX2.While the intrinsic isn't technically part of the
AVX2CPUID bit, it is logically part of the AVX2 family of instructions alongside BMI1, BMI2, F16C, and FMA. That is, the ISA bits were all introduced simultaneously and have not appeared in hardware independently. -- Formally, Intel definesPOPCNTis part ofSSE4.2andLZCNTis part ofBMI1, they exist as separate CPUID bits for compat with AMD which had them in their own earlierABMinstruction set.The JIT took several simplifications last release to "merge" various logical ISA groupings together. This was done to significantly simplify the testing, support, and general codegen complexity matrices. Such logical groupings are always disabled or enabled together and this flows with the new "unified" versioning scheme that is intended to exist moving forward under AVX10.
In .NET 10, we have:
- CMOV+CX8+SSE+SSE2 - baseline, logically x86-64-v1
- SSE3+SSSE3+SSE4.1+SSE4.2+POPCNT - logically x86-64-v2
- AVX - kept because enough real world AVX only hardware existed, making it significant enough to support
- AVX2+BMI1+BMI2+F16C+FMA+LZCNT+MOVBE - logically x86-64-v3
- AVX512F+BW+CD+DQ+VL - logically x86-64-v4
We then define some "pseudo-versions" for AVX512v2 (IFMA+VBMI) and AVX512v3 (BITALG+VBMI2+VNNI+VPOPCNTDQ), which also represent the real world logical groupings. After that, we just have
AVX10v1,AVX10v2, and will continue versioning this way in the future.In .NET 11 we have raised the baseline to x86-64-v2 and so
NI_SSE42_*no longer exists, it is simply part ofNI_X86Base_*@tannergooding The naming here does feel a bit misleading, especially considering the existence of vector leading-zero count instructions.
Look like the sign-extension for the
M2case was elided starting from .NET 8.It's no different than other scalar vs vector form instructions that exist in a given ISA. We differentiate where important/relevant
We can always do so here in the future if it becomes important. For now, it was the change that allowed simplification of the JIT without also being a massively in depth refactoring
I have looked into this and compiled the first
LeadingZeroCountclass.This is the corresponding GenTree:
[000004] ----------- * RETURN long [000003] ----------- \--* CAST long <- int [000002] ---------U- \--* CAST int <- ulong [000001] ----------- \--* HWINTRINSIC long 0 LeadingZeroCount [000000] ----------- \--* LCL_VAR long V00 arg0
Obviously this is the node that generates a
cdqe:CAST long <- int
Help appreciated here: not too sure about what is best?
-
we morph the double cast into a single
long <- ulongcast without sign extension IFFF we know that the intermediate type (inthere) is always positive (as is the case here withLeadingZeroCount)?
A bit similar to what is done here:
https://github.com/dotnet/runtime/blob/main/src/coreclr/jit/morph.cpp#L369 -
or maybe we should skip emitting the
cdqestraight at code generation time, and hardcodelzcnt/popcnt/tzcntinstructions here?
https://github.com/dotnet/runtime/blob/main/src/coreclr/jit/emitxarch.cpp#L1036 -
any other suggestion?
Also tagging @saucecontrol because this looks relatively similar to what we did with
popcntin my other PR.-
Option 1 sounds reasonable to me. You'd need to be able to prove not only that the value is non-negative but also that the value would fit in the intermediate type. The range assertions on PopCount and pals do that, so you should be good.
Reacted by LotendanI have looked into it this evening and there are a couple of things I'm unsure of:
I'm not sure to understand why the first
CASTis marked asUunsigned although the return type fromLeadingZeroCountislong.
It is important becauseIntegralRange::ForCastOutputpropagates the "non-negative" flag through a chain of casts:
runtime/src/coreclr/jit/assertionprop.cpp
Line 490 in 56ce9bd
/* static */ IntegralRange IntegralRange::ForCastOutput(GenTreeCast* cast, Compiler* compiler) Also, any idea about why this condition is limited to
TYP_INT?
runtime/src/coreclr/jit/assertionprop.cpp
Line 535 in 56ce9bd
if ((fromType == TYP_INT) && fromUnsigned) If I understand properly, propagating the
unsignedin the second cast would allow to use zero-extension instead of sign-extension.Thanks
The type you see on each node in the IR is what's called the JIT type, which is defined in the 3rd column of this table:
runtime/src/coreclr/jit/typelist.h
Lines 40 to 50 in 56ce9bd
DEF_TP(BYTE ,"byte" , TYP_INT, 1, 1, 4, 1, 1, VTR_INT, availableIntRegs, RBM_INT_CALLEE_SAVED, RBM_INT_CALLEE_TRASH, VTF_INT) DEF_TP(UBYTE ,"ubyte" , TYP_INT, 1, 1, 4, 1, 1, VTR_INT, availableIntRegs, RBM_INT_CALLEE_SAVED, RBM_INT_CALLEE_TRASH, VTF_INT|VTF_UNS) DEF_TP(SHORT ,"short" , TYP_INT, 2, 2, 4, 1, 2, VTR_INT, availableIntRegs, RBM_INT_CALLEE_SAVED, RBM_INT_CALLEE_TRASH, VTF_INT) DEF_TP(USHORT ,"ushort" , TYP_INT, 2, 2, 4, 1, 2, VTR_INT, availableIntRegs, RBM_INT_CALLEE_SAVED, RBM_INT_CALLEE_TRASH, VTF_INT|VTF_UNS) DEF_TP(INT ,"int" , TYP_INT, 4, 4, 4, 1, 4, VTR_INT, availableIntRegs, RBM_INT_CALLEE_SAVED, RBM_INT_CALLEE_TRASH, VTF_INT|VTF_I32) DEF_TP(UINT ,"uint" , TYP_INT, 4, 4, 4, 1, 4, VTR_INT, availableIntRegs, RBM_INT_CALLEE_SAVED, RBM_INT_CALLEE_TRASH, VTF_INT|VTF_UNS|VTF_I32) // Only used in GT_CAST nodes DEF_TP(LONG ,"long" , TYP_LONG, 8,EPS,EPS, 2, 8, VTR_INT, availableIntRegs, RBM_INT_CALLEE_SAVED, RBM_INT_CALLEE_TRASH, VTF_INT|VTF_I64) DEF_TP(ULONG ,"ulong" , TYP_LONG, 8,EPS,EPS, 2, 8, VTR_INT, availableIntRegs, RBM_INT_CALLEE_SAVED, RBM_INT_CALLEE_TRASH, VTF_INT|VTF_UNS|VTF_I64) // Only used in GT_CAST nodes You'll see that both
longandulonghave a JIT type ofTYP_LONG. Likewise, all primitive types 4 bytes or smaller have a JIT type ofTYP_INT.For HWIntrinsic nodes like
LeadingZeroCount, the real data type (ulongin this case) is stored in SimdBaseType, while the node itself is the JIT type (longin this case, because it returns a scalar value).Whether a cast is sign extending or not is determined by either the
fromtype of the cast (for small types, likeubyteandushort) or the unsigned flag on the node (foruintandulong), but the node itself will have the JIT type.Hope that clears it up 😄
I think you're looking in the right place, because it looks to me like we're ignoring the fact that your cast operand brought its own narrow range (0 to 127) assertion in, and it's leaving with a wider range (int.MinValue to int.MaxValue).
Reacted by LotendanLooks like this might be addressed by #128658
There's a few different range check things I'm working on atm, multiple have the chance to address this issue. We'll see which land or not.
- added 4 commits that reference this issue
on Aug 17, 2026
godbolt.org