Skip to content

Vector2.Dot and Vector2.LengthSquared slower on .NET 11. Also for Vector3 and Vector4. #133297

Description

@jwdj

Description

Vector2.Dot and Vector2.LengthSquared are considerably slower in .NET 11 compared to .NET 10
The same is true for Vector3 en Vector4.

Regression

Vector2.Dot and Vector2.LengthSquared about 16% slower on .NET 11
Vector3.Dot and Vector3.LengthSquared about 14% slower on .NET 11
Vector4.Dot and Vector4.LengthSquared about 25% slower on .NET 11
I ran the .NET 10 and .NET 11 benchmark separately so these numbers might not be exact, but at least it is a double digit performance regression.
This is compared to .NET 10 on a modern Intel Laptop processor.

.NET 11

BenchmarkDotNet v0.15.8, Windows 11 (10.0.26200.9168/25H2/2025Update/HudsonValley2)
Intel Core Ultra 9 275HX 2.70GHz, 1 CPU, 24 logical and 24 physical cores
.NET SDK 11.0.100-preview.7.26381.103
[Host] : .NET 11.0.0 (11.0.0-preview.7.26381.103, 11.0.26.38203), X64 RyuJIT x86-64-v3

Method Mean Error StdDev Allocated
Vector2_LengthSquared_Benchmark 582.8 ns 1.41 ns 1.32 ns -
Vector2_Dot_Benchmark 581.7 ns 1.99 ns 1.86 ns -
Manual2_Benchmark 379.6 ns 1.27 ns 0.99 ns -
Vector3_LengthSquared_Benchmark 691.8 ns 2.47 ns 2.31 ns -
Vector3_Dot_Benchmark 691.3 ns 0.99 ns 0.93 ns -
Manual3_Benchmark 716.5 ns 1.07 ns 1.00 ns -
Vector4_LengthSquared_Benchmark 496.1 ns 1.16 ns 1.09 ns -
Vector4_Dot_Benchmark 487.9 ns 6.21 ns 5.81 ns -
Manual4_Benchmark 673.2 ns 0.47 ns 0.44 ns -

// * Hints *
Outliers
UserQuery.Manual2_Benchmark: Default -> 3 outliers were removed (383.84 ns..390.57 ns)
UserQuery.Vector3_LengthSquared_Benchmark: Default -> 1 outlier was detected (684.36 ns)
UserQuery.Manual3_Benchmark: Default -> 1 outlier was detected (715.70 ns)

.NET 10

BenchmarkDotNet v0.15.8, Windows 11 (10.0.26200.9168/25H2/2025Update/HudsonValley2)
Intel Core Ultra 9 275HX 2.70GHz, 1 CPU, 24 logical and 24 physical cores
.NET SDK 11.0.100-preview.7.26381.103
[Host] : .NET 10.0.10 (10.0.10, 10.0.1026.32716), X64 RyuJIT x86-64-v3

Method Mean Error StdDev Allocated
Vector2_LengthSquared_Benchmark 491.0 ns 0.69 ns 0.58 ns -
Vector2_Dot_Benchmark 490.6 ns 0.42 ns 0.35 ns -
Manual2_Benchmark 378.0 ns 0.35 ns 0.31 ns -
Vector3_LengthSquared_Benchmark 596.4 ns 0.74 ns 0.69 ns -
Vector3_Dot_Benchmark 596.9 ns 0.41 ns 0.37 ns -
Manual3_Benchmark 584.4 ns 0.67 ns 0.63 ns -
Vector4_LengthSquared_Benchmark 390.5 ns 0.33 ns 0.31 ns -
Vector4_Dot_Benchmark 391.2 ns 0.28 ns 0.25 ns -
Manual4_Benchmark 672.9 ns 0.85 ns 0.80 ns -

// * Hints *
Outliers
UserQuery.Vector2_LengthSquared_Benchmark: Default -> 2 outliers were removed (493.76 ns, 494.16 ns)
UserQuery.Vector2_Dot_Benchmark: Default -> 2 outliers were removed (492.86 ns, 492.93 ns)
UserQuery.Manual2_Benchmark: Default -> 1 outlier was removed (379.72 ns)
UserQuery.Vector3_Dot_Benchmark: Default -> 1 outlier was removed, 2 outliers were detected (596.73 ns, 598.51 ns)
UserQuery.Vector4_Dot_Benchmark: Default -> 1 outlier was removed (392.86 ns)

Data

.NET 11 codegen

Vector2_LengthSquared_Benchmark ()
L0000	vxorps	xmm0, xmm0, xmm0
L0004	mov	rax, [rcx+0x10]
L0008	mov	ecx, [rax+8]
L000b	test	ecx, ecx
L000d	jle	short L0041
L000f	add	rax, 0x10
L0013	vmovsd	xmm1, [rax]
L0017	vinsertps	xmm1, xmm1, xmm1, 0x3c
L001d	vmulps	xmm1, xmm1, xmm1
L0021	vpermilps	xmm2, xmm1, 0xb1
L0027	vaddps	xmm1, xmm2, xmm1
L002b	vpermilps	xmm2, xmm1, 0x4e
L0031	vaddps	xmm1, xmm2, xmm1
L0035	vaddss	xmm0, xmm1, xmm0
L0039	add	rax, 8
L003d	dec	ecx
L003f	jne	short L0013
L0041	ret	
Vector2_Dot_Benchmark ()
L0000	vxorps	xmm0, xmm0, xmm0
L0004	mov	rax, [rcx+0x10]
L0008	mov	ecx, [rax+8]
L000b	test	ecx, ecx
L000d	jle	short L0041
L000f	add	rax, 0x10
L0013	vmovsd	xmm1, [rax]
L0017	vinsertps	xmm1, xmm1, xmm1, 0x3c
L001d	vmulps	xmm1, xmm1, xmm1
L0021	vpermilps	xmm2, xmm1, 0xb1
L0027	vaddps	xmm1, xmm2, xmm1
L002b	vpermilps	xmm2, xmm1, 0x4e
L0031	vaddps	xmm1, xmm2, xmm1
L0035	vaddss	xmm0, xmm1, xmm0
L0039	add	rax, 8
L003d	dec	ecx
L003f	jne	short L0013
L0041	ret	
Manual2_Benchmark ()
L0000	vxorps	xmm0, xmm0, xmm0
L0004	mov	rax, [rcx+0x10]
L0008	mov	ecx, [rax+8]
L000b	test	ecx, ecx
L000d	jle	short L0037
L000f	add	rax, 0x10
L0013	vmovsd	xmm1, [rax]
L0017	vmovaps	xmm2, xmm1
L001b	vmulss	xmm2, xmm2, xmm2
L001f	vmovshdup	xmm1, xmm1
L0023	vmulss	xmm1, xmm1, xmm1
L0027	vaddss	xmm1, xmm2, xmm1
L002b	vaddss	xmm0, xmm1, xmm0
L002f	add	rax, 8
L0033	dec	ecx
L0035	jne	short L0013
L0037	ret	
Vector3_LengthSquared_Benchmark ()
L0000	vxorps	xmm0, xmm0, xmm0
L0004	mov	rax, [rcx+0x18]
L0008	mov	ecx, [rax+8]
L000b	test	ecx, ecx
L000d	jle	short L0048
L000f	add	rax, 0x10
L0013	vmovsd	xmm1, [rax]
L0017	vinsertps	xmm1, xmm1, [rax+8], 0x28
L001e	vinsertps	xmm1, xmm1, xmm1, 0x38
L0024	vmulps	xmm1, xmm1, xmm1
L0028	vpermilps	xmm2, xmm1, 0xb1
L002e	vaddps	xmm1, xmm2, xmm1
L0032	vpermilps	xmm2, xmm1, 0x4e
L0038	vaddps	xmm1, xmm2, xmm1
L003c	vaddss	xmm0, xmm1, xmm0
L0040	add	rax, 0xc
L0044	dec	ecx
L0046	jne	short L0013
L0048	ret	
Vector3_Dot_Benchmark ()
L0000	vxorps	xmm0, xmm0, xmm0
L0004	mov	rax, [rcx+0x18]
L0008	mov	ecx, [rax+8]
L000b	test	ecx, ecx
L000d	jle	short L0048
L000f	add	rax, 0x10
L0013	vmovsd	xmm1, [rax]
L0017	vinsertps	xmm1, xmm1, [rax+8], 0x28
L001e	vinsertps	xmm1, xmm1, xmm1, 0x38
L0024	vmulps	xmm1, xmm1, xmm1
L0028	vpermilps	xmm2, xmm1, 0xb1
L002e	vaddps	xmm1, xmm2, xmm1
L0032	vpermilps	xmm2, xmm1, 0x4e
L0038	vaddps	xmm1, xmm2, xmm1
L003c	vaddss	xmm0, xmm1, xmm0
L0040	add	rax, 0xc
L0044	dec	ecx
L0046	jne	short L0013
L0048	ret	
Manual3_Benchmark ()
L0000	vxorps	xmm0, xmm0, xmm0
L0004	mov	rax, [rcx+0x18]
L0008	mov	ecx, [rax+8]
L000b	test	ecx, ecx
L000d	jle	short L004a
L000f	add	rax, 0x10
L0013	vmovsd	xmm1, [rax]
L0017	vinsertps	xmm1, xmm1, [rax+8], 0x28
L001e	vmovaps	xmm2, xmm1
L0022	vmulss	xmm2, xmm2, xmm2
L0026	vmovshdup	xmm3, xmm1
L002a	vmulss	xmm3, xmm3, xmm3
L002e	vaddss	xmm2, xmm2, xmm3
L0032	vunpckhps	xmm1, xmm1, xmm1
L0036	vmulss	xmm1, xmm1, xmm1
L003a	vaddss	xmm1, xmm2, xmm1
L003e	vaddss	xmm0, xmm1, xmm0
L0042	add	rax, 0xc
L0046	dec	ecx
L0048	jne	short L0013
L004a	ret	
Vector4_LengthSquared_Benchmark ()
L0000	vxorps	xmm0, xmm0, xmm0
L0004	mov	rax, [rcx+0x20]
L0008	mov	ecx, [rax+8]
L000b	test	ecx, ecx
L000d	jle	short L003b
L000f	add	rax, 0x10
L0013	vmovups	xmm1, [rax]
L0017	vmulps	xmm1, xmm1, xmm1
L001b	vpermilps	xmm2, xmm1, 0xb1
L0021	vaddps	xmm1, xmm2, xmm1
L0025	vpermilps	xmm2, xmm1, 0x4e
L002b	vaddps	xmm1, xmm2, xmm1
L002f	vaddss	xmm0, xmm1, xmm0
L0033	add	rax, 0x10
L0037	dec	ecx
L0039	jne	short L0013
L003b	ret	
Vector4_Dot_Benchmark ()
L0000	vxorps	xmm0, xmm0, xmm0
L0004	mov	rax, [rcx+0x20]
L0008	mov	ecx, [rax+8]
L000b	test	ecx, ecx
L000d	jle	short L003b
L000f	add	rax, 0x10
L0013	vmovups	xmm1, [rax]
L0017	vmulps	xmm1, xmm1, xmm1
L001b	vpermilps	xmm2, xmm1, 0xb1
L0021	vaddps	xmm1, xmm2, xmm1
L0025	vpermilps	xmm2, xmm1, 0x4e
L002b	vaddps	xmm1, xmm2, xmm1
L002f	vaddss	xmm0, xmm1, xmm0
L0033	add	rax, 0x10
L0037	dec	ecx
L0039	jne	short L0013
L003b	ret	
Manual4_Benchmark ()
L0000	vxorps	xmm0, xmm0, xmm0
L0004	mov	rax, [rcx+0x20]
L0008	mov	ecx, [rax+8]
L000b	test	ecx, ecx
L000d	jle	short L0050
L000f	add	rax, 0x10
L0013	vmovups	xmm1, [rax]
L0017	vmovaps	xmm2, xmm1
L001b	vmulss	xmm2, xmm2, xmm2
L001f	vmovshdup	xmm3, xmm1
L0023	vmulss	xmm3, xmm3, xmm3
L0027	vaddss	xmm2, xmm2, xmm3
L002b	vunpckhps	xmm3, xmm1, xmm1
L002f	vmulss	xmm3, xmm3, xmm3
L0033	vaddss	xmm2, xmm2, xmm3
L0037	vshufps	xmm1, xmm1, xmm1, 0xff
L003c	vmulss	xmm1, xmm1, xmm1
L0040	vaddss	xmm1, xmm2, xmm1
L0044	vaddss	xmm0, xmm1, xmm0
L0048	add	rax, 0x10
L004c	dec	ecx
L004e	jne	short L0013
L0050	ret

.NET 10 codegen

Vector2_LengthSquared_Benchmark ()
L0000	vxorps	xmm0, xmm0, xmm0
L0004	mov	rax, [rcx+0x10]
L0008	mov	ecx, [rax+8]
L000b	test	ecx, ecx
L000d	jle	short L003c
L000f	add	rax, 0x10
L0013	nop	[rax+rax]
L0018	nop	[rax+rax]
L0020	vmovsd	xmm1, [rax]
L0024	vinsertps	xmm1, xmm1, xmm1, 0x3c
L002a	vdpps	xmm1, xmm1, xmm1, 0xff
L0030	vaddss	xmm0, xmm1, xmm0
L0034	add	rax, 8
L0038	dec	ecx
L003a	jne	short L0020
L003c	ret	
Vector2_Dot_Benchmark ()
L0000	vxorps	xmm0, xmm0, xmm0
L0004	mov	rax, [rcx+0x10]
L0008	mov	ecx, [rax+8]
L000b	test	ecx, ecx
L000d	jle	short L003c
L000f	add	rax, 0x10
L0013	nop	[rax+rax]
L0018	nop	[rax+rax]
L0020	vmovsd	xmm1, [rax]
L0024	vinsertps	xmm1, xmm1, xmm1, 0x3c
L002a	vdpps	xmm1, xmm1, xmm1, 0xff
L0030	vaddss	xmm0, xmm1, xmm0
L0034	add	rax, 8
L0038	dec	ecx
L003a	jne	short L0020
L003c	ret	
Manual2_Benchmark ()
L0000	vxorps	xmm0, xmm0, xmm0
L0004	mov	rax, [rcx+0x10]
L0008	mov	ecx, [rax+8]
L000b	test	ecx, ecx
L000d	jle	short L0037
L000f	add	rax, 0x10
L0013	vmovsd	xmm1, [rax]
L0017	vmovaps	xmm2, xmm1
L001b	vmulss	xmm2, xmm2, xmm2
L001f	vmovshdup	xmm1, xmm1
L0023	vmulss	xmm1, xmm1, xmm1
L0027	vaddss	xmm1, xmm2, xmm1
L002b	vaddss	xmm0, xmm1, xmm0
L002f	add	rax, 8
L0033	dec	ecx
L0035	jne	short L0013
L0037	ret	
Vector3_LengthSquared_Benchmark ()
L0000	vxorps	xmm0, xmm0, xmm0
L0004	mov	rax, [rcx+0x18]
L0008	mov	ecx, [rax+8]
L000b	test	ecx, ecx
L000d	jle	short L0036
L000f	add	rax, 0x10
L0013	vmovsd	xmm1, [rax]
L0017	vinsertps	xmm1, xmm1, [rax+8], 0x28
L001e	vinsertps	xmm1, xmm1, xmm1, 0x38
L0024	vdpps	xmm1, xmm1, xmm1, 0xff
L002a	vaddss	xmm0, xmm1, xmm0
L002e	add	rax, 0xc
L0032	dec	ecx
L0034	jne	short L0013
L0036	ret	
Vector3_Dot_Benchmark ()
L0000	vxorps	xmm0, xmm0, xmm0
L0004	mov	rax, [rcx+0x18]
L0008	mov	ecx, [rax+8]
L000b	test	ecx, ecx
L000d	jle	short L0036
L000f	add	rax, 0x10
L0013	vmovsd	xmm1, [rax]
L0017	vinsertps	xmm1, xmm1, [rax+8], 0x28
L001e	vinsertps	xmm1, xmm1, xmm1, 0x38
L0024	vdpps	xmm1, xmm1, xmm1, 0xff
L002a	vaddss	xmm0, xmm1, xmm0
L002e	add	rax, 0xc
L0032	dec	ecx
L0034	jne	short L0013
L0036	ret	
Manual3_Benchmark ()
L0000	vxorps	xmm0, xmm0, xmm0
L0004	mov	rax, [rcx+0x18]
L0008	mov	ecx, [rax+8]
L000b	test	ecx, ecx
L000d	jle	short L004a
L000f	add	rax, 0x10
L0013	vmovsd	xmm1, [rax]
L0017	vinsertps	xmm1, xmm1, [rax+8], 0x28
L001e	vmovaps	xmm2, xmm1
L0022	vmulss	xmm2, xmm2, xmm2
L0026	vmovshdup	xmm3, xmm1
L002a	vmulss	xmm3, xmm3, xmm3
L002e	vaddss	xmm2, xmm2, xmm3
L0032	vunpckhps	xmm1, xmm1, xmm1
L0036	vmulss	xmm1, xmm1, xmm1
L003a	vaddss	xmm1, xmm2, xmm1
L003e	vaddss	xmm0, xmm1, xmm0
L0042	add	rax, 0xc
L0046	dec	ecx
L0048	jne	short L0013
L004a	ret	
Vector4_LengthSquared_Benchmark ()
L0000	vxorps	xmm0, xmm0, xmm0
L0004	mov	rax, [rcx+0x20]
L0008	mov	ecx, [rax+8]
L000b	test	ecx, ecx
L000d	jle	short L0036
L000f	add	rax, 0x10
L0013	nop	[rax+rax]
L0018	nop	[rax+rax]
L0020	vmovups	xmm1, [rax]
L0024	vdpps	xmm1, xmm1, xmm1, 0xff
L002a	vaddss	xmm0, xmm1, xmm0
L002e	add	rax, 0x10
L0032	dec	ecx
L0034	jne	short L0020
L0036	ret	
Vector4_Dot_Benchmark ()
L0000	vxorps	xmm0, xmm0, xmm0
L0004	mov	rax, [rcx+0x20]
L0008	mov	ecx, [rax+8]
L000b	test	ecx, ecx
L000d	jle	short L0036
L000f	add	rax, 0x10
L0013	nop	[rax+rax]
L0018	nop	[rax+rax]
L0020	vmovups	xmm1, [rax]
L0024	vdpps	xmm1, xmm1, xmm1, 0xff
L002a	vaddss	xmm0, xmm1, xmm0
L002e	add	rax, 0x10
L0032	dec	ecx
L0034	jne	short L0020
L0036	ret	
Manual4_Benchmark ()
L0000	vxorps	xmm0, xmm0, xmm0
L0004	mov	rax, [rcx+0x20]
L0008	mov	ecx, [rax+8]
L000b	test	ecx, ecx
L000d	jle	short L0050
L000f	add	rax, 0x10
L0013	vmovups	xmm1, [rax]
L0017	vmovaps	xmm2, xmm1
L001b	vmulss	xmm2, xmm2, xmm2
L001f	vmovshdup	xmm3, xmm1
L0023	vmulss	xmm3, xmm3, xmm3
L0027	vaddss	xmm2, xmm2, xmm3
L002b	vunpckhps	xmm3, xmm1, xmm1
L002f	vmulss	xmm3, xmm3, xmm3
L0033	vaddss	xmm2, xmm2, xmm3
L0037	vshufps	xmm1, xmm1, xmm1, 0xff
L003c	vmulss	xmm1, xmm1, xmm1
L0040	vaddss	xmm1, xmm2, xmm1
L0044	vaddss	xmm0, xmm1, xmm0
L0048	add	rax, 0x10
L004c	dec	ecx
L004e	jne	short L0013
L0050	ret

I have used a LINQPad script to measure with Benchmark.NET

#load "BenchmarkDotNet"

void Main()
{
	//BenchmarkSetup();
	//Vector2_LengthSquared_Benchmark().DumpTell();
	//Vector2_Dot_Benchmark().DumpTell();
	//Manual_Benchmark().DumpTell();

	RunBenchmark();
}

const int N = 1000;
Vector2[] vector2s;
Vector3[] vector3s;
Vector4[] vector4s;

[GlobalSetup]
public void BenchmarkSetup()
{
	vector2s = new System.Numerics.Vector2[N];
	vector3s = new System.Numerics.Vector3[N];
	vector4s = new System.Numerics.Vector4[N];
	Array.Fill(vector2s, Vector2.One);
	Array.Fill(vector3s, Vector3.One);
	Array.Fill(vector4s, Vector4.One);
}

[Benchmark]
public float Vector2_LengthSquared_Benchmark()
{
	float result = 0;
	foreach (var v in vector2s)
	{
		result += v.LengthSquared();
	}
	return result;
}

[Benchmark]
public float Vector2_Dot_Benchmark()
{
	float result = 0;
	foreach (var v in vector2s)
	{
		result += Vector2.Dot(v, v);
	}
	return result;
}

[Benchmark]
public float Manual2_Benchmark()
{
	float result = 0;
	foreach (var v in vector2s)
	{
		result += v.X * v.X + v.Y * v.Y;
	}
	return result;
}

[Benchmark]
public float Vector3_LengthSquared_Benchmark()
{
	float result = 0;
	foreach (var v in vector3s)
	{
		result += v.LengthSquared();
	}
	return result;
}

[Benchmark]
public float Vector3_Dot_Benchmark()
{
	float result = 0;
	foreach (var v in vector3s)
	{
		result += Vector3.Dot(v, v);
	}
	return result;
}

[Benchmark]
public float Manual3_Benchmark()
{
	float result = 0;
	foreach (var v in vector3s)
	{
		result += v.X * v.X + v.Y * v.Y + v.Z * v.Z;
	}
	return result;
}

[Benchmark]
public float Vector4_LengthSquared_Benchmark()
{
	float result = 0;
	foreach (var v in vector4s)
	{
		result += v.LengthSquared();
	}
	return result;
}

[Benchmark]
public float Vector4_Dot_Benchmark()
{
	float result = 0;
	foreach (var v in vector4s)
	{
		result += Vector4.Dot(v, v);
	}
	return result;
}

[Benchmark]
public float Manual4_Benchmark()
{
	float result = 0;
	foreach (var v in vector4s)
	{
		result += v.X * v.X + v.Y * v.Y + v.Z * v.Z + v.W * v.W;
	}
	return result;
}

Analysis

This is probably due to:

  • Faster DotProduct on AVX: Lowering for Vector128.Dot-style operations now emits a mul + permute + add sequence instead of vdpps/vdppd when AVX is available, which is consistently faster.

Activity

  1. added
    area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI
    on Sep 5, 2026
  2. dotnet-policy-service commented on Sep 5, 2026

    @dotnet-policy-service
    Contributor

    Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
    See info in area-owners.md if you want to be subscribed.

  3. tannergooding commented on Sep 5, 2026

    @tannergooding
    Member

    The decision to not use dpps is by design and is notably something that is going to be dependent on the CPU microarchitecture and code in question.

    The Intel and AMD optimization manuals alike call out the nuances of dpps/dppd and that while in some scenarios it can result in better perf due to shorter instruction sequences, there are many others where it slows down the pipeline and makes things worse instead, particularly due to the higher latency and lack of pipelining.

    Further, it is a "legacy" instruction and has not been brought forward or modernized as part of AVX512 or AVX10, the various horizontal (pairwise) instructions are in the same boat. This means they have concrete limitations and usability issues that can interfere with register allocation and hinder overall code throughput, particularly in more representative scenarios.

    On your CPU, which is an Arrow Lake based CPU, you will also see changes based on whether it gets scheduled on a P or E core, where you have:

    Instruction Latency P Reciprocal Throughput P Micro-ops P Latency E Reciprocal Throughput E Micro-ops E
    dpps 13 1.5 5 16-21 2.6 8
    haddps 4 1.5 3 7-8 1.7 5
    mulps 3 0.5 1 3 0.25 1
    addps 2 0.5 1 2 0.25 1
    permilps* 1 0.5 1 1 0.25 1

    * Noting that shuffle, unpack, and various other instructions all have the same timing and handling as permilps.


    Given that, there are multiple things at play here:

    1. .NET 10 is aligning the loop where-as .NET 11 is not, this is likely the primary reason for the perf difference
    2. Your benchmark is setup in a way that introduces a potential different bottleneck due to switching from vector to scalar prior to storage
    3. Your benchmark is covering a very simple loop and so not actually saturating the CPU as would be expected in perf sensitive code
    4. There is some unnecessary codegen in Vector2/3 that could be cleaned up -- I will work on fixing this
    5. BDN doesn't really handle P vs E cores and this can skew results, especially over many separate launches

    We also have a range of both Intel and AMD hardware and have measured consistent performance improvements on them from the associated change in real world scenarios, so I would say that these microbenches are not strictly representative. That being said, the perf being worse does not reproduce on my own AMD Zen 4, Intel Tiger Lake, or the Arrow Lake machine I have if restricted entirely to P cores or entirely to E cores.

  4. self-assigned this
    on Sep 5, 2026
  5. added this to the 11.0.0 milestone on Sep 11, 2026
  6. removed
    untriagedNew issue has not been triaged by the area owner
    on Sep 11, 2026
  7. tannergooding commented on Sep 21, 2026

    @tannergooding
    Member

    This was resolved in #133607

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMItenet-performancePerformance related issue

Type

No type

Projects

No projects

    Milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions