Skip to content

[JIT] Rationalize: Re-run pre-order processing on identity shuffle rewrite in PreOrderVisit - #134910

Open
adamperlin wants to merge 8 commits into
dotnet:mainfrom
adamperlin:adamperlin/fix-rewrite-hw-intrinsic-user-call
Open

adamperlin wants to merge 8 commits into
dotnet:mainfrom
adamperlin:adamperlin/fix-rewrite-hw-intrinsic-user-call

Conversation

@adamperlin

@adamperlin adamperlin commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

Resolves #133787; This case comes up with IR that looks like the following:

STOREIND simd16
├── LCL_VAR V00 RetBuf
└── HWINTRINSIC Shuffle                         [000022]
    ├── HWINTRINSIC ShiftRightLogical128BitLane [000020]
    │   ├── IND simd16
    │   │   └── LCL_VAR V01 arg0
    │   └── CAST int <- ubyte
    │       └── LCL_VAR V02 arg1
    └── CNS_VEC <0, 1, 2, ..., 15>              [000016]

The outer shuffle is the identity, so it is folded away as part of RewriteHwIntrinsicAsUserCall during PreOrderVisit. The inner ShiftRightLogical128BitLane -- which needs to be re-written back to a GT_CALL since it has a non-constant arg -- is the replacement node, but it is never processed by the preorder visitor. If we get back a different node from Rewrite*Intrinsic, we should re-do the processing logic on that new node.

adamperlin and others added 2 commits September 29, 2026 17:35
The replacement loop exits only after node and *use are equal, so the ARM64-specific reload is unnecessary. Also remove trailing whitespace from the adjacent comment.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
@github-actions github-actions Bot added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Sep 30, 2026
@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines:
Successfully started running 5 pipeline(s).
11 pipeline(s) were filtered out due to trigger conditions.
There may be pipelines that require an authorized user to comment /azp run to run.

@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

@adamperlin

adamperlin commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor Author

@dotnet/jit-contrib PTAL
cc @EgorBo

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

The reported Checked-build regression needs a targeted test to prevent recurrence.

Review effort: Balanced
Findings: 1 Medium severity

Open (1)
What changed in this PR

Fixes JIT rationalization so replacement intrinsic nodes receive pre-order processing.

Changes:

  • Repeats intrinsic rewriting until the node remains unchanged.
  • Ensures nested deferred hardware intrinsics are converted to user calls.
File Description
src/​coreclr/​jit/​rationalize.cpp Reprocesses rewritten intrinsic nodes during rationalization.

Comment thread src/coreclr/jit/rationalize.cpp Outdated
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Comment thread src/coreclr/jit/rationalize.cpp Outdated
// The below is a loop, because rewriting an intrinsic or HW intrinsic as a user call
// may replace the node with another node which also needs the same pre-order processing.
// Continue until no replacement is made.
do

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Is this a potential algorithmic complexity problem? Can we reduce the amount of looping/walking we need to do somehow as presumably this is uncommon and doesn't need everything to rewalk?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I agree, can the transformation which does this instead ensure that it does the necessary processing before returning?

@adamperlin adamperlin Oct 2, 2026 •

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I believe there is a potential algorithmic complexity problem for a parent-child chain of identity-shuffle's that are eliminated:

SHUFFLE_1 identity [deferred user call] 
 |____ SHUFFLE_2 identity 
 ...
       |___SHUFFLE_N identity 
            |__ T

Each iteration of RewriteHWIntrinsicAsUserCall strips one level away and re-runs fgSetTreeSeq over the child (unchanged) sub-tree chain at each iteration, leading to O(n^2) behavior where n is the chain length. Moving the re-processing down into RewriteHWIntrinsicAsUserCall doesn't necessarily fix the complexity issue in this worst case, but we could maybe optimize specifically the identity shuffle case to avoid re-processing unchanged sub-trees?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@tannergooding I believe the issue should be addressed now; I added an optimized path for when we have an identity shuffle that returns one of its operands unchanged.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can the path that creates the new operand or moves it into a position where it won't be visited naturally just call the walk recursively on it?

Modifying the root visit seems like a big hammer and it also has non trivial TP costs.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Got it, that makes sense. I updated this PR to pass the visitor through to RewriteHWIntrinsicAsUserCall so that we can re-invoke the PreOrderVisit on the path that moves the operand!

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@tannergooding can you take another look here?

@adamperlin adamperlin changed the title [JIT] Rationalize: Re-run pre-order processing if a node is re-written in PreOrderVisit [JIT] Rationalize: Re-run pre-order processing on identity shuffle rewrite in PreOrderVisit Oct 5, 2026

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

JIT: (bug) Assertion failed '!node->IsUserCall()' during 'Rationalize IR' with an identity Vector128.Shuffle

5 participants