Repository navigation
APX - Improve NDD on latest Intel Hardware - #132727
kendall1997 wants to merge 9 commits into
Conversation
…mance on latest Intel processors with APX
|
Azure Pipelines: Successfully started running 5 pipeline(s). 11 pipeline(s) were filtered out due to trigger conditions. There may be pipelines that require an authorized user to comment /azp run to run. |
|
Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch |
There was a problem hiding this comment.
Pull request overview
This PR adjusts x64 JIT instruction selection/emission to avoid generating the APX NDD (EVEX.ND) memory-source form, forcing a mov + binary-op sequence when the second source operand would come from memory.
Changes:
- Tightened NDD eligibility in
CodeGen::genCodeForBinaryto exclude cases whereop2is sourced from memory. - Updated
emitter::emitIns_BASE_R_R_RMto disable NDD when the RM operand is sourced from memory, falling back tomov+op.
Reviewed changes
Copilot reviewed 2 out of 2 changed files in this pull request and generated 2 comments.
| File | Description |
|---|---|
| src/coreclr/jit/codegenxarch.cpp | Prevents NDD selection when operand 2 is a memory source in genCodeForBinary. |
| src/coreclr/jit/emitxarch.cpp | Prevents NDD move-elision + NDD emission when the RM source operand is memory in emitIns_BASE_R_R_RM. |
Suppressed comments (1)
src/coreclr/jit/codegenxarch.cpp:1070
- The new comment is Intel-specific ("latest Intel processors") but the condition only checks DoJitUseApxNDD (JitConfig.EnableApxNDD + encodability) and operand placement, not CPU vendor. This can be misleading documentation if APX/NDD ever becomes relevant beyond Intel.
Consider rewording the comment to describe the policy (disable memory-source form) without hard-coding a vendor claim unless the code actually checks for it.
// On the latest Intel processors with APX, the memory form of NDD instructions has to be turned off by
// default for optimal hardware performance, so keep NDD only for register and immediate sources.
eligibleForNDD = emit->DoJitUseApxNDD(ins) && !op2->isUsedFromMemory();
| // On the latest Intel processors with APX, the memory form of NDD instructions has to be turned off by | ||
| // default for optimal hardware performance, so keep NDD only for register and immediate sources. | ||
| eligibleForNDD = emit->DoJitUseApxNDD(ins) && !op2->isUsedFromMemory(); |
Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
|
Azure Pipelines: Successfully started running 5 pipeline(s). 11 pipeline(s) were filtered out due to trigger conditions. There may be pipelines that require an authorized user to comment /azp run to run. |
Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
| // Disable the memory-source form of NDD (EVEX.ND) instructions for performance reasons; fall back to the | ||
| // mov+op sequence when the RM source is in memory. | ||
| bool useApxNdd = DoJitUseApxNDD(ins) && !rmOp->isUsedFromMemory(); | ||
|
|
||
| if (emitIns_Mov(INS_mov, attr, targetReg, regOp->GetRegNum(), true, useApxNdd) && useApxNdd) |
|
Hello @dotnet/intel, this PR is ready for your review! Thank you. |
| // For performance, avoid the memory-source form of EVEX.ND; keep NDD only for register and immediate sources. | ||
| eligibleForNDD = emit->DoJitUseApxNDD(ins) && !op2->isUsedFromMemory(); |
There was a problem hiding this comment.
Can this be elaborated on a bit?
Is this saying it's avoiding ins dst, op1, [op2]? Can you elaborate more on why this is necessary and why it is or isn't a problem for ins dst, [op2] or for SIMD instructions?
This is going to bloat codegen and rather seems like it should be impacting lowering/lsra to ensure preferencing is correct, as otherwise we shouldn't have marked the operand as contained at all.
There was a problem hiding this comment.
Hi @tannergooding, the tuning is based on performance results we observed on NVL. Similar tuning changes have also been upstreamed in both GCC and LLVM:
- GCC: https://gcc.gnu.org/git/?p=gcc.git;a=commit;h=a27731f6b52d94f4b1461709409057ab2a8a31ae
- LLVM: [X86][APX] Enable NDD tunings llvm/llvm-project#186049
Thanks for your comment, I'll update the PR soon with the latest version of the change for your review!
…-memory-operand-public
…-memory-operand-public
The memory-source form of APX NDD (`ins dst, src1, [mem]`) is not profitable, so it is never selected; the legacy `mov dst, src1` + `ins dst, [mem]` pair is used instead. Register and immediate NDD forms are unchanged. Move the decision out of emitIns_BASE_R_R_RM into its two callers, genCodeForBinary and genCodeForMul, through a new DoJitUseApxNDD(ins, rmOp) overload, and pass it in explicitly. The emitter asserts the caller's choice, and that the fallback `mov` cannot overwrite an address register of the memory operand. LSRA guarantees the latter by modelling these nodes as read-modify-write. No codegen change when APX NDD is disabled (the default).
For read-modify-write nodes, BuildRMWUses marks op2 delay-free because the legacy sequence writes the destination before reading op2. The APX NDD form reads both sources first, so the constraint is not needed when codegen emits NDD. For an integer sub with a register op2, skip the delay-free when NDD is enabled for sub and op1 is a register-candidate local that stays live after the node. The destination can't take op1's register there, so the JIT already emitted the three-register NDD form; now the destination may take op2's register. Codegen emits NDD `sub dst, op1, op2` with dst == op2, and the emitter allows it. The op1 preference is unchanged, and a contained op2 keeps its delay-free. No codegen change when APX NDD is disabled (the default).
…emory-operand-public
Move the LSRA check that drops the op2 delay-free for APX NDD sub out of the else-if chain into a separate check under TARGET_AMD64, keyed on delayUseOperand. The condition is unchanged. In genCodeForBinary, let a non-commutative op whose destination got op2's register fall through to the three-register path, which already emits NDD. On the non-NDD path, a noway_assert keeps release builds from emitting a mov that would overwrite op2 if LSRA and codegen ever disagree. No codegen change.
This pull request disables memory forms of NDD instructions for optimal performance on Intel CPUs with APX support.
Instruction selection and code generation improvements:
src/coreclr/jit/codegenxarch.cpp: Updated the eligibility check for NDD instructions ingenCodeForBinaryto exclude cases where the second operand is sourced from memory, ensuring NDD is used only for register and immediate sources.src/coreclr/jit/emitxarch.cpp: ModifiedemitIns_BASE_R_R_RMto disable NDD instructions when the RM source operand is in memory, defaulting to amov+opsequence instead.