Repository navigation
perf(optimizers): wire amsgrad to the fused compiled training path (#653) - #1653
Conversation
AMSGradOptimizer did not implement IFusedOptimizerSpec, so every Train() step fell back to the eager autograd tape — a multi-x perf cliff — even though the Tensors fused kernel already implements the AMSGrad variant (the same OptimizerType.AMSGrad that AdamOptimizer.UseAMSGrad selects). Implement IFusedOptimizerSpec on AMSGradOptimizer (Adam-shaped params: beta1/beta2/epsilon, no decoupled weight decay), mirroring the Adam/AdaMax specs and falling back when an adaptive LR or unsupported scheduler is set. Adds an AMSGrad case to FusedOptimizerParityTests proving the fused path engages (fusedSteps>0) and produces parameter-identical results to the eager tape over 40 steps (within the Adam control tolerance). Full parity suite green: 9/9. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
|
The latest updates on your projects. Learn more about Vercel for GitHub.
|
|
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: Path: .coderabbit.yaml Review profile: ASSERTIVE Plan: Pro Run ID: 📒 Files selected for processing (2)
Walkthrough
ChangesAMSGrad Fused Optimizer Integration
Estimated code review effort🎯 2 (Simple) | ⏱️ ~10 minutes Poem
🚥 Pre-merge checks | ✅ 4 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (4 passed)
✏️ Tip: You can configure your own custom pre-merge checks in the settings. ✨ Finishing Touches📝 Generate docstrings
🧪 Generate unit tests (beta)
Comment |
Summary
#653 core-saturation, training-perf lever: widen fused-optimizer coverage so AMSGrad no longer falls off the fused compiled training path into the eager tape fallback.
AMSGradOptimizer<T>now implementsIFusedOptimizerSpecand maps to the TensorsOptimizerType.AMSGradfused kernel viaTryMapToFusedOptimizerConfig, soTrainWithTapekeeps AMSGrad on the fused fast path (which eliminates the bulk of the tape-walk backward cost) instead of dropping to the eager fallback cliff. The map is guarded: it declines (returns false) whenUseAdaptiveLearningRateis set or the LR schedule isn't fused-representable, so those paths still run eager — correctness preserved.Tests
FusedOptimizerParityTests.AMSGrad_FusedMatchesEager_NoWorseThanAdam— runs the fused (compiled) vs eager training step for 40 steps and asserts:fusedSteps > 0),trainDelta > 1e-6),Green locally.
🤖 Generated with Claude Code
Summary by CodeRabbit
New Features
Tests