Fix Issue for Lion Optimizer - #381
Conversation
This commit implements the Lion (EvoLved Sign Momentum) optimizer as requested in issue #315. Key Features: - LionOptimizer.cs: Complete implementation of the Lion algorithm with sign-based gradient updates - LionOptimizerOptions.cs: Configuration options with sensible defaults (LR=1e-4, β1=0.9, β2=0.99) - Comprehensive unit tests covering all major functionality Implementation Details: - Inherits from GradientBasedOptimizerBase following project architecture - Uses single momentum state (50% memory reduction vs Adam) - Implements sign-based updates for improved generalization - Supports decoupled weight decay (similar to AdamW) - Provides UpdateParameters overloads for Vector, Matrix types - Full serialization/deserialization support - Generic type support via INumericOperations<T> Tests Include: - Parameter update correctness with various gradient configurations - Sign-based behavior verification (magnitude independence) - Weight decay functionality - Momentum accumulation across iterations - Both float and double type support - Serialization/deserialization preservation - Different beta1/beta2 configurations Resolves #315
There was a problem hiding this comment.
Pull Request Overview
This PR implements the Lion (Evolved Sign Momentum) optimization algorithm, a modern gradient-based optimizer that offers memory efficiency and strong performance characteristics compared to Adam.
Key Changes
- Adds
LionOptimizer<T, TInput, TOutput>class implementing the Lion algorithm with sign-based gradient updates - Adds
LionOptimizerOptions<T, TInput, TOutput>class for configuring Lion-specific parameters (learning rate, beta1, beta2, weight decay, and adaptive parameter settings) - Adds comprehensive unit tests covering constructor behavior, gradient updates (vectors and matrices), weight decay, momentum building, serialization, and multiple data types
Reviewed Changes
Copilot reviewed 3 out of 3 changed files in this pull request and generated 13 comments.
| File | Description |
|---|---|
src/Optimizers/LionOptimizer.cs |
Implements Lion optimizer with UpdateParameters methods for vectors/matrices, serialization support, and adaptive parameter handling |
src/Models/Options/LionOptimizerOptions.cs |
Defines configuration options with default values matching Lion algorithm specifications (learning rate 1e-4, beta1=0.9, beta2=0.99) |
tests/UnitTests/Optimizers/LionOptimizerTests.cs |
Provides 20 test cases covering optimizer initialization, parameter updates with various gradient scenarios, momentum behavior, serialization, and type support |
💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.
…ization - UpdateOptions: Add InitializeAdaptiveParameters() call after setting new options to ensure cached parameter values reflect the updated options - Deserialize: Add InitializeAdaptiveParameters() call after deserializing options to initialize cached values from deserialized data instead of keeping zero values Both fixes ensure _currentLearningRate, _currentBeta1, and _currentBeta2 are properly synchronized with the LionOptimizerOptions values. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude <noreply@anthropic.com>
|
Note Other AI code review bot(s) detectedCodeRabbit has detected other AI code review bot(s) in this pull request and will avoid duplicating their findings in the review comments. This may lead to a less comprehensive review. Summary by CodeRabbit
WalkthroughAdds a new Lion optimizer implementation, a corresponding options class with Lion-specific hyperparameters, and unit tests. Includes adaptive-parameter handling, momentum state, parameter update routines (vector/matrix), serialization/deserialization, reset, and gradient-cache key changes. Changes
Sequence Diagram(s)sequenceDiagram
participant Caller
participant Lion as LionOptimizer
participant Grad as GradientCalc
participant Eval as Evaluator
Caller->>Lion: Optimize(inputData)
activate Lion
Lion->>Lion: Initialize _m, _t, adaptive params
loop Iteration
Lion->>Grad: Compute gradients
Grad-->>Lion: gradients
rect rgb(230, 245, 255)
Note over Lion: Lion update (sign-based)
Lion->>Lion: Interpolate momentum: m' = β1·m + (1-β1)·grad
Lion->>Lion: Apply sign update: θ ← θ - lr·sign(m')
Lion->>Lion: Update momentum state m ← m'
end
Lion->>Lion: Optionally update β1, β2 (clamped)
Lion->>Eval: Evaluate loss / history
Eval-->>Lion: loss
alt Converged / Early stop
Lion->>Lion: Break loop
end
end
Lion-->>Caller: OptimizationResult
deactivate Lion
Estimated code review effort🎯 4 (Complex) | ⏱️ ~45 minutes
Poem
Pre-merge checks and finishing touches❌ Failed checks (1 warning, 1 inconclusive)
✅ Passed checks (3 passed)
✨ Finishing touches
🧪 Generate unit tests (beta)
Comment |
- Fix spelling: EvoLved -> Evolved in documentation - Use cached adaptive parameters (_currentBeta1, _currentBeta2, _currentLearningRate) in UpdateParameters methods instead of accessing options directly - Add proper null checks with guard clauses in test files instead of lazy null-forgiving operator All fixes ensure consistency with adaptive parameter adjustments and proper null safety without using the ! operator. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude <noreply@anthropic.com>
There was a problem hiding this comment.
Actionable comments posted: 5
🧹 Nitpick comments (2)
tests/UnitTests/Optimizers/LionOptimizerTests.cs (2)
278-318: Consider verifying momentum state preservation.The test validates that options are correctly serialized/deserialized but doesn't verify that the momentum state (
_m) and timestep (_t) are preserved. This is a critical part of the optimizer's state.Consider adding a third
UpdateParameterscall on both optimizers after deserialization and verifying they produce identical results, which would indirectly confirm momentum state preservation:// After deserialization and option verification, add: var gradient2 = new Vector<double>(new double[] { 0.3, -0.3, 0.7 }); // Both should produce same result if state was preserved var result1 = optimizer1.UpdateParameters(parameters, gradient2); var result2 = optimizer2.UpdateParameters(parameters, gradient2); for (int i = 0; i < parameters.Length; i++) { Assert.Equal(result1[i], result2[i], 1e-9); }
60-108: Consider adding precise numerical verification.While directional assertions (
>,<) are appropriate for some integration-style tests, at least one test should verify actual expected values calculated by hand or from a reference implementation to ensure algorithmic correctness.For example, for a simple scenario with known values:
- Parameters: [1.0]
- Gradient: [1.0]
- Beta1: 0.0, Beta2: 0.0 (no momentum)
- LearningRate: 0.1
- WeightDecay: 0.0
The expected result should be exactly:
1.0 - 0.1 * sign(1.0) = 1.0 - 0.1 = 0.9Adding such a test would provide stronger validation of the implementation.
📜 Review details
Configuration used: CodeRabbit UI
Review profile: CHILL
Plan: Pro
📒 Files selected for processing (3)
src/Models/Options/LionOptimizerOptions.cs(1 hunks)src/Optimizers/LionOptimizer.cs(1 hunks)tests/UnitTests/Optimizers/LionOptimizerTests.cs(1 hunks)
🚧 Files skipped from review as they are similar to previous changes (1)
- src/Models/Options/LionOptimizerOptions.cs
🧰 Additional context used
🧬 Code graph analysis (2)
src/Optimizers/LionOptimizer.cs (3)
src/Models/Options/LionOptimizerOptions.cs (1)
LionOptimizerOptions(18-134)src/Optimizers/OptimizerBase.cs (2)
UpdateBestSolution(641-669)UpdateIterationHistoryAndCheckEarlyStopping(818-834)src/Helpers/MathHelper.cs (1)
MathHelper(16-987)
tests/UnitTests/Optimizers/LionOptimizerTests.cs (2)
src/Optimizers/LionOptimizer.cs (7)
LionOptimizer(27-497)LionOptimizer(70-83)Vector(246-291)Matrix(304-357)Reset(367-371)Serialize(418-442)Deserialize(453-479)src/Models/Options/LionOptimizerOptions.cs (1)
LionOptimizerOptions(18-134)
⏰ Context from checks skipped due to timeout of 90000ms. You can increase the timeout in your CodeRabbit configuration to a maximum of 15 minutes (900000ms). (1)
- GitHub Check: Build All Frameworks
🔇 Additional comments (5)
tests/UnitTests/Optimizers/LionOptimizerTests.cs (1)
11-58: LGTM: Constructor tests verify correct initialization.The null-checking pattern (Assert.NotNull followed by explicit null checks with throws) effectively addresses the potential null reference warnings from past reviews while maintaining test clarity.
src/Optimizers/LionOptimizer.cs (4)
246-291: LGTM: UpdateParameters correctly uses cached adaptive parameters.The method now correctly uses
_currentBeta1,_currentBeta2, and_currentLearningRatethroughout, which ensures that adaptive parameter adjustments (if enabled) are properly respected. This addresses the past review concerns.
304-357: LGTM: Matrix UpdateParameters correctly implemented.Consistent with the Vector version, this method properly uses cached adaptive parameters and correctly implements the Lion algorithm for matrix parameters.
383-394: LGTM: State management methods correctly update adaptive parameters.Both
UpdateOptions(line 388) andDeserialize(line 468) now properly callInitializeAdaptiveParameters()to ensure the cached parameter values are synchronized with the options. This addresses the past review concerns.Also applies to: 453-479
195-233: Algorithm implementation appears correct.The Lion optimizer correctly implements the three-step algorithm:
- Interpolate between momentum and gradient (lines 204-207)
- Apply sign-based update with optional weight decay (lines 211-222)
- Update momentum for next iteration (lines 226-229)
The sign-based approach and decoupled weight decay align with the Lion algorithm specification.
- Add configurable Beta1/Beta2 increase/decrease factors to LionOptimizerOptions - Implement proper adaptive beta updates in UpdateAdaptiveParameters method - Add SetParameters call before return in UpdateSolution to persist changes - Fix test tolerance from 10 to 1e-9 for proper numerical precision - Add verification that different Beta1/Beta2 values produce different results Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude <noreply@anthropic.com>
There was a problem hiding this comment.
Actionable comments posted: 2
🧹 Nitpick comments (1)
src/Optimizers/LionOptimizer.cs (1)
70-83: Remove redundant field initialization.Lines 78-80 initialize
_currentLearningRate,_currentBeta1, and_currentBeta2toNumOps.Zero, but line 82 immediately callsInitializeAdaptiveParameters()which overwrites these values with the actual option values.Apply this diff to remove the redundant initialization:
{ _m = Vector<T>.Empty(); _t = 0; _options = options ?? new(); - _currentLearningRate = NumOps.Zero; - _currentBeta1 = NumOps.Zero; - _currentBeta2 = NumOps.Zero; InitializeAdaptiveParameters(); }
📜 Review details
Configuration used: CodeRabbit UI
Review profile: CHILL
Plan: Pro
📒 Files selected for processing (3)
src/Models/Options/LionOptimizerOptions.cs(1 hunks)src/Optimizers/LionOptimizer.cs(1 hunks)tests/UnitTests/Optimizers/LionOptimizerTests.cs(1 hunks)
🚧 Files skipped from review as they are similar to previous changes (1)
- tests/UnitTests/Optimizers/LionOptimizerTests.cs
🧰 Additional context used
🧬 Code graph analysis (1)
src/Optimizers/LionOptimizer.cs (2)
src/Models/Options/LionOptimizerOptions.cs (1)
LionOptimizerOptions(18-178)src/Helpers/MathHelper.cs (1)
MathHelper(16-987)
⏰ Context from checks skipped due to timeout of 90000ms. You can increase the timeout in your CodeRabbit configuration to a maximum of 15 minutes (900000ms). (1)
- GitHub Check: Build All Frameworks
🔇 Additional comments (12)
src/Models/Options/LionOptimizerOptions.cs (1)
1-18: LGTM! Spelling issue resolved.The class documentation correctly uses "Evolved Sign Momentum" throughout. This addresses the previous review feedback about the "EvoLved" spelling error.
src/Optimizers/LionOptimizer.cs (11)
1-27: LGTM! Class documentation is accurate.The optimizer correctly implements the Lion (Evolved Sign Momentum) algorithm with comprehensive documentation. The spelling issue from previous reviews has been resolved.
92-97: LGTM! Adaptive parameter initialization is correct.Properly converts hyperparameters from
doubleto generic typeTusingNumOps.FromDouble, following the architectural requirement to useINumericOperations<T>for all numeric operations.
109-146: LGTM! Main optimization loop is well-structured.The optimization process correctly implements the Lion algorithm flow: gradient calculation, parameter updates using the Lion three-step rule, fitness evaluation, adaptive parameter adjustments, and convergence checks. The implementation properly initializes momentum state and tracks the best solution.
212-251: LGTM! Lion update rule is correctly implemented.The three-step Lion algorithm is properly implemented:
- Interpolation:
c_t = beta1 * m_{t-1} + (1 - beta1) * g_t(lines 221-224)- Sign-based update with decoupled weight decay:
theta_t = theta_{t-1} - lr * (sign(c_t) + lambda * theta_{t-1})(lines 228-239)- Momentum update:
m_t = beta2 * m_{t-1} + (1 - beta2) * g_t(lines 243-246)The method now correctly calls
SetParameters()(line 249) before returning, which addresses the previous review concern about persisting parameter updates.
322-375: LGTM! Matrix parameter update correctly implements Lion algorithm.The matrix overload properly flattens the 2D parameter space into a 1D momentum vector and applies the Lion update rule element-wise. The implementation correctly uses the adaptive parameter fields (
_currentBeta1,_currentBeta2,_currentLearningRate) rather than directly accessing options, which addresses previous review feedback.Note: The momentum reset behavior on size change (lines 324-328) mirrors the vector version and should be verified as discussed in the previous comment.
385-389: LGTM! Reset functionality is correct.Properly clears the momentum state and resets the time step counter, allowing the optimizer to be reused for new training runs or re-initialized mid-training.
401-412: LGTM! Options update correctly reinitializes parameters.The method now properly calls
InitializeAdaptiveParameters()(line 406) after updating options, which addresses the previous review concern about cached parameter values not being synchronized with new options.
422-425: LGTM! Options accessor is straightforward.
436-497: LGTM! Serialization and deserialization are complete.Both methods correctly preserve and restore the optimizer's full state:
- Base class state via
base.Serialize()/base.Deserialize()- Lion-specific options via JSON serialization
- Time step counter
_t- Momentum vector
_mThe deserialization method now properly calls
InitializeAdaptiveParameters()(line 486) to synchronize cached adaptive parameter values with the deserialized options, which addresses previous review feedback.
510-514: LGTM! Gradient cache key generation is appropriate.Correctly extends the base cache key with Lion-specific parameters (
LearningRateandMaxIterations) to ensure gradient caching distinguishes between different optimizer configurations.
264-309: ****The momentum reset on parameter size change is an intentional, consistent pattern across all optimizers in the codebase—not a bug or oversight. The same logic appears in AdamOptimizer (lines 220-224), which checks if size changes and reinitializes momentum vectors, matching Lion's approach exactly. This defensive pattern supports flexible architectures and edge cases where parameter dimensions may shift, and there is no evidence in the codebase of problems with this behavior. Throwing an exception instead would break compatibility with legitimate use cases.
Likely an incorrect or invalid review comment.
| public double LearningRate { get; set; } = 1e-4; | ||
|
|
||
| /// <summary> | ||
| /// Gets or sets the exponential decay rate for the momentum interpolation (used for computing the update). | ||
| /// </summary> | ||
| /// <value>The beta1 value, defaulting to 0.9.</value> | ||
| /// <remarks> | ||
| /// <para><b>For Beginners:</b> Beta1 controls how much Lion blends the current gradient with past momentum | ||
| /// when deciding which direction to move. A value of 0.9 means it gives 90% weight to the past momentum | ||
| /// and 10% to the new gradient. This is like having inertia - you don't change direction immediately when | ||
| /// you get new information. Higher values (closer to 1) create smoother updates but slower adaptation, | ||
| /// while lower values respond more quickly to new gradients.</para> | ||
| /// </remarks> | ||
| public double Beta1 { get; set; } = 0.9; | ||
|
|
||
| /// <summary> | ||
| /// Gets or sets the exponential decay rate for updating the momentum state. | ||
| /// </summary> | ||
| /// <value>The beta2 value, defaulting to 0.99.</value> | ||
| /// <remarks> | ||
| /// <para><b>For Beginners:</b> Beta2 controls how much Lion remembers from its momentum history when | ||
| /// updating the momentum state for the next iteration. A value of 0.99 means it retains 99% of the old | ||
| /// momentum and incorporates 1% from the new gradient. This creates a long memory of past gradients, | ||
| /// helping smooth out noisy updates. Think of it like a heavy flywheel that doesn't change speed quickly - | ||
| /// it provides stability during training.</para> | ||
| /// </remarks> | ||
| public double Beta2 { get; set; } = 0.99; | ||
|
|
||
| /// <summary> | ||
| /// Gets or sets the weight decay (L2 regularization) coefficient. | ||
| /// </summary> | ||
| /// <value>The weight decay value, defaulting to 0.0.</value> | ||
| /// <remarks> | ||
| /// <para><b>For Beginners:</b> Weight decay helps prevent overfitting by penalizing large parameter values. | ||
| /// A value of 0.0 means no weight decay. When set to a small positive value (e.g., 0.01 or 0.1), it encourages | ||
| /// the model to keep weights small, which often improves generalization to new data. Think of it like a tax | ||
| /// on complexity - it encourages the model to be as simple as possible while still solving the problem. | ||
| /// Lion applies weight decay in a decoupled manner, similar to AdamW.</para> | ||
| /// </remarks> | ||
| public double WeightDecay { get; set; } = 0.0; | ||
|
|
||
| /// <summary> | ||
| /// Gets or sets whether to automatically adjust Beta1 during training. | ||
| /// </summary> | ||
| /// <value>False by default, as Lion typically uses fixed betas.</value> | ||
| /// <remarks> | ||
| /// <para><b>For Beginners:</b> When enabled, Beta1 will be automatically adjusted based on training progress. | ||
| /// However, Lion was designed to work well with fixed beta values, so this is disabled by default. | ||
| /// Unlike Adam, Lion is less sensitive to beta parameter choices due to its sign-based updates. | ||
| /// You typically don't need to enable this unless you're doing advanced experimentation.</para> | ||
| /// </remarks> | ||
| public bool UseAdaptiveBeta1 { get; set; } = false; | ||
|
|
||
| /// <summary> | ||
| /// Gets or sets whether to automatically adjust Beta2 during training. | ||
| /// </summary> | ||
| /// <value>False by default, as Lion typically uses fixed betas.</value> | ||
| /// <remarks> | ||
| /// <para><b>For Beginners:</b> When enabled, Beta2 will be automatically adjusted based on training progress. | ||
| /// However, Lion was designed to work well with fixed beta values, so this is disabled by default. | ||
| /// The sign-based nature of Lion makes it robust to beta parameter variations.</para> | ||
| /// </remarks> | ||
| public bool UseAdaptiveBeta2 { get; set; } = false; | ||
|
|
||
| /// <summary> | ||
| /// Gets or sets the minimum allowed value for Beta1. | ||
| /// </summary> | ||
| /// <value>The minimum Beta1 value, defaulting to 0.85.</value> | ||
| /// <remarks> | ||
| /// <para><b>For Beginners:</b> If adaptive Beta1 is enabled, this prevents it from dropping too low. | ||
| /// A minimum of 0.85 ensures some momentum is always maintained.</para> | ||
| /// </remarks> | ||
| public double MinBeta1 { get; set; } = 0.85; | ||
|
|
||
| /// <summary> | ||
| /// Gets or sets the maximum allowed value for Beta1. | ||
| /// </summary> | ||
| /// <value>The maximum Beta1 value, defaulting to 0.95.</value> | ||
| /// <remarks> | ||
| /// <para><b>For Beginners:</b> If adaptive Beta1 is enabled, this prevents it from becoming too high. | ||
| /// A maximum of 0.95 ensures some responsiveness to new gradients.</para> | ||
| /// </remarks> | ||
| public double MaxBeta1 { get; set; } = 0.95; | ||
|
|
||
| /// <summary> | ||
| /// Gets or sets the minimum allowed value for Beta2. | ||
| /// </summary> | ||
| /// <value>The minimum Beta2 value, defaulting to 0.95.</value> | ||
| /// <remarks> | ||
| /// <para><b>For Beginners:</b> If adaptive Beta2 is enabled, this prevents it from dropping too low. | ||
| /// A minimum of 0.95 ensures momentum state retains sufficient history.</para> | ||
| /// </remarks> | ||
| public double MinBeta2 { get; set; } = 0.95; | ||
|
|
||
| /// <summary> | ||
| /// Gets or sets the maximum allowed value for Beta2. | ||
| /// </summary> | ||
| /// <value>The maximum Beta2 value, defaulting to 0.999.</value> | ||
| /// <remarks> | ||
| /// <para><b>For Beginners:</b> If adaptive Beta2 is enabled, this prevents it from becoming too high. | ||
| /// A maximum of 0.999 ensures the momentum state can still adapt to changes.</para> | ||
| /// </remarks> | ||
| public double MaxBeta2 { get; set; } = 0.999; | ||
|
|
||
| /// <summary> | ||
| /// Gets or sets the factor by which Beta1 is increased when fitness improves. | ||
| /// </summary> | ||
| /// <value>The Beta1 increase factor, defaulting to 1.02.</value> | ||
| /// <remarks> | ||
| /// <para><b>For Beginners:</b> When adaptive Beta1 is enabled and the optimizer is improving, | ||
| /// Beta1 is multiplied by this factor. A value of 1.02 means Beta1 increases by 2% each time | ||
| /// fitness improves. Higher Beta1 values create smoother, more stable updates.</para> | ||
| /// </remarks> | ||
| public double Beta1IncreaseFactor { get; set; } = 1.02; | ||
|
|
||
| /// <summary> | ||
| /// Gets or sets the factor by which Beta1 is decreased when fitness does not improve. | ||
| /// </summary> | ||
| /// <value>The Beta1 decrease factor, defaulting to 0.98.</value> | ||
| /// <remarks> | ||
| /// <para><b>For Beginners:</b> When adaptive Beta1 is enabled and the optimizer is not improving, | ||
| /// Beta1 is multiplied by this factor. A value of 0.98 means Beta1 decreases by 2% each time | ||
| /// fitness doesn't improve. Lower Beta1 values make the optimizer more responsive to new gradients.</para> | ||
| /// </remarks> | ||
| public double Beta1DecreaseFactor { get; set; } = 0.98; | ||
|
|
||
| /// <summary> | ||
| /// Gets or sets the factor by which Beta2 is increased when fitness improves. | ||
| /// </summary> | ||
| /// <value>The Beta2 increase factor, defaulting to 1.02.</value> | ||
| /// <remarks> | ||
| /// <para><b>For Beginners:</b> When adaptive Beta2 is enabled and the optimizer is improving, | ||
| /// Beta2 is multiplied by this factor. A value of 1.02 means Beta2 increases by 2% each time | ||
| /// fitness improves. Higher Beta2 values create longer memory of past gradients for more stability.</para> | ||
| /// </remarks> | ||
| public double Beta2IncreaseFactor { get; set; } = 1.02; | ||
|
|
||
| /// <summary> | ||
| /// Gets or sets the factor by which Beta2 is decreased when fitness does not improve. | ||
| /// </summary> | ||
| /// <value>The Beta2 decrease factor, defaulting to 0.98.</value> | ||
| /// <remarks> | ||
| /// <para><b>For Beginners:</b> When adaptive Beta2 is enabled and the optimizer is not improving, | ||
| /// Beta2 is multiplied by this factor. A value of 0.98 means Beta2 decreases by 2% each time | ||
| /// fitness doesn't improve. Lower Beta2 values make the momentum state more responsive to recent changes.</para> | ||
| /// </remarks> | ||
| public double Beta2DecreaseFactor { get; set; } = 0.98; |
There was a problem hiding this comment.
🛠️ Refactor suggestion | 🟠 Major
Add parameter validation to ensure valid configurations.
The options class lacks validation for hyperparameter constraints. Invalid values could lead to runtime errors, non-convergence, or incorrect optimization behavior.
Consider adding validation in a constructor or via property setters:
public double LearningRate
{
get => _learningRate;
set
{
if (value <= 0)
throw new ArgumentException("LearningRate must be positive.", nameof(LearningRate));
_learningRate = value;
}
}Key constraints to validate:
LearningRate > 0Beta1 ∈ (0, 1)andBeta2 ∈ (0, 1)WeightDecay >= 0MinBeta1 < MaxBeta1andMinBeta2 < MaxBeta2Beta1/Beta2 ∈ [MinBeta1/MinBeta2, MaxBeta1/MaxBeta2]Beta1IncreaseFactor > 1andBeta1DecreaseFactor < 1(and similarly for Beta2)
Alternatively, perform validation in the LionOptimizer constructor if the framework convention is to validate at consumption rather than assignment.
| protected override void UpdateAdaptiveParameters(OptimizationStepData<T, TInput, TOutput> currentStepData, OptimizationStepData<T, TInput, TOutput> previousStepData) | ||
| { | ||
| base.UpdateAdaptiveParameters(currentStepData, previousStepData); | ||
|
|
||
| // Adaptive Beta1 updates (if enabled) | ||
| if (_options.UseAdaptiveBeta1) | ||
| { | ||
| // Increase Beta1 (more smoothing) if fitness is improving, decrease (faster adaptation) otherwise | ||
| if (NumOps.GreaterThan(currentStepData.FitnessScore, previousStepData.FitnessScore)) | ||
| { | ||
| _currentBeta1 = NumOps.Multiply(_currentBeta1, NumOps.FromDouble(_options.Beta1IncreaseFactor)); | ||
| } | ||
| else | ||
| { | ||
| _currentBeta1 = NumOps.Multiply(_currentBeta1, NumOps.FromDouble(_options.Beta1DecreaseFactor)); | ||
| } | ||
|
|
||
| // Clamp to configured bounds | ||
| _currentBeta1 = MathHelper.Max(NumOps.FromDouble(_options.MinBeta1), | ||
| MathHelper.Min(NumOps.FromDouble(_options.MaxBeta1), _currentBeta1)); | ||
| } | ||
|
|
||
| // Adaptive Beta2 updates (if enabled) | ||
| if (_options.UseAdaptiveBeta2) | ||
| { | ||
| // Increase Beta2 (more stability) if fitness is improving, decrease otherwise | ||
| if (NumOps.GreaterThan(currentStepData.FitnessScore, previousStepData.FitnessScore)) | ||
| { | ||
| _currentBeta2 = NumOps.Multiply(_currentBeta2, NumOps.FromDouble(_options.Beta2IncreaseFactor)); | ||
| } | ||
| else | ||
| { | ||
| _currentBeta2 = NumOps.Multiply(_currentBeta2, NumOps.FromDouble(_options.Beta2DecreaseFactor)); | ||
| } | ||
|
|
||
| // Clamp to configured bounds | ||
| _currentBeta2 = MathHelper.Max(NumOps.FromDouble(_options.MinBeta2), | ||
| MathHelper.Min(NumOps.FromDouble(_options.MaxBeta2), _currentBeta2)); | ||
| } | ||
| } |
There was a problem hiding this comment.
🧩 Analysis chain
Verify adaptive Beta logic handles minimization correctly.
The adaptive parameter logic on lines 165 and 183 checks currentStepData.FitnessScore > previousStepData.FitnessScore to determine if fitness is "improving." This assumes fitness is being maximized (higher = better).
If the framework minimizes a loss function without converting it to a maximization problem (e.g., by negating), the logic would be inverted:
- When loss decreases (improving), the code would decrease Beta (more responsive), which is backwards
- When loss increases (not improving), the code would increase Beta (more smoothing), which is also backwards
Run the following script to check how fitness/loss is handled in the base class:
If minimization mode exists, update the comparison logic to respect it, or confirm that the base class UpdateAdaptiveParameters call on line 159 handles the conversion before this method runs.
🏁 Script executed:
#!/bin/bash
# Description: Check if fitness is always maximized or if minimization is supported
# Search for FitnessScore assignments and transformations
rg -n --type=cs -C3 'FitnessScore\s*=' src/Optimizers/
# Search for loss-to-fitness conversions
rg -n --type=cs -C3 'Fitness.*[Ll]oss|[Ll]oss.*Fitness' src/Optimizers/
# Search for minimization/maximization mode configuration
rg -n --type=cs -C3 'Minimi[zs]|Maximi[zs]' src/Models/Options/Length of output: 4461
</script>
🏁 Script executed:
#!/bin/bash
# Search for IsHigherScoreBetter usage
rg -n --type=cs 'IsHigherScoreBetter' src/
# Check LionOptimizer options structure
rg -n --type=cs -B5 -A15 'class.*LionOptimizer' src/Optimizers/
# Search FitnessCalculator implementation
rg -n --type=cs -B3 -A10 'CalculateFitnessScore' src/Length of output: 14061
🏁 Script executed:
#!/bin/bash
# Check if _options in LionOptimizer has IsMaximization or similar
rg -n --type=cs 'class.*LionOptimizerOptions' src/Models/Options/Length of output: 151
🏁 Script executed:
#!/bin/bash
# Look at LionOptimizer class definition and options
ast-grep --pattern 'class LionOptimizer {
$$$
}'
# Check LionOptimizerOptions structure
rg -n --type=cs -A30 'class LionOptimizerOptions' src/Models/Options/Length of output: 3203
🏁 Script executed:
#!/bin/bash
# Check GradientBasedOptimizerOptions to see if it has IsMaximization
rg -n --type=cs -A50 'class GradientBasedOptimizerOptions' src/Models/Options/ | head -80
# Confirm access to FitnessCalculator in LionOptimizer
rg -n --type=cs 'FitnessCalculator' src/Optimizers/LionOptimizer.cs
# Check if any other adapting code checks IsHigherScoreBetter
rg -n --type=cs -B2 -A2 'UpdateAdaptiveParameters' src/Optimizers/*.cs | head -60Length of output: 9716
🏁 Script executed:
#!/bin/bash
# Check GradientBasedOptimizerOptions inheritance
rg -n --type=cs -A50 'class GradientBasedOptimizerOptions' src/Models/Options/
# Verify LionOptimizer has FitnessCalculator access
rg -n --type=cs 'FitnessCalculator' src/Optimizers/LionOptimizer.cs
# Check if other optimizers' adaptive parameter updates check direction
rg -n --type=cs -B3 -A8 'UpdateAdaptiveParameters' src/Optimizers/AdamOptimizer.csLength of output: 6338
Update adaptive parameter comparisons to respect optimization direction.
Lines 165 and 183 use hardcoded NumOps.GreaterThan() comparisons assuming fitness improvement when currentFitnessScore > previousFitnessScore. This assumes maximization, but the framework supports both maximization and minimization via FitnessCalculator.IsHigherScoreBetter.
When minimizing (e.g., loss functions where lower is better), the comparison logic is inverted:
- Condition increases Beta when loss worsens (backwards)
- Condition decreases Beta when loss improves (backwards)
Use the established pattern from the codebase (genetics algorithms, other optimizers):
bool fitnessImproved = FitnessCalculator.IsHigherScoreBetter
? NumOps.GreaterThan(currentStepData.FitnessScore, previousStepData.FitnessScore)
: NumOps.LessThan(currentStepData.FitnessScore, previousStepData.FitnessScore);
if (fitnessImproved) { /* increase Beta */ }
else { /* decrease Beta */ }Apply this fix to both Beta1 (line 165) and Beta2 (line 183) adaptive logic.
🤖 Prompt for AI Agents
In src/Optimizers/LionOptimizer.cs around lines 157 to 196, change the hardcoded
NumOps.GreaterThan comparisons for Beta1 and Beta2 to respect optimization
direction by computing a fitnessImproved boolean using
FitnessCalculator.IsHigherScoreBetter ? NumOps.GreaterThan(current, previous) :
NumOps.LessThan(current, previous); then use fitnessImproved to choose the
increase vs decrease factor for both _currentBeta1 and _currentBeta2, keeping
the existing multiplication and clamping logic unchanged.
This commit implements the Lion (EvoLved Sign Momentum) optimizer as requested in issue #315.
Key Features:
Implementation Details:
Tests Include:
Resolves #315
User Story / Context
merge-dev2-to-masterSummary
Verification
Copilot Review Loop (Outcome-Based)
Record counts before/after your last push:
Files Modified
Notes