Skip to content

Fix Issue for Lion Optimizer - #381

Merged
ooples merged 5 commits into
masterfrom
claude/fix-issue-011CUsoNxYGmTa6Gp8gBbLGi
Nov 7, 2025
Merged

ooples merged 5 commits into
masterfrom
claude/fix-issue-011CUsoNxYGmTa6Gp8gBbLGi

Conversation

@ooples

@ooples ooples commented Nov 7, 2025

Copy link
Copy Markdown
Owner

This commit implements the Lion (EvoLved Sign Momentum) optimizer as requested in issue #315.

Key Features:

  • LionOptimizer.cs: Complete implementation of the Lion algorithm with sign-based gradient updates
  • LionOptimizerOptions.cs: Configuration options with sensible defaults (LR=1e-4, β1=0.9, β2=0.99)
  • Comprehensive unit tests covering all major functionality

Implementation Details:

  • Inherits from GradientBasedOptimizerBase following project architecture
  • Uses single momentum state (50% memory reduction vs Adam)
  • Implements sign-based updates for improved generalization
  • Supports decoupled weight decay (similar to AdamW)
  • Provides UpdateParameters overloads for Vector, Matrix types
  • Full serialization/deserialization support
  • Generic type support via INumericOperations

Tests Include:

  • Parameter update correctness with various gradient configurations
  • Sign-based behavior verification (magnitude independence)
  • Weight decay functionality
  • Momentum accumulation across iterations
  • Both float and double type support
  • Serialization/deserialization preservation
  • Different beta1/beta2 configurations

Resolves #315

User Story / Context

  • Reference: [US-XXX] (if applicable)
  • Base branch: merge-dev2-to-master

Summary

  • What changed and why (scoped strictly to the user story / PR intent)

Verification

  • Builds succeed (scoped to changed projects)
  • Unit tests pass locally
  • Code coverage >= 90% for touched code
  • Codecov upload succeeded (if token configured)
  • TFM verification (net46, net6.0, net8.0) passes (if packaging)
  • No unresolved Copilot comments on HEAD

Copilot Review Loop (Outcome-Based)

Record counts before/after your last push:

  • Comments on HEAD BEFORE: [N]
  • Comments on HEAD AFTER (60s): [M]
  • Final HEAD SHA: [sha]

Files Modified

  • List files changed (must align with scope)

Notes

  • Any follow-ups, caveats, or migration details

This commit implements the Lion (EvoLved Sign Momentum) optimizer as requested in issue #315.

Key Features:
- LionOptimizer.cs: Complete implementation of the Lion algorithm with sign-based gradient updates
- LionOptimizerOptions.cs: Configuration options with sensible defaults (LR=1e-4, β1=0.9, β2=0.99)
- Comprehensive unit tests covering all major functionality

Implementation Details:
- Inherits from GradientBasedOptimizerBase following project architecture
- Uses single momentum state (50% memory reduction vs Adam)
- Implements sign-based updates for improved generalization
- Supports decoupled weight decay (similar to AdamW)
- Provides UpdateParameters overloads for Vector, Matrix types
- Full serialization/deserialization support
- Generic type support via INumericOperations<T>

Tests Include:
- Parameter update correctness with various gradient configurations
- Sign-based behavior verification (magnitude independence)
- Weight decay functionality
- Momentum accumulation across iterations
- Both float and double type support
- Serialization/deserialization preservation
- Different beta1/beta2 configurations

Resolves #315
Copilot AI review requested due to automatic review settings November 7, 2025 03:22

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull Request Overview

This PR implements the Lion (Evolved Sign Momentum) optimization algorithm, a modern gradient-based optimizer that offers memory efficiency and strong performance characteristics compared to Adam.

Key Changes

  • Adds LionOptimizer<T, TInput, TOutput> class implementing the Lion algorithm with sign-based gradient updates
  • Adds LionOptimizerOptions<T, TInput, TOutput> class for configuring Lion-specific parameters (learning rate, beta1, beta2, weight decay, and adaptive parameter settings)
  • Adds comprehensive unit tests covering constructor behavior, gradient updates (vectors and matrices), weight decay, momentum building, serialization, and multiple data types

Reviewed Changes

Copilot reviewed 3 out of 3 changed files in this pull request and generated 13 comments.

File Description
src/Optimizers/LionOptimizer.cs Implements Lion optimizer with UpdateParameters methods for vectors/matrices, serialization support, and adaptive parameter handling
src/Models/Options/LionOptimizerOptions.cs Defines configuration options with default values matching Lion algorithm specifications (learning rate 1e-4, beta1=0.9, beta2=0.99)
tests/UnitTests/Optimizers/LionOptimizerTests.cs Provides 20 test cases covering optimizer initialization, parameter updates with various gradient scenarios, momentum behavior, serialization, and type support

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment thread src/Optimizers/LionOptimizer.cs
Comment thread src/Optimizers/LionOptimizer.cs
Comment thread src/Optimizers/LionOptimizer.cs Outdated
Comment thread src/Models/Options/LionOptimizerOptions.cs Outdated
Comment thread src/Optimizers/LionOptimizer.cs
Comment thread tests/UnitTests/Optimizers/LionOptimizerTests.cs
Comment thread tests/UnitTests/Optimizers/LionOptimizerTests.cs
Comment thread tests/UnitTests/Optimizers/LionOptimizerTests.cs
Comment thread tests/UnitTests/Optimizers/LionOptimizerTests.cs
Comment thread tests/UnitTests/Optimizers/LionOptimizerTests.cs
…ization

- UpdateOptions: Add InitializeAdaptiveParameters() call after setting new options
  to ensure cached parameter values reflect the updated options
- Deserialize: Add InitializeAdaptiveParameters() call after deserializing options
  to initialize cached values from deserialized data instead of keeping zero values

Both fixes ensure _currentLearningRate, _currentBeta1, and _currentBeta2 are properly
synchronized with the LionOptimizerOptions values.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>
@coderabbitai

coderabbitai Bot commented Nov 7, 2025 •

Copy link
Copy Markdown
Contributor

Note

Other AI code review bot(s) detected

CodeRabbit has detected other AI code review bot(s) in this pull request and will avoid duplicating their findings in the review comments. This may lead to a less comprehensive review.

Summary by CodeRabbit

  • New Features
    • Introduced Lion optimization algorithm for model training with comprehensive customization options including learning rate configuration, momentum decay management, weight decay regularization, and adaptive parameter tuning capabilities to support diverse optimization scenarios and improve convergence performance.

Walkthrough

Adds a new Lion optimizer implementation, a corresponding options class with Lion-specific hyperparameters, and unit tests. Includes adaptive-parameter handling, momentum state, parameter update routines (vector/matrix), serialization/deserialization, reset, and gradient-cache key changes.

Changes

Cohort / File(s) Summary
Options Configuration
src/Models/Options/LionOptimizerOptions.cs
New generic LionOptimizerOptions<T, TInput, TOutput> extending GradientBasedOptimizerOptions with properties: LearningRate, Beta1, Beta2, WeightDecay, UseAdaptiveBeta1, UseAdaptiveBeta2, MinBeta1, MaxBeta1, MinBeta2, MaxBeta2, Beta1IncreaseFactor, Beta1DecreaseFactor, Beta2IncreaseFactor, Beta2DecreaseFactor (XML docs and defaults).
Core Optimizer Implementation
src/Optimizers/LionOptimizer.cs
New LionOptimizer<T, TInput, TOutput> extends gradient-based optimizer base. Implements momentum state _m, timestep _t, adaptive parameters, constructor/initialization, Optimize flow, UpdateSolution/UpdateParameters for vectors and matrices (sign-based updates), adaptive-parameter updates, Reset, UpdateOptions/GetOptions, Serialize/Deserialize (options via JSON), and gradient-cache key generation.
Unit Tests
tests/UnitTests/Optimizers/LionOptimizerTests.cs
New comprehensive tests covering default/custom option initialization, vector/matrix updates, momentum behavior across calls, sign-based updates, weight decay, serialization/deserialization, options update, reset behavior, and float/double variants.

Sequence Diagram(s)

sequenceDiagram
    participant Caller
    participant Lion as LionOptimizer
    participant Grad as GradientCalc
    participant Eval as Evaluator

    Caller->>Lion: Optimize(inputData)
    activate Lion
    Lion->>Lion: Initialize _m, _t, adaptive params

    loop Iteration
        Lion->>Grad: Compute gradients
        Grad-->>Lion: gradients

        rect rgb(230, 245, 255)
            Note over Lion: Lion update (sign-based)
            Lion->>Lion: Interpolate momentum: m' = β1·m + (1-β1)·grad
            Lion->>Lion: Apply sign update: θ ← θ - lr·sign(m')
            Lion->>Lion: Update momentum state m ← m'
        end

        Lion->>Lion: Optionally update β1, β2 (clamped)
        Lion->>Eval: Evaluate loss / history
        Eval-->>Lion: loss
        alt Converged / Early stop
            Lion->>Lion: Break loop
        end
    end

    Lion-->>Caller: OptimizationResult
    deactivate Lion
Loading

Estimated code review effort

🎯 4 (Complex) | ⏱️ ~45 minutes

  • Areas needing extra attention:
    • Generic numeric handling and INumericOperations usage across Lion update math.
    • Sign-based update math and momentum interpolation correctness vs. Lion paper.
    • Adaptive-parameter clamping and feature-flag behavior.
    • Serialization/deserialization fidelity for options and internal state (_m, _t).
    • Unit tests: coverage for float/double and vector vs matrix consistency.

Poem

🐰
Hop, hop—I tune the learning rate,
With signed momentum, steady and straight.
No second moments, just swift little hops,
I guard the weights while the optimizer stops.
Joyful gradients, off we race!

Pre-merge checks and finishing touches

❌ Failed checks (1 warning, 1 inconclusive)
Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 43.33% which is insufficient. The required threshold is 80.00%. You can run @coderabbitai generate docstrings to improve docstring coverage.
Title check ❓ Inconclusive Title is vague and generic, using a non-descriptive phrase 'Fix Issue for Lion Optimizer' that doesn't convey the specific implementation nature of adding the Lion optimizer. Consider a more specific title like 'Implement Lion Optimizer with sign-based gradient updates' to clearly indicate this adds new functionality rather than fixing an existing bug.
✅ Passed checks (3 passed)
Check name Status Explanation
Description check ✅ Passed The description provides comprehensive details about the Lion optimizer implementation, including key features, implementation details, test coverage, and resolves issue #315.
Linked Issues check ✅ Passed The implementation addresses most core requirements from #315 including LionOptimizer class with sign-based updates, LionOptimizerOptions, and comprehensive unit tests covering various scenarios.
Out of Scope Changes check ✅ Passed All changes are focused on implementing the Lion optimizer as specified in #315; no unrelated modifications or scope creep detected in the changeset.
✨ Finishing touches
  • 📝 Generate docstrings
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Post copyable unit tests in a comment
  • Commit unit tests in branch claude/fix-issue-011CUsoNxYGmTa6Gp8gBbLGi

Comment @coderabbitai help to get the list of available commands and usage tips.

ooples and others added 2 commits November 6, 2025 23:05
- Fix spelling: EvoLved -> Evolved in documentation
- Use cached adaptive parameters (_currentBeta1, _currentBeta2, _currentLearningRate)
  in UpdateParameters methods instead of accessing options directly
- Add proper null checks with guard clauses in test files instead of lazy
  null-forgiving operator

All fixes ensure consistency with adaptive parameter adjustments and proper
null safety without using the ! operator.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 5

🧹 Nitpick comments (2)
tests/UnitTests/Optimizers/LionOptimizerTests.cs (2)

278-318: Consider verifying momentum state preservation.

The test validates that options are correctly serialized/deserialized but doesn't verify that the momentum state (_m) and timestep (_t) are preserved. This is a critical part of the optimizer's state.

Consider adding a third UpdateParameters call on both optimizers after deserialization and verifying they produce identical results, which would indirectly confirm momentum state preservation:

// After deserialization and option verification, add:
var gradient2 = new Vector<double>(new double[] { 0.3, -0.3, 0.7 });

// Both should produce same result if state was preserved
var result1 = optimizer1.UpdateParameters(parameters, gradient2);
var result2 = optimizer2.UpdateParameters(parameters, gradient2);

for (int i = 0; i < parameters.Length; i++)
{
    Assert.Equal(result1[i], result2[i], 1e-9);
}

60-108: Consider adding precise numerical verification.

While directional assertions (>, <) are appropriate for some integration-style tests, at least one test should verify actual expected values calculated by hand or from a reference implementation to ensure algorithmic correctness.

For example, for a simple scenario with known values:

  • Parameters: [1.0]
  • Gradient: [1.0]
  • Beta1: 0.0, Beta2: 0.0 (no momentum)
  • LearningRate: 0.1
  • WeightDecay: 0.0

The expected result should be exactly: 1.0 - 0.1 * sign(1.0) = 1.0 - 0.1 = 0.9

Adding such a test would provide stronger validation of the implementation.

📜 Review details

Configuration used: CodeRabbit UI

Review profile: CHILL

Plan: Pro

📥 Commits

Reviewing files that changed from the base of the PR and between 0c165ae and e55bad0.

📒 Files selected for processing (3)
  • src/Models/Options/LionOptimizerOptions.cs (1 hunks)
  • src/Optimizers/LionOptimizer.cs (1 hunks)
  • tests/UnitTests/Optimizers/LionOptimizerTests.cs (1 hunks)
🚧 Files skipped from review as they are similar to previous changes (1)
  • src/Models/Options/LionOptimizerOptions.cs
🧰 Additional context used
🧬 Code graph analysis (2)
src/Optimizers/LionOptimizer.cs (3)
src/Models/Options/LionOptimizerOptions.cs (1)
  • LionOptimizerOptions (18-134)
src/Optimizers/OptimizerBase.cs (2)
  • UpdateBestSolution (641-669)
  • UpdateIterationHistoryAndCheckEarlyStopping (818-834)
src/Helpers/MathHelper.cs (1)
  • MathHelper (16-987)
tests/UnitTests/Optimizers/LionOptimizerTests.cs (2)
src/Optimizers/LionOptimizer.cs (7)
  • LionOptimizer (27-497)
  • LionOptimizer (70-83)
  • Vector (246-291)
  • Matrix (304-357)
  • Reset (367-371)
  • Serialize (418-442)
  • Deserialize (453-479)
src/Models/Options/LionOptimizerOptions.cs (1)
  • LionOptimizerOptions (18-134)
⏰ Context from checks skipped due to timeout of 90000ms. You can increase the timeout in your CodeRabbit configuration to a maximum of 15 minutes (900000ms). (1)
  • GitHub Check: Build All Frameworks
🔇 Additional comments (5)
tests/UnitTests/Optimizers/LionOptimizerTests.cs (1)

11-58: LGTM: Constructor tests verify correct initialization.

The null-checking pattern (Assert.NotNull followed by explicit null checks with throws) effectively addresses the potential null reference warnings from past reviews while maintaining test clarity.

src/Optimizers/LionOptimizer.cs (4)

246-291: LGTM: UpdateParameters correctly uses cached adaptive parameters.

The method now correctly uses _currentBeta1, _currentBeta2, and _currentLearningRate throughout, which ensures that adaptive parameter adjustments (if enabled) are properly respected. This addresses the past review concerns.


304-357: LGTM: Matrix UpdateParameters correctly implemented.

Consistent with the Vector version, this method properly uses cached adaptive parameters and correctly implements the Lion algorithm for matrix parameters.


383-394: LGTM: State management methods correctly update adaptive parameters.

Both UpdateOptions (line 388) and Deserialize (line 468) now properly call InitializeAdaptiveParameters() to ensure the cached parameter values are synchronized with the options. This addresses the past review concerns.

Also applies to: 453-479


195-233: Algorithm implementation appears correct.

The Lion optimizer correctly implements the three-step algorithm:

  1. Interpolate between momentum and gradient (lines 204-207)
  2. Apply sign-based update with optional weight decay (lines 211-222)
  3. Update momentum for next iteration (lines 226-229)

The sign-based approach and decoupled weight decay align with the Lion algorithm specification.

Comment thread src/Optimizers/LionOptimizer.cs
Comment thread tests/UnitTests/Optimizers/LionOptimizerTests.cs
Comment thread tests/UnitTests/Optimizers/LionOptimizerTests.cs
Comment thread tests/UnitTests/Optimizers/LionOptimizerTests.cs
- Add configurable Beta1/Beta2 increase/decrease factors to LionOptimizerOptions
- Implement proper adaptive beta updates in UpdateAdaptiveParameters method
- Add SetParameters call before return in UpdateSolution to persist changes
- Fix test tolerance from 10 to 1e-9 for proper numerical precision
- Add verification that different Beta1/Beta2 values produce different results

Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>
@ooples
ooples merged commit 82fe62a into master Nov 7, 2025
4 of 5 checks passed
@ooples
ooples deleted the claude/fix-issue-011CUsoNxYGmTa6Gp8gBbLGi branch November 7, 2025 04:30

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 2

🧹 Nitpick comments (1)
src/Optimizers/LionOptimizer.cs (1)

70-83: Remove redundant field initialization.

Lines 78-80 initialize _currentLearningRate, _currentBeta1, and _currentBeta2 to NumOps.Zero, but line 82 immediately calls InitializeAdaptiveParameters() which overwrites these values with the actual option values.

Apply this diff to remove the redundant initialization:

     {
         _m = Vector<T>.Empty();
         _t = 0;
         _options = options ?? new();
-        _currentLearningRate = NumOps.Zero;
-        _currentBeta1 = NumOps.Zero;
-        _currentBeta2 = NumOps.Zero;
 
         InitializeAdaptiveParameters();
     }
📜 Review details

Configuration used: CodeRabbit UI

Review profile: CHILL

Plan: Pro

📥 Commits

Reviewing files that changed from the base of the PR and between e55bad0 and 88ef396.

📒 Files selected for processing (3)
  • src/Models/Options/LionOptimizerOptions.cs (1 hunks)
  • src/Optimizers/LionOptimizer.cs (1 hunks)
  • tests/UnitTests/Optimizers/LionOptimizerTests.cs (1 hunks)
🚧 Files skipped from review as they are similar to previous changes (1)
  • tests/UnitTests/Optimizers/LionOptimizerTests.cs
🧰 Additional context used
🧬 Code graph analysis (1)
src/Optimizers/LionOptimizer.cs (2)
src/Models/Options/LionOptimizerOptions.cs (1)
  • LionOptimizerOptions (18-178)
src/Helpers/MathHelper.cs (1)
  • MathHelper (16-987)
⏰ Context from checks skipped due to timeout of 90000ms. You can increase the timeout in your CodeRabbit configuration to a maximum of 15 minutes (900000ms). (1)
  • GitHub Check: Build All Frameworks
🔇 Additional comments (12)
src/Models/Options/LionOptimizerOptions.cs (1)

1-18: LGTM! Spelling issue resolved.

The class documentation correctly uses "Evolved Sign Momentum" throughout. This addresses the previous review feedback about the "EvoLved" spelling error.

src/Optimizers/LionOptimizer.cs (11)

1-27: LGTM! Class documentation is accurate.

The optimizer correctly implements the Lion (Evolved Sign Momentum) algorithm with comprehensive documentation. The spelling issue from previous reviews has been resolved.


92-97: LGTM! Adaptive parameter initialization is correct.

Properly converts hyperparameters from double to generic type T using NumOps.FromDouble, following the architectural requirement to use INumericOperations<T> for all numeric operations.


109-146: LGTM! Main optimization loop is well-structured.

The optimization process correctly implements the Lion algorithm flow: gradient calculation, parameter updates using the Lion three-step rule, fitness evaluation, adaptive parameter adjustments, and convergence checks. The implementation properly initializes momentum state and tracks the best solution.


212-251: LGTM! Lion update rule is correctly implemented.

The three-step Lion algorithm is properly implemented:

  1. Interpolation: c_t = beta1 * m_{t-1} + (1 - beta1) * g_t (lines 221-224)
  2. Sign-based update with decoupled weight decay: theta_t = theta_{t-1} - lr * (sign(c_t) + lambda * theta_{t-1}) (lines 228-239)
  3. Momentum update: m_t = beta2 * m_{t-1} + (1 - beta2) * g_t (lines 243-246)

The method now correctly calls SetParameters() (line 249) before returning, which addresses the previous review concern about persisting parameter updates.


322-375: LGTM! Matrix parameter update correctly implements Lion algorithm.

The matrix overload properly flattens the 2D parameter space into a 1D momentum vector and applies the Lion update rule element-wise. The implementation correctly uses the adaptive parameter fields (_currentBeta1, _currentBeta2, _currentLearningRate) rather than directly accessing options, which addresses previous review feedback.

Note: The momentum reset behavior on size change (lines 324-328) mirrors the vector version and should be verified as discussed in the previous comment.


385-389: LGTM! Reset functionality is correct.

Properly clears the momentum state and resets the time step counter, allowing the optimizer to be reused for new training runs or re-initialized mid-training.


401-412: LGTM! Options update correctly reinitializes parameters.

The method now properly calls InitializeAdaptiveParameters() (line 406) after updating options, which addresses the previous review concern about cached parameter values not being synchronized with new options.


422-425: LGTM! Options accessor is straightforward.


436-497: LGTM! Serialization and deserialization are complete.

Both methods correctly preserve and restore the optimizer's full state:

  • Base class state via base.Serialize()/base.Deserialize()
  • Lion-specific options via JSON serialization
  • Time step counter _t
  • Momentum vector _m

The deserialization method now properly calls InitializeAdaptiveParameters() (line 486) to synchronize cached adaptive parameter values with the deserialized options, which addresses previous review feedback.


510-514: LGTM! Gradient cache key generation is appropriate.

Correctly extends the base cache key with Lion-specific parameters (LearningRate and MaxIterations) to ensure gradient caching distinguishes between different optimizer configurations.


264-309: ****

The momentum reset on parameter size change is an intentional, consistent pattern across all optimizers in the codebase—not a bug or oversight. The same logic appears in AdamOptimizer (lines 220-224), which checks if size changes and reinitializes momentum vectors, matching Lion's approach exactly. This defensive pattern supports flexible architectures and edge cases where parameter dimensions may shift, and there is no evidence in the codebase of problems with this behavior. Throwing an exception instead would break compatibility with legitimate use cases.

Likely an incorrect or invalid review comment.

Comment on lines +31 to +177
public double LearningRate { get; set; } = 1e-4;

/// <summary>
/// Gets or sets the exponential decay rate for the momentum interpolation (used for computing the update).
/// </summary>
/// <value>The beta1 value, defaulting to 0.9.</value>
/// <remarks>
/// <para><b>For Beginners:</b> Beta1 controls how much Lion blends the current gradient with past momentum
/// when deciding which direction to move. A value of 0.9 means it gives 90% weight to the past momentum
/// and 10% to the new gradient. This is like having inertia - you don't change direction immediately when
/// you get new information. Higher values (closer to 1) create smoother updates but slower adaptation,
/// while lower values respond more quickly to new gradients.</para>
/// </remarks>
public double Beta1 { get; set; } = 0.9;

/// <summary>
/// Gets or sets the exponential decay rate for updating the momentum state.
/// </summary>
/// <value>The beta2 value, defaulting to 0.99.</value>
/// <remarks>
/// <para><b>For Beginners:</b> Beta2 controls how much Lion remembers from its momentum history when
/// updating the momentum state for the next iteration. A value of 0.99 means it retains 99% of the old
/// momentum and incorporates 1% from the new gradient. This creates a long memory of past gradients,
/// helping smooth out noisy updates. Think of it like a heavy flywheel that doesn't change speed quickly -
/// it provides stability during training.</para>
/// </remarks>
public double Beta2 { get; set; } = 0.99;

/// <summary>
/// Gets or sets the weight decay (L2 regularization) coefficient.
/// </summary>
/// <value>The weight decay value, defaulting to 0.0.</value>
/// <remarks>
/// <para><b>For Beginners:</b> Weight decay helps prevent overfitting by penalizing large parameter values.
/// A value of 0.0 means no weight decay. When set to a small positive value (e.g., 0.01 or 0.1), it encourages
/// the model to keep weights small, which often improves generalization to new data. Think of it like a tax
/// on complexity - it encourages the model to be as simple as possible while still solving the problem.
/// Lion applies weight decay in a decoupled manner, similar to AdamW.</para>
/// </remarks>
public double WeightDecay { get; set; } = 0.0;

/// <summary>
/// Gets or sets whether to automatically adjust Beta1 during training.
/// </summary>
/// <value>False by default, as Lion typically uses fixed betas.</value>
/// <remarks>
/// <para><b>For Beginners:</b> When enabled, Beta1 will be automatically adjusted based on training progress.
/// However, Lion was designed to work well with fixed beta values, so this is disabled by default.
/// Unlike Adam, Lion is less sensitive to beta parameter choices due to its sign-based updates.
/// You typically don't need to enable this unless you're doing advanced experimentation.</para>
/// </remarks>
public bool UseAdaptiveBeta1 { get; set; } = false;

/// <summary>
/// Gets or sets whether to automatically adjust Beta2 during training.
/// </summary>
/// <value>False by default, as Lion typically uses fixed betas.</value>
/// <remarks>
/// <para><b>For Beginners:</b> When enabled, Beta2 will be automatically adjusted based on training progress.
/// However, Lion was designed to work well with fixed beta values, so this is disabled by default.
/// The sign-based nature of Lion makes it robust to beta parameter variations.</para>
/// </remarks>
public bool UseAdaptiveBeta2 { get; set; } = false;

/// <summary>
/// Gets or sets the minimum allowed value for Beta1.
/// </summary>
/// <value>The minimum Beta1 value, defaulting to 0.85.</value>
/// <remarks>
/// <para><b>For Beginners:</b> If adaptive Beta1 is enabled, this prevents it from dropping too low.
/// A minimum of 0.85 ensures some momentum is always maintained.</para>
/// </remarks>
public double MinBeta1 { get; set; } = 0.85;

/// <summary>
/// Gets or sets the maximum allowed value for Beta1.
/// </summary>
/// <value>The maximum Beta1 value, defaulting to 0.95.</value>
/// <remarks>
/// <para><b>For Beginners:</b> If adaptive Beta1 is enabled, this prevents it from becoming too high.
/// A maximum of 0.95 ensures some responsiveness to new gradients.</para>
/// </remarks>
public double MaxBeta1 { get; set; } = 0.95;

/// <summary>
/// Gets or sets the minimum allowed value for Beta2.
/// </summary>
/// <value>The minimum Beta2 value, defaulting to 0.95.</value>
/// <remarks>
/// <para><b>For Beginners:</b> If adaptive Beta2 is enabled, this prevents it from dropping too low.
/// A minimum of 0.95 ensures momentum state retains sufficient history.</para>
/// </remarks>
public double MinBeta2 { get; set; } = 0.95;

/// <summary>
/// Gets or sets the maximum allowed value for Beta2.
/// </summary>
/// <value>The maximum Beta2 value, defaulting to 0.999.</value>
/// <remarks>
/// <para><b>For Beginners:</b> If adaptive Beta2 is enabled, this prevents it from becoming too high.
/// A maximum of 0.999 ensures the momentum state can still adapt to changes.</para>
/// </remarks>
public double MaxBeta2 { get; set; } = 0.999;

/// <summary>
/// Gets or sets the factor by which Beta1 is increased when fitness improves.
/// </summary>
/// <value>The Beta1 increase factor, defaulting to 1.02.</value>
/// <remarks>
/// <para><b>For Beginners:</b> When adaptive Beta1 is enabled and the optimizer is improving,
/// Beta1 is multiplied by this factor. A value of 1.02 means Beta1 increases by 2% each time
/// fitness improves. Higher Beta1 values create smoother, more stable updates.</para>
/// </remarks>
public double Beta1IncreaseFactor { get; set; } = 1.02;

/// <summary>
/// Gets or sets the factor by which Beta1 is decreased when fitness does not improve.
/// </summary>
/// <value>The Beta1 decrease factor, defaulting to 0.98.</value>
/// <remarks>
/// <para><b>For Beginners:</b> When adaptive Beta1 is enabled and the optimizer is not improving,
/// Beta1 is multiplied by this factor. A value of 0.98 means Beta1 decreases by 2% each time
/// fitness doesn't improve. Lower Beta1 values make the optimizer more responsive to new gradients.</para>
/// </remarks>
public double Beta1DecreaseFactor { get; set; } = 0.98;

/// <summary>
/// Gets or sets the factor by which Beta2 is increased when fitness improves.
/// </summary>
/// <value>The Beta2 increase factor, defaulting to 1.02.</value>
/// <remarks>
/// <para><b>For Beginners:</b> When adaptive Beta2 is enabled and the optimizer is improving,
/// Beta2 is multiplied by this factor. A value of 1.02 means Beta2 increases by 2% each time
/// fitness improves. Higher Beta2 values create longer memory of past gradients for more stability.</para>
/// </remarks>
public double Beta2IncreaseFactor { get; set; } = 1.02;

/// <summary>
/// Gets or sets the factor by which Beta2 is decreased when fitness does not improve.
/// </summary>
/// <value>The Beta2 decrease factor, defaulting to 0.98.</value>
/// <remarks>
/// <para><b>For Beginners:</b> When adaptive Beta2 is enabled and the optimizer is not improving,
/// Beta2 is multiplied by this factor. A value of 0.98 means Beta2 decreases by 2% each time
/// fitness doesn't improve. Lower Beta2 values make the momentum state more responsive to recent changes.</para>
/// </remarks>
public double Beta2DecreaseFactor { get; set; } = 0.98;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🛠️ Refactor suggestion | 🟠 Major

Add parameter validation to ensure valid configurations.

The options class lacks validation for hyperparameter constraints. Invalid values could lead to runtime errors, non-convergence, or incorrect optimization behavior.

Consider adding validation in a constructor or via property setters:

public double LearningRate
{
    get => _learningRate;
    set
    {
        if (value <= 0)
            throw new ArgumentException("LearningRate must be positive.", nameof(LearningRate));
        _learningRate = value;
    }
}

Key constraints to validate:

  • LearningRate > 0
  • Beta1 ∈ (0, 1) and Beta2 ∈ (0, 1)
  • WeightDecay >= 0
  • MinBeta1 < MaxBeta1 and MinBeta2 < MaxBeta2
  • Beta1/Beta2 ∈ [MinBeta1/MinBeta2, MaxBeta1/MaxBeta2]
  • Beta1IncreaseFactor > 1 and Beta1DecreaseFactor < 1 (and similarly for Beta2)

Alternatively, perform validation in the LionOptimizer constructor if the framework convention is to validate at consumption rather than assignment.

Comment on lines +157 to +196
protected override void UpdateAdaptiveParameters(OptimizationStepData<T, TInput, TOutput> currentStepData, OptimizationStepData<T, TInput, TOutput> previousStepData)
{
base.UpdateAdaptiveParameters(currentStepData, previousStepData);

// Adaptive Beta1 updates (if enabled)
if (_options.UseAdaptiveBeta1)
{
// Increase Beta1 (more smoothing) if fitness is improving, decrease (faster adaptation) otherwise
if (NumOps.GreaterThan(currentStepData.FitnessScore, previousStepData.FitnessScore))
{
_currentBeta1 = NumOps.Multiply(_currentBeta1, NumOps.FromDouble(_options.Beta1IncreaseFactor));
}
else
{
_currentBeta1 = NumOps.Multiply(_currentBeta1, NumOps.FromDouble(_options.Beta1DecreaseFactor));
}

// Clamp to configured bounds
_currentBeta1 = MathHelper.Max(NumOps.FromDouble(_options.MinBeta1),
MathHelper.Min(NumOps.FromDouble(_options.MaxBeta1), _currentBeta1));
}

// Adaptive Beta2 updates (if enabled)
if (_options.UseAdaptiveBeta2)
{
// Increase Beta2 (more stability) if fitness is improving, decrease otherwise
if (NumOps.GreaterThan(currentStepData.FitnessScore, previousStepData.FitnessScore))
{
_currentBeta2 = NumOps.Multiply(_currentBeta2, NumOps.FromDouble(_options.Beta2IncreaseFactor));
}
else
{
_currentBeta2 = NumOps.Multiply(_currentBeta2, NumOps.FromDouble(_options.Beta2DecreaseFactor));
}

// Clamp to configured bounds
_currentBeta2 = MathHelper.Max(NumOps.FromDouble(_options.MinBeta2),
MathHelper.Min(NumOps.FromDouble(_options.MaxBeta2), _currentBeta2));
}
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🔴 Critical

🧩 Analysis chain

Verify adaptive Beta logic handles minimization correctly.

The adaptive parameter logic on lines 165 and 183 checks currentStepData.FitnessScore > previousStepData.FitnessScore to determine if fitness is "improving." This assumes fitness is being maximized (higher = better).

If the framework minimizes a loss function without converting it to a maximization problem (e.g., by negating), the logic would be inverted:

  • When loss decreases (improving), the code would decrease Beta (more responsive), which is backwards
  • When loss increases (not improving), the code would increase Beta (more smoothing), which is also backwards

Run the following script to check how fitness/loss is handled in the base class:

If minimization mode exists, update the comparison logic to respect it, or confirm that the base class UpdateAdaptiveParameters call on line 159 handles the conversion before this method runs.


🏁 Script executed:

#!/bin/bash
# Description: Check if fitness is always maximized or if minimization is supported

# Search for FitnessScore assignments and transformations
rg -n --type=cs -C3 'FitnessScore\s*=' src/Optimizers/

# Search for loss-to-fitness conversions
rg -n --type=cs -C3 'Fitness.*[Ll]oss|[Ll]oss.*Fitness' src/Optimizers/

# Search for minimization/maximization mode configuration
rg -n --type=cs -C3 'Minimi[zs]|Maximi[zs]' src/Models/Options/

Length of output: 4461


</script>

🏁 Script executed:

#!/bin/bash
# Search for IsHigherScoreBetter usage
rg -n --type=cs 'IsHigherScoreBetter' src/

# Check LionOptimizer options structure
rg -n --type=cs -B5 -A15 'class.*LionOptimizer' src/Optimizers/

# Search FitnessCalculator implementation
rg -n --type=cs -B3 -A10 'CalculateFitnessScore' src/

Length of output: 14061


🏁 Script executed:

#!/bin/bash
# Check if _options in LionOptimizer has IsMaximization or similar
rg -n --type=cs 'class.*LionOptimizerOptions' src/Models/Options/

Length of output: 151


🏁 Script executed:

#!/bin/bash
# Look at LionOptimizer class definition and options
ast-grep --pattern 'class LionOptimizer {
  $$$
}'

# Check LionOptimizerOptions structure
rg -n --type=cs -A30 'class LionOptimizerOptions' src/Models/Options/

Length of output: 3203



🏁 Script executed:

#!/bin/bash
# Check GradientBasedOptimizerOptions to see if it has IsMaximization
rg -n --type=cs -A50 'class GradientBasedOptimizerOptions' src/Models/Options/ | head -80

# Confirm access to FitnessCalculator in LionOptimizer
rg -n --type=cs 'FitnessCalculator' src/Optimizers/LionOptimizer.cs

# Check if any other adapting code checks IsHigherScoreBetter
rg -n --type=cs -B2 -A2 'UpdateAdaptiveParameters' src/Optimizers/*.cs | head -60

Length of output: 9716


🏁 Script executed:

#!/bin/bash
# Check GradientBasedOptimizerOptions inheritance
rg -n --type=cs -A50 'class GradientBasedOptimizerOptions' src/Models/Options/

# Verify LionOptimizer has FitnessCalculator access
rg -n --type=cs 'FitnessCalculator' src/Optimizers/LionOptimizer.cs

# Check if other optimizers' adaptive parameter updates check direction
rg -n --type=cs -B3 -A8 'UpdateAdaptiveParameters' src/Optimizers/AdamOptimizer.cs

Length of output: 6338


Update adaptive parameter comparisons to respect optimization direction.

Lines 165 and 183 use hardcoded NumOps.GreaterThan() comparisons assuming fitness improvement when currentFitnessScore > previousFitnessScore. This assumes maximization, but the framework supports both maximization and minimization via FitnessCalculator.IsHigherScoreBetter.

When minimizing (e.g., loss functions where lower is better), the comparison logic is inverted:

  • Condition increases Beta when loss worsens (backwards)
  • Condition decreases Beta when loss improves (backwards)

Use the established pattern from the codebase (genetics algorithms, other optimizers):

bool fitnessImproved = FitnessCalculator.IsHigherScoreBetter 
    ? NumOps.GreaterThan(currentStepData.FitnessScore, previousStepData.FitnessScore)
    : NumOps.LessThan(currentStepData.FitnessScore, previousStepData.FitnessScore);

if (fitnessImproved) { /* increase Beta */ } 
else { /* decrease Beta */ }

Apply this fix to both Beta1 (line 165) and Beta2 (line 183) adaptive logic.

🤖 Prompt for AI Agents
In src/Optimizers/LionOptimizer.cs around lines 157 to 196, change the hardcoded
NumOps.GreaterThan comparisons for Beta1 and Beta2 to respect optimization
direction by computing a fitnessImproved boolean using
FitnessCalculator.IsHigherScoreBetter ? NumOps.GreaterThan(current, previous) :
NumOps.LessThan(current, previous); then use fitnessImproved to choose the
increase vs decrease factor for both _currentBeta1 and _currentBeta2, keeping
the existing multiplication and clamping logic unchanged.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Gap Analysis] Implement Lion Optimizer

3 participants