Skip to content

feat: add 8-bit Adam optimizer for memory efficiency - #814

Merged
ooples merged 2 commits into
masterfrom
feat/8bit-adam-optimizer-314
Feb 3, 2026
Merged

ooples merged 2 commits into
masterfrom
feat/8bit-adam-optimizer-314

Conversation

@ooples

@ooples ooples commented Feb 2, 2026

Copy link
Copy Markdown
Owner

Summary

  • Implements 8-bit Adam optimizer that reduces optimizer state memory by ~4x
  • Uses block-wise quantization with configurable block size for accurate representation
  • Provides options for dynamic vs static quantization, stochastic rounding, and partial compression
  • Includes comprehensive integration tests covering various configurations and edge cases

Motivation

Closes #314

Large models can have billions of parameters, and standard Adam stores 2 full-precision values per parameter (first and second moment estimates). This can require 16GB+ of optimizer memory alone. 8-bit Adam quantizes these states, reducing memory usage to approximately 2 bytes per parameter plus a small overhead for scaling factors.

Implementation Details

  • Adam8BitOptimizerOptions<T, TInput, TOutput>: Configuration class with options for:

    • BlockSize (default 2048): Number of elements per quantization block
    • UseDynamicQuantization (default true): Whether to adapt scales during training
    • CompressBothMoments (default true): Whether to quantize both m and v
    • QuantizationPercentile (default 99.9): Percentile for outlier-aware scaling
    • UseStochasticRounding (default false): Use probabilistic rounding
  • Adam8BitOptimizer<T, TInput, TOutput>: Main optimizer implementation with:

    • Signed 8-bit quantization for first moment (values can be negative)
    • Unsigned 8-bit quantization for second moment (always positive)
    • Block-wise scaling factors for better precision
    • Full serialization/deserialization support
    • GetMemoryUsage() method for tracking memory savings

Test plan

  • All 28 Adam8Bit integration tests pass
  • All 100 Adam-related tests pass (no regressions)
  • Tests cover:
    • Simple and complex optimization targets (quadratic, Rosenbrock)
    • Comparison with standard Adam
    • Large parameter counts (10K parameters)
    • Edge cases (zero gradients, large gradients, small gradients)
    • Various configurations (dynamic/static, compress both/one moment, stochastic/standard rounding)
    • Serialization/deserialization roundtrip
    • Float and double types

🤖 Generated with Claude Code

Implements Adam optimizer with 8-bit quantized state storage, reducing
optimizer memory usage by approximately 4x compared to standard Adam.

Key features:
- Block-wise quantization with configurable block size (default 2048)
- Signed quantization for first moment (m), unsigned for second moment (v)
- Dynamic quantization that adapts scales during training
- Option to keep first moment in full precision (CompressBothMoments=false)
- Stochastic rounding option for reduced quantization bias
- Percentile-based scaling for outlier robustness
- Full serialization/deserialization support
- Memory usage tracking via GetMemoryUsage()

This is particularly useful for:
- Training large models where optimizer memory is a bottleneck
- GPU training with limited VRAM
- Distributed training with memory constraints

Closes #314

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
Copilot AI review requested due to automatic review settings February 2, 2026 20:19
@vercel

vercel Bot commented Feb 2, 2026 •

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated (UTC)
aidotnet-playground-api Ready Ready Preview, Comment Feb 3, 2026 6:26pm

@coderabbitai

coderabbitai Bot commented Feb 2, 2026 •

Copy link
Copy Markdown
Contributor

Summary by CodeRabbit

  • New Features

    • Added new 8-bit Adam optimizer variant with fully customizable configuration parameters for training workflows
    • Supports state persistence through serialization and deserialization, enabling checkpoint management
  • Tests

    • Added comprehensive integration test suite validating convergence behavior, robustness across edge cases, and various configuration scenarios

Walkthrough

Introduces an 8-bit quantized Adam optimizer with configurable options and comprehensive integration tests. Stores optimizer states in 8-bit quantized format with per-block scaling, dynamic quantization, and optional full-precision updates for memory efficiency.

Changes

Cohort / File(s) Summary
Options Configuration
src/Models/Options/Adam8BitOptimizerOptions.cs
New generic options class extending AdamOptimizerOptions with 6 configurable properties: BlockSize (default 2048), UseDynamicQuantization (default true), QuantizationPercentile (default 99.9), FullPrecisionUpdateFrequency (default 0), UseStochasticRounding (default false), and CompressBothMoments (default true). Each property includes detailed XML documentation.
Optimizer Implementation
src/Optimizers/Adam8BitOptimizer.cs
New optimizer class implementing 8-bit quantization of Adam states with per-block scaling. Features include quantize/dequantize helpers, adaptive parameter updates with bias correction, mixed-precision support, stochastic rounding, serialization/deserialization, memory tracking, and batch-wise optimization workflow with early stopping.
Integration Tests
tests/AiDotNet.Tests/IntegrationTests/Optimizers/Adam8BitOptimizerIntegrationTests.cs
Comprehensive test suite with 14 test methods covering convergence on quadratic and Rosenbrock functions, comparison with standard Adam, large-scale parameter handling, memory statistics, edge cases (zero/large/small/mixed-sign gradients), serialization round-trips, configuration variants (dynamic quantization, moment compression, rounding modes), and multi-type support (float).

Sequence Diagram

sequenceDiagram
    participant Client
    participant Optimizer as Adam8BitOptimizer
    participant State as Quantized State
    participant Compute as Computation Engine
    
    Client->>Optimizer: Initialize(parameters, options)
    Optimizer->>State: InitializeQuantizedState()
    State-->>Optimizer: _mQuantized, _vQuantized, scales
    
    loop Each Batch
        Client->>Optimizer: UpdateParameters(params, gradient)
        
        Optimizer->>State: Dequantize(_mQuantized, _mScales)
        State-->>Optimizer: m (full precision)
        
        Optimizer->>State: Dequantize(_vQuantized, _vScales)
        State-->>Optimizer: v (full precision)
        
        Optimizer->>Compute: Adam Update with bias correction<br/>(m, v, gradient, β₁, β₂, α)
        Compute-->>Optimizer: Updated m, v, parameters
        
        Optimizer->>State: Quantize(m, blockSize)<br/>with optional stochastic rounding
        State-->>Optimizer: _mQuantized, _mScales
        
        Optimizer->>State: Quantize(v, blockSize)<br/>with optional stochastic rounding
        State-->>Optimizer: _vQuantized, _vScales
        
        Optimizer-->>Client: Updated parameters
    end
Loading

Estimated code review effort

🎯 4 (Complex) | ⏱️ ~60 minutes

Suggested labels

feature

Poem

🐰 Hops with glee ✨

Eight bits of wisdom, memory reduced,
Through quantum tunnels, states deduced,
The Adam optimizer hops so fast,
Where smaller bits make memory last,
More models now fit in our warren!

🚥 Pre-merge checks | ✅ 4 | ❌ 1
❌ Failed checks (1 warning)
Check name Status Explanation Resolution
Linked Issues check ⚠️ Warning The PR partially addresses #314 requirements: AC 1.2 (Adam8BitOptimizer implementation) is complete, AC 1.3 (tests) are included; however, AC 1.1 (quantization scheme research documentation) is not evidenced and critical architectural requirements regarding INumericOperations usage, PredictionModelBuilder integration, and interface/base-class patterns are not met. Verify INumericOperations is properly used throughout, add interfaces to src/Interfaces/, create base classes following inheritance patterns, integrate with PredictionModelBuilder, add XML documentation with beginner-friendly explanations, and ensure all critical architectural requirements are met.
✅ Passed checks (4 passed)
Check name Status Explanation
Title check ✅ Passed The title 'feat: add 8-bit Adam optimizer for memory efficiency' clearly summarizes the main change: implementing an 8-bit Adam optimizer focused on memory efficiency.
Description check ✅ Passed The PR description comprehensively explains the implementation of 8-bit Adam optimizer, motivation (closes #314), design details, and test coverage, all directly related to the changeset.
Out of Scope Changes check ✅ Passed All changes (Adam8BitOptimizer, Adam8BitOptimizerOptions, and integration tests) are directly in scope for implementing the 8-bit Adam optimizer as specified in issue #314.
Docstring Coverage ✅ Passed Docstring coverage is 100.00% which is sufficient. The required threshold is 80.00%.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing touches
  • 📝 Generate docstrings
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Post copyable unit tests in a comment
  • Commit unit tests in branch feat/8bit-adam-optimizer-314

Comment @coderabbitai help to get the list of available commands and usage tips.

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR implements an 8-bit quantized Adam optimizer to reduce memory consumption for optimizer states. The implementation uses block-wise quantization to compress the momentum and variance estimates from full precision to 8-bit representation, achieving approximately 4x memory savings while maintaining optimization quality.

Changes:

  • Added Adam8BitOptimizer<T, TInput, TOutput> class with quantization/dequantization logic for optimizer states
  • Created Adam8BitOptimizerOptions<T, TInput, TOutput> with configurable block size, quantization modes, and compression settings
  • Implemented comprehensive integration tests covering various optimization scenarios, edge cases, and configurations

Reviewed changes

Copilot reviewed 3 out of 3 changed files in this pull request and generated 3 comments.

File Description
Adam8BitOptimizer.cs Core optimizer implementation with 8-bit quantization for Adam's moment estimates
Adam8BitOptimizerOptions.cs Configuration options for block size, dynamic/static quantization, and compression modes
Adam8BitOptimizerIntegrationTests.cs Integration tests verifying optimizer behavior across different scenarios and edge cases

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment thread src/Optimizers/Adam8BitOptimizer.cs Outdated
Comment thread src/Optimizers/Adam8BitOptimizer.cs Outdated
@coderabbitai coderabbitai Bot added the feature Feature work item label Feb 2, 2026

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 2

🤖 Fix all issues with AI agents
In `@src/Optimizers/Adam8BitOptimizer.cs`:
- Around line 667-676: Serialize currently writes v state only when _vQuantized
!= null but Deserialize unconditionally reads vLength, causing mismatch; modify
Serialize() to write an explicit null/hasValue marker (e.g., a bool or byte)
before writing v state when handling _vQuantized and _vScales, and update
Deserialize() to first read that marker and only read vLength, the _vQuantized
byte[] and the _vScales floats when the marker indicates presence; reference the
Serialize(), Deserialize(), _vQuantized, _vScales and the vLength read logic
when applying the change.
- Around line 461-467: UpdateParameters currently uses fixed
_options.Beta1/_options.Beta2 while UpdateSolution uses adaptive
_currentBeta1/_currentBeta2; change UpdateParameters to use the adaptive values
when UseAdaptiveBetas is true by replacing references to
_options.Beta1/_options.Beta2 with the current values
(_currentBeta1/_currentBeta2) when computing beta1, beta2, oneMinusBeta1,
oneMinusBeta2 and biasCorrection1/2 (compute biasCorrection using Math.Pow on
the underlying double adaptive betas or their converted values and _t), and fall
back to _options.* when adaptation is disabled so both methods behave
consistently (look for the block computing beta1, beta2, oneMinusBeta1,
oneMinusBeta2, epsilon, biasCorrection1, biasCorrection2).
🧹 Nitpick comments (6)
src/Models/Options/Adam8BitOptimizerOptions.cs (2)

46-46: Consider adding validation for BlockSize.

A BlockSize of zero or negative would cause division-by-zero or invalid array allocations in the optimizer. Consider adding validation either here with a property backing field, or document that validation occurs in the optimizer.

💡 Optional: Add validation in a backing field
+    private int _blockSize = 2048;
+
     /// <summary>
     /// Gets or sets the block size for block-wise quantization.
     /// </summary>
-    public int BlockSize { get; set; } = 2048;
+    public int BlockSize
+    {
+        get => _blockSize;
+        set => _blockSize = value > 0 ? value : throw new ArgumentOutOfRangeException(nameof(value), "BlockSize must be positive.");
+    }

81-81: Consider validating QuantizationPercentile range.

Values outside 0-100 could cause index-out-of-bounds errors in the percentile calculation logic.

tests/AiDotNet.Tests/IntegrationTests/Optimizers/Adam8BitOptimizerIntegrationTests.cs (1)

492-525: Good float type coverage.

Consider adding tests for other numeric types if supported (e.g., decimal) to increase type coverage per the issue requirements.

src/Optimizers/Adam8BitOptimizer.cs (3)

118-118: Fixed random seed may affect stochastic rounding reproducibility across instances.

The fixed seed 42 ensures reproducibility within a single optimizer instance, but multiple optimizers will share the same random sequence. Consider making the seed configurable or using Random.Shared for non-deterministic behavior when UseStochasticRounding is enabled.


706-725: Deserialization reads compressBothMoments separately from options.

The serialized compressBothMoments flag (line 706) should match _options.CompressBothMoments (from deserialized options). If they ever diverge, the state restoration could be incorrect. Consider adding a consistency check.

💡 Optional: Add consistency validation
             // Deserialize first moment
             bool compressBothMoments = reader.ReadBoolean();
+            if (compressBothMoments != _options.CompressBothMoments)
+            {
+                throw new InvalidOperationException(
+                    $"Serialized state compression mode ({compressBothMoments}) does not match options ({_options.CompressBothMoments}).");
+            }
             if (compressBothMoments)

606-608: Type size assumption may not cover all numeric types.

The code assumes T is either float (4 bytes) or double (8 bytes). Other numeric types like decimal (16 bytes) or Half (2 bytes) would report incorrect memory usage.

Comment thread src/Optimizers/Adam8BitOptimizer.cs Outdated
Comment thread src/Optimizers/Adam8BitOptimizer.cs
- Clarify comment about [-127, 127] to [1, 255] mapping (0 is unused)
- Rename typeSize2 to bytesPerElement and compute once
- Use adaptive betas (_currentBeta1, _currentBeta2) instead of fixed _options values
- Fix serialization/deserialization mismatch by adding hasVQuantized marker
- Document the alignment buffer constant in test

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

@github-advanced-security github-advanced-security AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

CodeQL found more than 20 potential problems in the proposed changes. Check the Files changed tab for more details.

@sonarqubecloud

sonarqubecloud Bot commented Feb 3, 2026

Copy link
Copy Markdown

Quality Gate Failed Quality Gate failed

Failed conditions
0.0% Coverage on New Code (required ≥ 80%)
10.1% Duplication on New Code (required ≤ 3%)

See analysis details on SonarQube Cloud

This branch was successfully deployed

1 active deployment
Preview — 6f190f97 Deployed Feb 3, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

feature Feature work item

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Gap Analysis] Implement 8-bit Adam Optimizer for Memory Efficiency

4 participants