Skip to content

perf(training): make BF16-Adam fused-compatible (proper bf16 moment kernel, not a gate) - #1745

Merged
ooples merged 8 commits into
masterfrom
perf/bf16-adam-fused-compat-memory-gate
Jul 2, 2026
Merged

ooples merged 8 commits into
masterfrom
perf/bf16-adam-fused-compat-memory-gate

Conversation

@ooples

@ooples ooples commented Jun 29, 2026 •

Copy link
Copy Markdown
Owner

What changed

Replaces the interim memory-gate workaround with the real fix: BF16-Adam now keeps the fused fast path, so large models get both the fused speed and the halved optimizer-state footprint — no tradeoff.

The problem (why the gate was lazy)

Adam8BitOptimizer in BF16 moment mode was not fused-kernel-compatible, so selecting it dropped the whole model off the compiled fused-training path onto the eager autograd tape (~10× slower). The first version of this PR just gated BF16 off under memory pressure to keep the fused path — trading the memory saving away instead of solving it.

The real fix — proper fused BF16 moment kernel

Pairs with AiDotNet.Tensors PR ooples/AiDotNet.Tensors#713 (fused bf16 moment Adam/AdamW kernel + ICompiledTrainingPlan.RequestBf16MomentStorage):

  • Adam8BitOptimizer implements IFusedOptimizerSpec — in BF16 moment mode it maps to the fused Adam kernel with UseBf16Moments=true. True int8 block-quant mode (and adaptive-LR / AMSGrad) still has no fused kernel and correctly falls back to eager.
  • FusedOptimizerConfig.UseBf16Moments flows through TryMapToFusedOptimizerConfig; CompiledTapeTrainingStep calls plan.RequestBf16MomentStorage(true) before ConfigureOptimizer, so the plan allocates half-size ushort[] m/v buffers and dispatches the bf16 kernel.
  • ShouldUseBFloat16Optimizer reverts to a plain size threshold — the memory gate existed only to avoid losing the fused path, which no longer happens.

Verification

  • AiDotNet lib builds against the patched Tensors; Adam8BitFusedSpecTests (3/3) prove BF16 mode → fused Adam (UseBf16Moments=true), block-quant → no map, AMSGrad/adaptive-LR → fall back.
  • Tensors side: bf16-moment Adam and AdamW track an fp32 reference within a tight bound, bit-deterministic; 58 existing fused-optimizer tests stay green; builds net471/net8.0/net10.0.

⚠️ Merge ordering (cross-repo)

This depends on AiDotNet.Tensors#713 releasing first:

  1. Merge + release Tensors#713 → bump the AiDotNet.Tensors package here → CI goes green.
  2. Locally verified via a temporary ProjectReference to the patched Tensors (kept out of the committed diff); CI will be red until the package bump lands.

Closes the BF16-fused-compat half of the optimizer-memory-ladder work.

Summary by CodeRabbit

  • New Features

    • Improved fused training support for Adam/AdamW-style optimization with optional BF16 moment storage.
    • Expanded 8-bit Adam behavior so BF16 mode can use the faster fused execution path when eligible.
  • Bug Fixes

    • BF16 optimizer settings now stay on the fused path instead of falling back to the slower eager path.
    • Better handling of fused optimizer configuration for supported training setups.
  • Tests

    • Added integration coverage for fused optimizer selection and BF16-related training paths.
  • Chores

    • Updated AiDotNet tensor and native backend package versions.

…ps ≥50M models on the fused path)

ShouldUseBFloat16Optimizer engaged BF16 moment storage proactively for ANY model with
≥50M parameters. But the BF16/8-bit Adam (Adam8BitOptimizer) is NOT fused-kernel-
compatible — TryMapToFusedOptimizerConfig only accepts plain Adam/AdamW/SGD — so
selecting it silently drops the ENTIRE model off the compiled fused-training fast path
onto the eager autograd tape, ~10x slower per step. Net effect: every ≥50M-param model
was training on the slow path even when it fit in memory with room to spare.

Measured on ViT-Base (86.5M): the proactive BF16 forced the eager tape at ~5.0 s/step;
gating it off (the model fits) keeps it on the fused path at ~3.4 s/step (~1.5x), with no
change to small (<50M) models.

Fix: only engage proactive BF16 when the fp32 moment state would actually consume a
meaningful fraction of available memory (> 25%), mirroring the fits-in-memory guard used
for weight streaming. Models that genuinely don't fit still get BF16. The reactive memory
ladder (_memoryLeversForced, set on an actual OOM) is untouched and still engages
BF16/8-bit on demand. AIDOTNET_BF16_ADAM=1/0 still force/disable explicitly.

This is the broad half of the "memory-lever optimizers break fused training" finding in #1743.

Refs #1743, #1706.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
@vercel

vercel Bot commented Jun 29, 2026 •

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

2 Skipped Deployments
Project Deployment Actions Updated (UTC)
aidotnet_website Ignored Ignored Preview Jul 2, 2026 1:07am
aidotnet-playground-api Ignored Ignored Preview Jul 2, 2026 1:07am

@coderabbitai

coderabbitai Bot commented Jun 29, 2026 •

Copy link
Copy Markdown
Contributor

Review Change Stack

Warning

Review limit reached

You’ve reached a temporary PR review limit under our Fair Usage Limits Policy.

Your recent review volume is higher than typical usage, so adaptive limits are currently applied.

Next review available in: 52 minutes

Enable usage-based reviews in Billing to review now. Otherwise, wait until the next included review is available.
You're only billed for reviews past your plan's rate limits ($0.25/file).

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro

Run ID: 3e66eeba-84ed-4e0c-8cc4-03075e692c94

📥 Commits

Reviewing files that changed from the base of the PR and between 687d824 and ca388b4.

📒 Files selected for processing (4)
  • Directory.Packages.props
  • src/Optimizers/Adam8BitOptimizer.cs
  • src/Optimizers/Fused/IFusedOptimizerSpec.cs
  • src/Training/CompiledTapeTrainingStep.cs

Walkthrough

Adds a UseBf16Moments flag to FusedOptimizerConfig, implemented by Adam8BitOptimizer via IFusedOptimizerSpec, threaded through NeuralNetworkBase's fused-config mapping into CompiledTapeTrainingStep.TryStepWithFusedOptimizer, which requests BF16 moment storage on plan initialization. Includes new integration tests and a package version bump.

Changes

BF16 moment storage for fused Adam8Bit optimizer

Layer / File(s) Summary
FusedOptimizerConfig contract
src/Optimizers/Fused/IFusedOptimizerSpec.cs
Adds a UseBf16Moments boolean field (default false) with documentation describing BF16 moment buffer storage.
Adam8BitOptimizer fused spec
src/Optimizers/Adam8BitOptimizer.cs
Implements IFusedOptimizerSpec.TryGetFusedOptimizerConfig, returning fused Adam config with UseBf16Moments: true only when BF16 moment storage is on, adaptive LR/AMSGrad are off, and a valid LR schedule exists.
NeuralNetworkBase fused-config mapping and wiring
src/NeuralNetworks/NeuralNetworkBase.cs
Extends TryMapToFusedOptimizerConfig with a useBf16Moments out-parameter, threads it through TryTrainWithFusedOptimizer into TryStepWithFusedOptimizer, and updates comments in ShouldUseBFloat16Optimizer.
CompiledTapeTrainingStep BF16 allocation
src/Training/CompiledTapeTrainingStep.cs
TryStepWithFusedOptimizer gains a useBf16Moments parameter and calls plan.RequestBf16MomentStorage(true) during first-time plan initialization when enabled.
Tests and package bump
tests/AiDotNet.Tests/IntegrationTests/Optimizers/Adam8BitFusedSpecTests.cs, Directory.Packages.props
Adds integration tests validating fused-config mapping across BF16/int8/AMSGrad/adaptive-LR combinations, and bumps AiDotNet.Tensors/native backend packages from 0.105.0 to 0.106.0.

Estimated code review effort: 3 (Moderate) | ~25 minutes

Sequence Diagram(s)

sequenceDiagram
  participant Adam8BitOptimizer
  participant NeuralNetworkBase
  participant CompiledTapeTrainingStep
  participant FusedPlan

  Adam8BitOptimizer->>NeuralNetworkBase: TryGetFusedOptimizerConfig (UseBf16Moments=true)
  NeuralNetworkBase->>NeuralNetworkBase: TryMapToFusedOptimizerConfig extracts useBf16Moments
  NeuralNetworkBase->>CompiledTapeTrainingStep: TryStepWithFusedOptimizer(useBf16Moments)
  CompiledTapeTrainingStep->>FusedPlan: RequestBf16MomentStorage(true)
  FusedPlan-->>CompiledTapeTrainingStep: plan configured with BF16 moment buffers
Loading

Possibly related PRs

  • ooples/AiDotNet#1653: Both PRs hook into the fused optimizer configuration path, extending FusedOptimizerConfig/IFusedOptimizerSpec for a new optimizer's fused kernel selection.
  • ooples/AiDotNet#1635: Both PRs modify CompiledTapeTrainingStep.TryStepWithFusedOptimizer's fused compile-time path with additional configuration flags.
  • ooples/AiDotNet#1641: Both PRs modify TryTrainWithFusedOptimizer in NeuralNetworkBase.cs at the same method level.

Poem

A rabbit hops through bits of eight,
Trims moments down to lighten weight,
BF16 whispers through the plan,
"Fuse me fast, if optimizer can!"
Tests confirm the flag holds true —
🐇 0.106.0, hop through! 🥕

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title clearly reflects the main change: BF16 Adam training now stays on the fused path via proper fused BF16 moment support.
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch perf/bf16-adam-fused-compat-memory-gate

Comment @coderabbitai help to get the list of available commands.

ooples and others added 2 commits June 29, 2026 21:56
…te (#1745)

Replaces the interim "gate proactive BF16 on memory pressure" workaround with
the real fix: BF16-Adam now keeps the fused fast path instead of dropping to the
eager autograd tape, so large models get BOTH the fused speed AND the halved
optimizer-state footprint — no tradeoff.

Pairs with AiDotNet.Tensors PR #713 (fused bf16 moment kernel +
ICompiledTrainingPlan.RequestBf16MomentStorage):
- Adam8BitOptimizer implements IFusedOptimizerSpec: in BFloat16 moment-storage
  mode it maps to the fused Adam kernel with UseBf16Moments=true. The true 8-bit
  block-quant mode (and adaptive-LR / AMSGrad) still has no fused kernel and
  correctly falls back to eager.
- FusedOptimizerConfig carries UseBf16Moments; TryMapToFusedOptimizerConfig
  surfaces it; CompiledTapeTrainingStep calls plan.RequestBf16MomentStorage
  before ConfigureOptimizer so the plan allocates half-size m/v buffers.
- ShouldUseBFloat16Optimizer reverts to a plain size threshold — the memory
  gate existed only to avoid losing the fused path, which no longer happens.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
@ooples ooples changed the title perf(training): gate proactive BF16-Adam on memory pressure — keeps ≥50M models on the fused path perf(training): make BF16-Adam fused-compatible (proper bf16 moment kernel, not a gate) Jun 30, 2026
ooples added 2 commits July 1, 2026 16:17
0.106.0 ships the fused bf16 adam moment kernel (tensors #713:
adamupdatebf16simd / requestbf16momentstorage) this pr wires into
compiledtapetrainingstep, plus the weightregistry dead-owner sweep
(tensors #716).
Copilot AI review requested due to automatic review settings July 1, 2026 20:22

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (1)
src/Training/CompiledTapeTrainingStep.cs (1)

775-775: 🗄️ Data Integrity & Integration | 🟡 Minor | ⚡ Quick win

Drift-detection tuple omits useBf16Moments.

_configuredOptimizerConfig tracks type/LR/beta/eps/weight-decay for the single-plan drift check, but not useBf16Moments. If the optimizer's BF16-moment flag changes between steps on the same configured plan (e.g. via Adam8BitOptimizer.UpdateOptions toggling UseBFloat16MomentStorage mid-run without a shape change), the drift check won't catch it and training would silently continue with the plan's original moment-buffer layout.

🔧 Proposed fix to include the flag in drift detection
-    private static (int OptType, float Lr, float B1, float B2, float Eps, float Wd)? _configuredOptimizerConfig;
+    private static (int OptType, float Lr, float B1, float B2, float Eps, float Wd, bool UseBf16Moments)? _configuredOptimizerConfig;
-            var currentConfig = ((int)optimizerType, learningRate, beta1, beta2, epsilon, weightDecay);
+            var currentConfig = ((int)optimizerType, learningRate, beta1, beta2, epsilon, weightDecay, useBf16Moments);

Also applies to: 839-847

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@src/Training/CompiledTapeTrainingStep.cs` at line 775, The drift-detection
tuple in `CompiledTapeTrainingStep` is missing `useBf16Moments`, so changes to
the optimizer’s BF16 moment storage can slip past the single-plan config check.
Update the optimizer config snapshot used by the drift check (including the
tuple around `currentConfig` and the matching `_configuredOptimizerConfig`
comparison logic in the same training step flow) to include `useBf16Moments`, so
any toggle in `Adam8BitOptimizer.UpdateOptions` or similar is treated as a
config drift.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@Directory.Packages.props`:
- Around line 205-208: Add a changelog-style comment for the AiDotNet.Tensors
0.106.0 version bump in Directory.Packages.props, matching the existing
convention used by other PackageVersion entries. Place the note alongside the
AiDotNet.Tensors, AiDotNet.Native.OneDNN, AiDotNet.Native.OpenBLAS, and
AiDotNet.Native.CLBlast updates, and reference the upstream driver such as
AiDotNet.Tensors#713 so the rationale is preserved for future readers.

---

Outside diff comments:
In `@src/Training/CompiledTapeTrainingStep.cs`:
- Line 775: The drift-detection tuple in `CompiledTapeTrainingStep` is missing
`useBf16Moments`, so changes to the optimizer’s BF16 moment storage can slip
past the single-plan config check. Update the optimizer config snapshot used by
the drift check (including the tuple around `currentConfig` and the matching
`_configuredOptimizerConfig` comparison logic in the same training step flow) to
include `useBf16Moments`, so any toggle in `Adam8BitOptimizer.UpdateOptions` or
similar is treated as a config drift.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro

Run ID: 67292955-7d62-4baf-b8fe-d457dd1ae9e4

📥 Commits

Reviewing files that changed from the base of the PR and between e50bd03 and 687d824.

📒 Files selected for processing (6)
  • Directory.Packages.props
  • src/NeuralNetworks/NeuralNetworkBase.cs
  • src/Optimizers/Adam8BitOptimizer.cs
  • src/Optimizers/Fused/IFusedOptimizerSpec.cs
  • src/Training/CompiledTapeTrainingStep.cs
  • tests/AiDotNet.Tests/IntegrationTests/Optimizers/Adam8BitFusedSpecTests.cs

Comment thread Directory.Packages.props Outdated

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Note

Copilot couldn't run its full agentic review because no GitHub Actions runner was available. Make sure your repository has a runner available to run Copilot's review, or add a copilot-setup-steps.yml file specifying one with the runs-on attribute. See the docs for more details.

This PR restores BF16 moment storage for Adam8BitOptimizer without sacrificing the compiled fused-training fast path by plumbing a UseBf16Moments signal through the fused optimizer mapping and compiled training plan.

Changes:

  • Add UseBf16Moments to FusedOptimizerConfig and propagate it through fused optimizer mapping into CompiledTapeTrainingStep.
  • Implement IFusedOptimizerSpec on Adam8BitOptimizer so BF16 moment mode maps to fused Adam with bf16 moment buffers (while block-quant / AMSGrad / adaptive LR correctly fall back).
  • Add integration tests covering the optimizer→fused-config mapping; bump AiDotNet.Tensors + native packages.

Reviewed changes

Copilot reviewed 6 out of 6 changed files in this pull request and generated 3 comments.

Show a summary per file
File Description
tests/AiDotNet.Tests/IntegrationTests/Optimizers/Adam8BitFusedSpecTests.cs Adds assertions for BF16 mapping to fused Adam and correct fallbacks for unsupported modes.
src/Training/CompiledTapeTrainingStep.cs Requests bf16 moment storage on the compiled plan prior to optimizer configuration when requested.
src/Optimizers/Fused/IFusedOptimizerSpec.cs Extends fused optimizer config with a UseBf16Moments flag (default false).
src/Optimizers/Adam8BitOptimizer.cs Implements fused optimizer spec for BF16 moment mode and wires UseBf16Moments=true.
src/NeuralNetworks/NeuralNetworkBase.cs Plumbs UseBf16Moments from optimizer mapping into the fused training step; updates BF16 threshold rationale comment.
Directory.Packages.props Updates AiDotNet.Tensors + native package versions to 0.106.0.

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment thread src/Training/CompiledTapeTrainingStep.cs Outdated
Comment thread src/Optimizers/Fused/IFusedOptimizerSpec.cs Outdated
Comment thread Directory.Packages.props Outdated
- FusedOptimizerConfig: move UseBf16Moments from the primary constructor to an
  init-only property so Deconstruct arity and positional construction sites are
  unchanged (only Adam8Bit sets it, now via object initializer); still part of
  record value equality.
- TryStepWithFusedOptimizer: append useBf16Moments after eagerOptimizer instead
  of inserting it before, so positional call sites aren't shifted (sole caller
  uses named args).
- Directory.Packages.props: document the 0.104.6 -> 0.106.0 bump (Tensors #713
  fused bf16 moment kernel) per the file's changelog convention; note 0.106.0 is
  already published so CI isn't gated on an unreleased dependency.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Copilot AI review requested due to automatic review settings July 2, 2026 00:01

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

@ooples
ooples merged commit fbe7fe3 into master Jul 2, 2026
71 of 85 checks passed
@ooples
ooples deleted the perf/bf16-adam-fused-compat-memory-gate branch July 2, 2026 11:51
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants