Skip to content

fix(#1405): moe default optimizer overshoots — use amsgrad adam(1e-4) - #1409

Merged
ooples merged 9 commits into
masterfrom
fix/issue-1405-moe-overshoot
May 20, 2026
Merged

ooples merged 9 commits into
masterfrom
fix/issue-1405-moe-overshoot

Conversation

@ooples

@ooples ooples commented May 20, 2026 •

Copy link
Copy Markdown
Owner

Summary

MixtureOfExpertsNeuralNetworkTests.MoreData_ShouldNotDegrade was failing because the network's built-in default optimizer (vanilla AdamOptimizer<T,Tensor,Tensor>(this)) diverges in the long-training regime: short training fits, but loss climbs again as training proceeds.

This is the same signature as the DenseNet divergence fixed in #1393, and the same recipe applies.

Root cause

Two reinforcing problems with the previous default:

  1. LR too aggressive for MoE gradient paths. In MoE, the gating network plus each expert all touch shared parameters on every step, so the effective per-parameter update is larger than for a typical feed-forward net. Vanilla Adam's default LR makes step sizes grow late in training.
  2. Vanilla Adam's v̂ denominator drifts late in training (Reddi et al. 2018, On the Convergence of Adam and Beyond). AMSGrad fixes this by keeping the running max of v instead of the EMA, which prevents the denominator from shrinking and the step from blowing up.

Fix

Switch the built-in default in MixtureOfExpertsNeuralNetwork to AMSGrad-mode Adam at lr = 1e-4:

_optimizer = optimizer ?? new AdamOptimizer<T, Tensor<T>, Tensor<T>>(
    this,
    new AdamOptimizerOptions<T, Tensor<T>, Tensor<T>>
    {
        InitialLearningRate = 1e-4,
        UseAMSGrad = true
    });

Users who pass an explicit optimizer are unaffected.

Test plan

  • MixtureOfExpertsNeuralNetworkTests.MoreData_ShouldNotDegrade passes on net10.0
  • Release build clean (0 errors)

Closes #1405

🤖 Generated with Claude Code

Summary by CodeRabbit

Release Notes

  • Bug Fixes
    • Enhanced default optimizer configuration for Mixture of Experts Neural Network with optimized learning rate and AMSGrad mode to improve training stability and convergence.

Review Change Stack

franklinic and others added 9 commits May 19, 2026 12:19
…v deserialize fallback

PR #1290 CI Cluster 1: 25 of 25 DCGANTests failing post-master with one
of two errors:

  1. Most (23 tests): "Invalid layer configuration: The last layer's
     output shape [3, -1, -1] must match the architecture output size
     (12288)."
  2. Clone tests (2): "Input spatial dims after padding (1+2*1, 1+2*1)
     must be >= kernelSize (4)" raised inside DeserializationHelper's
     pre-resolve of the discriminator's first conv layer.

Plus 1 SparseNN test (intermittent mode-collapse) that re-runs pass
without code change — flaky, not a regression target.

## Root causes

(1) NeuralNetworkBase.IsLastLayerShapeCompatible: PR #1329 (commit
969977d) added a `outputShape.Any(d => d < 0)` early-return so the
validator defers the flat-OutputSize check when any output-shape dim
is deferred — DCGAN's last transposed-conv emits [3, -1, -1] until
its first Forward resolves H/W. That guard was inadvertently deleted
by the grafprint PR (c8cac23, May 16) one day later. Restoring it
unblocks all 23 validator-rejection cases at once.

(2) DeserializationHelper conv path: when the saved layer record's
inputShape carries -1 sentinels (a lazy conv layer serialized before
its first Forward — DCGAN's discriminator on a Predict-only probe
sees only the generator), the pre-existing code coerced all -1 dims
to 1 and called conv.ResolveShapesOnly(...). For DCGAN's first conv
(kernel=4, padding=1) this fails OnFirstForward's kernel-size check
(1 + 2 < 4). Coercing to Math.Max(1, KernelSize) fixes that
specific check, but locks InputDepth at 1 — then the real Forward
with the [3, 64, 64] RGB image throws "Expected input depth 1, but
got 3". The correct fix is to skip pre-resolve entirely when
InputDepth is deferred — ConvolutionalLayer.SetParameters has its
own auto-resolve fallback at line ~1598 that derives InputDepth from
the saved parameter vector's length, and uses KernelSize as the
spatial placeholder. Pre-resolve still runs (and uses
Math.Max(1, KernelSize) for any deferred spatial dim) when
InputDepth is concrete — that's the original PR #1329 contract for
the auto-resolve-disambiguation case.

## Verification

  $ dotnet test --filter "FullyQualifiedName~DCGANTests|FullyQualifiedName~SparseNeuralNetworkTests" --framework net10.0
  Failed!  - Failed: 2, Passed: 44, Skipped: 0, Total: 46

26 → 2 failures. The remaining two are NOT cluster-1 shape-contract
issues:

  - DCGANTests.MoreData_ShouldNotDegrade — `Test execution timed
    out after 120000 milliseconds`. Pre-existing GAN training-path
    perf gap; the deep deconv+conv chain in tape mode is ~5-10×
    slower than PyTorch CPU baseline. Substep profile (Release):
    Generator.Predict 19 ms, Discriminator.Train 187 ms, Generator
    adversarial 313 ms — 519 ms/step × 250 iters = 130 s vs 120 s
    timeout. Filed separately so this PR ships the actual
    cluster-1 root causes (validator + conv-deserialize) without
    bundling a multi-week perf project.

  - SparseNeuralNetworkTests.DifferentInputs_AfterTraining_ShouldProduceDifferentOutputs
    — intermittent mode-collapse, passes on re-runs. Separate
    flaky-test issue, not a shape-contract bug.

Closes #1309 partially (cluster-1 shape-contract root causes).
The MoreData_ShouldNotDegrade timeout + SparseNN mode-collapse
flakiness are tracked separately.

🤖 Generated with [Claude Code](https://claude.com/claude-code)
…ses DCGAN MoreData timeout

Previously GenerativeAdversarialNetwork.Train ran the generator forward TWICE
per training step:
  1. Generator.Predict(input)  (eval mode, NoGradScope) → detached fake images
     for the combined real+fake discriminator step.
  2. ForwardForTraining(input) (train mode, on tape) inside
     TrainWithCustomLoss — duplicate of the same forward, just for the
     gen-adversarial backward.

On the DCGAN MoreData fixture (250 iters, double-precision, batch=2, 64×64
RGB) this duplicate forward contributed ~19 ms of the 519 ms / step
profiled in #1390 — pushing the test 10 s over its 120 s budget.

Refactor:
  - Open a single GradientTape at the start of the step.
  - Run ForwardForTraining(input) ONCE on that tape → fakeTapeTracked.
  - Take a value-copy detached snapshot (fakeImages) for the disc step;
    fresh Tensor<T> with no GradNode chain so disc.Train (which opens
    its own nested tape) can not leak gradients back into the generator.
  - Walk the discriminator layer-by-layer on the existing gen tape for
    the adversarial loss (unchanged from the prior closure semantics).
  - Drive the gen optimizer step via the new
    NeuralNetworkBase.BackwardAndStepOnPrecomputedLoss helper, which
    reuses the open tape instead of TrainWithCustomLoss opening a fresh
    one + re-running ForwardForTraining.

Behavior note: the disc step now sees train-mode generator output
(batch BN stats) instead of eval-mode (running BN stats). This matches
PyTorch's standard DCGAN training pattern (fake = G(z); fake_detached =
fake.detach()) and the existing gen step's own train-mode forward.
DCGAN has no Dropout, so the only distribution shift is BN stats, which
is the conventional adversarial behavior.

Verified locally with the canonical Tensors 0.81.3 dependency:
  - DCGANTests.MoreData_ShouldNotDegrade: 1 m 47 s (was timing out at
    > 120 s) — closes the test's perf gap.
  - Full DCGANTests class: 25 / 25 passing.
  - ConditionalGANTests + InfoGANTests (other GAN.Train consumers):
    50 / 50 passing.
  - Full SparseNeuralNetworkTests: 21 / 21 passing (previously
    "intermittent mode-collapse" in PR #1389 description — appears
    stable now, may have been transient).

Closes #1390.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
addresses three coderabbit comments on backwardandsteponprecomputed
loss in pr #1389:

1. visibility narrowed public -> internal. the codebase contract is
   "users should only interact with aimodelbuilder / aimodelresult"
   and this helper is training plumbing for in-assembly callers
   (currently generativeadversarialnetwork.train); no reason for it
   to live on the public surface. only caller is in same assembly.

2. added using var __reentrancyguard = acquiretrainsentinel() at the
   top, mirroring trainwithtape's sentinel discipline. without it,
   concurrent callers on the same model race on lastloss + optimizer
   internal state.

3. trainableparams now concats getextratrainabletensors() with the
   layer params, matching trainwithtape's parameter set. without this
   models that expose raw tensors via getextratrainabletensors (rather
   than layer-resident params) silently skipped updates on the
   precomputed-loss path -- divergent semantics between the two
   training entry points.

build passes.
vanilla adam default at the recommended learning rate diverges on the
moredata long-training regression for mixtureofexpertsneuralnetwork
(short=ok, long=overshoot) — same signature as the densenet fix in
#1393.

two root causes:
  1. lr too aggressive for the moe gating + per-expert summed-gradient
     path, where multiple experts touch each parameter per step.
  2. vanilla adam's v denominator drifts late in training, so step size
     grows even after loss has bottomed out.

switching the built-in default to amsgrad-mode adam at 1e-4 keeps the
running max of v (reddi et al. 2018), which removes the late-training
denominator decay and matches the recipe already proven on densenet.
users passing an explicit optimizer are unaffected.

Closes #1405
Copilot AI review requested due to automatic review settings May 20, 2026 18:48
@vercel

vercel Bot commented May 20, 2026 •

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated (UTC)
aidotnet_website Ready Ready Preview, Comment May 20, 2026 6:48pm
aidotnet-playground-api Ready Ready Preview, Comment May 20, 2026 6:48pm

@coderabbitai

coderabbitai Bot commented May 20, 2026 •

Copy link
Copy Markdown
Contributor

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro

Run ID: d58217a1-1db5-413a-bae3-8db84caaafe9

📥 Commits

Reviewing files that changed from the base of the PR and between 330380e and 6bdcd34.

📒 Files selected for processing (1)
  • src/NeuralNetworks/MixtureOfExpertsNeuralNetwork.cs

Walkthrough

MixtureOfExpertsNeuralNetwork<T> adds using AiDotNet.Models.Options; and updates the default optimizer initialization in its parameterized constructor to instantiate AdamOptimizer with explicit AdamOptimizerOptions configuration: InitialLearningRate = 1e-4 and UseAMSGrad = true. Existing optimizer overrides remain unchanged.

Changes

Default AMSGrad optimizer configuration

Layer / File(s) Summary
Default AMSGrad optimizer configuration
src/NeuralNetworks/MixtureOfExpertsNeuralNetwork.cs
Adds using AiDotNet.Models.Options; import and replaces parameterless AdamOptimizer default with explicit AdamOptimizerOptions configuration (learning rate 1e-4, AMSGrad enabled) in the constructor's _optimizer initialization.

Estimated code review effort

🎯 2 (Simple) | ⏱️ ~10 minutes

Possibly related PRs

  • ooples/AiDotNet#1403: Applies identical AMSGrad optimizer default configuration (learning rate 1e-4, UseAMSGrad = true) to DenseNetNetwork constructor, following the same pattern as this MixtureOfExpertsNeuralNetwork change.

Poem

🧠 The mixture finds its balance true,
With AMSGrad's steadying hand to guide,
A learning rate of 1e-4 sets the view—
No more divergence, just adaptive stride. ✨

🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 0.00% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title accurately reflects the main change: switching MoE default optimizer to AMSGrad Adam with 1e-4 learning rate to fix overshooting behavior.
Linked Issues check ✅ Passed The PR fully addresses issue #1405 by implementing the exact fix specified: changing default optimizer to AMSGrad Adam (lr=1e-4) to prevent loss divergence in late training, directly resolving the failing MoreData_ShouldNotDegrade test.
Out of Scope Changes check ✅ Passed All changes are strictly scoped to the MixtureOfExpertsNeuralNetwork default optimizer configuration; no unrelated modifications or ancillary changes are present.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch fix/issue-1405-moe-overshoot

Comment @coderabbitai help to get the list of available commands and usage tips.

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot encountered an error and was unable to review this pull request. You can try again by re-requesting a review.

@ooples
ooples merged commit 2ce0b30 into master May 20, 2026
37 of 50 checks passed
@ooples
ooples deleted the fix/issue-1405-moe-overshoot branch May 20, 2026 22:37

This branch was successfully deployed

2 active deployments
Preview – aidotnet_website — 6bdcd349 Deployed May 20, 2026 by vercel[bot]
Preview – aidotnet-playground-api — 6bdcd349 Deployed May 20, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[CI] Tests (net10.0) - Unit - 08d NN-Adapters/Other: MoE MoreData failing

3 participants