Skip to content

chore(deps): bump AiDotNet.Tensors 0.92.1 -> 0.92.5 (closes #1221) - #1563

Merged
ooples merged 1 commit into
masterfrom
chore/bump-tensors-0.92.5
Jun 11, 2026
Merged

ooples merged 1 commit into
masterfrom
chore/bump-tensors-0.92.5

Conversation

@ooples

@ooples ooples commented Jun 11, 2026 •

Copy link
Copy Markdown
Owner

Bumps AiDotNet.Tensors (and the lockstep AiDotNet.Native.OneDNN/OpenBLAS/CLBlast) from 0.92.1 → 0.92.5.

Why

0.92.5 is the first release carrying Tensors #583 — validate the prepacked-B GEMM cache against in-place weight mutation. The SgemmWithCachedB inference path reused a stale pre-packed B panel when a layer's weights were mutated in place without a MarkWeightDirty, so FusedLinear inference returned pre-update results. That was the root cause of #1221 (transformer production-scale convergence failure).

Verification

  • dotnet restore clean against 0.92.5 (no NU1102); all four lockstep packages confirmed published at 0.92.5.
  • Build clean (net10.0, Release).
  • Issue1221 suite: 11/11 pass (TransformerProductionScaleConvergenceIssue1221Tests + TransformerProductionScaleHarnessIssue1221Tests).

Closes #1221.

🤖 Generated with Claude Code

Summary by CodeRabbit

  • Chores
    • Updated package dependencies to latest compatible versions for improved stability and performance.

0.92.5 is the first release carrying Tensors #583 (validate the prepacked-B GEMM
cache against in-place weight mutation). The SgemmWithCachedB inference path reused
a stale pre-packed B panel when a layer's weights were mutated in place without a
MarkWeightDirty, so FusedLinear inference returned pre-update results — the root
cause of the #1221 transformer production-scale convergence failure. Bumps the
AiDotNet.Native.* packages (OneDNN/OpenBLAS/CLBlast) in lockstep, all published at
0.92.5. Restores clean (no NU1102); the 11 Issue1221 convergence/harness tests pass.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
@vercel

vercel Bot commented Jun 11, 2026 •

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated (UTC)
aidotnet_website Ready Ready Preview, Comment Jun 11, 2026 12:11am
aidotnet-playground-api Ready Ready Preview, Comment Jun 11, 2026 12:11am

@coderabbitai

coderabbitai Bot commented Jun 11, 2026 •

Copy link
Copy Markdown
Contributor

Review Change Stack

Walkthrough

Centralized package version declarations for the AiDotNet tensor and native backend ecosystem are bumped from 0.92.1 to 0.92.5. The inline XML documentation comment describing the version delta is updated to align with the new version.

Changes

AiDotNet Tensor and Native Dependencies

Layer / File(s) Summary
AiDotNet.Tensors and native packages version bump
Directory.Packages.props
AiDotNet.Tensors, AiDotNet.Native.OneDNN, AiDotNet.Native.OpenBLAS, and AiDotNet.Native.CLBlast updated from 0.92.1 to 0.92.5, with version comment documentation updated to match.

Estimated Code Review Effort

🎯 1 (Trivial) | ⏱️ ~3 minutes

Possibly Related PRs

  • ooples/AiDotNet#1519: Both PRs update Directory.Packages.props to bump AiDotNet.Tensors and related native packages to newer versions.
  • ooples/AiDotNet#1255: Both PRs modify the same centralized props file to update AiDotNet.Tensors and native backend package versions (OneDNN, OpenBLAS).
  • ooples/AiDotNet#955: Both PRs directly update Directory.Packages.props version entries for AiDotNet.Tensors and its native dependencies.

Poem

🚀 From 0.92.1 to 0.92.5 we climb,
Tensors and natives in perfect time,
Props aligned, dependencies flow,
The ecosystem's ready to grow! 📊

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed Title accurately summarizes the main change: bumping AiDotNet.Tensors from 0.92.1 to 0.92.5 and references the closed issue.
Linked Issues check ✅ Passed The PR successfully addresses issue #1221 by updating to AiDotNet.Tensors 0.92.5, which includes the critical GEMM cache validation fix (Tensors issue #583) that resolves the transformer uniform output bug.
Out of Scope Changes check ✅ Passed All changes are narrowly scoped to version bumps in Directory.Packages.props for AiDotNet.Tensors and lockstep native packages, directly supporting the bug fix objective.
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch chore/bump-tensors-0.92.5

Comment @coderabbitai help to get the list of available commands and usage tips.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@Directory.Packages.props`:
- Around line 142-145: The PR currently pins four AiDotNet package versions
(AiDotNet.Tensors, AiDotNet.Native.OneDNN, AiDotNet.Native.OpenBLAS,
AiDotNet.Native.CLBlast) to 0.92.5 in Directory.Packages.props and the reviewer
confirmed they exist on nuget.org; to prevent future mismatches add a
CI/pre-merge step that validates each package/version from
Directory.Packages.props is available on nuget.org (or the configured feed) and
fails fast if any PackageVersion (e.g., the AiDotNet.Tensors or
AiDotNet.Native.* entries) is not published so NU1102 regressions are caught
before merge.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro

Run ID: 41949f1b-c548-4bdc-a648-73d9a34f3f2d

📥 Commits

Reviewing files that changed from the base of the PR and between 51de880 and 11127b5.

📒 Files selected for processing (1)
  • Directory.Packages.props

Comment thread Directory.Packages.props
@ooples
ooples merged commit 811ef20 into master Jun 11, 2026
41 of 63 checks passed
@ooples
ooples deleted the chore/bump-tensors-0.92.5 branch June 11, 2026 13:16
ooples added a commit that referenced this pull request Jun 12, 2026
…1563)

0.92.5 is the first release carrying Tensors #583 (validate the prepacked-B GEMM
cache against in-place weight mutation). The SgemmWithCachedB inference path reused
a stale pre-packed B panel when a layer's weights were mutated in place without a
MarkWeightDirty, so FusedLinear inference returned pre-update results — the root
cause of the #1221 transformer production-scale convergence failure. Bumps the
AiDotNet.Native.* packages (OneDNN/OpenBLAS/CLBlast) in lockstep, all published at
0.92.5. Restores clean (no NU1102); the 11 Issue1221 convergence/harness tests pass.

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
ooples added a commit that referenced this pull request Jun 13, 2026
…ors, continual-learning + batchnorm tests (#1565)

* fix(diffusion): resolve lazy shapes in clone roundtrip for video/editing models

The lazy-init clone bug surfaced in ModelFamily Clone_ShouldProduceIdenticalOutput
for RealESRGAN, SeedEdit3, and StableVideoDiffusion: the predictor/VAE time-
embedding MLPs (and the VideoUNet image-condition projection) report
ParameterCount=0 until their first forward, so reconstructing a fresh model and
calling SetParameters(GetParameters()) copied an under-counted snapshot and let
the clone re-resolve with fresh random weights on its first Predict.

- RealESRGAN/SeedEdit3 Clone now delegate to the predictor/VAE's own Clone, which
  trigger lazy shape resolution on both source and clone before copying weights.
- VideoUNetPredictor.Clone gains TriggerLazyShapeResolution (mirroring
  UNetNoisePredictor), and it also resolves the image-condition projection since
  image-to-video models (SVD) always forward through that path. TemporalVAE has
  no lazy layers, so StableVideoDiffusion's existing delegating Clone is now
  correct end-to-end.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore: ignore local .ci-failures tracking dir

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(diffusion): round-trip VideoUNet FiLM projection + image-mode lazy trigger

StableVideoDiffusion's Clone_ShouldProduceIdenticalOutput still diverged after
the initial lazy-resolution fix because VideoUNetPredictor's parameter walk
omitted the per-block FiLM TimeCondProjection (the AdaGN timestep-conditioning
DenseLayer applied in ApplyVideoBlock). It was excluded from AddBlockParameters,
SetBlockParameters, AND AddBlockGradients, so its weights never round-tripped
(the clone kept fresh-random FiLM weights) and never trained.

- Add TimeCondProjection to all three block walks, in the same position
  (right after SpatialResBlock, matching the forward order).
- TriggerLazyShapeResolution now mirrors the real Predict path exactly: a tiny
  4D image-mode forward with no conditioning (the denoising loop never runs the
  5D video / image-condition / cross-attention paths, and the temporal-mixing
  DenseLayer is sized strictly [_numFrames -> _numFrames]). Lazy weights are
  spatial-independent, so a minimal dummy resolves identical shapes on source
  and clone cheaply, keeping the GetParameters/SetParameters copy index-aligned.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(rl): correct WorldModelsAgent network sizing, deterministic clone, pure predict

Three bugs surfaced by the Generated Layers N-Z shard
(WorldModelsAgentTests.ActionSelection_ShouldBeFinite / Clone_ShouldProduceSamePolicy):

1. Malformed, multi-GB networks. Each VAE/RNN network was built from an
   architecture with NO explicit layers, so NeuralNetwork.InitializeLayers
   auto-generated default hidden layers (~2x the flattened input) and the
   agent's AddLayer calls then stacked on top. For a 64x64x3 = 12,288-wide
   observation that default layer alone is a single 12,288x24,576 weight
   (~2.4 GB at FP64), which tripped weight-streaming's per-tensor byte cap and
   threw on the first forward. Now each network is built with its EXACT layer
   list passed to the architecture, so no defaults are generated.

2. Non-deterministic clone. NeuralNetwork.InitializeLayers does not wire
   explicitly-supplied custom layers, so their lazy weights initialized from
   the shared non-deterministic RNG — a clone re-initialized a different
   policy. Each layer now gets a deterministic per-layer RandomSeed, and Clone
   reproduces the networks via the constructor (they are never trained; only
   the controller is, via the evolution strategy) and copies the trained
   controller weights + rollout state directly. A serialization round-trip is
   NOT used for cloning because it rebuilds the layers and drops the seed pins.

3. Non-pure Predict. SelectAction advances the RNN hidden state for sequential
   rollout, making repeated Predict calls non-deterministic. Predict now
   snapshots and restores the hidden state so inference is side-effect-free.

All 7 WorldModelsAgentTests pass.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(physics): correct HamiltonianNeuralNetwork parameterless ctor input size

The parameterless constructor built its architecture with
inputSize = DefaultStateDim * 2 (= 4) while passing stateDim = DefaultStateDim
(= 2). The constructor enforces CalculatedInputSize == stateDim (the network
input IS the full phase-space state vector (q, p), which it splits into q/p
halves), so every `new HamiltonianNeuralNetwork<T>()` threw
"Hamiltonian network input size (4) must match state dimension (2)".

The source-generated ModelFamily test scaffold builds the model via the
parameterless ctor, so the throw failed ALL 21 HamiltonianNeuralNetworkTests
(Architecture/Clone/ForwardPass/Training/Metadata/Parameters/...).

Fix: set inputSize = DefaultStateDim so input size matches stateDim. DefaultStateDim
= 2 is the canonical 1-DOF Hamiltonian phase space (one position + one momentum).
All 21 tests pass.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(continual-learning): declare IParameterizable on EWC/SI test mock model

The interface-segregation refactor moved GetParameters / SetParameters /
ParameterCount / WithParameters off IFullModel and onto the separate
IParameterizable interface. EWC and Synaptic Intelligence guard-require it via
InterfaceGuard.Parameterizable (correctly — EWC's Fisher information matrix and
SI's per-parameter importance are both computed OVER the model's parameters, so
a model without trainable parameters cannot participate).

The CLMockModel test fixture still declared only IFullModel, so the runtime
`model is IParameterizable<...>` check failed and threw
"CLMockModel does not implement IParameterizable<...>" from PrepareForTask —
failing all 14 EWC/SI deep-math integration tests. The mock already implements
every IParameterizable member (GetParameters, SetParameters, ParameterCount,
SupportsParameterInitialization, WithParameters, SanitizeParameters); it just
needed to declare the interface so the cast succeeds.

No assertion or expected value changed. All 36 ContinualLearningDeepMath tests pass.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(batchnorm): use eager ctor for init-contract assertions in deep-math tests

Eight BatchNorm deep-math tests constructed the layer via the parameterless ctor
`new BatchNormalizationLayer<double>()`, which is intentionally LAZY — gamma /
beta / running stats are sized on the first forward (a separate eager ctor
`(numFeatures)` was added in #1370 for immediate allocation). The tests then read
GetGamma() / ParameterCount / ran ZeroInitGamma BEFORE any forward, so they saw
length/count 0 and failed. They were also mutually inconsistent (Init_* assumed
3 features while ParameterCount assumed 5), so no single lazy/eager default could
satisfy them, and making the parameterless ctor eager-default would regress lazy
BatchNorm resolution across the codebase.

Fix (test-only, assertions unchanged):
- Init_* (Gamma=1, Beta=0, RunningMean=0, RunningVar=1) and ZeroInitGamma use the
  eager ctor `(3)` — this exercises the documented eager init, which matches the
  PyTorch BatchNorm(num_features) standard (weight=1, bias=0, running_mean=0,
  running_var=1, eps default).
- ParameterCount uses `(5)` so 2 * 5 = 10.
- CustomEpsilon / CustomMomentum pass `epsilon: 1.0` / `momentum: 0.5` explicitly
  (their math assumed those values, not the 1e-5 / 0.9 defaults); they forward an
  input, so lazy feature resolution still applies there.

No production code changed. All 46 NormalizationLayersDeepMath tests pass.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(finance): Kairos size-router operates on the full look-back window

KairosMultiSizePatchLayer.Forward received a rank-3 [B, lookback, channels]
tensor from Kairos.ForwardNative but only handled rank-1/rank-2. The DenseLayer
router therefore bound to the trailing (channel) axis on the first forward,
while SetParameters re-resolves it from contextLength on deserialize — so the
serialization round-trip threw "expected 225 parameters across sublayers, got
218" (router 9 vs 2 params, the contextLength=8 vs channels=1 difference).

Per Kairos (Mixture-of-Size Encoder) the size-router weights the patch scales
from the FULL look-back window, so Forward now collapses any rank>2 input to
[B, contextLength] before routing. This is both paper-faithful (router sees the
whole series, not one channel) and makes the forward and SetParameters router
shapes agree, fixing the round-trip.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(finance): MOIRAI decoder-only training projects mixture head to forecast

ForwardDecoderOnlyForTraining handled only a rank-2 mixture head, but a
per-position head emerges as rank-3 [B, seqLen, numMixtures*3]. The encoder
training path (ForwardNativeForTraining) already bridges rank-3 via ReduceMean
over the context axis before extracting point predictions; the decoder-only
path was missing that step, so it returned the raw [1, 8, 30] head and the MSE
training loss threw a shape mismatch against the [1, 4, 1] forecast target.

Mirror the encoder path: average the per-position predictive distributions
across the sequence axis (paper-faithful pooling per Woo et al. 2024) before
ExtractPointPredictionsTapeSafe, yielding the [B, forecastHorizon, 1] forecast.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(finance): CSDI routes packed denoiser input through the input projection

CSDI pre-projected the conditioning through _inputProjection (consuming its
resolved input width on the conditioning shape) and then fed the raw packed
per-step denoiser input [noisy | cond | sin(t)] (e.g. [1, 33]) straight into the
residual stack, whose BatchNorm channels are sized to hiddenDimension (24) —
[1, 33] vs [1, 24] broadcast crash on the first reverse step.

Per Tashiro et al. 2021 ("CSDI") and the layer-helper layout (input projection
-> residual blocks -> output projection), _inputProjection is meant to project
the WHOLE packed per-step input to hidden width. Both ForwardNative (inference)
and ComputeDenoisingPairTape (training) now pack the RAW conditioning and route
the packed input through _inputProjection -> residual stack -> output
projection, and the training path slices the flat score to the denoised target
length so it aligns with the true-noise tensor. Inference and training denoiser
graphs are now identical.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(finance): TSDiff/TimeDiff/CCDM route packed denoiser input through projection

Same root cause as CSDI: each pre-projected the conditioning through
_inputProjection and then fed the raw packed per-step denoiser input
[noisy | cond | time/sigma] straight into the residual/transformer/denoising
stack, whose BatchNorm channels are sized to hiddenDimension — a [1, N] vs
[1, hiddenDimension] broadcast crash on the first reverse step.

Per each model's paper (TSDiff: Kollovieh et al. 2024; TimeDiff; CCDM) and the
layer-helper layout, _inputProjection projects the WHOLE packed per-step input
to hidden width. All three now pack the RAW conditioning and route the packed
input through _inputProjection before their denoiser stack (TSDiff also for its
classifier-free-guidance unconditional pass).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* feat(nn): add eager FullyConnectedLayer(inputSize, outputSize) constructor

FullyConnectedLayer only had an output-only ctor that resolves the input size
lazily on first forward, so ParameterCount reports 0 weight params (bias only)
before any forward. The FullyConnectedLayer_ParameterCount deep-math test
constructed the layer and asserted 3*2+2=8 params without a forward.

Add an eager (inputSize, outputSize, activation) ctor that allocates and
initializes the weight/bias tensors immediately — the PyTorch
nn.Linear(in_features, out_features) convention — and point the test at it. The
new overload is additive (existing output-only / output+activation call-sites
are unaffected). No assertion changed.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(nn): AttentionLayer cross-attn/masked first-forward init + block test ctors

AttentionLayer (Integration N-O):
- Cross-attention disambiguation compared the K/V feature dim against _inputSize,
  which is -1 on a lazy layer's first forward, so every first-forward cross-
  attention call threw "ambiguous shape". Compare against the QUERY's feature
  dim (K/V shares the query embedding width) instead.
- ForwardCrossAttention / ForwardMaskedAttention did not resolve _inputSize or
  allocate the Q/K/V projections when cross-/masked-attention was the layer's
  FIRST forward (only the self-attention Forward did), so the Q reshape to
  [batch*seqLen, _inputSize] failed. Both now call EnsureInitializedFromInput
  on the query/input first.

SpecializedBlocks tests (Integration N-O), test-only:
- BottleneckBlock is parameterized by base width (PyTorch Bottleneck(inplanes,
  planes): output = planes*4, inChannels inferred from the input). The test
  passed (inChannels, outChannels) which set baseChannels=64/stride=32; fixed to
  pass the base width so output = 32*4 = 128.
- DenseBlock.OutputChannels resolves inputChannels on first forward, so the test
  now runs one forward before reading the property (it reports -1 until then).

All 32 AttentionLayer integration tests and both SpecializedBlocks tests pass.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(nn): Autoencoder.EncodedSize and SpiralNet.ParameterCount report eagerly

Both reported a non-positive value on a freshly-constructed (un-forwarded)
network because their default layer stacks are lazy:

- Autoencoder.EncodedSize read the middle layer's INPUT shape (the latent
  handoff), which is -1 until the first forward. It now takes the narrowest
  Dense OUTPUT width (the bottleneck), which is concrete at construction.
- SpiralNet.ParameterCount summed lazy SpiralConvLayer params (0 until forward).
  Added an eager SpiralConvLayer(inputChannels, outputChannels, spiralLength)
  ctor that allocates weights immediately, and CreateDefaultSpiralNetLayers now
  passes the chained input channel count so the conv weights exist at
  construction. OnFirstForward skips re-allocation when weights are already
  present (no double-register, no re-init).

Autoencoder_EncodedSize_IsPositive and SpiralNet_GetParameterCount_ReturnsPositiveValue pass.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(nn): ConditionalGAN generator architecture accounts for the condition input

The conditional generator's runtime input is [noise | condition], but
CreateConditionalGeneratorArchitecture returned the generator architecture
unchanged, so its first layer was sized for the noise-only width and Predict
threw a shape mismatch (expected [32], got [42] for noise=32 + 10 classes).
Mirror the discriminator's input expansion: grow the flat input width by
numConditionClasses (or add numConditionClasses channels for spatial
generators) so the generator's first layer matches the concatenated input.

ConditionalGAN_Predict_ProducesOutput and ConditionalGAN_GetParameterCount pass.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(nn): GRU optional stateful mode + GraphConv requires adjacency

GRULayer: add a Keras-style `stateful` constructor flag (default false). When
false the layer keeps the standard stateless contract (h0 = zeros each Forward,
matching PyTorch nn.GRU) that Clone-after-training parity and repeated-Predict
determinism depend on. When true the hidden state carries over between Forward
calls (reset only on first call / batch change / ResetState). The
GRU_HiddenStateCarriesOver test opts into stateful:true.

GraphConvolutionalLayer: Forward now throws InvalidOperationException when no
adjacency matrix has been set, instead of silently substituting an identity
(which degrades the layer to a per-node MLP). Per Kipf & Welling 2017 the
propagation rule H' = A-hat*H*W is undefined without a graph. Audited all
in-repo GCN users (GraphNeuralNetwork, GraphSAGENetwork, GraphAttentionNetwork,
GraphNeuralOperator) — every one sets the adjacency before Forward — so this is
safe. The set-but-shape-stale rebuild path is preserved.

All 50 GRU + GraphLayers integration tests pass.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(nested-learning): align AssociativeMemory tests with Ramsauer Hopfield

AssociativeMemory implements modern continuous Hopfield (Ramsauer et al. 2021):
Associate/Update store (key, value) pairs, GetAssociationMatrix returns the
unscaled outer-product sum W = Σ value ⊗ key, and Retrieve is
softmax(β · Kᵀq) · V. The 8 deep-math tests asserted classic Hebbian behavior
instead (lr=0.01-scaled outer products, matrix-multiply + cosine-gated blending
retrieval), so they failed against the modern-Hopfield implementation.

Rewrite the assertions to the paper-faithful Ramsauer values (all hand-derived):
- matrix tests: unscaled outer products (W[1,0]=1, W=I, etc.)
- 2-memory retrieve with an equidistant query: softmax [0.5,0.5] (β-independent)
- single-memory retrieve: returns the stored value (softmax=[1]); 1e-9 local
  tolerance absorbs the /(sumExp+1e-10) stability epsilon
- identity-basis retrieve: assert the softmax-distribution invariants (sums to 1,
  components in (0,1), ordering follows the query-key scores)

Implementation unchanged (kept Ramsauer per design decision). All 11 pass.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(nn): DenseNet adds batch dim for rank-3 image input

DenseNet's conv and dense-block layers preserve the input rank, and its
per-channel BatchNorm broadcasts [1, C, 1] against the channel axis at dim 1.
A single image supplied as rank-3 [C, H, W] therefore propagates with channels
at dim 0, so the BatchNorm broadcast misaligns ([64, 32, 32] vs [1, 64, 1]).
Forward now prepends a batch dim for rank-3 input so the whole stack stays 4-D
with channels at dim 1, matching the standard image-model input contract.

DenseNetNetwork_Predict_ProducesOutput passes.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(nn): pool-aware memory-budget estimator bounds the actual peak

PredictWithMemoryBudget's slope-fit estimator probed only
GC.GetAllocatedBytesForCurrentThread, which counts NEW allocations. The tensor
allocator pools its buffers, so once a probe primes the pool at a batch size a
same-size forward reuses buffers and allocates ~0 — collapsing the per-sample
slope β toward zero. The chunk size (budget − alpha)/β then grew unbounded, and
a single oversized forward ballooned the pool to ~4.9 GB retained against a 1 GB
budget.

Add MeasureForwardRetainedBytes: measure the RETAINED managed-heap growth of a
forward at probeLarge, taken BEFORE the allocated probes prime the pool and with
a B=1-only warm-up, so the delta reflects the pooled footprint. Use it as a
per-sample floor (β = max(allocated-slope, retained-per-sample)) so the chunk is
sized against the memory that actually persists.

All 21 Issue1296LargeXTrainBatchingTests pass (peak now within budget).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(code): paper-faithful residual encoder for CodeBERT/GraphCodeBERT + clone round-trip

CodeBERT/GraphCodeBERT (BERT-class encoders, Feng et al. 2020 / Devlin 2018) were
built from a hand-stacked MHA→LN→FFN→LN sequence with NO residual connections, so
gradients could not flow and Training_ShouldReduceLoss failed. Replace the encoder
stack with the residual TransformerEncoderBlock (attention+residual+LN,
GELU-FFN+residual+LN, Vaswani 2017 §3).

Supporting fixes so the residual block and graph layer survive Clone:
- TransformerEncoderBlock: add an optional FFN activation arg (default ReLU =
  Vaswani; GELU = BERT) and persist it in GetMetadata; DeserializationHelper now
  restores it. Without this a cloned block fell back to the ctor default and
  diverged from the original (Clone_ShouldProduceIdenticalOutput).
- GraphConvolutionalLayer: add an opt-in implicitIdentityWhenUnset flag (default
  false = throw, Kipf & Welling 2017). DeserializationHelper reconstructs cloned
  graph layers with it enabled so a model whose adjacency is only set at runtime
  survives serialization, while direct construction keeps the strict contract the
  layer unit tests assert. The code-synthesis factory opts in (its token stream
  has no explicit data-flow graph).

CodeBERT/GraphCodeBERT Training + Clone (+ all 4 clone tests) pass; the layer's
require-adjacency unit test still passes.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(code): final LayerNorm + AdamW stabilize CodeBERT/GraphCodeBERT training

Per-layer instrumentation across training iterations showed the parameters
staying bounded (max ~1.0) while the activations grew ~400x over 100 iterations
to NaN — the signature of a Pre-LN transformer with no final LayerNorm.
TransformerEncoderBlock is Pre-LN (each sublayer adds to the residual stream
without normalizing it), so without a final LayerNorm the unnormalized residual
fed the output projection and the logits — and the training loss — blew up.

- Add the final LayerNorm after the encoder stack (GPT-2 / BERT-class standard).
- Switch the default optimizer to AdamW (Loshchilov & Hutter 2019) at lr=1e-5
  with weight decay 0.01 + AMSGrad — the BERT/RoBERTa setup. Decoupled weight
  decay and the lower (BERT fine-tuning range) LR remove the residual
  post-convergence drift the LayerNorm alone left behind.

CodeBERT is now fully green (Training, Clone, MoreData); GraphCodeBERT Training +
Clone pass. GraphCodeBERT MoreData/LossStrictly remain perf timeouts (separate).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(code): reduced-scale GraphCodeBERT scaffold (fix MoreData perf timeout)

The source-generated GraphCodeBERTTests instantiated the model via its
parameterless ctor, which builds the CodeSynthesisArchitecture defaults —
BERT-base scale (6 layers, 512 dim, 512 seq, 50000 vocab). Profiling showed the
small smoke config trains at ~8 ms/step, but the default config's 512->50000
output projection (~25M params, ~205 MB output tensor per forward) makes the
~250-step training invariants overflow the 120s budget (MoreData_ShouldNotDegrade
timed out). The per-layer cost is fine — the bottleneck is purely the
production-scale default the generator picked.

Exclude GraphCodeBERT from the auto-generator (TestScaffoldGenerator
.ExcludedClassNames, same precedent as Janus/Helix/Donut) and add a manual
reduced-scale GraphCodeBERTTests in ModelFamilyTests/CodeModel mirroring
CodeBERTTests (2 layers, 64 dim, 4 heads, 32 seq, 128 vocab) with
UseDataFlow=true so the data-flow graph-convolution path is still exercised. All
21 invariants now pass in ~6s.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(causal): paper-canonical time-series fixtures for information-flow algorithms

The shared causal-discovery fixture was a NOISELESS straight-line ramp
(X0 = 0.1·i with exact linear functions of it). For time-series information-flow
estimators that data is mathematically degenerate: a ramp is perfectly
self-predictable (Y[t] = 2Y[t-1] - Y[t-2]) and its transitions are constant, so
transfer entropy (Schreiber 2000) and causation entropy (Sun et al. 2015) are
EXACTLY ZERO by the papers' own definitions — the paper-faithful algorithms
correctly found no edges and the >=1-edge assertion failed. Equally, the
generated-suite fixture used i.i.d. SEM rows, where sample t is independent of
sample t-1 and the true lagged-transfer score is zero for every pair.

- InformationTheoreticCausalDiscoveryTests: coupled stochastic AR system with
  seeded noise and lagged drive X0 -> X1 -> X2 (the canonical validation setup
  for these estimators). Assertions unchanged.
- CausalDiscoveryTestBase.CreateKnownStructureData: emits a LAGGED variant of
  the same SEM (same true edges, same coefficients, same X1 _|_ X3 | X0
  conditional independence) when the algorithm declares SupportsTimeSeries.

Verified: TransferEntropy/OCSE integration tests and the generated
TransferEntropyAlgorithm RecoversTrueEdges now pass; the 200-test causal
namespace run shows no regressions among previously-passing algorithms.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(causal,vision): paper-faithful fixes for SegMamba, BCDNets, IGCI, NOTEARSLowRank + degenerate fixtures

SegMamba (Xing et al. 2024): decoder upsamples compute 2*ceil(n/2) >= n, so on
non-2^depth-divisible inputs the skip-concat threw a shape mismatch. Forward now
center-crops the upsampled tensor to the skip / input spatial dims (U-Net
convention, Ronneberger 2015) via tape-safe TensorNarrow + Contiguous.

BCDNetsAlgorithm (Cundy et al. 2021): three compounding optimization bugs made
the posterior collapse to an empty or reversed graph —
1. dual ascent ran after EVERY gradient step against an ABSOLUTE h > 0.25 test,
   so rho exploded 10x/step to 1e16 (NOTEARS schedule: outer rounds, escalate
   only when h > 0.25 * h_prev);
2. unclamped edge logits saturated so deep under the early symmetric acyclicity
   pressure that sigmoid'(Z/tau) ~ 0 froze ALL gradients permanently (standard
   logit clamping to [-4,4] keeps the gate responsive);
3. OLS init turned BOTH directions of every correlated pair ON, starting the
   search in a dense 2-cycle graph whose pruning is orientation-blind. Init is
   now direction-aware via the equal-variance identifiability criterion (Peters
   & Buhlmann 2014): only the lower-conditional-residual direction starts ON.
Verified: recovers all 3 true edges with correct orientations and near-true
weights (0->1 = 0.785 vs true 0.8) on the 4-var SEM.

IGCIAlgorithm (Daniusis 2010 / Janzing 2012): the score rank-transformed BOTH
variables, which maps every deterministic monotone relationship to the identity
(all slopes 1, score exactly 0) — IGCI could never orient any edge. The paper's
estimator rescales each variable's RANGE to [0,1] (uniform reference measure);
replaced the rank transform with affine min-max rescaling.

NOTEARSLowRank: guard the L-BFGS objective against non-finite W from
line-search overshoot (tr(exp(W.W)) overflow propagated NaN into Math.Sign,
throwing ArithmeticException); a large finite objective makes the line search
backtrack — standard augmented-Lagrangian handling.

Degenerate-fixture fixes (same approved principle as the TE/oCSE commit: the
noiseless collinear ramp is mathematically outside these algorithms' stated
domains):
- CCM: coupled chaotic logistic maps — the CCM paper's own validation system;
  convergence of cross-map skill cannot be observed on a saturated 1-D ramp.
- tsFCI: lagged stochastic AR chain; Fisher-z CI tests are 0/0-degenerate on
  exact collinearity.
- FastIAMB/CDNOD: noisy i.i.d. SEM; |r|=1 makes Fisher-z infinite, and CD-NOD's
  time-index surrogate fully explains a deterministic ramp (empty X-X graph is
  its paper-correct output there).
- IGCI: nonlinear deterministic invertible chain (x^3, tanh) — the paper states
  linear f is its unidentifiable boundary case.

All verified green in the 168/172 causal-namespace run + slice 2B.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(nn): SpikingNeuralNetwork trains its composite-core topology via the tape

The default SNN topology (LayerHelper.CreateDefaultSpikingLayers) is a single
composite SpikingNetworkCore + ActivationLayer. The hand-rolled spike-history
delta rule in Train() assumes a MULTI-LAYER stack: it walks back to the first
trainable layer and updates layers via `for (layerIndex = outputLayerIndex;
layerIndex > 0; layerIndex--)` — with the composite core the first trainable
layer IS index 0, so the loop body never executed and Train silently updated
NOTHING (diagnosed: parameter L1-delta was exactly 0.0 after Train at every
input scale; Training_ShouldChangeParameters / GradientFlow / TrainingError
failed with "parameters did not change").

SpikingNetworkCore is a BPTT-unrolled surrogate-gradient SNN built entirely
from Engine ops (straight-through Heaviside with a fast-sigmoid surrogate —
Neftci, Mostafa & Zenke 2019 §III), so it is tape-differentiable end-to-end:
route the composite topology through TrainWithTape. The legacy spike-history
path remains for custom multi-layer SpikingLayer stacks supplied via
Architecture.Layers.

All 21 SpikingNeuralNetworkTests invariants pass (verified parameter delta now
nonzero and prediction converges toward the target).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(causal): stabilize GraN-DAG augmented Lagrangian (NaN weight explosion)

GraN-DAG (Lachapelle et al. 2020) found 0 edges at every learning rate because
its raw path-norm adjacency was NaN: the dual update compared h against an
ABSOLUTE 0.25 threshold and escalated rho 10x per epoch (the NOTEARS schedule
is relative — escalate only when h > 0.25 * h_prev), and unlike BCD-Nets'
clamped edge logits the MLP weights here are unbounded, so the ~1e17 penalty
coefficient launched them into the tr(exp(A.A)) overflow regime and every
weight went NaN.

- Dual ascent moved to the relative NOTEARS stall rule with an h < 1e-8 stop.
- The augmented coefficient alpha + rho*h is clamped at 1e6 (constraint force
  stays strong but finite) and the inner/outer loops bail to the last finite
  state if h goes non-finite.
- Per-element gradient clipping (|g| <= 10) bounds the per-step weight change —
  the standard guards in NOTEARS-family reference implementations.

The integration test now passes with true edges recovered (0->1 found at the
default lr) and the fixture uses a noisy i.i.d. SEM (GraN-DAG fits a Gaussian
likelihood with per-node noise variances — the noiseless rank-1 ramp is its
degenerate zero-variance limit; the paper's experiments use stochastic SEMs).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(diffusion): CatVTON clone delegation + reduced-scale test scaffold

CatVTONModel.Clone() constructed a fresh model — which rebuilds the DEFAULT
SD-1.5-scale predictor/VAE — then called SetParameters with this model's
weights: for a model built with custom-sized components the parameter counts
differ and SetParameters throws, and even at default scale the fresh
components re-hit the lazy-init divergence bug. Clone now delegates to the
predictor/VAE's own Clone implementations (RealESRGAN / SeedEdit3 /
ImprovedConsistencyModel pattern, PR #1555/#1565).

The test scaffold previously instantiated the production default — the full
SD-1.5-inpainting UNet (~860M params): a single Predict is a multi-step
reverse-sampling loop over ~15s/forward CPU passes, which hung the shared
diffusion-test host (the crash-test inactivity dump on
Scheduler_ShouldBeNonNull). Per the AnimateDiff/Janus reduced-scale precedent
the scaffold now supplies a small UNet + VAE with the same architecture shape,
critically keeping CatVTON's defining 12-channel inflated input convolution
(the person/garment/mask latent concatenation of Chong et al. 2024,
"Concatenation Is All You Need"). The smoke latent is 32x32, not 64x64: at
64x64 the resolution-1 self-attention spans 4096 tokens and the per-step
attention tensors accumulating on the tape paged the host into a >14 GB swap
stall (blame-hang isolated Clone_ShouldProduceIdenticalOutput); at 32x32 the
full 13-test suite passes in ~90s.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: franklinic <franklin@ivorycloud.com>

This branch was successfully deployed

2 active deployments
Preview – aidotnet_website — 11127b50 Deployed Jun 11, 2026 by vercel[bot]
Preview – aidotnet-playground-api — 11127b50 Deployed Jun 11, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Transformer<T>.Train() produces uniform output at production scale (v0.166.0) — top-1=0% / PPL=|V| on 1MB byte-LM

1 participant