chore(deps): bump AiDotNet.Tensors 0.92.1 -> 0.92.5 (closes #1221) - #1563
Conversation
0.92.5 is the first release carrying Tensors #583 (validate the prepacked-B GEMM cache against in-place weight mutation). The SgemmWithCachedB inference path reused a stale pre-packed B panel when a layer's weights were mutated in place without a MarkWeightDirty, so FusedLinear inference returned pre-update results — the root cause of the #1221 transformer production-scale convergence failure. Bumps the AiDotNet.Native.* packages (OneDNN/OpenBLAS/CLBlast) in lockstep, all published at 0.92.5. Restores clean (no NU1102); the 11 Issue1221 convergence/harness tests pass. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
|
The latest updates on your projects. Learn more about Vercel for GitHub.
|
WalkthroughCentralized package version declarations for the AiDotNet tensor and native backend ecosystem are bumped from ChangesAiDotNet Tensor and Native Dependencies
Estimated Code Review Effort🎯 1 (Trivial) | ⏱️ ~3 minutes Possibly Related PRs
Poem
🚥 Pre-merge checks | ✅ 5✅ Passed checks (5 passed)
✏️ Tip: You can configure your own custom pre-merge checks in the settings. ✨ Finishing Touches🧪 Generate unit tests (beta)
Comment |
There was a problem hiding this comment.
Actionable comments posted: 1
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@Directory.Packages.props`:
- Around line 142-145: The PR currently pins four AiDotNet package versions
(AiDotNet.Tensors, AiDotNet.Native.OneDNN, AiDotNet.Native.OpenBLAS,
AiDotNet.Native.CLBlast) to 0.92.5 in Directory.Packages.props and the reviewer
confirmed they exist on nuget.org; to prevent future mismatches add a
CI/pre-merge step that validates each package/version from
Directory.Packages.props is available on nuget.org (or the configured feed) and
fails fast if any PackageVersion (e.g., the AiDotNet.Tensors or
AiDotNet.Native.* entries) is not published so NU1102 regressions are caught
before merge.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Path: .coderabbit.yaml
Review profile: ASSERTIVE
Plan: Pro
Run ID: 41949f1b-c548-4bdc-a648-73d9a34f3f2d
📒 Files selected for processing (1)
Directory.Packages.props
…1563) 0.92.5 is the first release carrying Tensors #583 (validate the prepacked-B GEMM cache against in-place weight mutation). The SgemmWithCachedB inference path reused a stale pre-packed B panel when a layer's weights were mutated in place without a MarkWeightDirty, so FusedLinear inference returned pre-update results — the root cause of the #1221 transformer production-scale convergence failure. Bumps the AiDotNet.Native.* packages (OneDNN/OpenBLAS/CLBlast) in lockstep, all published at 0.92.5. Restores clean (no NU1102); the 11 Issue1221 convergence/harness tests pass. Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
…ors, continual-learning + batchnorm tests (#1565) * fix(diffusion): resolve lazy shapes in clone roundtrip for video/editing models The lazy-init clone bug surfaced in ModelFamily Clone_ShouldProduceIdenticalOutput for RealESRGAN, SeedEdit3, and StableVideoDiffusion: the predictor/VAE time- embedding MLPs (and the VideoUNet image-condition projection) report ParameterCount=0 until their first forward, so reconstructing a fresh model and calling SetParameters(GetParameters()) copied an under-counted snapshot and let the clone re-resolve with fresh random weights on its first Predict. - RealESRGAN/SeedEdit3 Clone now delegate to the predictor/VAE's own Clone, which trigger lazy shape resolution on both source and clone before copying weights. - VideoUNetPredictor.Clone gains TriggerLazyShapeResolution (mirroring UNetNoisePredictor), and it also resolves the image-condition projection since image-to-video models (SVD) always forward through that path. TemporalVAE has no lazy layers, so StableVideoDiffusion's existing delegating Clone is now correct end-to-end. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore: ignore local .ci-failures tracking dir Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(diffusion): round-trip VideoUNet FiLM projection + image-mode lazy trigger StableVideoDiffusion's Clone_ShouldProduceIdenticalOutput still diverged after the initial lazy-resolution fix because VideoUNetPredictor's parameter walk omitted the per-block FiLM TimeCondProjection (the AdaGN timestep-conditioning DenseLayer applied in ApplyVideoBlock). It was excluded from AddBlockParameters, SetBlockParameters, AND AddBlockGradients, so its weights never round-tripped (the clone kept fresh-random FiLM weights) and never trained. - Add TimeCondProjection to all three block walks, in the same position (right after SpatialResBlock, matching the forward order). - TriggerLazyShapeResolution now mirrors the real Predict path exactly: a tiny 4D image-mode forward with no conditioning (the denoising loop never runs the 5D video / image-condition / cross-attention paths, and the temporal-mixing DenseLayer is sized strictly [_numFrames -> _numFrames]). Lazy weights are spatial-independent, so a minimal dummy resolves identical shapes on source and clone cheaply, keeping the GetParameters/SetParameters copy index-aligned. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(rl): correct WorldModelsAgent network sizing, deterministic clone, pure predict Three bugs surfaced by the Generated Layers N-Z shard (WorldModelsAgentTests.ActionSelection_ShouldBeFinite / Clone_ShouldProduceSamePolicy): 1. Malformed, multi-GB networks. Each VAE/RNN network was built from an architecture with NO explicit layers, so NeuralNetwork.InitializeLayers auto-generated default hidden layers (~2x the flattened input) and the agent's AddLayer calls then stacked on top. For a 64x64x3 = 12,288-wide observation that default layer alone is a single 12,288x24,576 weight (~2.4 GB at FP64), which tripped weight-streaming's per-tensor byte cap and threw on the first forward. Now each network is built with its EXACT layer list passed to the architecture, so no defaults are generated. 2. Non-deterministic clone. NeuralNetwork.InitializeLayers does not wire explicitly-supplied custom layers, so their lazy weights initialized from the shared non-deterministic RNG — a clone re-initialized a different policy. Each layer now gets a deterministic per-layer RandomSeed, and Clone reproduces the networks via the constructor (they are never trained; only the controller is, via the evolution strategy) and copies the trained controller weights + rollout state directly. A serialization round-trip is NOT used for cloning because it rebuilds the layers and drops the seed pins. 3. Non-pure Predict. SelectAction advances the RNN hidden state for sequential rollout, making repeated Predict calls non-deterministic. Predict now snapshots and restores the hidden state so inference is side-effect-free. All 7 WorldModelsAgentTests pass. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(physics): correct HamiltonianNeuralNetwork parameterless ctor input size The parameterless constructor built its architecture with inputSize = DefaultStateDim * 2 (= 4) while passing stateDim = DefaultStateDim (= 2). The constructor enforces CalculatedInputSize == stateDim (the network input IS the full phase-space state vector (q, p), which it splits into q/p halves), so every `new HamiltonianNeuralNetwork<T>()` threw "Hamiltonian network input size (4) must match state dimension (2)". The source-generated ModelFamily test scaffold builds the model via the parameterless ctor, so the throw failed ALL 21 HamiltonianNeuralNetworkTests (Architecture/Clone/ForwardPass/Training/Metadata/Parameters/...). Fix: set inputSize = DefaultStateDim so input size matches stateDim. DefaultStateDim = 2 is the canonical 1-DOF Hamiltonian phase space (one position + one momentum). All 21 tests pass. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(continual-learning): declare IParameterizable on EWC/SI test mock model The interface-segregation refactor moved GetParameters / SetParameters / ParameterCount / WithParameters off IFullModel and onto the separate IParameterizable interface. EWC and Synaptic Intelligence guard-require it via InterfaceGuard.Parameterizable (correctly — EWC's Fisher information matrix and SI's per-parameter importance are both computed OVER the model's parameters, so a model without trainable parameters cannot participate). The CLMockModel test fixture still declared only IFullModel, so the runtime `model is IParameterizable<...>` check failed and threw "CLMockModel does not implement IParameterizable<...>" from PrepareForTask — failing all 14 EWC/SI deep-math integration tests. The mock already implements every IParameterizable member (GetParameters, SetParameters, ParameterCount, SupportsParameterInitialization, WithParameters, SanitizeParameters); it just needed to declare the interface so the cast succeeds. No assertion or expected value changed. All 36 ContinualLearningDeepMath tests pass. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(batchnorm): use eager ctor for init-contract assertions in deep-math tests Eight BatchNorm deep-math tests constructed the layer via the parameterless ctor `new BatchNormalizationLayer<double>()`, which is intentionally LAZY — gamma / beta / running stats are sized on the first forward (a separate eager ctor `(numFeatures)` was added in #1370 for immediate allocation). The tests then read GetGamma() / ParameterCount / ran ZeroInitGamma BEFORE any forward, so they saw length/count 0 and failed. They were also mutually inconsistent (Init_* assumed 3 features while ParameterCount assumed 5), so no single lazy/eager default could satisfy them, and making the parameterless ctor eager-default would regress lazy BatchNorm resolution across the codebase. Fix (test-only, assertions unchanged): - Init_* (Gamma=1, Beta=0, RunningMean=0, RunningVar=1) and ZeroInitGamma use the eager ctor `(3)` — this exercises the documented eager init, which matches the PyTorch BatchNorm(num_features) standard (weight=1, bias=0, running_mean=0, running_var=1, eps default). - ParameterCount uses `(5)` so 2 * 5 = 10. - CustomEpsilon / CustomMomentum pass `epsilon: 1.0` / `momentum: 0.5` explicitly (their math assumed those values, not the 1e-5 / 0.9 defaults); they forward an input, so lazy feature resolution still applies there. No production code changed. All 46 NormalizationLayersDeepMath tests pass. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(finance): Kairos size-router operates on the full look-back window KairosMultiSizePatchLayer.Forward received a rank-3 [B, lookback, channels] tensor from Kairos.ForwardNative but only handled rank-1/rank-2. The DenseLayer router therefore bound to the trailing (channel) axis on the first forward, while SetParameters re-resolves it from contextLength on deserialize — so the serialization round-trip threw "expected 225 parameters across sublayers, got 218" (router 9 vs 2 params, the contextLength=8 vs channels=1 difference). Per Kairos (Mixture-of-Size Encoder) the size-router weights the patch scales from the FULL look-back window, so Forward now collapses any rank>2 input to [B, contextLength] before routing. This is both paper-faithful (router sees the whole series, not one channel) and makes the forward and SetParameters router shapes agree, fixing the round-trip. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(finance): MOIRAI decoder-only training projects mixture head to forecast ForwardDecoderOnlyForTraining handled only a rank-2 mixture head, but a per-position head emerges as rank-3 [B, seqLen, numMixtures*3]. The encoder training path (ForwardNativeForTraining) already bridges rank-3 via ReduceMean over the context axis before extracting point predictions; the decoder-only path was missing that step, so it returned the raw [1, 8, 30] head and the MSE training loss threw a shape mismatch against the [1, 4, 1] forecast target. Mirror the encoder path: average the per-position predictive distributions across the sequence axis (paper-faithful pooling per Woo et al. 2024) before ExtractPointPredictionsTapeSafe, yielding the [B, forecastHorizon, 1] forecast. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(finance): CSDI routes packed denoiser input through the input projection CSDI pre-projected the conditioning through _inputProjection (consuming its resolved input width on the conditioning shape) and then fed the raw packed per-step denoiser input [noisy | cond | sin(t)] (e.g. [1, 33]) straight into the residual stack, whose BatchNorm channels are sized to hiddenDimension (24) — [1, 33] vs [1, 24] broadcast crash on the first reverse step. Per Tashiro et al. 2021 ("CSDI") and the layer-helper layout (input projection -> residual blocks -> output projection), _inputProjection is meant to project the WHOLE packed per-step input to hidden width. Both ForwardNative (inference) and ComputeDenoisingPairTape (training) now pack the RAW conditioning and route the packed input through _inputProjection -> residual stack -> output projection, and the training path slices the flat score to the denoised target length so it aligns with the true-noise tensor. Inference and training denoiser graphs are now identical. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(finance): TSDiff/TimeDiff/CCDM route packed denoiser input through projection Same root cause as CSDI: each pre-projected the conditioning through _inputProjection and then fed the raw packed per-step denoiser input [noisy | cond | time/sigma] straight into the residual/transformer/denoising stack, whose BatchNorm channels are sized to hiddenDimension — a [1, N] vs [1, hiddenDimension] broadcast crash on the first reverse step. Per each model's paper (TSDiff: Kollovieh et al. 2024; TimeDiff; CCDM) and the layer-helper layout, _inputProjection projects the WHOLE packed per-step input to hidden width. All three now pack the RAW conditioning and route the packed input through _inputProjection before their denoiser stack (TSDiff also for its classifier-free-guidance unconditional pass). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * feat(nn): add eager FullyConnectedLayer(inputSize, outputSize) constructor FullyConnectedLayer only had an output-only ctor that resolves the input size lazily on first forward, so ParameterCount reports 0 weight params (bias only) before any forward. The FullyConnectedLayer_ParameterCount deep-math test constructed the layer and asserted 3*2+2=8 params without a forward. Add an eager (inputSize, outputSize, activation) ctor that allocates and initializes the weight/bias tensors immediately — the PyTorch nn.Linear(in_features, out_features) convention — and point the test at it. The new overload is additive (existing output-only / output+activation call-sites are unaffected). No assertion changed. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(nn): AttentionLayer cross-attn/masked first-forward init + block test ctors AttentionLayer (Integration N-O): - Cross-attention disambiguation compared the K/V feature dim against _inputSize, which is -1 on a lazy layer's first forward, so every first-forward cross- attention call threw "ambiguous shape". Compare against the QUERY's feature dim (K/V shares the query embedding width) instead. - ForwardCrossAttention / ForwardMaskedAttention did not resolve _inputSize or allocate the Q/K/V projections when cross-/masked-attention was the layer's FIRST forward (only the self-attention Forward did), so the Q reshape to [batch*seqLen, _inputSize] failed. Both now call EnsureInitializedFromInput on the query/input first. SpecializedBlocks tests (Integration N-O), test-only: - BottleneckBlock is parameterized by base width (PyTorch Bottleneck(inplanes, planes): output = planes*4, inChannels inferred from the input). The test passed (inChannels, outChannels) which set baseChannels=64/stride=32; fixed to pass the base width so output = 32*4 = 128. - DenseBlock.OutputChannels resolves inputChannels on first forward, so the test now runs one forward before reading the property (it reports -1 until then). All 32 AttentionLayer integration tests and both SpecializedBlocks tests pass. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(nn): Autoencoder.EncodedSize and SpiralNet.ParameterCount report eagerly Both reported a non-positive value on a freshly-constructed (un-forwarded) network because their default layer stacks are lazy: - Autoencoder.EncodedSize read the middle layer's INPUT shape (the latent handoff), which is -1 until the first forward. It now takes the narrowest Dense OUTPUT width (the bottleneck), which is concrete at construction. - SpiralNet.ParameterCount summed lazy SpiralConvLayer params (0 until forward). Added an eager SpiralConvLayer(inputChannels, outputChannels, spiralLength) ctor that allocates weights immediately, and CreateDefaultSpiralNetLayers now passes the chained input channel count so the conv weights exist at construction. OnFirstForward skips re-allocation when weights are already present (no double-register, no re-init). Autoencoder_EncodedSize_IsPositive and SpiralNet_GetParameterCount_ReturnsPositiveValue pass. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(nn): ConditionalGAN generator architecture accounts for the condition input The conditional generator's runtime input is [noise | condition], but CreateConditionalGeneratorArchitecture returned the generator architecture unchanged, so its first layer was sized for the noise-only width and Predict threw a shape mismatch (expected [32], got [42] for noise=32 + 10 classes). Mirror the discriminator's input expansion: grow the flat input width by numConditionClasses (or add numConditionClasses channels for spatial generators) so the generator's first layer matches the concatenated input. ConditionalGAN_Predict_ProducesOutput and ConditionalGAN_GetParameterCount pass. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(nn): GRU optional stateful mode + GraphConv requires adjacency GRULayer: add a Keras-style `stateful` constructor flag (default false). When false the layer keeps the standard stateless contract (h0 = zeros each Forward, matching PyTorch nn.GRU) that Clone-after-training parity and repeated-Predict determinism depend on. When true the hidden state carries over between Forward calls (reset only on first call / batch change / ResetState). The GRU_HiddenStateCarriesOver test opts into stateful:true. GraphConvolutionalLayer: Forward now throws InvalidOperationException when no adjacency matrix has been set, instead of silently substituting an identity (which degrades the layer to a per-node MLP). Per Kipf & Welling 2017 the propagation rule H' = A-hat*H*W is undefined without a graph. Audited all in-repo GCN users (GraphNeuralNetwork, GraphSAGENetwork, GraphAttentionNetwork, GraphNeuralOperator) — every one sets the adjacency before Forward — so this is safe. The set-but-shape-stale rebuild path is preserved. All 50 GRU + GraphLayers integration tests pass. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(nested-learning): align AssociativeMemory tests with Ramsauer Hopfield AssociativeMemory implements modern continuous Hopfield (Ramsauer et al. 2021): Associate/Update store (key, value) pairs, GetAssociationMatrix returns the unscaled outer-product sum W = Σ value ⊗ key, and Retrieve is softmax(β · Kᵀq) · V. The 8 deep-math tests asserted classic Hebbian behavior instead (lr=0.01-scaled outer products, matrix-multiply + cosine-gated blending retrieval), so they failed against the modern-Hopfield implementation. Rewrite the assertions to the paper-faithful Ramsauer values (all hand-derived): - matrix tests: unscaled outer products (W[1,0]=1, W=I, etc.) - 2-memory retrieve with an equidistant query: softmax [0.5,0.5] (β-independent) - single-memory retrieve: returns the stored value (softmax=[1]); 1e-9 local tolerance absorbs the /(sumExp+1e-10) stability epsilon - identity-basis retrieve: assert the softmax-distribution invariants (sums to 1, components in (0,1), ordering follows the query-key scores) Implementation unchanged (kept Ramsauer per design decision). All 11 pass. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(nn): DenseNet adds batch dim for rank-3 image input DenseNet's conv and dense-block layers preserve the input rank, and its per-channel BatchNorm broadcasts [1, C, 1] against the channel axis at dim 1. A single image supplied as rank-3 [C, H, W] therefore propagates with channels at dim 0, so the BatchNorm broadcast misaligns ([64, 32, 32] vs [1, 64, 1]). Forward now prepends a batch dim for rank-3 input so the whole stack stays 4-D with channels at dim 1, matching the standard image-model input contract. DenseNetNetwork_Predict_ProducesOutput passes. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(nn): pool-aware memory-budget estimator bounds the actual peak PredictWithMemoryBudget's slope-fit estimator probed only GC.GetAllocatedBytesForCurrentThread, which counts NEW allocations. The tensor allocator pools its buffers, so once a probe primes the pool at a batch size a same-size forward reuses buffers and allocates ~0 — collapsing the per-sample slope β toward zero. The chunk size (budget − alpha)/β then grew unbounded, and a single oversized forward ballooned the pool to ~4.9 GB retained against a 1 GB budget. Add MeasureForwardRetainedBytes: measure the RETAINED managed-heap growth of a forward at probeLarge, taken BEFORE the allocated probes prime the pool and with a B=1-only warm-up, so the delta reflects the pooled footprint. Use it as a per-sample floor (β = max(allocated-slope, retained-per-sample)) so the chunk is sized against the memory that actually persists. All 21 Issue1296LargeXTrainBatchingTests pass (peak now within budget). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(code): paper-faithful residual encoder for CodeBERT/GraphCodeBERT + clone round-trip CodeBERT/GraphCodeBERT (BERT-class encoders, Feng et al. 2020 / Devlin 2018) were built from a hand-stacked MHA→LN→FFN→LN sequence with NO residual connections, so gradients could not flow and Training_ShouldReduceLoss failed. Replace the encoder stack with the residual TransformerEncoderBlock (attention+residual+LN, GELU-FFN+residual+LN, Vaswani 2017 §3). Supporting fixes so the residual block and graph layer survive Clone: - TransformerEncoderBlock: add an optional FFN activation arg (default ReLU = Vaswani; GELU = BERT) and persist it in GetMetadata; DeserializationHelper now restores it. Without this a cloned block fell back to the ctor default and diverged from the original (Clone_ShouldProduceIdenticalOutput). - GraphConvolutionalLayer: add an opt-in implicitIdentityWhenUnset flag (default false = throw, Kipf & Welling 2017). DeserializationHelper reconstructs cloned graph layers with it enabled so a model whose adjacency is only set at runtime survives serialization, while direct construction keeps the strict contract the layer unit tests assert. The code-synthesis factory opts in (its token stream has no explicit data-flow graph). CodeBERT/GraphCodeBERT Training + Clone (+ all 4 clone tests) pass; the layer's require-adjacency unit test still passes. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(code): final LayerNorm + AdamW stabilize CodeBERT/GraphCodeBERT training Per-layer instrumentation across training iterations showed the parameters staying bounded (max ~1.0) while the activations grew ~400x over 100 iterations to NaN — the signature of a Pre-LN transformer with no final LayerNorm. TransformerEncoderBlock is Pre-LN (each sublayer adds to the residual stream without normalizing it), so without a final LayerNorm the unnormalized residual fed the output projection and the logits — and the training loss — blew up. - Add the final LayerNorm after the encoder stack (GPT-2 / BERT-class standard). - Switch the default optimizer to AdamW (Loshchilov & Hutter 2019) at lr=1e-5 with weight decay 0.01 + AMSGrad — the BERT/RoBERTa setup. Decoupled weight decay and the lower (BERT fine-tuning range) LR remove the residual post-convergence drift the LayerNorm alone left behind. CodeBERT is now fully green (Training, Clone, MoreData); GraphCodeBERT Training + Clone pass. GraphCodeBERT MoreData/LossStrictly remain perf timeouts (separate). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(code): reduced-scale GraphCodeBERT scaffold (fix MoreData perf timeout) The source-generated GraphCodeBERTTests instantiated the model via its parameterless ctor, which builds the CodeSynthesisArchitecture defaults — BERT-base scale (6 layers, 512 dim, 512 seq, 50000 vocab). Profiling showed the small smoke config trains at ~8 ms/step, but the default config's 512->50000 output projection (~25M params, ~205 MB output tensor per forward) makes the ~250-step training invariants overflow the 120s budget (MoreData_ShouldNotDegrade timed out). The per-layer cost is fine — the bottleneck is purely the production-scale default the generator picked. Exclude GraphCodeBERT from the auto-generator (TestScaffoldGenerator .ExcludedClassNames, same precedent as Janus/Helix/Donut) and add a manual reduced-scale GraphCodeBERTTests in ModelFamilyTests/CodeModel mirroring CodeBERTTests (2 layers, 64 dim, 4 heads, 32 seq, 128 vocab) with UseDataFlow=true so the data-flow graph-convolution path is still exercised. All 21 invariants now pass in ~6s. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(causal): paper-canonical time-series fixtures for information-flow algorithms The shared causal-discovery fixture was a NOISELESS straight-line ramp (X0 = 0.1·i with exact linear functions of it). For time-series information-flow estimators that data is mathematically degenerate: a ramp is perfectly self-predictable (Y[t] = 2Y[t-1] - Y[t-2]) and its transitions are constant, so transfer entropy (Schreiber 2000) and causation entropy (Sun et al. 2015) are EXACTLY ZERO by the papers' own definitions — the paper-faithful algorithms correctly found no edges and the >=1-edge assertion failed. Equally, the generated-suite fixture used i.i.d. SEM rows, where sample t is independent of sample t-1 and the true lagged-transfer score is zero for every pair. - InformationTheoreticCausalDiscoveryTests: coupled stochastic AR system with seeded noise and lagged drive X0 -> X1 -> X2 (the canonical validation setup for these estimators). Assertions unchanged. - CausalDiscoveryTestBase.CreateKnownStructureData: emits a LAGGED variant of the same SEM (same true edges, same coefficients, same X1 _|_ X3 | X0 conditional independence) when the algorithm declares SupportsTimeSeries. Verified: TransferEntropy/OCSE integration tests and the generated TransferEntropyAlgorithm RecoversTrueEdges now pass; the 200-test causal namespace run shows no regressions among previously-passing algorithms. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(causal,vision): paper-faithful fixes for SegMamba, BCDNets, IGCI, NOTEARSLowRank + degenerate fixtures SegMamba (Xing et al. 2024): decoder upsamples compute 2*ceil(n/2) >= n, so on non-2^depth-divisible inputs the skip-concat threw a shape mismatch. Forward now center-crops the upsampled tensor to the skip / input spatial dims (U-Net convention, Ronneberger 2015) via tape-safe TensorNarrow + Contiguous. BCDNetsAlgorithm (Cundy et al. 2021): three compounding optimization bugs made the posterior collapse to an empty or reversed graph — 1. dual ascent ran after EVERY gradient step against an ABSOLUTE h > 0.25 test, so rho exploded 10x/step to 1e16 (NOTEARS schedule: outer rounds, escalate only when h > 0.25 * h_prev); 2. unclamped edge logits saturated so deep under the early symmetric acyclicity pressure that sigmoid'(Z/tau) ~ 0 froze ALL gradients permanently (standard logit clamping to [-4,4] keeps the gate responsive); 3. OLS init turned BOTH directions of every correlated pair ON, starting the search in a dense 2-cycle graph whose pruning is orientation-blind. Init is now direction-aware via the equal-variance identifiability criterion (Peters & Buhlmann 2014): only the lower-conditional-residual direction starts ON. Verified: recovers all 3 true edges with correct orientations and near-true weights (0->1 = 0.785 vs true 0.8) on the 4-var SEM. IGCIAlgorithm (Daniusis 2010 / Janzing 2012): the score rank-transformed BOTH variables, which maps every deterministic monotone relationship to the identity (all slopes 1, score exactly 0) — IGCI could never orient any edge. The paper's estimator rescales each variable's RANGE to [0,1] (uniform reference measure); replaced the rank transform with affine min-max rescaling. NOTEARSLowRank: guard the L-BFGS objective against non-finite W from line-search overshoot (tr(exp(W.W)) overflow propagated NaN into Math.Sign, throwing ArithmeticException); a large finite objective makes the line search backtrack — standard augmented-Lagrangian handling. Degenerate-fixture fixes (same approved principle as the TE/oCSE commit: the noiseless collinear ramp is mathematically outside these algorithms' stated domains): - CCM: coupled chaotic logistic maps — the CCM paper's own validation system; convergence of cross-map skill cannot be observed on a saturated 1-D ramp. - tsFCI: lagged stochastic AR chain; Fisher-z CI tests are 0/0-degenerate on exact collinearity. - FastIAMB/CDNOD: noisy i.i.d. SEM; |r|=1 makes Fisher-z infinite, and CD-NOD's time-index surrogate fully explains a deterministic ramp (empty X-X graph is its paper-correct output there). - IGCI: nonlinear deterministic invertible chain (x^3, tanh) — the paper states linear f is its unidentifiable boundary case. All verified green in the 168/172 causal-namespace run + slice 2B. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(nn): SpikingNeuralNetwork trains its composite-core topology via the tape The default SNN topology (LayerHelper.CreateDefaultSpikingLayers) is a single composite SpikingNetworkCore + ActivationLayer. The hand-rolled spike-history delta rule in Train() assumes a MULTI-LAYER stack: it walks back to the first trainable layer and updates layers via `for (layerIndex = outputLayerIndex; layerIndex > 0; layerIndex--)` — with the composite core the first trainable layer IS index 0, so the loop body never executed and Train silently updated NOTHING (diagnosed: parameter L1-delta was exactly 0.0 after Train at every input scale; Training_ShouldChangeParameters / GradientFlow / TrainingError failed with "parameters did not change"). SpikingNetworkCore is a BPTT-unrolled surrogate-gradient SNN built entirely from Engine ops (straight-through Heaviside with a fast-sigmoid surrogate — Neftci, Mostafa & Zenke 2019 §III), so it is tape-differentiable end-to-end: route the composite topology through TrainWithTape. The legacy spike-history path remains for custom multi-layer SpikingLayer stacks supplied via Architecture.Layers. All 21 SpikingNeuralNetworkTests invariants pass (verified parameter delta now nonzero and prediction converges toward the target). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(causal): stabilize GraN-DAG augmented Lagrangian (NaN weight explosion) GraN-DAG (Lachapelle et al. 2020) found 0 edges at every learning rate because its raw path-norm adjacency was NaN: the dual update compared h against an ABSOLUTE 0.25 threshold and escalated rho 10x per epoch (the NOTEARS schedule is relative — escalate only when h > 0.25 * h_prev), and unlike BCD-Nets' clamped edge logits the MLP weights here are unbounded, so the ~1e17 penalty coefficient launched them into the tr(exp(A.A)) overflow regime and every weight went NaN. - Dual ascent moved to the relative NOTEARS stall rule with an h < 1e-8 stop. - The augmented coefficient alpha + rho*h is clamped at 1e6 (constraint force stays strong but finite) and the inner/outer loops bail to the last finite state if h goes non-finite. - Per-element gradient clipping (|g| <= 10) bounds the per-step weight change — the standard guards in NOTEARS-family reference implementations. The integration test now passes with true edges recovered (0->1 found at the default lr) and the fixture uses a noisy i.i.d. SEM (GraN-DAG fits a Gaussian likelihood with per-node noise variances — the noiseless rank-1 ramp is its degenerate zero-variance limit; the paper's experiments use stochastic SEMs). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(diffusion): CatVTON clone delegation + reduced-scale test scaffold CatVTONModel.Clone() constructed a fresh model — which rebuilds the DEFAULT SD-1.5-scale predictor/VAE — then called SetParameters with this model's weights: for a model built with custom-sized components the parameter counts differ and SetParameters throws, and even at default scale the fresh components re-hit the lazy-init divergence bug. Clone now delegates to the predictor/VAE's own Clone implementations (RealESRGAN / SeedEdit3 / ImprovedConsistencyModel pattern, PR #1555/#1565). The test scaffold previously instantiated the production default — the full SD-1.5-inpainting UNet (~860M params): a single Predict is a multi-step reverse-sampling loop over ~15s/forward CPU passes, which hung the shared diffusion-test host (the crash-test inactivity dump on Scheduler_ShouldBeNonNull). Per the AnimateDiff/Janus reduced-scale precedent the scaffold now supplies a small UNet + VAE with the same architecture shape, critically keeping CatVTON's defining 12-channel inflated input convolution (the person/garment/mask latent concatenation of Chong et al. 2024, "Concatenation Is All You Need"). The smoke latent is 32x32, not 64x64: at 64x64 the resolution-1 self-attention spans 4096 tokens and the per-step attention tensors accumulating on the tape paged the host into a >14 GB swap stall (blame-hang isolated Clone_ShouldProduceIdenticalOutput); at 32x32 the full 13-test suite passes in ~90s. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> Co-authored-by: franklinic <franklin@ivorycloud.com>
Bumps
AiDotNet.Tensors(and the lockstepAiDotNet.Native.OneDNN/OpenBLAS/CLBlast) from 0.92.1 → 0.92.5.Why
0.92.5 is the first release carrying Tensors #583 — validate the prepacked-B GEMM cache against in-place weight mutation. The
SgemmWithCachedBinference path reused a stale pre-packed B panel when a layer's weights were mutated in place without aMarkWeightDirty, soFusedLinearinference returned pre-update results. That was the root cause of #1221 (transformer production-scale convergence failure).Verification
dotnet restoreclean against 0.92.5 (no NU1102); all four lockstep packages confirmed published at 0.92.5.Issue1221suite: 11/11 pass (TransformerProductionScaleConvergenceIssue1221Tests+TransformerProductionScaleHarnessIssue1221Tests).Closes #1221.
🤖 Generated with Claude Code
Summary by CodeRabbit