test: add ProgramSynthesis integration tests - #763
Conversation
Add 50 comprehensive integration tests for the ProgramSynthesis module covering: - Models: CodePosition, CodeSpan, CodeLocation, CodeIssue, Program<T> - Enums: ProgramLanguage, CodeTask, CodeIssueSeverity, SqlDialect, etc. - Execution: CompilationDiagnostic, SqlValue, ProgramExecuteResponse - Results: CodeGenerationResult All tests validate model construction, property access, and default values. Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
|
The latest updates on your projects. Learn more about Vercel for GitHub.
|
Summary by CodeRabbit
✏️ Tip: You can customize this high-level summary in your review settings. WalkthroughA project reference to CLBlast native bindings was removed from the test project, and a comprehensive integration test suite was added for the ProgramSynthesis module, covering type instantiation, property handling, enum values, and cross-component interactions. Changes
Estimated code review effort🎯 3 (Moderate) | ⏱️ ~25 minutes Suggested labels
Poem
🚥 Pre-merge checks | ✅ 2 | ❌ 3❌ Failed checks (2 warnings, 1 inconclusive)
✅ Passed checks (2 passed)
✏️ Tip: You can configure your own custom pre-merge checks in the settings. ✨ Finishing touches
🧪 Generate unit tests (beta)
Comment |
|
Add 50 comprehensive integration tests for the ProgramSynthesis module covering: - Models: CodePosition, CodeSpan, CodeLocation, CodeIssue, Program<T> - Enums: ProgramLanguage, CodeTask, CodeIssueSeverity, SqlDialect, etc. - Execution: CompilationDiagnostic, SqlValue, ProgramExecuteResponse - Results: CodeGenerationResult All tests validate model construction, property access, and default values. Co-authored-by: franklinic <franklin@ivorycloud.com> Co-authored-by: Claude Opus 4.5 <noreply@anthropic.com>
Fixed critical bugs found during production-readiness review: 1. PredictionTypeInference (PR #765): Integer overflow when calculating class label range. maxClass - minClass overflows for extreme values. Fixed by using long for range calculation. 2. GeneticOptimizer (PR #762): IndexOf bug in tournament selection. When population contains duplicate prompts, IndexOf returns first occurrence index, causing wrong fitness selection. Fixed by tracking index directly. 3. NeuralProgramSynthesizer (PR #763): Absolute error comparison fails for large numbers. 1e12 + 0.5 vs 1e12 incorrectly fails with 1e-6 absolute tolerance. Fixed with relative error comparison for large numbers and absolute for small. 4. TrainingMonitor (PR #755): CSV export shows "0" for missing metrics instead of empty string. FirstOrDefault returns default(T) which is 0 for numerics, not null. Fixed by checking if match exists. Added comprehensive tests that expose all bugs and verify fixes. Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
* test: add comprehensive Serving module integration tests (#670) Add 62 integration tests covering: - ContinuousBatcherConfig (defaults, configuration, model presets) - BatchSchedulerConfig (defaults, model-specific configs) - BatchScheduler<T> (scheduling, preemption, cancellation, statistics) - SequenceState<T> (lifecycle, token management, stop conditions) - GenerationRequest<T> (configuration options) - GenerationResult<T> (result data) - BatcherStatistics and SchedulerStatistics - SequenceStatus, StopReason, SchedulingPolicy enums - Full integration workflows Also remove stale CLBlast project reference from test project. Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * test: add comprehensive FineTuning module integration tests (#640) (#753) Add 49 integration tests covering all 16 fine-tuning methods: - SFT (Supervised Fine-Tuning) - DPO (Direct Preference Optimization) - SimPO (Simple Preference Optimization) - ORPO (Odds Ratio Preference Optimization) - IPO (Identity Preference Optimization) - RDPO (Robust Direct Preference Optimization) - KTO (Kahneman-Tversky Optimization) - CPO (Contrastive Preference Optimization) - RLHF-PPO (Reinforcement Learning Human Feedback) - GRPO (Group Relative Policy Optimization) - PRO (Pairwise Ranking Optimization) - RRHF (Rank Responses Human Feedback) - RSO (Statistical Rejection Sampling) - SPIN (Self-Play Fine-Tuning) - CAI (Constitutional AI) Tests cover: - Constructor initialization - Fine-tuning workflow completion - Data validation (SFT, Preference, RL, Ranking) - Serialization/deserialization - Edge cases and parameter handling - Cancellation support Also removes stale CLBlast project reference from test project. Co-authored-by: franklinic <franklin@ivorycloud.com> Co-authored-by: Claude Opus 4.5 <noreply@anthropic.com> * test: add Evaluation integration tests (#655) (#765) Add 51 comprehensive integration tests for the Evaluation module covering: - Enums: PredictionType (BinaryClassification, Regression, MultiClass, MultiLabel) - DefaultModelEvaluator: Construction, options handling - PredictionStatsOptions: Default values, configuration - PredictionTypeInference: Inference from vectors, matrices, tensors Tests validate prediction type inference logic including: - Binary classification detection (0/1 labels) - Multi-class detection (integer labels with low unique ratio) - Regression detection (continuous values, NaN, infinity, high unique ratio) - Multi-label detection (matrix with multiple positives per row) Also fixes stale CLBlast project reference in test project. Co-authored-by: franklinic <franklin@ivorycloud.com> Co-authored-by: Claude Opus 4.5 <noreply@anthropic.com> * test: add ProgramSynthesis integration tests (#665) (#763) Add 50 comprehensive integration tests for the ProgramSynthesis module covering: - Models: CodePosition, CodeSpan, CodeLocation, CodeIssue, Program<T> - Enums: ProgramLanguage, CodeTask, CodeIssueSeverity, SqlDialect, etc. - Execution: CompilationDiagnostic, SqlValue, ProgramExecuteResponse - Results: CodeGenerationResult All tests validate model construction, property access, and default values. Co-authored-by: franklinic <franklin@ivorycloud.com> Co-authored-by: Claude Opus 4.5 <noreply@anthropic.com> * test: add comprehensive integration tests for PromptEngineering module (#762) Add 81 integration tests covering Templates, Analysis, FewShot selectors, Compression, Chains, and Tools components. Tests include: - SimplePromptTemplate: construction, variable extraction, formatting, validation - PromptMetrics: property values and timestamp behavior - PromptIssue and IssueSeverity: issue tracking and severity levels - ValidationOptions: default, strict, and lenient configurations - CompressionResult: token savings, compression ratio, success detection - CompressionOptions: default, aggressive, conservative presets - FixedExampleSelector: example management and selection behavior - ToolRegistry: registration, execution, case-insensitivity, descriptions - SequentialChain: step management, sync/async execution, cancellation - FewShotExample: model properties Also removes stale CLBlast project reference from test project. Closes #666 Co-authored-by: franklinic <franklin@ivorycloud.com> Co-authored-by: Claude Opus 4.5 <noreply@anthropic.com> * test: add comprehensive integration tests for Prototypes module (#761) Add 58 integration tests covering PrototypeVector, PrototypeAdamOptimizer, SimpleLinearRegression, and SimpleNeuralNetwork classes. Tests include: - PrototypeVector: construction, indexing, arithmetic operations, factory methods - PrototypeAdamOptimizer: construction, parameter updates, convergence behavior - SimpleLinearRegression: training, prediction, MSE/R2 computation - SimpleNeuralNetwork: forward/backward passes, XOR training scenario - Cross-component integration scenarios Also removes stale CLBlast project reference from test project. Closes #667 Co-authored-by: franklinic <franklin@ivorycloud.com> Co-authored-by: Claude Opus 4.5 <noreply@anthropic.com> * test: add comprehensive TrainingMonitoring module integration tests (#673) (#755) Add 92 integration tests for the TrainingMonitoring module covering: - TrainingMonitor<T>: session lifecycle, metric logging, progress tracking, speed stats, resource usage, issue detection, and data export - ResourceMonitor: lifecycle, snapshots, history, threshold alerts, events - ExperimentTracker: experiments, runs, parameters, metrics, artifacts, tags, deletion/restoration, persistence, and search - NotificationManager: services, sending, filtering, buffering, events - Dashboard implementations (ConsoleDashboard, HtmlDashboard, LiveDashboard): scalar/histogram logging, hyperparameters, confusion matrix, reports - Cross-module integration tests combining multiple components Also removes stale CLBlast project reference from test project. Co-authored-by: franklinic <franklin@ivorycloud.com> Co-authored-by: Claude Opus 4.5 <noreply@anthropic.com> * Add comprehensive integration tests for AiDotNet.Serving module Added 42 integration tests covering: - ServingOptions configuration defaults and custom values - StartupModel property validation - PerformanceMetrics recording, percentiles, and statistics - ContinuousBatchingStrategy concurrency and adaptive behavior - TimeoutBatchingStrategy timeout-based processing - SizeBatchingStrategy fixed-size batching - AdaptiveBatchingStrategy latency-based adaptation - BucketBatchingStrategy bucket index handling - Padding strategies (Minimal, Bucket, Fixed) data preservation - Configuration enum value verification Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * fix: replace thread.sleep with deterministic timestamp in queue time test Instead of using Thread.Sleep(50) which is flaky under CI load, directly set GenerationStartedAt to a known offset from CreatedAt. This makes the test deterministic and asserts the exact expected duration instead of using >= comparison. Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * fix: use net471-compatible Enum.GetValues in serving tests Replace Enum.GetValues<T>() (NET 5+ only) with Enum.GetValues(typeof(T)).Cast<T>() for .NET Framework 4.7.1 compatibility. Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> --------- Co-authored-by: franklinic <franklin@ivorycloud.com> Co-authored-by: Claude Opus 4.5 <noreply@anthropic.com>
* fix: production bugs in recently merged PRs Fixed critical bugs found during production-readiness review: 1. PredictionTypeInference (PR #765): Integer overflow when calculating class label range. maxClass - minClass overflows for extreme values. Fixed by using long for range calculation. 2. GeneticOptimizer (PR #762): IndexOf bug in tournament selection. When population contains duplicate prompts, IndexOf returns first occurrence index, causing wrong fitness selection. Fixed by tracking index directly. 3. NeuralProgramSynthesizer (PR #763): Absolute error comparison fails for large numbers. 1e12 + 0.5 vs 1e12 incorrectly fails with 1e-6 absolute tolerance. Fixed with relative error comparison for large numbers and absolute for small. 4. TrainingMonitor (PR #755): CSV export shows "0" for missing metrics instead of empty string. FirstOrDefault returns default(T) which is 0 for numerics, not null. Fixed by checking if match exists. Added comprehensive tests that expose all bugs and verify fixes. Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * fix: critical FineTuning bugs - log probability and SFT constructor FineTuningBase.cs: - Implemented ComputeLogProbabilityFromPrediction that was returning 0.0 - Added proper handling for probability distributions (cross-entropy) - Added cosine similarity for embeddings - Added scalar comparison for numeric values - Added string similarity using Levenshtein distance - Fixed single-element array bug (cosine similarity is always 1.0) SupervisedFineTuning.cs: - Added single-parameter constructor for Activator.CreateInstance compatibility - This enables reflection-based instantiation used by test frameworks MergedPRBugFixTests.cs: - Added tests for log probability computation - Added test verifying SFT can be instantiated with single-parameter constructor - All 14 tests passing Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * fix: critical Diagnostics bugs - memory leak, thread safety, variance MemoryTracker.cs: - Added MaxHistorySize property with default of 10000 - Enforces limit to prevent unbounded memory growth in long-running apps - Removes oldest snapshots when limit is reached (FIFO) ProfilerSession.cs: - Fixed thread-unsafe System.Random by using RandomHelper.ThreadSafeRandom - System.Random is NOT thread-safe; concurrent access corrupts internal state - Fixed variance calculation: was population (_m2/_count), now sample (_m2/(_count-1)) - Fixed call stack cleanup: handles out-of-order Stop() calls (e.g., due to exceptions) - Now searches stack for timer instead of only checking top, prevents memory leak MergedPRBugFixTests.cs: - Added 5 tests for Diagnostics bug fixes - All 19 tests passing Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * fix: ModelRegistry production bugs (PR #774) Fixed 5 critical production-readiness bugs: 1. Mutable internal state returned: GetModel, GetModelByStage, and SearchModels now return clones to prevent external modification of internal state 2. Console.WriteLine in library code: Replaced with LoadErrors property for proper diagnostic exposure without polluting stdout 3. TOCTOU race conditions: Fixed file operations in DeleteModelVersion and GetModelCard to use try-catch pattern instead of File.Exists checks 4. DeleteModelVersion didn't validate modelName: Added ValidateModelName call for consistency with other methods 5. Lineage tracking didn't work: _lineage dictionary is now populated when model versions are created via RegisterModel and CreateModelVersion Also added GetInternalModel helper method for internal mutation operations. Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * fix: PR #773 Metrics - add null checks, empty tensor handling, and validation Bugs fixed in Metrics module: - PSNR.ComputeBatch: Added null argument validation - STOI.Compute: Added null argument validation - STOI.ComputeNormalizedCorrelation: Fixed bounds check to include both arrays - SI-SDR.Compute: Added null validation and empty tensor handling - SNR.Compute: Added null argument validation - IoU3D.ComputeBoxIoU: Added null validation and coordinate validation (min <= max) - ChamferDistance.ComputeOneWay: Throws for empty target with non-empty source Added 9 new tests to verify fixes. Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * fix: PR #772 Logging - add null checks, validation, and TOCTOU fix Bugs fixed in Logging module: - SummaryWriter.AddScalars: Added null check for tagScalarDict - SummaryWriter.AddHistogram: Added null check for values arrays - SummaryWriter.AddImage: Validate dataformats parameter (must be CHW or HWC) - SummaryWriter.AddImages: Added null check and parameter validation - SummaryWriter.AddPrCurve: Added null checks, length validation - SummaryWriter.LogWeights: Added null check, handle empty weights array - TensorBoardWriter.WriteEmbedding: Added null check, metadata length validation - TensorBoardWriter.EncodePng: Validate pixels array length - TensorBoardWriter.WriteProjectorConfig: Fixed TOCTOU race condition Added 7 new tests to verify fixes. Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * fix(LanguageModels): add null/empty validation for model names and API version Bug fixes for PR #771 LanguageModels: - OpenAIChatModel: Validate modelName not null/empty (was NullReferenceException) - OpenAIChatModel: Validate maxTokens > 0 early with clear error message - AnthropicChatModel: Validate modelName not null/empty (was NullReferenceException) - AzureOpenAIChatModel: Validate apiVersion not null/empty (caused invalid URL) - AzureOpenAIChatModel: Validate maxTokens > 0 early with clear error message Added 12 tests covering these validation scenarios. Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * fix(JitCompiler): add null checks and remove side effect from Validate Bug fixes for PR #770 JitCompiler: - IRGraph.Validate(): Remove side effect that modified TensorShapes (validation should be read-only) - TensorShapeExtensions.GetElementCount(): Add null check - TensorShapeExtensions.ShapeToString(): Add null check - TensorShapeExtensions.GetShapeHashCode(): Add null check - TensorShapeExtensions.GetShape(): Add null tensor check - IRTypeExtensions.FromSystemType(): Add null Type check Added 10 tests covering these validation scenarios. Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * fix(Interpretability): add null argument validation to helper methods Bug fixes for PR #769 Interpretability: - InterpretabilityMetricsHelper.GetUniqueGroups: Add null check for sensitiveFeature - InterpretabilityMetricsHelper.GetGroupIndices: Add null check for sensitiveFeature - InterpretabilityMetricsHelper.GetSubset: Add null checks for vector and indices - InterpretabilityMetricsHelper.ComputePositiveRate: Add null check for predictions - InterpretabilityMetricsHelper.ComputeTruePositiveRate: Add null checks for predictions and actualLabels - InterpretabilityMetricsHelper.ComputeFalsePositiveRate: Add null checks for predictions and actualLabels - InterpretabilityMetricsHelper.ComputePrecision: Add null checks for predictions and actualLabels - InterpretableModelHelper: Add null checks for model, enabledMethods, and input parameters across all async methods (GetGlobalFeatureImportanceAsync, GetLocalFeatureImportanceAsync, GetShapValuesAsync, GetLimeExplanationAsync, GetPartialDependenceAsync, GetCounterfactualAsync, GetModelSpecificInterpretabilityAsync, GenerateTextExplanationAsync, GetFeatureInteractionAsync, ValidateFairnessAsync, GetAnchorExplanationAsync) Added 10 tests covering null argument validation scenarios. Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * fix(InferenceOptimization): add null validation to optimization graph and node methods PR #768 production-readiness fixes: - OptimizationNode.AddInput: add null check for inputNode - OptimizationNode.RemoveInput: add null check for inputNode - OptimizationNode.ReplaceInput: add null checks for oldInput/newInput - OptimizationGraph.FindNodeById: add null check for id - OptimizationGraph.FindNodesByName: add null check for name - IRDataTypeExtensions.FromSystemType: add null check for type - TensorType.IsBroadcastCompatible: add null check for other - GraphOptimizer.Optimize: add null check for graph - GraphOptimizer.AddPass: add null check for pass Added 14 tests covering all 9 bug fixes and valid input scenarios. Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * fix(HyperparameterOptimization): add null validation to trial pruner and distributions PR #767 production-readiness fixes: - TrialPruner.ReportAndCheckPrune(trial): add null check for trial - TrialPruner.ReportAndCheckPrune(trialId): add null check for trialId - TrialPruner.MarkComplete: add null check for trialId - ContinuousDistribution.Sample: add null check for random - IntegerDistribution.Sample: add null check for random - CategoricalDistribution.Sample: add null check for random - HyperparameterOptimizerBase.FindBestTrial: add null check for completedTrials - HyperparameterOptimizerBase.EvaluateTrialSafely: add null checks for trial, objectiveFunction, parameters Added 13 tests covering all 8 bug fixes and valid input scenarios. Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * fix(ExperimentTracking): fix production bugs in experiment tracking Bug fixes: - StartRun: persist experiment timestamp after Touch() to maintain disk consistency - SearchRuns: validate maxResults > 0 to prevent invalid queries - GetRunDirectory: use TryGetValue with descriptive InvalidOperationException - DeleteExperiment/DeleteRun: handle IOException gracefully when directory deletion fails - LogArtifact: validate extracted filename isn't empty for root paths - LogArtifacts: wrap UnauthorizedAccessException with descriptive message - Add null validation to GetExperiment, GetRun, ListRuns, DeleteExperiment, DeleteRun - Add null validation to SerializeToJson, DeserializeFromJson, GetLatestMetric Added 11 tests covering: - Null argument validation - MaxResults validation - Timestamp persistence verification - Thread safety for concurrent metric logging - Graceful error handling for directory operations Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * fix(DistributedTraining): fix production bugs in distributed communication Bug fixes: - CommunicationManager.Broadcast: add root range validation to prevent invalid operations - CommunicationManager.Scatter: add root range validation to prevent invalid operations - CommunicationManager.ReduceScatter: add null validation for data parameter - ParameterAnalyzer.CalculateDistributionStats: throw for null/empty groups (consistent API) - InMemoryCommunicationBackend.PerformReduction: validate all vectors have same length - InMemoryCommunicationBackend.Receive: validate message size BEFORE dequeuing (prevents data loss) - ShardingConfiguration factory methods: add null validation for better error messages Added 10 tests covering: - Root validation in Broadcast/Scatter - Null data validation in ReduceScatter - Null/empty groups in ParameterAnalyzer - Null backend in ShardingConfiguration factory methods - Single-process AllReduce optimization path Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * fix(Reasoning): fix production bugs in reasoning module Bug fixes: - ReasoningChain.AddStep: add null validation for step parameter - DepthFirstSearch.SearchAsync: add null validation for generator, evaluator, config (consistent with other search algorithms) - MonteCarloTreeSearch constructor: validate numSimulations >= 1 and explorationConstant >= 0 - BreadthFirstSearch.CollectAllNodes: use iterative approach instead of recursion to prevent StackOverflow on deep trees Added 9 tests covering: - Null step validation in ReasoningChain - Step number auto-increment - MCTS constructor parameter validation - ThoughtNode path reconstruction - ThoughtNode leaf/root detection - ReasoningConfig default values Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * fix(Serialization): add validation for negative dimensions in JSON converters - VectorJsonConverter: validate that length is non-negative - TensorJsonConverter: validate that shape array is not empty - TensorJsonConverter: validate that all shape dimensions are non-negative - Added 8 tests for serialization validation Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * fix(Serving): add parameter validation to batching strategies and padding Add production bug fixes for PR #758 (Serving module): Batching Strategies: - ContinuousBatchingStrategy: validate maxConcurrency >= 1, minWaitMs >= 0, targetLatencyMs > 0 - TimeoutBatchingStrategy: validate timeoutMs >= 0, maxBatchSize >= 1 - AdaptiveBatchingStrategy: validate minBatchSize >= 1, maxBatchSize >= minBatchSize, maxWaitMs >= 0, targetLatencyMs > 0, latencyToleranceFactor > 0 - SizeBatchingStrategy: validate batchSize >= 1, maxWaitMs >= 0 - BucketBatchingStrategy: validate maxBatchSize >= 1, maxWaitMs >= 0, bucket values > 0 Monitoring: - PerformanceMetrics: validate maxSamples >= 1, maxQueueDepthSamples >= 1 Padding Strategies (MinimalPaddingStrategy, BucketPaddingStrategy, FixedSizePaddingStrategy): - PadBatch: validate no null vectors in input array - UnpadBatch: validate originalLengths are non-negative - FixedSizePaddingStrategy: validate fixedLength > 0 - BucketPaddingStrategy: validate bucketSizes not null/empty Added 25 validation tests across BatchingStrategyTests, PaddingStrategyTests, and PerformanceMetricsTests. Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * fix(Tokenization): add parameter validation to tokenizers and vocabulary Add production bug fixes for PR #757 (Tokenization module): Vocabulary: - AddTokens: validate tokens parameter is not null - Constructor(Dictionary): validate tokenToId parameter is not null Tokenizers: - BpeTokenizer.Train: validate corpus not null, vocabSize >= 1 - WordPieceTokenizer.Train: validate corpus not null, vocabSize >= 1 - WordPieceTokenizer constructor: validate maxInputCharsPerWord >= 1 - CharacterTokenizer.Train: validate corpus not null, minFrequency >= 1 MidiTokenizer: - Constructor: validate ticksPerBeat >= 1, numVelocityBins >= 1 - CreateREMI: validate ticksPerBeat >= 1, numVelocityBins >= 1 - CreateCPWord: validate ticksPerBeat >= 1, numVelocityBins >= 1 - CreateSimpleNote: validate ticksPerBeat >= 1 TokenizationConfig: - ParallelBatchThreshold: validate value >= 1 via property setter Added 17 validation tests across BpeTokenizerTests, CharacterTokenizerTests, WordPieceTokenizerTests, SpecializedTokenizerTests, and VocabularyTests. Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * fix(Tools): add parameter validation and robust type conversion - Add upper bound validation (100) for topK in VectorSearchTool and RAGTool - Add validation that topKAfterRerank cannot exceed topK in RAGTool - Make ToolBase TryGetInt/TryGetDouble/TryGetBool handle type conversion errors gracefully - Add 34 unit tests covering parameter validation and edge cases PR #756 bug fixes - prevent performance issues from excessive topK values and improve robustness when receiving invalid JSON property types. Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * fix(DistributedTraining): fix initialization order and add parameter validation Bugs fixed: 1. Initialization order bug in derived classes (DivideByZeroException): - Base class was calling InitializeSharding() before derived class fields were set - Fix: Remove InitializeSharding() from base constructor, each derived class now calls it explicitly at the end of their constructor 2. ShardingConfiguration missing learningRate validation: - Added validation that learningRate > 0 3. PipelineParallelModel missing microBatchSize validation: - Added validation that microBatchSize >= 1 4. HybridShardedModel missing parallelism size validation: - Added validation that pipelineParallelSize >= 1 - Added validation that tensorParallelSize >= 1 5. Inconsistent learning rate usage: - DDPModel, FSDPModel, PipelineParallelModel, HybridShardedModel were using hardcoded 0.01 instead of Config.LearningRate - Fixed to use Config.LearningRate consistently Affected files: - ShardedModelBase.cs - removed InitializeSharding() call from constructor - DDPModel.cs, FSDPModel.cs, ZeRO1Model.cs, ZeRO2Model.cs - added InitializeSharding() call - TensorParallelModel.cs, PipelineParallelModel.cs, HybridShardedModel.cs - same + validation - ShardingConfiguration.cs - added learningRate validation Tests: Added 26 validation tests in DistributedTrainingValidationTests.cs Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * fix(LoRA): fix matrix indexing bugs and initialization order issues - Add input/output size validation in LoRALayer constructor - Add pruningInterval validation in AdaLoRAAdapter to prevent division by zero - Fix matrix indexing in MergeWeights (returns [inputSize, outputSize]): - LoRAAdapterBase.MergeToDenseOrFullyConnected - StandardLoRAAdapter.MergeToOriginalLayer - QLoRAAdapter.MergeToOriginalLayer - DoRAAdapter.Forward() and MergeToOriginalLayer - Fix DefaultLoRAConfiguration.CreateAdapter to handle different constructor signatures - Fix VeRAAdapter initialization order bug: - Move scaling vector init to CreateLoRALayer (called before ParameterCount) - Add UpdateParametersFromLayers override for VeRA-specific parameter sync - Add 22 validation tests covering all bug fixes Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * fix(Autodiff): add numerical stability validation to tensor operations - Add division by zero validation in TensorOperations.Divide - Add non-positive value validation in TensorOperations.Log - Add negative value validation in TensorOperations.Sqrt - Handle sqrt(0) edge case in backward pass (use 0 instead of infinity) - Fix null axes handling in TensorOperations.Sum OperationParams - Add segmentSize validation in GradientCheckpointing.SequentialCheckpoint - Add 19 validation tests covering all bug fixes Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * fix: address pr review comments and code scanning alerts - ExperimentTracker: add logging to IOException catch block as comment indicated - MonteCarloTreeSearch: add NaN/Infinity guards to explorationConstant validation - TensorJsonConverter: remove validation rejecting empty shapes (scalars are valid) - BucketBatchingStrategy: add validation for empty bucket boundaries array - ToolBase: update exception handling to catch JsonSerializationException - LoRALayer: update XML docs to reflect actual exception types thrown - MergedPRBugFixTests: update size-mismatch test to actually verify behavior - FineTuningBase: fix generic catch clauses and collection equality check - ShardedModelBase: use lazy initialization to avoid virtual calls in constructor Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * fix: correct misleading comment and tighten assertion in sum operation - Update comment in TensorOperations.cs to accurately describe OperationParams behavior - Tighten Sum_NullAxes_OperationParamsHandledCorrectly test assertion to properly verify that Axes key is NOT present when axes parameter is null Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * fix: remove virtual calls from distributed training model constructors - Remove direct InitializeSharding() calls from DDPModel, FSDPModel, ZeRO1Model, and ZeRO2Model constructors - Remove redundant InitializeSharding() call from TensorParallelModel's OnBeforeInitializeSharding() method - All models now rely on lazy initialization via EnsureShardingInitialized() in the base class to avoid virtual calls in constructors Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * fix: address coderabbit review comments for pr #783 - TensorOperations.cs: convert if/else to ternary for operationparams assignment - HybridShardedModel.cs: move pendingconfig.value = null to onbeforeinitializesharding where it is consumed (lazy init compatibility) - FineTuningBase.cs: move string check before ienumerable check since strings implement ienumerable<char> Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * fix: address additional coderabbit review comments - TensorOperations.cs: only emit "Axes" metadata when non-null AND non-empty (empty array means sum-all like null) - FineTuningBase.cs: treat null-null elements as matches in sequence matching fallback Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * fix: guard numeric conversion against mixed runtime types Check both prediction and target are numeric before calling Convert.ToDouble to avoid throwing when TOutput is object and types don't match. Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * chore: remove accidentally committed _playground_publish folder Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * feat: add comprehensive ActiveLearning integration tests Add 62 integration tests covering: - EntropySampling (8 tests) - UncertaintySampling (5 tests) - BALD (4 tests) - RandomSampling (3 tests) - MarginSampling (2 tests) - LeastConfidenceSampling (2 tests) - VariationRatios (2 tests) - DiversitySampling (5 tests with all methods/metrics) - CoreSetSelection (2 tests) - HybridSampling (3 tests with all combination methods) - InformationDensity (2 tests) - DensityWeightedSampling (1 test) - ExpectedModelChange (2 tests) - BatchBALD (2 tests) - QueryByCommittee (2 tests) - Edge cases and mathematical validation (6 tests) Tests include: - Correct batch size validation - Null argument handling - Score range validation - Mathematical properties (entropy of uniform/certain distributions) - Diversity selection across clusters Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * feat: add comprehensive ContinualLearning integration tests Add 53 integration tests covering: - ElasticWeightConsolidation (EWC): 12 tests - SynapticIntelligence (SI): 5 tests - MemoryAwareSynapses (MAS): 4 tests - GradientEpisodicMemory (GEM): 4 tests - LearningWithoutForgetting (LwF): 4 tests - OnlineEWC: 3 tests - ExperienceReplay: 3 tests - PackNet: 3 tests - ProgressiveNeuralNetworks: 2 tests - GenerativeReplay: 2 tests - AveragedGEM (A-GEM): 3 tests - Edge cases and cross-strategy validation: 8 tests Tests verify: - Constructor initialization - BeforeTask/AfterTask lifecycle - ComputeLoss returns non-negative values - ModifyGradients produces valid output - Reset clears stored data - Multiple sequential tasks work correctly - Lambda property can be modified - Null argument handling Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * feat: add comprehensive curriculum learning integration tests - Add 35 integration tests for CurriculumLearning components - Test schedulers: LinearScheduler, SelfPacedScheduler, CompetenceBasedScheduler - Test difficulty estimators: LossBased, ConfidenceBased, TransferBased, ExpertDefined, Ensemble - Test edge cases: empty arrays, zero epochs, reset behavior - Fix bug in CurriculumSchedulerBase.GetIndicesAtPhase that crashed on empty arrays Note: Tests document a bug with generic T? default values - for unconstrained generics, T? with default value is 0.0 for value types, not null. Tests work around this by providing explicit values for optional parameters. Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * test: add comprehensive selfsupervisedlearning integration tests Add 52 integration tests covering the SelfSupervisedLearning module: - NT-Xent Loss tests (4 tests): temperature effects, gradient computation - InfoNCE Loss tests (5 tests): memory bank integration, accuracy metrics - BYOL Loss tests (5 tests): cosine similarity, symmetric loss computation - Linear Projector tests (6 tests): shape validation, gradient backprop - MLP Projector tests (5 tests): batch norm, training mode, backward pass - Symmetric Projector tests (6 tests): predictor head, combined operations - Memory Bank tests (11 tests): FIFO queue, momentum updates, sampling - Momentum Encoder static method tests (3 tests): cosine schedule - Edge cases and error handling tests (6 tests) Uses RandomHelper for secure random number generation and proper Tensor API patterns for cross-framework compatibility. Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * test: add nestedlearning integration tests for associativememory and contextflow Add 30 integration tests covering the NestedLearning module: AssociativeMemory tests (12 tests): - Constructor validation and initialization - Associate/Retrieve with dimension validation - Capacity limit enforcement (FIFO) - Association matrix updates with Hebbian learning - Clear/Reset functionality - Multiple associations and large capacity handling ContextFlow tests (15 tests): - Constructor and matrix initialization - PropagateContext with level validation - ComputeContextGradients backpropagation - UpdateFlow transformation matrix updates - GetContextState and CompressContext operations - Reset clears all context states - Independent states across multiple levels Integration tests (3 tests): - Combined AssociativeMemory + ContextFlow workflow - Large capacity stress testing - Sequential propagation state accumulation Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * test: add mixedprecision integration tests (39 tests) Adds comprehensive integration tests for MixedPrecision module: - LossScaler: scaling, unscaling, overflow detection, dynamic scaling - MixedPrecisionConfig: defaults follow NVIDIA recommendations - MixedPrecisionContext: FP32/FP16 weight management, gradient preparation - Full workflow tests: training iterations, overflow recovery Closes #642 Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * test: add comprehensive integration tests for adversarial robustness module - Add 83 integration tests for AdversarialRobustness module: - FGSM attack tests (8 tests) - PGD attack tests (4 tests) - CW attack tests (3 tests) - AutoAttack tests (2 tests) - Adversarial training defense tests (7 tests) - Randomized smoothing certification tests (8 tests) - Interval bound propagation tests (7 tests) - CROWN verification tests (6 tests) - Safety filter tests (11 tests) - Rule-based content classifier tests (11 tests) - Integration scenarios (6 tests) - Edge cases and error handling (10 tests) - Fix bug in FGSMAttack: add null check for trueLabel parameter to throw ArgumentNullException instead of InvalidOperationException Closes #631 Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * test: fix randomhelper usage in finetuning integration tests Replace new Random(42) with RandomHelper.CreateSeededRandom(42) to follow project security standards for random number generation. Closes #640 Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * test: fix lora integration tests and loralayer exception consistency - Fix LoRALayer to throw ArgumentOutOfRangeException consistently for all invalid rank values (was throwing ArgumentException for rank > min(in, out)) - Update test to expect ArgumentOutOfRangeException - Replace new Random(42) with RandomHelper.CreateSeededRandom(42) Closes #641 Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * test: add knowledgedistillation integration tests Comprehensive integration tests covering: - DistillationLoss: constructor, compute loss/gradient, edge cases - All distillation strategies: Feature, Attention, Contrastive, Probabilistic, Hybrid, Curriculum, Adaptive, Variational, NeuronSelectivity, Relational, SimilarityPreserving, FlowBased, FactorTransfer - DistillationStrategyFactory: all strategy types - DistillationForwardResult and DistillationCheckpointConfig - IntermediateActivations: add/get/count - Numerical stability and edge case testing Total: 85 tests, 0 bugs found Closes #636 Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * test: add comprehensive physicsinformed integration tests Add 94 integration tests covering the PhysicsInformed module: - PhysicsInformedLoss tests (loss computation, gradients, edge cases) - PDE tests (HeatEquation, WaveEquation, PoissonEquation, BurgersEquation, AllenCahnEquation, KdV, AdvectionDiffusion) - PINN tests (PhysicsInformedNeuralNetwork, VariationalPINN, DeepRitzMethod) - Neural Operator tests (FourierNeuralOperator, FourierLayer) - ScientificML tests (HamiltonianNeuralNetwork) - TrainingHistory, PDEDerivatives, PDEResidualGradient tests - Edge cases and numerical stability tests - Integration workflow tests - Serialization tests Closes #637 Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * fix: move tensor type check before text fullname check in datamodalitydetector The Tensor check was happening after the fullName-based Text check, which caused Tensor<T> types to be incorrectly detected as Text modality since the fullName could contain substrings matching other checks. Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * test: add comprehensive augmentation integration tests with 106 tests Coverage includes: - Image augmentations (GaussianNoise, Brightness, Contrast, Cutout, RandomCrop, etc.) - Audio augmentations (AudioNoise, TimeStretch, PitchShift, etc.) - Video augmentations (TemporalFlip, FrameDropout, SpeedChange) - Text augmentations (RandomDeletion, RandomInsertion, SynonymReplacement, etc.) - Object detection augmentations (BoundingBox, Keypoint transformations) - Compose and auto-augment pipelines - DataModalityDetector type detection Tests verify correct behavior for each augmentation category including: - Apply with probability, deterministic behavior, edge cases - Parameter validation, composition, and context management Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * test: add comprehensive federated learning integration tests Adds 64 integration tests for the FederatedLearning module covering: - Aggregation strategies: FedAvg, FedProx, FedBN with weighted averaging - Byzantine-robust aggregators: Krum, MultiKrum, Bulyan, TrimmedMean, WinsorizedMean, GeometricMedian, RFA - Client selection strategies: UniformRandom, WeightedRandom, Stratified, AvailabilityAware, PerformanceAware, Clustered - Privacy mechanisms: GaussianDifferentialPrivacy with clipping and noise - Privacy accounting: BasicComposition and RDP privacy accountants - Cryptography: HKDF key derivation - Server optimizers: FedAdam, FedAdagrad, FedYogi, FedAvgM - Heterogeneity corrections: SCAFFOLD, FedNova, FedDyn - Secure aggregation: SecureAggregationVector, ThresholdSecureAggregationVector - Additional tests: GaussianDifferentialPrivacyVector Tests verify mathematical correctness and proper API behavior. Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * test: add checkpoint management integration tests Adds 38 integration tests for the CheckpointManagement module covering: - CheckpointManager construction and directory creation - Auto-checkpointing configuration (save frequency, keep last, save on improvement) - ShouldAutoSaveCheckpoint logic (frequency-based and improvement-based triggers) - UpdateAutoSaveState for tracking last save step and best metric values - AutoCheckpointState properties and ToString formatting - Thread-safe concurrent configuration updates and state reads - ListCheckpoints, LoadLatestCheckpoint, LoadBestCheckpoint edge cases - CleanupOldCheckpoints and CleanupKeepBest cleanup strategies - MetricOptimizationDirection enum values - Path validation and nested directory support Tests verify proper state tracking for minimization and maximization scenarios. Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * test: add 64 integration tests for Configuration module - AutoMLBudgetOptions: default values, property setting, all presets - RLTrainingOptions: default values, property setting, callbacks - RLCheckpointConfig: default values, property setting - RLEarlyStoppingConfig: default values, generic type support - ExplorationScheduleConfig: default values, all decay types - InferenceOptimizationConfig: default values, validation, all enum values - ResNetConfiguration: variants, block counts, expansion, factory methods - BenchmarkingOptions: default values, federated configs - CurriculumLearningOptions: schedule types, difficulty estimators - Supporting options classes: SelfPacedOptions, CompetenceBasedOptions Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * feat: add dataprocessor integration tests and fix empty matrix bug - Add 24 comprehensive integration tests for DataProcessor module - Tests cover DefaultDataPreprocessor, DataProcessorOptions, SplitData - Tests validate preprocessing pipeline with Matrix, Vector, and Tensor types - Fix DivideByZeroException in FeatureSelectorHelper.CreateFilteredData when handling empty matrices (0 rows) - The fix returns a properly dimensioned empty matrix instead of attempting to call FromColumns on empty column vectors Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * feat: add dataversioncontrol integration tests (62 tests) - Add comprehensive integration tests for DataVersionControl module - Tests cover versioning, hashing, integrity verification, run linking - Tests cover tagging, lineage tracking, snapshots, and persistence - Tests verify thread safety with concurrent version creation - Tests model classes: DatasetVersion, DatasetLineage, DatasetStatistics, etc. Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * test: add 65 integration tests for dataversioning module Tests cover: - Constructor and directory structure creation - CreateDataset with validation, metadata, and duplicate handling - AddVersion with files, directories, and deduplication - GetVersion by ID, version number, and "latest" - ListVersions and ListDatasets with ordering - GetDataPath with directory structure preservation - CompareVersions detecting additions, removals, modifications - DeleteVersion and DeleteDataset with file cleanup - RecordLineage and GetLineage with recursive upstream resolution - Persistence across instance restarts - Model classes (DatasetInfo, DataVersion, DataFileInfo, DataVersionDiff, DataLineage) - Thread safety for concurrent operations - Edge cases (large file count, empty directories, special characters) Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * fix: correct json parsing bug in agent response extraction The ExtractJsonFromResponse method in all agent classes used a non-greedy regex pattern that incorrectly matched nested JSON objects. For example, with input: {"reasoning_steps": [...], "tool_calls": [{"tool_name": "X"}]} The regex would match up to the first "}" (end of tool_calls inner object) instead of the outer closing brace. Fixed by implementing proper brace-balancing algorithm that: - Tracks brace count while respecting string boundaries - Handles escape sequences within strings - Returns the complete outermost JSON object Also added comprehensive integration tests for all agent types: - Agent (ReAct pattern): 15 tests - ChainOfThoughtAgent: 8 tests - PlanAndExecuteAgent: 7 tests - RAGAgent: 12 tests - AgentBase: 3 tests - JSON parsing edge cases: 4 tests - Thread safety: 1 test Total: 50 new tests for the Agents module Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * fix: correct training data dimensions in vision and tabular benchmarking tests The tests were training KNN models on 1-feature data but then running benchmarks that generate test data with different feature dimensions: - CIFAR10/CIFAR100: 3072 features (32x32x3 pixels flattened) - TabularNonIID: FeatureCount features (3 in this test) Fixed by providing training data that matches the benchmark feature dimensions: - CIFAR tests: 2 samples with 3072 features each (normalized pixel values) - TabularNonIID test: 3 samples with 3 features each Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * test: add comprehensive benchmarking integration tests Add 98 integration tests for the Benchmarking module covering: - BenchmarkSuiteRegistry: display names, suite categories, mappings - BenchmarkReport: construction, serialization, aggregation, statistics - BenchmarkMetricValue: creation, edge cases, formatting - BenchmarkExecutionStatus: all status values and transitions - BenchmarkSuite enum: all 23 benchmark suites - BenchmarkMetric enum: all 8 metric types - BenchmarkSuiteKind enum: all 3 suite kinds - Model classes: BenchmarkSuiteReport, BenchmarkDataSelectionSummary Tests verify behavior for: - Enum value coverage for all benchmark-related enums - Display name generation and formatting - Report creation and metric aggregation - Suite category classification (ReasoningSuite vs DatasetSuite) - Edge cases like empty reports and default values Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * test: add comprehensive computervision integration tests (66 tests) Tests cover: - BoundingBox format conversions (XYXY, XYWH, CXCYWH, YOLO) - BoundingBox IoU, Area, Clip, IsValid operations - Detection and DetectionResult classes - DetectionStatistics and BatchDetectionResult - NMS (standard, class-aware, batched) - IoU variants (IoU, GIoU, DIoU, CIoU) - GIoULoss for bounding box regression - SORT tracker with Kalman filtering - Track, TrackingOptions, TrackingResult classes - End-to-end integration scenarios Closes #647 Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * fix: net471 compatibility for string.split and enum.getvalues - Use char array overload for string.Split with StringSplitOptions - Use typeof() overload for Enum.GetValues instead of generic version Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * fix: address pr review comments for integration tests - Fix divide-by-zero guard in ActiveLearning CreatePoolWithKnownUncertainty - Add missing assertion in RandomSampling_InformativenessScores test - Add meaningful assertion in Rotation_ApplyWithTargets test - Rename AllRegularizationStrategies to CoreRegularizationStrategies for accuracy - Fix reflection test to assert if type exists but methods don't - Fix greedy regex in ChainOfThoughtAgent JSON extraction - Add BindingFlags for non-public property reflection in Benchmarking - Add cleanup for default checkpoints directory - Fix culture-invariant decimal comparison in AutoCheckpointState test Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * fix: address additional pr review comments - Rename misleading test names in DataVersionControl - Make MockChatModel thread-safe with Interlocked and ConcurrentBag - Guard against deleting pre-existing default directories in DataVersioning - Add delays to prevent timestamp-tie flakes in ordered tests Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * fix: replace generic catch clauses with specific exception types in finetuning - Replace `catch (Exception ex) when (...)` patterns with specific exception type catches (InvalidCastException, FormatException, OverflowException) - Extract ComputeSequenceMatchLogProbability helper to reduce code duplication - Fix Equals on collections issue by using EqualityComparer<TOutput>.Default Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * fix: replace generic catch clause with specific exception types in ssl tests Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * fix: guard against single-class divide-by-zero in mock model Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * fix: update lora tests to expect argumentoutofrangeexception The code correctly throws ArgumentOutOfRangeException (more specific than ArgumentException) for invalid rank values. Tests now expect the correct exception type. Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * fix: update vblora test to expect argumentoutofrangeexception The test Constructor_WithRankExceedingBankSizeA_ThrowsArgumentException was expecting ArgumentException but the code correctly throws ArgumentOutOfRangeException which is more specific for range validation. Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> --------- Co-authored-by: franklinic <franklin@ivorycloud.com> Co-authored-by: Claude Opus 4.5 <noreply@anthropic.com>
…s (4F/4G/4H) Consumes Tensors PR ooples/AiDotNet.Tensors#763 (compiled backward with createGraph=true, DP-SGD helper, multi-slot persistent inputs). ## 4G — WGAN-GP correctness fix + fused disc steps (5 files) Root cause fixed for AiDotNet issue #1844: every ComputeGradientPenalty implementation used the inner GradientTape's ComputeGradients WITHOUT createGraph=true. The inner backward's ops didn't record on the outer tape, so inputGradients had no GradFn chain back to the discriminator weights. Effect: the gradient penalty was included in the loss VALUE but NOT in the disc's gradient — WGAN-GP silently degraded to plain WGAN with no 1-Lipschitz enforcement. Fix (5 files: CTGAN, DPCTGAN, CopulaGAN, CausalGAN, TableGAN): add createGraph: true to the inner ComputeGradients so the outer tape can differentiate the penalty through the disc weights. Also wired 4 disc steps through the fused compiled plan (CausalGAN, CTGAN, CopulaGAN, TableGAN — DPCTGAN's disc is DP so goes via 4H). Packs (real, fake) into one persistent input along axis 0; loss closure splits scores back out and computes wasserstein + λ·GP. The fused path activates once Tensors PR #763 lands (its 4C change removes the !createGraph gate at GradientTape.cs:727); until then these wire-ups fall through to the (now-corrected) eager tape path. ## 4H — DP-SGD wire-ups (2 files) Local AiDotNet.Training.DpSgdStep<T> — drop-in mirror of the same helper in Tensors PR #763 (Engines.Training.DpSgdStep<T>). Enables the wire-ups to land in this PR without waiting on the Tensors NuGet publish. Both implementations enforce the Abadi 2016 §3 Algorithm 1 clip-BEFORE-aggregate order via their structure. * DPCTGAN.TrainDiscriminatorStepBatchedDP — routes the per-example WGAN-GP + DP-SGD critic through DpSgdStep<T>. Objective: Wasserstein + λ·GP per example, clipped per-example, aggregated, noised, averaged. * MedSynth.TrainDiscriminatorStepPerExampleDPSGD — routes the per- example non-saturating BCE + DP-SGD critic through the same helper. Once Tensors PR #763 merges + NuGet publishes, the local mirror can be swapped for the Tensors version by changing the `using` — the API surface is identical by design. ## 4F — Batched-per-element diffusion foundation DiffusionModelBase.Train samples per-batch-element timesteps for rank ≥ 2 inputs (Ho et al. 2020 canonical pattern; HuggingFace diffusers reference), routing through NoiseSchedulerBase.AddNoiseBatched + DiffusionModelBase.PredictNoiseBatched. Rank-1 unbatched inputs keep the scalar path for backward compat. net471 doesn't support default interface implementations, so AddNoiseBatched lives on NoiseSchedulerBase<T> (as virtual with a default per-element delegate) rather than on the INoiseScheduler<T> interface. DiffusionModelBase's fused-training wire-up (via Tensors PR #763's PersistentInputRegistry) is queued for the NuGet-bump follow-up commit — the API foundation is already in place. Builds green net471 / net8.0 / net10.0. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
… + consumer wire-ups Adds three fused-plan training primitives (mirrors of AiDotNet.Tensors PR #763) and wires the DP-SGD consumers through them: * DpSgdFusedStep<T> — Abadi 2016 §3 Algorithm 1 per-example clip-before-aggregate. Runs each per-example forward+backward through a compiled plan (LR=0 so weights don't drift between replays), then computes the global L2 norm, clips, noises, and aggregates via vectorized IEngine ops (TensorMultiply/ReduceSum for the norm, TensorMultiplyScalar/TensorAdd for the accumulator, TensorRandomNormalInto for on-device Gaussian noise). Returns aggregated gradients so the caller's configured optimizer (Adam/AdamW/SGD/...) applies the update — no hardcoded SGD step inside. * WganGpFusedStep<T> — Gulrajani 2017 WGAN-GP critic step. Composes E[D(fake)] − E[D(real)] + λ·(‖∇_x̃ D(x̃)‖₂ − 1)² inside one compiled plan with the inner ∇_x̃ D(x̃) recorded via createGraph=true so it differentiates into disc weights (issue #1844 fix). OnesLike uses vectorized Engine.TensorFill. * MultiSlotFusedStep<T> — N-slot persistent input mechanism with plan cache keyed by composite shape + parameter identity. Slots refreshed via AsSpan().CopyTo(AsWritableSpan()) — no per-element loops on the hot path. Consumer wire-ups (this PR): * DPCTGANGenerator.TrainDiscriminatorStepBatchedDP — primary path routes through DpSgdFusedStep.TryStep, falls back to the existing ComputePerExampleNoisedGradients when the fused path can't engage (non-GPU host / compilation disabled). * MedSynthGenerator.TrainDiscriminatorStepPerExampleDPSGD — same primary/fallback layering. * Both legacy fallback loops (ComputePerExampleNoisedGradients + MedSynth eager) are now themselves vectorized: global-L2-norm via ReduceSum(g·g), clipped accumulation via TensorMultiplyScalar+TensorAdd, on-device Gaussian noise via TensorRandomNormalInto+TensorAdd — no scalar per-element loops anywhere on the DP-SGD path. Codebase convention adopted: * Class-scope `private static IEngine Engine => AiDotNetEngine.Current;` and `private static readonly INumericOperations<T> Ops = MathHelper.GetNumericOperations<T>();` on each fused-step class (matches the ~20 activation/optimizer/etc. bases). No IEngine or INumericOperations<T> threaded through method signatures. Skipped (follow-ups filed): * WGAN-GP consumer rewire → #1845 (blocked on IGradientBasedOptimizer<T> config accessor extension; consumers already correct via GpuResidentFusedStep + local createGraph=true ComputeGradientPenalty). * Diffusion consumer rewire → #1846 (blocked on per-generator forward-refactor to move _timestepProjection inside compiled plan while treating raw timestep as a persistent slot). Verification: * net8.0 + net471 + net10.0 all build clean on both AiDotNet and AiDotNet.Tensors. * Old DpSgdStep.cs deleted (replaced by DpSgdFusedStep). Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
…s (4F/4G/4H) Consumes Tensors PR ooples/AiDotNet.Tensors#763 (compiled backward with createGraph=true, DP-SGD helper, multi-slot persistent inputs). ## 4G — WGAN-GP correctness fix + fused disc steps (5 files) Root cause fixed for AiDotNet issue #1844: every ComputeGradientPenalty implementation used the inner GradientTape's ComputeGradients WITHOUT createGraph=true. The inner backward's ops didn't record on the outer tape, so inputGradients had no GradFn chain back to the discriminator weights. Effect: the gradient penalty was included in the loss VALUE but NOT in the disc's gradient — WGAN-GP silently degraded to plain WGAN with no 1-Lipschitz enforcement. Fix (5 files: CTGAN, DPCTGAN, CopulaGAN, CausalGAN, TableGAN): add createGraph: true to the inner ComputeGradients so the outer tape can differentiate the penalty through the disc weights. Also wired 4 disc steps through the fused compiled plan (CausalGAN, CTGAN, CopulaGAN, TableGAN — DPCTGAN's disc is DP so goes via 4H). Packs (real, fake) into one persistent input along axis 0; loss closure splits scores back out and computes wasserstein + λ·GP. The fused path activates once Tensors PR #763 lands (its 4C change removes the !createGraph gate at GradientTape.cs:727); until then these wire-ups fall through to the (now-corrected) eager tape path. ## 4H — DP-SGD wire-ups (2 files) Local AiDotNet.Training.DpSgdStep<T> — drop-in mirror of the same helper in Tensors PR #763 (Engines.Training.DpSgdStep<T>). Enables the wire-ups to land in this PR without waiting on the Tensors NuGet publish. Both implementations enforce the Abadi 2016 §3 Algorithm 1 clip-BEFORE-aggregate order via their structure. * DPCTGAN.TrainDiscriminatorStepBatchedDP — routes the per-example WGAN-GP + DP-SGD critic through DpSgdStep<T>. Objective: Wasserstein + λ·GP per example, clipped per-example, aggregated, noised, averaged. * MedSynth.TrainDiscriminatorStepPerExampleDPSGD — routes the per- example non-saturating BCE + DP-SGD critic through the same helper. Once Tensors PR #763 merges + NuGet publishes, the local mirror can be swapped for the Tensors version by changing the `using` — the API surface is identical by design. ## 4F — Batched-per-element diffusion foundation DiffusionModelBase.Train samples per-batch-element timesteps for rank ≥ 2 inputs (Ho et al. 2020 canonical pattern; HuggingFace diffusers reference), routing through NoiseSchedulerBase.AddNoiseBatched + DiffusionModelBase.PredictNoiseBatched. Rank-1 unbatched inputs keep the scalar path for backward compat. net471 doesn't support default interface implementations, so AddNoiseBatched lives on NoiseSchedulerBase<T> (as virtual with a default per-element delegate) rather than on the INoiseScheduler<T> interface. DiffusionModelBase's fused-training wire-up (via Tensors PR #763's PersistentInputRegistry) is queued for the NuGet-bump follow-up commit — the API foundation is already in place. Builds green net471 / net8.0 / net10.0. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
… + consumer wire-ups Adds three fused-plan training primitives (mirrors of AiDotNet.Tensors PR #763) and wires the DP-SGD consumers through them: * DpSgdFusedStep<T> — Abadi 2016 §3 Algorithm 1 per-example clip-before-aggregate. Runs each per-example forward+backward through a compiled plan (LR=0 so weights don't drift between replays), then computes the global L2 norm, clips, noises, and aggregates via vectorized IEngine ops (TensorMultiply/ReduceSum for the norm, TensorMultiplyScalar/TensorAdd for the accumulator, TensorRandomNormalInto for on-device Gaussian noise). Returns aggregated gradients so the caller's configured optimizer (Adam/AdamW/SGD/...) applies the update — no hardcoded SGD step inside. * WganGpFusedStep<T> — Gulrajani 2017 WGAN-GP critic step. Composes E[D(fake)] − E[D(real)] + λ·(‖∇_x̃ D(x̃)‖₂ − 1)² inside one compiled plan with the inner ∇_x̃ D(x̃) recorded via createGraph=true so it differentiates into disc weights (issue #1844 fix). OnesLike uses vectorized Engine.TensorFill. * MultiSlotFusedStep<T> — N-slot persistent input mechanism with plan cache keyed by composite shape + parameter identity. Slots refreshed via AsSpan().CopyTo(AsWritableSpan()) — no per-element loops on the hot path. Consumer wire-ups (this PR): * DPCTGANGenerator.TrainDiscriminatorStepBatchedDP — primary path routes through DpSgdFusedStep.TryStep, falls back to the existing ComputePerExampleNoisedGradients when the fused path can't engage (non-GPU host / compilation disabled). * MedSynthGenerator.TrainDiscriminatorStepPerExampleDPSGD — same primary/fallback layering. * Both legacy fallback loops (ComputePerExampleNoisedGradients + MedSynth eager) are now themselves vectorized: global-L2-norm via ReduceSum(g·g), clipped accumulation via TensorMultiplyScalar+TensorAdd, on-device Gaussian noise via TensorRandomNormalInto+TensorAdd — no scalar per-element loops anywhere on the DP-SGD path. Codebase convention adopted: * Class-scope `private static IEngine Engine => AiDotNetEngine.Current;` and `private static readonly INumericOperations<T> Ops = MathHelper.GetNumericOperations<T>();` on each fused-step class (matches the ~20 activation/optimizer/etc. bases). No IEngine or INumericOperations<T> threaded through method signatures. Skipped (follow-ups filed): * WGAN-GP consumer rewire → #1845 (blocked on IGradientBasedOptimizer<T> config accessor extension; consumers already correct via GpuResidentFusedStep + local createGraph=true ComputeGradientPenalty). * Diffusion consumer rewire → #1846 (blocked on per-generator forward-refactor to move _timestepProjection inside compiled plan while treating raw timestep as a persistent slot). Verification: * net8.0 + net471 + net10.0 all build clean on both AiDotNet and AiDotNet.Tensors. * Old DpSgdStep.cs deleted (replaced by DpSgdFusedStep). Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
…mirrors 0.113.0 publishes Tensors PR #763 — this branch's gating dependency. #763 makes the compiled/persistent backward honor createGraph=true (GradientTape.Compute- Gradients previously gated the compiled path on !createGraph), so the WGAN-GP gradient penalty's inner backward now differentiates into the disc weights through the fused compiled plan instead of silently returning zeros (issue #1844). Against 0.111.2 the fused GPU-resident WGAN-GP path would have silently degraded to plain WGAN. Also carries #765 (compiled ReduceMax axis fill) and #764 (resident-param fused-Adam fix). Now that 0.113.0 is published to NuGet, retire the local mirror classes that stood in until it shipped (they were 1:1 copies with identical public APIs), and point every consumer at the Engines.Training.* versions in the package: - Delete src/Training/{DpSgdFusedStep,WganGpFusedStep,MultiSlotFusedStep}.cs. - Repoint all call sites (added by #1847's fused-primitive centralization) to AiDotNet.Tensors.Engines.Training.*: * WganGpFusedStep — CausalGAN, CTGAN, CopulaGAN, TableGAN * MultiSlotFusedStep — DiffusionModelBase, CSDI, TabDDPM, TabSyn * DpSgdFusedStep — DPCTGAN, MedSynth Signatures are identical, so this is a pure namespace swap. GpuResidentFusedStep stays consumer-side (not part of #763; it composes the loss — including the createGraph=true gradient penalty — and drives the generic fused plan). Bumps AiDotNet.Native.OneDNN / OpenBLAS to 0.113.0 in lockstep (both published). Builds green on net471 / net8.0 / net10.0. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
…for fused paths) 0.113.0 publishes Tensors PR #763 — this branch's gating dependency. #763's key fix (4C): the compiled/persistent backward now honors createGraph=true (GradientTape. ComputeGradients previously gated the compiled path on !createGraph), so the WGAN-GP gradient penalty's inner backward differentiates into the disc weights through the fused compiled plan instead of silently returning zeros (issue #1844). Against 0.111.2 the fused GPU-resident WGAN-GP path would have silently degraded to plain WGAN. The src/Training fused-primitive mirrors (WganGpFusedStep, MultiSlotFusedStep, DpSgdFusedStep) stay as the consumer-side primitives that #1847/#1848 centralize every consumer through; they call the public engine API, so the bump alone routes them onto #763's fixed compiled backward. Also picks up #765 (compiled ReduceMax axis fill) and #764 (resident-param fused-Adam fix). Native OneDNN/OpenBLAS/CLBlast bumped in lockstep (all published). Builds green on net8.0. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
…1843) * feat(training): gPU-resident fused step for non-TS single-net models Kicks off the non-time-series GPU-residency sweep. Adds a shared helper for classes that don't inherit from NeuralNetworkBase or TimeSeriesModelBase but still want their Train() to route forward + backward + optimizer through the compiled fused plan (weights / activations / Adam moments resident on-device across the whole step). * src/Training/GpuResidentFusedStep.cs — shared helper. TryResolveOptimizerConfig maps a runtime IGradientBasedOptimizer to the fused-plan OptimizerType + hyperparameters (Adam / AdamW / SGD supported via case-insensitive class-name match + reflection over Options.InitialLearningRate / Beta1 / Beta2 / Epsilon / WeightDecay). TryStep is the one-shot entry. Wired the following single-net models through the helper (mirrors NeuralNetworkBase.TrainWithFusedStep): * GraphClassificationModel (NN base + GNN + pooling + cross-entropy) — routes tape training through the fused plan when float + DirectGpu + compilation are live; falls through to the existing eager tape+optimizer loop on any failure. * LinkPredictionModel (NN base + GNN + node embeddings + BCE) — same wire-up as GraphClassificationModel. * FourierNeuralOperator (lift → FourierLayers → project, hardcoded SGD + learning rate 0.001) — captures the whole spectral-conv chain in a fused SGD plan; falls through to the in-place SGD loop below. GAN-family generators (CTGAN, DPCTGAN, PATEGAN, ...), CLAPModel (learned scalar outside layer hierarchy), AutoDiffTabGenerator (per-sample variable shape defeats plan replay) and other non-conforming shapes are queued for follow-up commits. The shared helper is the common seam so they all wire the same way once their per-step structure fits the fused-plan contract (constant shape + Adam/AdamW/SGD optimizer + ITrainableLayer-carried params). Builds green net471 / net8.0 / net10.0. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> * feat(training): gPU-resident fused step for Finance foundation forecasters Wires TFC, TOTEM, CSDI, MQCNN through the compiled fused SGD plan (all four use in-place SGD with lr=0.001 as their eager path). Extends the shared helper with an inline IsGpuResidentAvailable check so callers don't need to duplicate the availability gate. * GpuResidentFusedStep — adds IsGpuResidentAvailable static property (mirrors TimeSeriesModelBase.CanTrainOnGpu / NeuralNetworkBase.CanTrainOnGpu). TryStep short-circuits when unavailable so callers stay on their eager path. * TFC — supervised + contrastive branches fused into one closure; the fused plan captures both losses in a single backward. * TOTEM — reconstruction + VQ commitment terms fused; the recompute-loss closure re-derives the commitment loss on each replay for graph parity. * CSDI — denoising-score-matching: the forward closure re-samples timestep + noise each step (via ComputeDenoisingPairTape) so replay produces fresh training data even though the plan's captured graph shape is fixed. * MQCNN — multi-quantile pinball loss captured directly via ComputeMultiQuantilePinballLossTape. All four fall through to the existing in-place SGD path on any fused-plan failure (unsupported op, non-compilable graph, etc). Builds green net471 / net8.0 / net10.0. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> * feat(training): gPU-resident fused generator step for GAN-family synthetic generators + TVAE Wires the generator side of the dual-network training loop through the fused compiled plan for 9 SyntheticData generators. Each covers a distinct pattern: * PATEGAN — per-sample loop; fused plan compiles on first sample, replays across the batch with fresh noise per iteration. * MedSynth — batched noise -> decoder -> discriminator-frozen -> non-saturating log-sigmoid loss. * CausalGAN — batched noise -> generator -> optional causal-structure -> output activations -> disc-frozen -> -avgFake. * OCTGAN — per-sample loop; noise -> gen -> disc-frozen embedding -> SVDD dist². * TableGAN — noise + real batch -> gen -> composite loss (fake-scores + information-loss + optional classification). * CTGAN — the tricky one: pack (noise, cond, mask) into one persistent input tensor so the closure can slice cond and mask back out on replay; captures both -avgFake AND conditional cross-entropy on the fused plan. * DPCTGAN — same pattern as CTGAN but simpler (no mask, no CE). * CopulaGAN — same pattern as DPCTGAN. * TVAE — encoder + reparam + decoder + composite ELBO (recon + KL) all fused; Reparameterize re-samples inside the closure so training stays stochastic. Discriminator/critic/student layers are NOT passed to the fused step (their weights stay frozen on the gen step, matching the eager path semantics). Falls through to the eager tape+optimizer path on any failure. The discriminator STEPS of these GANs are not yet wired — they need a separate pass because their loss involves gradient penalty (WGAN-GP) or per-example DP clipping (DPCTGAN) which don't compose cleanly with the single-closure fused step yet. Builds green net471 / net8.0 / net10.0. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> * feat(training): gPU-resident fused step for TabSyn VAE + MisGAN mask/data gens Two more SyntheticData generators covered — each covers a distinct pattern: * TabSyn.TrainVAEBatch — per-row VAE ELBO training. Fused plan compiles on first row, replays across the batch with refreshed input per iteration. Reparameterize re-samples inside the closure so training stays stochastic. * MisGAN.TrainDataGeneratorStep — dual-noise generator (data noise + mask noise) packed into a single persistent input tensor; closure slices them back out on replay. Loss = -E[D_x(fakeRow ⊙ fakeMask)]. * MisGAN.TrainMaskGeneratorStep — simpler single-noise version; per-sample loop with fused plan replay across the batch. The DataDiscriminator step (WGAN-style critic + weight clipping) and mask discriminator step aren't wired — they need separate plans and the ClipWeights post-step interacts poorly with the fused optimizer's in-place update. Follow-up. Builds green net471 / net8.0 / net10.0. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> * feat(training): gPU-resident fused step for TimeGAN P1/P2 + MedSynth non-DP disc * TimeGAN Phase 1 (TrainReconstructionStepBatched) — embedder + recovery reconstruction MSE. Straight single-network fused-step conversion. * TimeGAN Phase 2 (TrainSupervisedStepBatched) — supervisor's next-step MSE in the embedded space, with the embedder frozen (its layers not in the trainable set for this step). Two-tensor closure (ht, htNext). * MedSynth non-DP disc step — pack (real, fake) into a single input tensor along axis 0; slice back in the loss for BCE-real + BCE-fake. Generator runs OUTSIDE the fused plan so its weights stay frozen on the critic step. TimeGAN Phase 3 (adversarial + supervised joint) and MedSynth DP-SGD paths still use their eager tape — the WGAN-GP gradient penalty and per-example DP clipping don't compose with the single-closure fused optimizer step. Builds green net471 / net8.0 / net10.0. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> * feat(training): extra-tensors parameter on fused-step API Extends the fused compiled-plan training step to accept raw trainable tensors that aren't naturally carried by an ITrainableLayer (e.g. GOGGLE's soft adjacency A; CLAP's learned _logTemperature scalar). Solves the SOLID concerns raised in review of the ITrainableLayer-wrapper alternative: * ISP — no forcing raw tensors to implement the full ILayer surface with no-op Forward / SetTrainingMode / GetParameterGradients members * LSP — no risk of "layer" wrappers with identity Forward diverging from the layer contract's behavioral expectations elsewhere in the codebase Wire-up: * CompiledTapeTrainingStep.TryStepWithFusedOptimizer — new optional extraTensors param. Threaded into the dedup-aware parameter collection (CollectDeduplicatedParametersWithExtras) so extras get moment buffers, gradient accumulation, and GPU-residency in exactly the same code path that layer-carried params do. * GpuResidentFusedStep.TryStep — plumbs extras through. A callsite with only extras (no layers) is a valid config for models whose whole trainable surface is raw tensors. Dedup is by Tensor<T> reference across BOTH sources, so an extra tensor that also happens to be layer-carried is registered exactly once — the same shared/tied-weight protection the layer-only collector provides. Enables Phase 4E (GOGGLE + CLAP wire-ups). Builds green net471 / net8.0 / net10.0. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> * feat(training): wire GOGGLE + CLAP through fused-step extra-tensors API Consumes Phase 4A's extraTensors parameter to route models with raw trainable state (outside the ITrainableLayer contract) through the fused compiled plan. * GOGGLE — soft adjacency A (initialised in InitializeModel, projected after each optimizer step via ProjectAdjacencyConstraints) is passed through extraTensors. The fused optimizer allocates its moment buffer, accumulates gradients, and applies the update in-place. Reparameterize re-samples inside the forward closure so training stays stochastic; ProjectAdjacencyConstraints runs after the fused step to keep the adjacency on the valid soft-adjacency manifold. * CLAP — the learned _logTemperature scalar (Radford 2021 / Wu 2023 contrastive-alignment temperature) goes through extraTensors. Both audio and text encoders are in the layers list; the fused step handles all three parameter classes uniformly. Both fall through to the existing eager tape+optimizer path on any failure (unsupported optimizer, non-compilable graph, etc). Builds green net471 / net8.0 / net10.0. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> * feat(diffusion): batched per-element timesteps for DDPM-canonical training Adds industry-standard batched-per-element timestep API to the diffusion stack (Ho et al. 2020 §3, HuggingFace diffusers reference). Previously DiffusionModelBase.Train sampled ONE timestep for the whole batch — reducing the signal-to-noise diversity of the training gradient. Canonical DDPM training samples a distinct timestep per batch element. * INoiseScheduler.AddNoiseBatched(cleanBatch, noiseBatch, timesteps) — default interface implementation delegates to the scalar AddNoise per element for backward compatibility. Concrete schedulers can override with a fused batched implementation. * NoisePredictorBase.PredictNoiseBatched(noisyBatch, timesteps, conditioning) — virtual with default slice-then-call-scalar implementation. Subclasses that want a fused batched forward override this to keep the training loop on-device. * DiffusionModelBase.PredictNoiseBatched — model-level counterpart with the same default slice-then-call-scalar behavior. * DiffusionModelBase.Train detects batched vs rank-1 input and routes through the batched noise scheduler + batched predictor when input.Rank >= 2. Rank-1 unbatched inputs stay on the scalar path for backward compatibility. This is the foundation for the industry-exceeding "batched-per-element + fused-resident" diffusion training path — see follow-up commits that wire DiffusionModelBase.Train through the fused compiled plan. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> * feat(training): wGAN-GP correctness fix + fused disc + DP-SGD wire-ups (4F/4G/4H) Consumes Tensors PR ooples/AiDotNet.Tensors#763 (compiled backward with createGraph=true, DP-SGD helper, multi-slot persistent inputs). ## 4G — WGAN-GP correctness fix + fused disc steps (5 files) Root cause fixed for AiDotNet issue #1844: every ComputeGradientPenalty implementation used the inner GradientTape's ComputeGradients WITHOUT createGraph=true. The inner backward's ops didn't record on the outer tape, so inputGradients had no GradFn chain back to the discriminator weights. Effect: the gradient penalty was included in the loss VALUE but NOT in the disc's gradient — WGAN-GP silently degraded to plain WGAN with no 1-Lipschitz enforcement. Fix (5 files: CTGAN, DPCTGAN, CopulaGAN, CausalGAN, TableGAN): add createGraph: true to the inner ComputeGradients so the outer tape can differentiate the penalty through the disc weights. Also wired 4 disc steps through the fused compiled plan (CausalGAN, CTGAN, CopulaGAN, TableGAN — DPCTGAN's disc is DP so goes via 4H). Packs (real, fake) into one persistent input along axis 0; loss closure splits scores back out and computes wasserstein + λ·GP. The fused path activates once Tensors PR #763 lands (its 4C change removes the !createGraph gate at GradientTape.cs:727); until then these wire-ups fall through to the (now-corrected) eager tape path. ## 4H — DP-SGD wire-ups (2 files) Local AiDotNet.Training.DpSgdStep<T> — drop-in mirror of the same helper in Tensors PR #763 (Engines.Training.DpSgdStep<T>). Enables the wire-ups to land in this PR without waiting on the Tensors NuGet publish. Both implementations enforce the Abadi 2016 §3 Algorithm 1 clip-BEFORE-aggregate order via their structure. * DPCTGAN.TrainDiscriminatorStepBatchedDP — routes the per-example WGAN-GP + DP-SGD critic through DpSgdStep<T>. Objective: Wasserstein + λ·GP per example, clipped per-example, aggregated, noised, averaged. * MedSynth.TrainDiscriminatorStepPerExampleDPSGD — routes the per- example non-saturating BCE + DP-SGD critic through the same helper. Once Tensors PR #763 merges + NuGet publishes, the local mirror can be swapped for the Tensors version by changing the `using` — the API surface is identical by design. ## 4F — Batched-per-element diffusion foundation DiffusionModelBase.Train samples per-batch-element timesteps for rank ≥ 2 inputs (Ho et al. 2020 canonical pattern; HuggingFace diffusers reference), routing through NoiseSchedulerBase.AddNoiseBatched + DiffusionModelBase.PredictNoiseBatched. Rank-1 unbatched inputs keep the scalar path for backward compat. net471 doesn't support default interface implementations, so AddNoiseBatched lives on NoiseSchedulerBase<T> (as virtual with a default per-element delegate) rather than on the INoiseScheduler<T> interface. DiffusionModelBase's fused-training wire-up (via Tensors PR #763's PersistentInputRegistry) is queued for the NuGet-bump follow-up commit — the API foundation is already in place. Builds green net471 / net8.0 / net10.0. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> * fix(review): resolve 12 CodeRabbit findings on PR #1843 Addresses every unresolved review comment on the non-TS GPU-residency PR. ## Correctness fixes * DiffusionModelBase.Train — RecomputeForward closure now dispatches isBatched ? PredictNoiseBatched(inp, timesteps) : PredictNoise(inp, timestep), so optimizer re-evaluation matches the recorded batched forward objective. * NoisePredictorBase.PredictNoiseBatched — detect batch-aligned conditioning (leading dim == batchSize) and slice per element; preserve shared conditioning when the leading dim doesn't match. Fixes shape-mismatch / conditioning-shared bugs in classifier-free-guidance-style predictors. * CSDI.Train — remove the fused-plan branch. ComputeDenoisingPairTape samples fresh (timestep, noise) each call, and calling it in both Fwd and Loss produces independent samples that don't match. The compiled plan can't refresh the RNG per replay either. Path stays on the eager tape until Tensors' PersistentInputRegistry (PR ooples/AiDotNet.Tensors#763) lands. * TFC.Train — ComputeContrastiveLossTape now runs INSIDE ForwardCombined so it consumes the current-step persistent input (`inp`), not the outer `input` which would freeze at compile time. Closure-captured local; Loss reads it with a null-guard covering the Fwd-then-Loss ordering invariant. * TOTEM.Train — capture the commitment tensor from ForwardNativeForTrainingWithCommitment's first call; reuse it in Loss instead of running the quantizer a second time. Without this, the compiled path performs an EMA SetCodebookValue update TWICE per step and diverges from the eager path. * NoiseSchedulerBase.AddNoiseBatched — validate full shape parity (rank + every dim), not only the leading batch dim. Prevents indexing beyond noiseBatch's span when a caller passes [B, smaller...] shapes. ## Correctness cleanup * CTGAN.TrainGeneratorStepBatched — capture (act, condFromInput, maskFromInput) from Fwd's single generator pass; reuse in Loss so the conditional-CE term doesn't re-run GeneratorForwardWithResidualBatched + ApplyOutputActivationsBatched (was doubling the per-step generator forward cost). * TabSynGenerator.TrainVAEBatch — capture (mean, logVar) from Fwd's single encoder pass; reuse in Loss so the encoder doesn't run twice per row. Also: if the fused step fails mid-batch after fusedEngaged=true, drop to the eager path for the remaining rows (no row is silently skipped). * FourierNeuralOperator.TapeTrainStep — remove the duplicate _fourierLayers loop in both the fused-path allTrainable collection AND the eager paramList. Fourier layers are already registered into Layers at construction, so iterating them again double-registered each Fourier parameter and drove the eager SGD to apply the update twice per step. ## API cleanup * GpuResidentFusedStep<T> — public → internal. Aligns with CompiledTapeTrainingStep<T> which is already internal; keeps in-assembly training plumbing out of the public API surface. Also fixed an inadvertent null-forgiving operator (`!`) usage that slipped into the review-fix edits. All new nullable annotations use explicit null-guards with descriptive InvalidOperationException on invariant violation (per CLAUDE.md's "never use null-forgiving operators" rule). INoiseScheduler default-interface issue (net471 incompatibility) was already fixed in an earlier commit (AddNoiseBatched moved to NoiseSchedulerBase). Builds green net471 / net8.0 / net10.0. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> * feat(training): vectorized fused DP-SGD/WGAN-GP/multi-slot primitives + consumer wire-ups Adds three fused-plan training primitives (mirrors of AiDotNet.Tensors PR #763) and wires the DP-SGD consumers through them: * DpSgdFusedStep<T> — Abadi 2016 §3 Algorithm 1 per-example clip-before-aggregate. Runs each per-example forward+backward through a compiled plan (LR=0 so weights don't drift between replays), then computes the global L2 norm, clips, noises, and aggregates via vectorized IEngine ops (TensorMultiply/ReduceSum for the norm, TensorMultiplyScalar/TensorAdd for the accumulator, TensorRandomNormalInto for on-device Gaussian noise). Returns aggregated gradients so the caller's configured optimizer (Adam/AdamW/SGD/...) applies the update — no hardcoded SGD step inside. * WganGpFusedStep<T> — Gulrajani 2017 WGAN-GP critic step. Composes E[D(fake)] − E[D(real)] + λ·(‖∇_x̃ D(x̃)‖₂ − 1)² inside one compiled plan with the inner ∇_x̃ D(x̃) recorded via createGraph=true so it differentiates into disc weights (issue #1844 fix). OnesLike uses vectorized Engine.TensorFill. * MultiSlotFusedStep<T> — N-slot persistent input mechanism with plan cache keyed by composite shape + parameter identity. Slots refreshed via AsSpan().CopyTo(AsWritableSpan()) — no per-element loops on the hot path. Consumer wire-ups (this PR): * DPCTGANGenerator.TrainDiscriminatorStepBatchedDP — primary path routes through DpSgdFusedStep.TryStep, falls back to the existing ComputePerExampleNoisedGradients when the fused path can't engage (non-GPU host / compilation disabled). * MedSynthGenerator.TrainDiscriminatorStepPerExampleDPSGD — same primary/fallback layering. * Both legacy fallback loops (ComputePerExampleNoisedGradients + MedSynth eager) are now themselves vectorized: global-L2-norm via ReduceSum(g·g), clipped accumulation via TensorMultiplyScalar+TensorAdd, on-device Gaussian noise via TensorRandomNormalInto+TensorAdd — no scalar per-element loops anywhere on the DP-SGD path. Codebase convention adopted: * Class-scope `private static IEngine Engine => AiDotNetEngine.Current;` and `private static readonly INumericOperations<T> Ops = MathHelper.GetNumericOperations<T>();` on each fused-step class (matches the ~20 activation/optimizer/etc. bases). No IEngine or INumericOperations<T> threaded through method signatures. Skipped (follow-ups filed): * WGAN-GP consumer rewire → #1845 (blocked on IGradientBasedOptimizer<T> config accessor extension; consumers already correct via GpuResidentFusedStep + local createGraph=true ComputeGradientPenalty). * Diffusion consumer rewire → #1846 (blocked on per-generator forward-refactor to move _timestepProjection inside compiled plan while treating raw timestep as a persistent slot). Verification: * net8.0 + net471 + net10.0 all build clean on both AiDotNet and AiDotNet.Tensors. * Old DpSgdStep.cs deleted (replaced by DpSgdFusedStep). Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> * fix(review): tFC + TOTEM preprocessing rewritten with traceable engine ops Resolves the two BLOCKING CodeRabbit findings on PR #1843: both fused-plan training paths were freezing preprocessing into the compiled plan because their inner ops used host-side .Data.Span loops, so later replays reused the first batch's normalized values / spectrum / argmin decisions instead of recomputing per batch. TFC — ApplyInstanceNormalization (RevIN) + ComputeFrequencyRepresentation: * ApplyInstanceNormalization now delegates to a new stateless NormalizeWithStats that returns (normalized, mean, std) as tensors. Under the hood: ReduceMean + ReduceVariance + TensorSqrt + TensorBroadcastSubtract + TensorBroadcastDivide. All ops record on the tape and re-execute per replay under the compiled plan. * DenormalizeForecast likewise delegates to DenormalizeForecastWithStats, which takes mean/std as explicit tensor parameters. ForwardNative threads them through as locals so the compile-mode replay uses CURRENT-step stats instead of frozen trace-time values. * _revinMean/_revinStd (Vector<T> scalars) → _revinMeanTensor/_revinStdTensor (Tensor<T>? nullable) — kept for the abstract override's external callers, but the fused path never reads them. * ComputeFrequencyRepresentation now uses Engine.RFFT for batched real FFT → reshape to [B, halfN, 2] pairs → TensorMultiply + ReduceSum(axis=2) for magSquared → TensorSqrt + TensorMultiplyScalar(1/n) for the one-sided magnitude spectrum → TensorSlice + TensorFlip + TensorConcatenate to mirror bins [1..n-halfN] into the tail. Handles even and odd n identically to the old scalar impl. * Fused-plan fast path in TFC.Train restored — now safe because both preprocessing methods trace correctly. TOTEM — VectorQuantize (VQ-VAE) traceable rewrite + post-Step EMA: * New VectorQuantizeTraceable returns (quantized, commitmentLoss, argmin, head) using Engine ops end-to-end: TensorBroadcastSubtract + TensorMultiply + ReduceSum for distances, TensorArgMin along the codebookSize axis, per-c TensorSliceAxis + TensorIndexSelectDiff + TensorStack for the gather, Engine.StopGradient for the straight-through estimator, ReduceSum-based commitment loss weighted by β/totalLen. * EMA moved OUT of the compiled forward into a new UpdateCodebookEMA(head, argmin). Called POST-Step by the fused path with the trace-time graph-node references — their .Data reflects the LAST replay so the update lands exactly once per batch (matches CodeRabbit's "EMA must execute exactly once per batch" contract). Under the compiled plan, argmin/head are refreshed by every _plan.Step() so post-Step reads see the current batch's values. * UpdateCodebookEMA expresses the per-codebook scatter as: current codebook slice + TensorScatterAdd((1-decay)·(head - gathered), argmin) → new slice, then TensorConcatenate across codebook axis into the full [numCodebooks, codebookSize, codebookDim] tensor, then Engine.TensorCopy back into the _codebooks tensor object to preserve identity (future reads via the same reference see the update). * Legacy VectorQuantize is now a thin adapter around VectorQuantizeTraceable for callers that don't need the extras. * ForwardNativeForTrainingWithCommitment delegates to a new ForwardNativeForTrainingWithVQExtras that exposes the argmin/head; the original (forecast, commitmentLoss) contract is preserved for callers that don't need EMA state. * Fused-plan fast path in TOTEM.Train restored — now safe because VQ is fully traceable AND the EMA runs exactly once per batch in post-Step eager code. No new Tensors primitives were needed — the engine already has RFFT, ReduceMean/Variance, TensorSqrt, TensorBroadcastSubtract/Divide, TensorFlip, TensorConcatenate, TensorArgMin, TensorIndexSelectDiff, TensorSliceAxis, TensorStack, StopGradient, TensorScatterAdd, and TensorCopy, all in the autodiff/compile registry per OpRegistry.cs. Verification: * net8.0 + net471 + net10.0 all build clean. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> * feat(training): centralize WGAN-GP + diffusion consumers through fused primitives (#1847) Closes #1845 and #1846. Routes all four WGAN-GP critics through WganGpFusedStep (the fused-plan primitive from PR #1843) and adds MultiSlotFusedStep wire-ups to the three tabular diffusion consumers plus a base-class opt-in hook for DiffusionModelBase subclasses. ## Optimizer config plumbing Zero new API surface: IFusedOptimizerSpec.TryGetFusedOptimizerConfig already exposes (OptimizerType, LR, Beta1, Beta2, Epsilon, WeightDecay, Schedule, UseBf16Moments). Widened NeuralNetworkBase<T>.TryMapToFusedOptimizerConfig from private to internal so the sibling generators in the same assembly can reuse the existing helper. ## WGAN-GP consumers (#1845) * CTGANGenerator, CopulaGANGenerator, TableGANGenerator, CausalGANGenerator — each critic training method now attempts WganGpFusedStep.TryStep FIRST with the discriminator's optimizer hyperparameters extracted via TryMapToFusedOptimizerConfig, falls back to the existing GpuResidentFusedStep path (secondary fused), then the eager tape (final fallback). The ε ∈ [0, 1]^B epsilon sampler uses Engine.TensorRandomUniformRange to match each critic's local ComputeGradientPenalty behavior. * Non-Adam optimizers (Lion, LBFGS) that don't implement IFusedOptimizerSpec cleanly fall through — TryMapToFusedOptimizerConfig returns false and the code path skips to GpuResidentFusedStep as before. ## Diffusion consumers (#1846) * TabDDPMGenerator — refactored to expose a slot-based forward (BuildTabDDPMSlots + DenoiserForwardFromTensors + ComputeDiffusionLossTapeFromTensors). Per-row TrainBatch loop now attempts MultiSlotFusedStep with (numNoisy, actualNoise, catNoisy, catClean, rawSinusoidalTimeEmbed) as persistent slots. The learnable _timestepProjection stays INSIDE the compiled forward closure so its weights participate in the backward pass. Plan is compiled once on the first row and replayed via slot-data refresh for subsequent rows. * TabSynGenerator — TrainDiffusionBatch's per-row loop wired with MultiSlotFusedStep on (noisyLatent, actualNoise, projectedTimeEmbed). Matches the existing eager path's semantic that _timestepProjection is NOT in _diffMLPLayers (kept detached in the eager path too), so the projected embedding is precomputed host-side per row and passed as slot data. * Finance/Forecasting/Foundation/CSDI — - ApplyInstanceNormalization rewritten with traceable engine ops (ReduceMean + ReduceVariance + TensorSqrt + broadcast subtract/divide) — same pattern as the TFC RevIN fix. The previous `.Data.Span` per-batch loop froze at trace time. - New BuildCsdiSlots + DenoiserForwardFromSlots express the DDPM x_t formation and packed denoising input via TensorConcatenate + engine scalar multiplies. Replaces the `.Data.Span[i] = xt[0, i]` fill that baked the trace batch's x_t into the compiled plan. - Train() attempts MultiSlotFusedStep first, falls back to the existing eager ComputeDenoisingPairTape path when the fused path can't engage. * Diffusion/DiffusionModelBase — - New opt-in `protected virtual bool SupportsFusedDenoising => false;` property. Base default is false so no existing subclass changes behavior. - Train() attempts MultiSlotFusedStep when SupportsFusedDenoising is true AND the training optimizer maps cleanly to a fused config. Slots: (noisySample, noise). Loss = MSE(pred, noise). QAT shadow restoration is preserved on the fused-success path. - Subclasses with fully-traceable PredictNoise / PredictNoiseBatched (e.g. after auditing to remove `.Data.Span` host loops) can opt in via a single-line override; no infrastructure changes needed elsewhere. ## Verification * net8.0, net471, net10.0 all build clean. * No API surface changes on IGradientBasedOptimizer<T> — the existing IFusedOptimizerSpec interface (already implemented by all fuse-able optimizers) provided everything needed. * All consumers preserve eager fallback path for non-fuse-able optimizers and non-GPU hosts. Co-authored-by: franklinic <franklin@ivorycloud.com> Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com> * feat(training): wire additional WGAN-GP + diffusion consumers through fused primitives (#1848) Extends PR #1847 with three additional consumer wire-ups discovered during a comprehensive audit against the primitives added in PR #1843. Each consumer attempts the fused path first via the existing IFusedOptimizerSpec / TryMapToFusedOptimizerConfig plumbing and falls back to eager tape training when the fused path can't engage. ## WGAN-GP consumers (additional beyond #1845) * NeuralNetworks/WGANGP.cs — standalone WGAN-GP class. Previously used the legacy flat-vector round-trip (three Critic.Predict calls per step, GetParameterGradients + host-side vector combine + UpdateCriticWithOptimizer). TrainCriticBatchWithGP now attempts WganGpFusedStep.TryStep first, using Critic.Layers as the parameter source and a per-batch ε ~ U(0, 1) sampler matching each critic's local ComputeGradientPenalty behavior. Falls back to the legacy path when the critic's optimizer has no fused-kernel mapping. ## Diffusion consumers (additional beyond #1846) * NeuralNetworks/SyntheticData/FinDiffGenerator.cs — per-row DDPM training (Sattarov et al. 2023). TrainBatch caches a MultiSlotFusedStep across rows so the compiled plan is built once and replayed via slot-data refresh for subsequent rows. Slots: (packedDenoiserInput, targetNoise). Falls back to the existing Train(input, targetNoise) call on miss. * NeuralNetworks/SyntheticData/AutoDiffTabGenerator.cs — per-row DDPM training with a custom TapeStepOver optimizer step. Same MultiSlotFusedStep caching pattern as FinDiff. Slots: (denoiserInput, targetNoise). Falls back to the existing tape-based TapeStepOver path on miss. ## Explicitly not wired in this PR (documented) * NeuralNetworks/GenerativeAdversarialNetwork.cs — uses BCE loss with optional GP regularization (a separate auxiliary optimizer step), NOT Wasserstein + GP. Wiring WganGpFusedStep would change training semantics from BCE to Wasserstein. Requires a separate design decision. * Finance/Forecasting/Foundation/{CCDM,MGTSD,TSDiff,TimeDiff,TimeGrad}.cs and Finance/Probabilistic/DiffusionTS.cs — each uses .Data.Span host-side packs for the denoising input (same trace-freeze issue TFC/CSDI hit in PR #1843). Each needs a TFC/CSDI-scale traceable rewrite BEFORE fused wiring is safe. Deferred to individual per-consumer PRs. * MetaLearning/Algorithms/{MetaDDPMAlgorithm,MetaDMAlgorithm,MetaDiffAlgorithm}.cs — hand-rolled Vector<T> params with index arithmetic, NOT the ILayer/Engine/tape infrastructure. Would need full rewrite to fit MultiSlotFusedStep. * DiffusionModelBase subclasses (DDPMModel, DiffWaveModel, LatentDiffusionModelBase) — the SupportsFusedDenoising opt-in hook (added in PR #1847) is available, but flipping it safely requires per-class audit of PredictNoise → UNet traceability. Left as follow-up work; the mechanism is in place. ## Verification * net8.0, net471, net10.0 all build clean. * No API surface changes. Co-authored-by: franklinic <franklin@ivorycloud.com> Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com> * chore(deps): bump AiDotNet.Tensors 0.111.2 -> 0.113.0 (activates #763 for fused paths) 0.113.0 publishes Tensors PR #763 — this branch's gating dependency. #763's key fix (4C): the compiled/persistent backward now honors createGraph=true (GradientTape. ComputeGradients previously gated the compiled path on !createGraph), so the WGAN-GP gradient penalty's inner backward differentiates into the disc weights through the fused compiled plan instead of silently returning zeros (issue #1844). Against 0.111.2 the fused GPU-resident WGAN-GP path would have silently degraded to plain WGAN. The src/Training fused-primitive mirrors (WganGpFusedStep, MultiSlotFusedStep, DpSgdFusedStep) stay as the consumer-side primitives that #1847/#1848 centralize every consumer through; they call the public engine API, so the bump alone routes them onto #763's fixed compiled backward. Also picks up #765 (compiled ReduceMax axis fill) and #764 (resident-param fused-Adam fix). Native OneDNN/OpenBLAS/CLBlast bumped in lockstep (all published). Builds green on net8.0. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: franklinic <franklin@ivorycloud.com> Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>



Summary
Test Coverage
Linked Issue
Closes #665
Test plan
🤖 Generated with Claude Code