Skip to content

test: add Evaluation integration tests - #765

Merged
ooples merged 1 commit into
masterfrom
test/evaluation-integration-tests
Jan 24, 2026
Merged

ooples merged 1 commit into
masterfrom
test/evaluation-integration-tests

Conversation

@ooples

@ooples ooples commented Jan 23, 2026

Copy link
Copy Markdown
Owner

Summary

  • Adds 51 comprehensive integration tests for the Evaluation module
  • Tests cover PredictionType enum, DefaultModelEvaluator, PredictionStatsOptions, and PredictionTypeInference
  • Validates prediction type inference logic for various input types

Test Coverage

  • PredictionType enum: All 4 values (BinaryClassification, Regression, MultiClass, MultiLabel)
  • DefaultModelEvaluator: Default construction, custom options, null options handling
  • PredictionStatsOptions: Default values, ConfidenceLevel, LearningCurveSteps
  • PredictionTypeInference (via reflection):
    • Binary classification: 0/1 labels, all zeros, all ones, near-integer values
    • Multi-class: integer labels with contiguous range and low unique ratio
    • Regression: continuous values, NaN, infinity, high unique ratio
    • Multi-label: matrix with multiple positives per row
    • Matrix inputs: single column, multi-column, one-hot encoding
    • Tensor inputs: rank 1, rank 2, empty/null handling

Bug Fix

  • Removes stale AiDotNet.Native.CLBlast.csproj project reference from test project

Linked Issue

Closes #655

Test plan

  • All 51 tests pass on net10.0
  • All 51 tests pass on net471
  • Build succeeds with no warnings

🤖 Generated with Claude Code

Add 51 comprehensive integration tests for the Evaluation module covering:
- Enums: PredictionType (BinaryClassification, Regression, MultiClass, MultiLabel)
- DefaultModelEvaluator: Construction, options handling
- PredictionStatsOptions: Default values, configuration
- PredictionTypeInference: Inference from vectors, matrices, tensors

Tests validate prediction type inference logic including:
- Binary classification detection (0/1 labels)
- Multi-class detection (integer labels with low unique ratio)
- Regression detection (continuous values, NaN, infinity, high unique ratio)
- Multi-label detection (matrix with multiple positives per row)

Also fixes stale CLBlast project reference in test project.

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
@vercel

vercel Bot commented Jan 23, 2026 •

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Review Updated (UTC)
aidotnet-playground-api Ready Ready Preview, Comment Jan 23, 2026 10:04pm

@coderabbitai

coderabbitai Bot commented Jan 23, 2026 •

Copy link
Copy Markdown
Contributor

Summary by CodeRabbit

Release Notes

  • Tests

    • Added comprehensive integration tests for the Evaluation module that validate prediction type inference across various data shapes and configurations, verify model evaluator construction with multiple options, and confirm prediction statistics functionality.
  • Chores

    • Simplified test project build dependencies.

✏️ Tip: You can customize this high-level summary in your review settings.

Walkthrough

This PR removes a CLBlast project reference from the test project configuration and adds comprehensive integration tests for the Evaluation module, covering PredictionType enum validation, DefaultModelEvaluator constructors, PredictionStatsOptions properties, and PredictionTypeInference logic across various data shapes.

Changes

Cohort / File(s) Summary
Project Reference Cleanup
tests/AiDotNet.Tests/AiDotNetTests.csproj
Removed ProjectReference to AiDotNet.Native.CLBlast.csproj
Evaluation Integration Tests
tests/AiDotNet.Tests/IntegrationTests/Evaluation/EvaluationIntegrationTests.cs
Added new test suite validating PredictionType enum, DefaultModelEvaluator constructors/options, PredictionStatsOptions properties, and PredictionTypeInference behavior via reflection across multiple data configurations (Vector, Matrix, Tensor) and target types

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~20 minutes

Possibly related PRs

Suggested labels

feature

Poem

🐰 A cottontail's delight, with test cases so bright,
The Evaluation module now shines in the light!
From PredictionType checks to inference so wise,
We hop through the data with quantitative eyes,
With 651 lines of assurance on test,
This integration suite puts our code to the best!

🚥 Pre-merge checks | ✅ 4 | ❌ 1
❌ Failed checks (1 warning)
Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 2.56% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Title check ✅ Passed The title clearly and concisely summarizes the primary change: adding integration tests for the Evaluation module.
Description check ✅ Passed The description is directly related to the changeset, providing detailed test coverage information and explaining the bug fix included.
Linked Issues check ✅ Passed The PR adds 51 integration tests covering PredictionType, DefaultModelEvaluator, PredictionStatsOptions, and PredictionTypeInference, addressing the core requirement to test the Evaluation module [#655].
Out of Scope Changes check ✅ Passed All changes are within scope: integration tests for the Evaluation module and removal of a stale project reference from the test project.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing touches
  • 📝 Generate docstrings
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Post copyable unit tests in a comment
  • Commit unit tests in branch test/evaluation-integration-tests

Comment @coderabbitai help to get the list of available commands and usage tips.

@coderabbitai coderabbitai Bot added the feature Feature work item label Jan 23, 2026
@sonarqubecloud

Copy link
Copy Markdown

@ooples
ooples merged commit 1cd26a3 into master Jan 24, 2026
41 checks passed
@ooples
ooples deleted the test/evaluation-integration-tests branch January 24, 2026 02:16
ooples added a commit that referenced this pull request Jan 24, 2026
Add 51 comprehensive integration tests for the Evaluation module covering:
- Enums: PredictionType (BinaryClassification, Regression, MultiClass, MultiLabel)
- DefaultModelEvaluator: Construction, options handling
- PredictionStatsOptions: Default values, configuration
- PredictionTypeInference: Inference from vectors, matrices, tensors

Tests validate prediction type inference logic including:
- Binary classification detection (0/1 labels)
- Multi-class detection (integer labels with low unique ratio)
- Regression detection (continuous values, NaN, infinity, high unique ratio)
- Multi-label detection (matrix with multiple positives per row)

Also fixes stale CLBlast project reference in test project.

Co-authored-by: franklinic <franklin@ivorycloud.com>
Co-authored-by: Claude Opus 4.5 <noreply@anthropic.com>
ooples pushed a commit that referenced this pull request Jan 24, 2026
Fixed critical bugs found during production-readiness review:

1. PredictionTypeInference (PR #765): Integer overflow when calculating
   class label range. maxClass - minClass overflows for extreme values.
   Fixed by using long for range calculation.

2. GeneticOptimizer (PR #762): IndexOf bug in tournament selection.
   When population contains duplicate prompts, IndexOf returns first
   occurrence index, causing wrong fitness selection. Fixed by tracking
   index directly.

3. NeuralProgramSynthesizer (PR #763): Absolute error comparison fails
   for large numbers. 1e12 + 0.5 vs 1e12 incorrectly fails with 1e-6
   absolute tolerance. Fixed with relative error comparison for large
   numbers and absolute for small.

4. TrainingMonitor (PR #755): CSV export shows "0" for missing metrics
   instead of empty string. FirstOrDefault returns default(T) which is
   0 for numerics, not null. Fixed by checking if match exists.

Added comprehensive tests that expose all bugs and verify fixes.

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
@ooples ooples mentioned this pull request Jan 24, 2026
3 tasks done
ooples added a commit that referenced this pull request Jan 26, 2026
* test: add comprehensive Serving module integration tests (#670)

Add 62 integration tests covering:
- ContinuousBatcherConfig (defaults, configuration, model presets)
- BatchSchedulerConfig (defaults, model-specific configs)
- BatchScheduler<T> (scheduling, preemption, cancellation, statistics)
- SequenceState<T> (lifecycle, token management, stop conditions)
- GenerationRequest<T> (configuration options)
- GenerationResult<T> (result data)
- BatcherStatistics and SchedulerStatistics
- SequenceStatus, StopReason, SchedulingPolicy enums
- Full integration workflows

Also remove stale CLBlast project reference from test project.

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* test: add comprehensive FineTuning module integration tests (#640) (#753)

Add 49 integration tests covering all 16 fine-tuning methods:
- SFT (Supervised Fine-Tuning)
- DPO (Direct Preference Optimization)
- SimPO (Simple Preference Optimization)
- ORPO (Odds Ratio Preference Optimization)
- IPO (Identity Preference Optimization)
- RDPO (Robust Direct Preference Optimization)
- KTO (Kahneman-Tversky Optimization)
- CPO (Contrastive Preference Optimization)
- RLHF-PPO (Reinforcement Learning Human Feedback)
- GRPO (Group Relative Policy Optimization)
- PRO (Pairwise Ranking Optimization)
- RRHF (Rank Responses Human Feedback)
- RSO (Statistical Rejection Sampling)
- SPIN (Self-Play Fine-Tuning)
- CAI (Constitutional AI)

Tests cover:
- Constructor initialization
- Fine-tuning workflow completion
- Data validation (SFT, Preference, RL, Ranking)
- Serialization/deserialization
- Edge cases and parameter handling
- Cancellation support

Also removes stale CLBlast project reference from test project.

Co-authored-by: franklinic <franklin@ivorycloud.com>
Co-authored-by: Claude Opus 4.5 <noreply@anthropic.com>

* test: add Evaluation integration tests (#655) (#765)

Add 51 comprehensive integration tests for the Evaluation module covering:
- Enums: PredictionType (BinaryClassification, Regression, MultiClass, MultiLabel)
- DefaultModelEvaluator: Construction, options handling
- PredictionStatsOptions: Default values, configuration
- PredictionTypeInference: Inference from vectors, matrices, tensors

Tests validate prediction type inference logic including:
- Binary classification detection (0/1 labels)
- Multi-class detection (integer labels with low unique ratio)
- Regression detection (continuous values, NaN, infinity, high unique ratio)
- Multi-label detection (matrix with multiple positives per row)

Also fixes stale CLBlast project reference in test project.

Co-authored-by: franklinic <franklin@ivorycloud.com>
Co-authored-by: Claude Opus 4.5 <noreply@anthropic.com>

* test: add ProgramSynthesis integration tests (#665) (#763)

Add 50 comprehensive integration tests for the ProgramSynthesis module covering:
- Models: CodePosition, CodeSpan, CodeLocation, CodeIssue, Program<T>
- Enums: ProgramLanguage, CodeTask, CodeIssueSeverity, SqlDialect, etc.
- Execution: CompilationDiagnostic, SqlValue, ProgramExecuteResponse
- Results: CodeGenerationResult

All tests validate model construction, property access, and default values.

Co-authored-by: franklinic <franklin@ivorycloud.com>
Co-authored-by: Claude Opus 4.5 <noreply@anthropic.com>

* test: add comprehensive integration tests for PromptEngineering module (#762)

Add 81 integration tests covering Templates, Analysis, FewShot selectors,
Compression, Chains, and Tools components.

Tests include:
- SimplePromptTemplate: construction, variable extraction, formatting, validation
- PromptMetrics: property values and timestamp behavior
- PromptIssue and IssueSeverity: issue tracking and severity levels
- ValidationOptions: default, strict, and lenient configurations
- CompressionResult: token savings, compression ratio, success detection
- CompressionOptions: default, aggressive, conservative presets
- FixedExampleSelector: example management and selection behavior
- ToolRegistry: registration, execution, case-insensitivity, descriptions
- SequentialChain: step management, sync/async execution, cancellation
- FewShotExample: model properties

Also removes stale CLBlast project reference from test project.

Closes #666

Co-authored-by: franklinic <franklin@ivorycloud.com>
Co-authored-by: Claude Opus 4.5 <noreply@anthropic.com>

* test: add comprehensive integration tests for Prototypes module (#761)

Add 58 integration tests covering PrototypeVector, PrototypeAdamOptimizer,
SimpleLinearRegression, and SimpleNeuralNetwork classes.

Tests include:
- PrototypeVector: construction, indexing, arithmetic operations, factory methods
- PrototypeAdamOptimizer: construction, parameter updates, convergence behavior
- SimpleLinearRegression: training, prediction, MSE/R2 computation
- SimpleNeuralNetwork: forward/backward passes, XOR training scenario
- Cross-component integration scenarios

Also removes stale CLBlast project reference from test project.

Closes #667

Co-authored-by: franklinic <franklin@ivorycloud.com>
Co-authored-by: Claude Opus 4.5 <noreply@anthropic.com>

* test: add comprehensive TrainingMonitoring module integration tests (#673) (#755)

Add 92 integration tests for the TrainingMonitoring module covering:

- TrainingMonitor<T>: session lifecycle, metric logging, progress tracking,
  speed stats, resource usage, issue detection, and data export
- ResourceMonitor: lifecycle, snapshots, history, threshold alerts, events
- ExperimentTracker: experiments, runs, parameters, metrics, artifacts,
  tags, deletion/restoration, persistence, and search
- NotificationManager: services, sending, filtering, buffering, events
- Dashboard implementations (ConsoleDashboard, HtmlDashboard, LiveDashboard):
  scalar/histogram logging, hyperparameters, confusion matrix, reports
- Cross-module integration tests combining multiple components

Also removes stale CLBlast project reference from test project.

Co-authored-by: franklinic <franklin@ivorycloud.com>
Co-authored-by: Claude Opus 4.5 <noreply@anthropic.com>

* Add comprehensive integration tests for AiDotNet.Serving module

Added 42 integration tests covering:
- ServingOptions configuration defaults and custom values
- StartupModel property validation
- PerformanceMetrics recording, percentiles, and statistics
- ContinuousBatchingStrategy concurrency and adaptive behavior
- TimeoutBatchingStrategy timeout-based processing
- SizeBatchingStrategy fixed-size batching
- AdaptiveBatchingStrategy latency-based adaptation
- BucketBatchingStrategy bucket index handling
- Padding strategies (Minimal, Bucket, Fixed) data preservation
- Configuration enum value verification

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* fix: replace thread.sleep with deterministic timestamp in queue time test

Instead of using Thread.Sleep(50) which is flaky under CI load,
directly set GenerationStartedAt to a known offset from CreatedAt.
This makes the test deterministic and asserts the exact expected
duration instead of using >= comparison.

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* fix: use net471-compatible Enum.GetValues in serving tests

Replace Enum.GetValues<T>() (NET 5+ only) with
Enum.GetValues(typeof(T)).Cast<T>() for .NET Framework 4.7.1 compatibility.

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

---------

Co-authored-by: franklinic <franklin@ivorycloud.com>
Co-authored-by: Claude Opus 4.5 <noreply@anthropic.com>
ooples added a commit that referenced this pull request Jan 27, 2026
* fix: production bugs in recently merged PRs

Fixed critical bugs found during production-readiness review:

1. PredictionTypeInference (PR #765): Integer overflow when calculating
   class label range. maxClass - minClass overflows for extreme values.
   Fixed by using long for range calculation.

2. GeneticOptimizer (PR #762): IndexOf bug in tournament selection.
   When population contains duplicate prompts, IndexOf returns first
   occurrence index, causing wrong fitness selection. Fixed by tracking
   index directly.

3. NeuralProgramSynthesizer (PR #763): Absolute error comparison fails
   for large numbers. 1e12 + 0.5 vs 1e12 incorrectly fails with 1e-6
   absolute tolerance. Fixed with relative error comparison for large
   numbers and absolute for small.

4. TrainingMonitor (PR #755): CSV export shows "0" for missing metrics
   instead of empty string. FirstOrDefault returns default(T) which is
   0 for numerics, not null. Fixed by checking if match exists.

Added comprehensive tests that expose all bugs and verify fixes.

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* fix: critical FineTuning bugs - log probability and SFT constructor

FineTuningBase.cs:
- Implemented ComputeLogProbabilityFromPrediction that was returning 0.0
- Added proper handling for probability distributions (cross-entropy)
- Added cosine similarity for embeddings
- Added scalar comparison for numeric values
- Added string similarity using Levenshtein distance
- Fixed single-element array bug (cosine similarity is always 1.0)

SupervisedFineTuning.cs:
- Added single-parameter constructor for Activator.CreateInstance compatibility
- This enables reflection-based instantiation used by test frameworks

MergedPRBugFixTests.cs:
- Added tests for log probability computation
- Added test verifying SFT can be instantiated with single-parameter constructor
- All 14 tests passing

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* fix: critical Diagnostics bugs - memory leak, thread safety, variance

MemoryTracker.cs:
- Added MaxHistorySize property with default of 10000
- Enforces limit to prevent unbounded memory growth in long-running apps
- Removes oldest snapshots when limit is reached (FIFO)

ProfilerSession.cs:
- Fixed thread-unsafe System.Random by using RandomHelper.ThreadSafeRandom
- System.Random is NOT thread-safe; concurrent access corrupts internal state
- Fixed variance calculation: was population (_m2/_count), now sample (_m2/(_count-1))
- Fixed call stack cleanup: handles out-of-order Stop() calls (e.g., due to exceptions)
- Now searches stack for timer instead of only checking top, prevents memory leak

MergedPRBugFixTests.cs:
- Added 5 tests for Diagnostics bug fixes
- All 19 tests passing

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* fix: ModelRegistry production bugs (PR #774)

Fixed 5 critical production-readiness bugs:

1. Mutable internal state returned: GetModel, GetModelByStage, and SearchModels
   now return clones to prevent external modification of internal state

2. Console.WriteLine in library code: Replaced with LoadErrors property for
   proper diagnostic exposure without polluting stdout

3. TOCTOU race conditions: Fixed file operations in DeleteModelVersion and
   GetModelCard to use try-catch pattern instead of File.Exists checks

4. DeleteModelVersion didn't validate modelName: Added ValidateModelName call
   for consistency with other methods

5. Lineage tracking didn't work: _lineage dictionary is now populated when
   model versions are created via RegisterModel and CreateModelVersion

Also added GetInternalModel helper method for internal mutation operations.

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* fix: PR #773 Metrics - add null checks, empty tensor handling, and validation

Bugs fixed in Metrics module:
- PSNR.ComputeBatch: Added null argument validation
- STOI.Compute: Added null argument validation
- STOI.ComputeNormalizedCorrelation: Fixed bounds check to include both arrays
- SI-SDR.Compute: Added null validation and empty tensor handling
- SNR.Compute: Added null argument validation
- IoU3D.ComputeBoxIoU: Added null validation and coordinate validation (min <= max)
- ChamferDistance.ComputeOneWay: Throws for empty target with non-empty source

Added 9 new tests to verify fixes.

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* fix: PR #772 Logging - add null checks, validation, and TOCTOU fix

Bugs fixed in Logging module:
- SummaryWriter.AddScalars: Added null check for tagScalarDict
- SummaryWriter.AddHistogram: Added null check for values arrays
- SummaryWriter.AddImage: Validate dataformats parameter (must be CHW or HWC)
- SummaryWriter.AddImages: Added null check and parameter validation
- SummaryWriter.AddPrCurve: Added null checks, length validation
- SummaryWriter.LogWeights: Added null check, handle empty weights array
- TensorBoardWriter.WriteEmbedding: Added null check, metadata length validation
- TensorBoardWriter.EncodePng: Validate pixels array length
- TensorBoardWriter.WriteProjectorConfig: Fixed TOCTOU race condition

Added 7 new tests to verify fixes.

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* fix(LanguageModels): add null/empty validation for model names and API version

Bug fixes for PR #771 LanguageModels:
- OpenAIChatModel: Validate modelName not null/empty (was NullReferenceException)
- OpenAIChatModel: Validate maxTokens > 0 early with clear error message
- AnthropicChatModel: Validate modelName not null/empty (was NullReferenceException)
- AzureOpenAIChatModel: Validate apiVersion not null/empty (caused invalid URL)
- AzureOpenAIChatModel: Validate maxTokens > 0 early with clear error message

Added 12 tests covering these validation scenarios.

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* fix(JitCompiler): add null checks and remove side effect from Validate

Bug fixes for PR #770 JitCompiler:
- IRGraph.Validate(): Remove side effect that modified TensorShapes
  (validation should be read-only)
- TensorShapeExtensions.GetElementCount(): Add null check
- TensorShapeExtensions.ShapeToString(): Add null check
- TensorShapeExtensions.GetShapeHashCode(): Add null check
- TensorShapeExtensions.GetShape(): Add null tensor check
- IRTypeExtensions.FromSystemType(): Add null Type check

Added 10 tests covering these validation scenarios.

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* fix(Interpretability): add null argument validation to helper methods

Bug fixes for PR #769 Interpretability:
- InterpretabilityMetricsHelper.GetUniqueGroups: Add null check for sensitiveFeature
- InterpretabilityMetricsHelper.GetGroupIndices: Add null check for sensitiveFeature
- InterpretabilityMetricsHelper.GetSubset: Add null checks for vector and indices
- InterpretabilityMetricsHelper.ComputePositiveRate: Add null check for predictions
- InterpretabilityMetricsHelper.ComputeTruePositiveRate: Add null checks for predictions and actualLabels
- InterpretabilityMetricsHelper.ComputeFalsePositiveRate: Add null checks for predictions and actualLabels
- InterpretabilityMetricsHelper.ComputePrecision: Add null checks for predictions and actualLabels
- InterpretableModelHelper: Add null checks for model, enabledMethods, and input parameters
  across all async methods (GetGlobalFeatureImportanceAsync, GetLocalFeatureImportanceAsync,
  GetShapValuesAsync, GetLimeExplanationAsync, GetPartialDependenceAsync,
  GetCounterfactualAsync, GetModelSpecificInterpretabilityAsync,
  GenerateTextExplanationAsync, GetFeatureInteractionAsync, ValidateFairnessAsync,
  GetAnchorExplanationAsync)

Added 10 tests covering null argument validation scenarios.

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* fix(InferenceOptimization): add null validation to optimization graph and node methods

PR #768 production-readiness fixes:
- OptimizationNode.AddInput: add null check for inputNode
- OptimizationNode.RemoveInput: add null check for inputNode
- OptimizationNode.ReplaceInput: add null checks for oldInput/newInput
- OptimizationGraph.FindNodeById: add null check for id
- OptimizationGraph.FindNodesByName: add null check for name
- IRDataTypeExtensions.FromSystemType: add null check for type
- TensorType.IsBroadcastCompatible: add null check for other
- GraphOptimizer.Optimize: add null check for graph
- GraphOptimizer.AddPass: add null check for pass

Added 14 tests covering all 9 bug fixes and valid input scenarios.

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* fix(HyperparameterOptimization): add null validation to trial pruner and distributions

PR #767 production-readiness fixes:
- TrialPruner.ReportAndCheckPrune(trial): add null check for trial
- TrialPruner.ReportAndCheckPrune(trialId): add null check for trialId
- TrialPruner.MarkComplete: add null check for trialId
- ContinuousDistribution.Sample: add null check for random
- IntegerDistribution.Sample: add null check for random
- CategoricalDistribution.Sample: add null check for random
- HyperparameterOptimizerBase.FindBestTrial: add null check for completedTrials
- HyperparameterOptimizerBase.EvaluateTrialSafely: add null checks for trial, objectiveFunction, parameters

Added 13 tests covering all 8 bug fixes and valid input scenarios.

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* fix(ExperimentTracking): fix production bugs in experiment tracking

Bug fixes:
- StartRun: persist experiment timestamp after Touch() to maintain disk consistency
- SearchRuns: validate maxResults > 0 to prevent invalid queries
- GetRunDirectory: use TryGetValue with descriptive InvalidOperationException
- DeleteExperiment/DeleteRun: handle IOException gracefully when directory deletion fails
- LogArtifact: validate extracted filename isn't empty for root paths
- LogArtifacts: wrap UnauthorizedAccessException with descriptive message
- Add null validation to GetExperiment, GetRun, ListRuns, DeleteExperiment, DeleteRun
- Add null validation to SerializeToJson, DeserializeFromJson, GetLatestMetric

Added 11 tests covering:
- Null argument validation
- MaxResults validation
- Timestamp persistence verification
- Thread safety for concurrent metric logging
- Graceful error handling for directory operations

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* fix(DistributedTraining): fix production bugs in distributed communication

Bug fixes:
- CommunicationManager.Broadcast: add root range validation to prevent invalid operations
- CommunicationManager.Scatter: add root range validation to prevent invalid operations
- CommunicationManager.ReduceScatter: add null validation for data parameter
- ParameterAnalyzer.CalculateDistributionStats: throw for null/empty groups (consistent API)
- InMemoryCommunicationBackend.PerformReduction: validate all vectors have same length
- InMemoryCommunicationBackend.Receive: validate message size BEFORE dequeuing (prevents data loss)
- ShardingConfiguration factory methods: add null validation for better error messages

Added 10 tests covering:
- Root validation in Broadcast/Scatter
- Null data validation in ReduceScatter
- Null/empty groups in ParameterAnalyzer
- Null backend in ShardingConfiguration factory methods
- Single-process AllReduce optimization path

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* fix(Reasoning): fix production bugs in reasoning module

Bug fixes:
- ReasoningChain.AddStep: add null validation for step parameter
- DepthFirstSearch.SearchAsync: add null validation for generator, evaluator, config (consistent with other search algorithms)
- MonteCarloTreeSearch constructor: validate numSimulations >= 1 and explorationConstant >= 0
- BreadthFirstSearch.CollectAllNodes: use iterative approach instead of recursion to prevent StackOverflow on deep trees

Added 9 tests covering:
- Null step validation in ReasoningChain
- Step number auto-increment
- MCTS constructor parameter validation
- ThoughtNode path reconstruction
- ThoughtNode leaf/root detection
- ReasoningConfig default values

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* fix(Serialization): add validation for negative dimensions in JSON converters

- VectorJsonConverter: validate that length is non-negative
- TensorJsonConverter: validate that shape array is not empty
- TensorJsonConverter: validate that all shape dimensions are non-negative
- Added 8 tests for serialization validation

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* fix(Serving): add parameter validation to batching strategies and padding

Add production bug fixes for PR #758 (Serving module):

Batching Strategies:
- ContinuousBatchingStrategy: validate maxConcurrency >= 1, minWaitMs >= 0, targetLatencyMs > 0
- TimeoutBatchingStrategy: validate timeoutMs >= 0, maxBatchSize >= 1
- AdaptiveBatchingStrategy: validate minBatchSize >= 1, maxBatchSize >= minBatchSize,
  maxWaitMs >= 0, targetLatencyMs > 0, latencyToleranceFactor > 0
- SizeBatchingStrategy: validate batchSize >= 1, maxWaitMs >= 0
- BucketBatchingStrategy: validate maxBatchSize >= 1, maxWaitMs >= 0, bucket values > 0

Monitoring:
- PerformanceMetrics: validate maxSamples >= 1, maxQueueDepthSamples >= 1

Padding Strategies (MinimalPaddingStrategy, BucketPaddingStrategy, FixedSizePaddingStrategy):
- PadBatch: validate no null vectors in input array
- UnpadBatch: validate originalLengths are non-negative
- FixedSizePaddingStrategy: validate fixedLength > 0
- BucketPaddingStrategy: validate bucketSizes not null/empty

Added 25 validation tests across BatchingStrategyTests, PaddingStrategyTests,
and PerformanceMetricsTests.

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* fix(Tokenization): add parameter validation to tokenizers and vocabulary

Add production bug fixes for PR #757 (Tokenization module):

Vocabulary:
- AddTokens: validate tokens parameter is not null
- Constructor(Dictionary): validate tokenToId parameter is not null

Tokenizers:
- BpeTokenizer.Train: validate corpus not null, vocabSize >= 1
- WordPieceTokenizer.Train: validate corpus not null, vocabSize >= 1
- WordPieceTokenizer constructor: validate maxInputCharsPerWord >= 1
- CharacterTokenizer.Train: validate corpus not null, minFrequency >= 1

MidiTokenizer:
- Constructor: validate ticksPerBeat >= 1, numVelocityBins >= 1
- CreateREMI: validate ticksPerBeat >= 1, numVelocityBins >= 1
- CreateCPWord: validate ticksPerBeat >= 1, numVelocityBins >= 1
- CreateSimpleNote: validate ticksPerBeat >= 1

TokenizationConfig:
- ParallelBatchThreshold: validate value >= 1 via property setter

Added 17 validation tests across BpeTokenizerTests, CharacterTokenizerTests,
WordPieceTokenizerTests, SpecializedTokenizerTests, and VocabularyTests.

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* fix(Tools): add parameter validation and robust type conversion

- Add upper bound validation (100) for topK in VectorSearchTool and RAGTool
- Add validation that topKAfterRerank cannot exceed topK in RAGTool
- Make ToolBase TryGetInt/TryGetDouble/TryGetBool handle type conversion errors gracefully
- Add 34 unit tests covering parameter validation and edge cases

PR #756 bug fixes - prevent performance issues from excessive topK values and
improve robustness when receiving invalid JSON property types.

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* fix(DistributedTraining): fix initialization order and add parameter validation

Bugs fixed:

1. Initialization order bug in derived classes (DivideByZeroException):
   - Base class was calling InitializeSharding() before derived class fields were set
   - Fix: Remove InitializeSharding() from base constructor, each derived class now
     calls it explicitly at the end of their constructor

2. ShardingConfiguration missing learningRate validation:
   - Added validation that learningRate > 0

3. PipelineParallelModel missing microBatchSize validation:
   - Added validation that microBatchSize >= 1

4. HybridShardedModel missing parallelism size validation:
   - Added validation that pipelineParallelSize >= 1
   - Added validation that tensorParallelSize >= 1

5. Inconsistent learning rate usage:
   - DDPModel, FSDPModel, PipelineParallelModel, HybridShardedModel were using
     hardcoded 0.01 instead of Config.LearningRate
   - Fixed to use Config.LearningRate consistently

Affected files:
- ShardedModelBase.cs - removed InitializeSharding() call from constructor
- DDPModel.cs, FSDPModel.cs, ZeRO1Model.cs, ZeRO2Model.cs - added InitializeSharding() call
- TensorParallelModel.cs, PipelineParallelModel.cs, HybridShardedModel.cs - same + validation
- ShardingConfiguration.cs - added learningRate validation

Tests: Added 26 validation tests in DistributedTrainingValidationTests.cs

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* fix(LoRA): fix matrix indexing bugs and initialization order issues

- Add input/output size validation in LoRALayer constructor
- Add pruningInterval validation in AdaLoRAAdapter to prevent division by zero
- Fix matrix indexing in MergeWeights (returns [inputSize, outputSize]):
  - LoRAAdapterBase.MergeToDenseOrFullyConnected
  - StandardLoRAAdapter.MergeToOriginalLayer
  - QLoRAAdapter.MergeToOriginalLayer
  - DoRAAdapter.Forward() and MergeToOriginalLayer
- Fix DefaultLoRAConfiguration.CreateAdapter to handle different constructor signatures
- Fix VeRAAdapter initialization order bug:
  - Move scaling vector init to CreateLoRALayer (called before ParameterCount)
  - Add UpdateParametersFromLayers override for VeRA-specific parameter sync
- Add 22 validation tests covering all bug fixes

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* fix(Autodiff): add numerical stability validation to tensor operations

- Add division by zero validation in TensorOperations.Divide
- Add non-positive value validation in TensorOperations.Log
- Add negative value validation in TensorOperations.Sqrt
- Handle sqrt(0) edge case in backward pass (use 0 instead of infinity)
- Fix null axes handling in TensorOperations.Sum OperationParams
- Add segmentSize validation in GradientCheckpointing.SequentialCheckpoint
- Add 19 validation tests covering all bug fixes

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* fix: address pr review comments and code scanning alerts

- ExperimentTracker: add logging to IOException catch block as comment indicated
- MonteCarloTreeSearch: add NaN/Infinity guards to explorationConstant validation
- TensorJsonConverter: remove validation rejecting empty shapes (scalars are valid)
- BucketBatchingStrategy: add validation for empty bucket boundaries array
- ToolBase: update exception handling to catch JsonSerializationException
- LoRALayer: update XML docs to reflect actual exception types thrown
- MergedPRBugFixTests: update size-mismatch test to actually verify behavior
- FineTuningBase: fix generic catch clauses and collection equality check
- ShardedModelBase: use lazy initialization to avoid virtual calls in constructor

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* fix: correct misleading comment and tighten assertion in sum operation

- Update comment in TensorOperations.cs to accurately describe OperationParams behavior
- Tighten Sum_NullAxes_OperationParamsHandledCorrectly test assertion to properly
  verify that Axes key is NOT present when axes parameter is null

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* fix: remove virtual calls from distributed training model constructors

- Remove direct InitializeSharding() calls from DDPModel, FSDPModel,
  ZeRO1Model, and ZeRO2Model constructors
- Remove redundant InitializeSharding() call from TensorParallelModel's
  OnBeforeInitializeSharding() method
- All models now rely on lazy initialization via EnsureShardingInitialized()
  in the base class to avoid virtual calls in constructors

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* fix: address coderabbit review comments for pr #783

- TensorOperations.cs: convert if/else to ternary for operationparams assignment
- HybridShardedModel.cs: move pendingconfig.value = null to onbeforeinitializesharding where it is consumed (lazy init compatibility)
- FineTuningBase.cs: move string check before ienumerable check since strings implement ienumerable<char>

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* fix: address additional coderabbit review comments

- TensorOperations.cs: only emit "Axes" metadata when non-null AND non-empty (empty array means sum-all like null)
- FineTuningBase.cs: treat null-null elements as matches in sequence matching fallback

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* fix: guard numeric conversion against mixed runtime types

Check both prediction and target are numeric before calling Convert.ToDouble
to avoid throwing when TOutput is object and types don't match.

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* chore: remove accidentally committed _playground_publish folder

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* feat: add comprehensive ActiveLearning integration tests

Add 62 integration tests covering:
- EntropySampling (8 tests)
- UncertaintySampling (5 tests)
- BALD (4 tests)
- RandomSampling (3 tests)
- MarginSampling (2 tests)
- LeastConfidenceSampling (2 tests)
- VariationRatios (2 tests)
- DiversitySampling (5 tests with all methods/metrics)
- CoreSetSelection (2 tests)
- HybridSampling (3 tests with all combination methods)
- InformationDensity (2 tests)
- DensityWeightedSampling (1 test)
- ExpectedModelChange (2 tests)
- BatchBALD (2 tests)
- QueryByCommittee (2 tests)
- Edge cases and mathematical validation (6 tests)

Tests include:
- Correct batch size validation
- Null argument handling
- Score range validation
- Mathematical properties (entropy of uniform/certain distributions)
- Diversity selection across clusters

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* feat: add comprehensive ContinualLearning integration tests

Add 53 integration tests covering:
- ElasticWeightConsolidation (EWC): 12 tests
- SynapticIntelligence (SI): 5 tests
- MemoryAwareSynapses (MAS): 4 tests
- GradientEpisodicMemory (GEM): 4 tests
- LearningWithoutForgetting (LwF): 4 tests
- OnlineEWC: 3 tests
- ExperienceReplay: 3 tests
- PackNet: 3 tests
- ProgressiveNeuralNetworks: 2 tests
- GenerativeReplay: 2 tests
- AveragedGEM (A-GEM): 3 tests
- Edge cases and cross-strategy validation: 8 tests

Tests verify:
- Constructor initialization
- BeforeTask/AfterTask lifecycle
- ComputeLoss returns non-negative values
- ModifyGradients produces valid output
- Reset clears stored data
- Multiple sequential tasks work correctly
- Lambda property can be modified
- Null argument handling

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* feat: add comprehensive curriculum learning integration tests

- Add 35 integration tests for CurriculumLearning components
- Test schedulers: LinearScheduler, SelfPacedScheduler, CompetenceBasedScheduler
- Test difficulty estimators: LossBased, ConfidenceBased, TransferBased, ExpertDefined, Ensemble
- Test edge cases: empty arrays, zero epochs, reset behavior
- Fix bug in CurriculumSchedulerBase.GetIndicesAtPhase that crashed on empty arrays

Note: Tests document a bug with generic T? default values - for unconstrained generics,
T? with default value is 0.0 for value types, not null. Tests work around this by
providing explicit values for optional parameters.

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* test: add comprehensive selfsupervisedlearning integration tests

Add 52 integration tests covering the SelfSupervisedLearning module:
- NT-Xent Loss tests (4 tests): temperature effects, gradient computation
- InfoNCE Loss tests (5 tests): memory bank integration, accuracy metrics
- BYOL Loss tests (5 tests): cosine similarity, symmetric loss computation
- Linear Projector tests (6 tests): shape validation, gradient backprop
- MLP Projector tests (5 tests): batch norm, training mode, backward pass
- Symmetric Projector tests (6 tests): predictor head, combined operations
- Memory Bank tests (11 tests): FIFO queue, momentum updates, sampling
- Momentum Encoder static method tests (3 tests): cosine schedule
- Edge cases and error handling tests (6 tests)

Uses RandomHelper for secure random number generation and proper
Tensor API patterns for cross-framework compatibility.

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* test: add nestedlearning integration tests for associativememory and contextflow

Add 30 integration tests covering the NestedLearning module:

AssociativeMemory tests (12 tests):
- Constructor validation and initialization
- Associate/Retrieve with dimension validation
- Capacity limit enforcement (FIFO)
- Association matrix updates with Hebbian learning
- Clear/Reset functionality
- Multiple associations and large capacity handling

ContextFlow tests (15 tests):
- Constructor and matrix initialization
- PropagateContext with level validation
- ComputeContextGradients backpropagation
- UpdateFlow transformation matrix updates
- GetContextState and CompressContext operations
- Reset clears all context states
- Independent states across multiple levels

Integration tests (3 tests):
- Combined AssociativeMemory + ContextFlow workflow
- Large capacity stress testing
- Sequential propagation state accumulation

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* test: add mixedprecision integration tests (39 tests)

Adds comprehensive integration tests for MixedPrecision module:
- LossScaler: scaling, unscaling, overflow detection, dynamic scaling
- MixedPrecisionConfig: defaults follow NVIDIA recommendations
- MixedPrecisionContext: FP32/FP16 weight management, gradient preparation
- Full workflow tests: training iterations, overflow recovery

Closes #642

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* test: add comprehensive integration tests for adversarial robustness module

- Add 83 integration tests for AdversarialRobustness module:
  - FGSM attack tests (8 tests)
  - PGD attack tests (4 tests)
  - CW attack tests (3 tests)
  - AutoAttack tests (2 tests)
  - Adversarial training defense tests (7 tests)
  - Randomized smoothing certification tests (8 tests)
  - Interval bound propagation tests (7 tests)
  - CROWN verification tests (6 tests)
  - Safety filter tests (11 tests)
  - Rule-based content classifier tests (11 tests)
  - Integration scenarios (6 tests)
  - Edge cases and error handling (10 tests)

- Fix bug in FGSMAttack: add null check for trueLabel parameter
  to throw ArgumentNullException instead of InvalidOperationException

Closes #631

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* test: fix randomhelper usage in finetuning integration tests

Replace new Random(42) with RandomHelper.CreateSeededRandom(42) to follow
project security standards for random number generation.

Closes #640

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* test: fix lora integration tests and loralayer exception consistency

- Fix LoRALayer to throw ArgumentOutOfRangeException consistently for all
  invalid rank values (was throwing ArgumentException for rank > min(in, out))
- Update test to expect ArgumentOutOfRangeException
- Replace new Random(42) with RandomHelper.CreateSeededRandom(42)

Closes #641

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* test: add knowledgedistillation integration tests

Comprehensive integration tests covering:
- DistillationLoss: constructor, compute loss/gradient, edge cases
- All distillation strategies: Feature, Attention, Contrastive,
  Probabilistic, Hybrid, Curriculum, Adaptive, Variational,
  NeuronSelectivity, Relational, SimilarityPreserving, FlowBased,
  FactorTransfer
- DistillationStrategyFactory: all strategy types
- DistillationForwardResult and DistillationCheckpointConfig
- IntermediateActivations: add/get/count
- Numerical stability and edge case testing

Total: 85 tests, 0 bugs found

Closes #636

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* test: add comprehensive physicsinformed integration tests

Add 94 integration tests covering the PhysicsInformed module:
- PhysicsInformedLoss tests (loss computation, gradients, edge cases)
- PDE tests (HeatEquation, WaveEquation, PoissonEquation, BurgersEquation,
  AllenCahnEquation, KdV, AdvectionDiffusion)
- PINN tests (PhysicsInformedNeuralNetwork, VariationalPINN, DeepRitzMethod)
- Neural Operator tests (FourierNeuralOperator, FourierLayer)
- ScientificML tests (HamiltonianNeuralNetwork)
- TrainingHistory, PDEDerivatives, PDEResidualGradient tests
- Edge cases and numerical stability tests
- Integration workflow tests
- Serialization tests

Closes #637

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* fix: move tensor type check before text fullname check in datamodalitydetector

The Tensor check was happening after the fullName-based Text check, which
caused Tensor<T> types to be incorrectly detected as Text modality since
the fullName could contain substrings matching other checks.

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* test: add comprehensive augmentation integration tests with 106 tests

Coverage includes:
- Image augmentations (GaussianNoise, Brightness, Contrast, Cutout, RandomCrop, etc.)
- Audio augmentations (AudioNoise, TimeStretch, PitchShift, etc.)
- Video augmentations (TemporalFlip, FrameDropout, SpeedChange)
- Text augmentations (RandomDeletion, RandomInsertion, SynonymReplacement, etc.)
- Object detection augmentations (BoundingBox, Keypoint transformations)
- Compose and auto-augment pipelines
- DataModalityDetector type detection

Tests verify correct behavior for each augmentation category including:
- Apply with probability, deterministic behavior, edge cases
- Parameter validation, composition, and context management

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* test: add comprehensive federated learning integration tests

Adds 64 integration tests for the FederatedLearning module covering:

- Aggregation strategies: FedAvg, FedProx, FedBN with weighted averaging
- Byzantine-robust aggregators: Krum, MultiKrum, Bulyan, TrimmedMean,
  WinsorizedMean, GeometricMedian, RFA
- Client selection strategies: UniformRandom, WeightedRandom, Stratified,
  AvailabilityAware, PerformanceAware, Clustered
- Privacy mechanisms: GaussianDifferentialPrivacy with clipping and noise
- Privacy accounting: BasicComposition and RDP privacy accountants
- Cryptography: HKDF key derivation
- Server optimizers: FedAdam, FedAdagrad, FedYogi, FedAvgM
- Heterogeneity corrections: SCAFFOLD, FedNova, FedDyn
- Secure aggregation: SecureAggregationVector, ThresholdSecureAggregationVector
- Additional tests: GaussianDifferentialPrivacyVector

Tests verify mathematical correctness and proper API behavior.

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* test: add checkpoint management integration tests

Adds 38 integration tests for the CheckpointManagement module covering:

- CheckpointManager construction and directory creation
- Auto-checkpointing configuration (save frequency, keep last, save on improvement)
- ShouldAutoSaveCheckpoint logic (frequency-based and improvement-based triggers)
- UpdateAutoSaveState for tracking last save step and best metric values
- AutoCheckpointState properties and ToString formatting
- Thread-safe concurrent configuration updates and state reads
- ListCheckpoints, LoadLatestCheckpoint, LoadBestCheckpoint edge cases
- CleanupOldCheckpoints and CleanupKeepBest cleanup strategies
- MetricOptimizationDirection enum values
- Path validation and nested directory support

Tests verify proper state tracking for minimization and maximization scenarios.

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* test: add 64 integration tests for Configuration module

- AutoMLBudgetOptions: default values, property setting, all presets
- RLTrainingOptions: default values, property setting, callbacks
- RLCheckpointConfig: default values, property setting
- RLEarlyStoppingConfig: default values, generic type support
- ExplorationScheduleConfig: default values, all decay types
- InferenceOptimizationConfig: default values, validation, all enum values
- ResNetConfiguration: variants, block counts, expansion, factory methods
- BenchmarkingOptions: default values, federated configs
- CurriculumLearningOptions: schedule types, difficulty estimators
- Supporting options classes: SelfPacedOptions, CompetenceBasedOptions

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* feat: add dataprocessor integration tests and fix empty matrix bug

- Add 24 comprehensive integration tests for DataProcessor module
- Tests cover DefaultDataPreprocessor, DataProcessorOptions, SplitData
- Tests validate preprocessing pipeline with Matrix, Vector, and Tensor types
- Fix DivideByZeroException in FeatureSelectorHelper.CreateFilteredData when
  handling empty matrices (0 rows)
- The fix returns a properly dimensioned empty matrix instead of attempting
  to call FromColumns on empty column vectors

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* feat: add dataversioncontrol integration tests (62 tests)

- Add comprehensive integration tests for DataVersionControl module
- Tests cover versioning, hashing, integrity verification, run linking
- Tests cover tagging, lineage tracking, snapshots, and persistence
- Tests verify thread safety with concurrent version creation
- Tests model classes: DatasetVersion, DatasetLineage, DatasetStatistics, etc.

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* test: add 65 integration tests for dataversioning module

Tests cover:
- Constructor and directory structure creation
- CreateDataset with validation, metadata, and duplicate handling
- AddVersion with files, directories, and deduplication
- GetVersion by ID, version number, and "latest"
- ListVersions and ListDatasets with ordering
- GetDataPath with directory structure preservation
- CompareVersions detecting additions, removals, modifications
- DeleteVersion and DeleteDataset with file cleanup
- RecordLineage and GetLineage with recursive upstream resolution
- Persistence across instance restarts
- Model classes (DatasetInfo, DataVersion, DataFileInfo, DataVersionDiff, DataLineage)
- Thread safety for concurrent operations
- Edge cases (large file count, empty directories, special characters)

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* fix: correct json parsing bug in agent response extraction

The ExtractJsonFromResponse method in all agent classes used a non-greedy
regex pattern that incorrectly matched nested JSON objects. For example,
with input:
  {"reasoning_steps": [...], "tool_calls": [{"tool_name": "X"}]}
The regex would match up to the first "}" (end of tool_calls inner object)
instead of the outer closing brace.

Fixed by implementing proper brace-balancing algorithm that:
- Tracks brace count while respecting string boundaries
- Handles escape sequences within strings
- Returns the complete outermost JSON object

Also added comprehensive integration tests for all agent types:
- Agent (ReAct pattern): 15 tests
- ChainOfThoughtAgent: 8 tests
- PlanAndExecuteAgent: 7 tests
- RAGAgent: 12 tests
- AgentBase: 3 tests
- JSON parsing edge cases: 4 tests
- Thread safety: 1 test

Total: 50 new tests for the Agents module

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* fix: correct training data dimensions in vision and tabular benchmarking tests

The tests were training KNN models on 1-feature data but then running
benchmarks that generate test data with different feature dimensions:
- CIFAR10/CIFAR100: 3072 features (32x32x3 pixels flattened)
- TabularNonIID: FeatureCount features (3 in this test)

Fixed by providing training data that matches the benchmark feature dimensions:
- CIFAR tests: 2 samples with 3072 features each (normalized pixel values)
- TabularNonIID test: 3 samples with 3 features each

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* test: add comprehensive benchmarking integration tests

Add 98 integration tests for the Benchmarking module covering:
- BenchmarkSuiteRegistry: display names, suite categories, mappings
- BenchmarkReport: construction, serialization, aggregation, statistics
- BenchmarkMetricValue: creation, edge cases, formatting
- BenchmarkExecutionStatus: all status values and transitions
- BenchmarkSuite enum: all 23 benchmark suites
- BenchmarkMetric enum: all 8 metric types
- BenchmarkSuiteKind enum: all 3 suite kinds
- Model classes: BenchmarkSuiteReport, BenchmarkDataSelectionSummary

Tests verify behavior for:
- Enum value coverage for all benchmark-related enums
- Display name generation and formatting
- Report creation and metric aggregation
- Suite category classification (ReasoningSuite vs DatasetSuite)
- Edge cases like empty reports and default values

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* test: add comprehensive computervision integration tests (66 tests)

Tests cover:
- BoundingBox format conversions (XYXY, XYWH, CXCYWH, YOLO)
- BoundingBox IoU, Area, Clip, IsValid operations
- Detection and DetectionResult classes
- DetectionStatistics and BatchDetectionResult
- NMS (standard, class-aware, batched)
- IoU variants (IoU, GIoU, DIoU, CIoU)
- GIoULoss for bounding box regression
- SORT tracker with Kalman filtering
- Track, TrackingOptions, TrackingResult classes
- End-to-end integration scenarios

Closes #647

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* fix: net471 compatibility for string.split and enum.getvalues

- Use char array overload for string.Split with StringSplitOptions
- Use typeof() overload for Enum.GetValues instead of generic version

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* fix: address pr review comments for integration tests

- Fix divide-by-zero guard in ActiveLearning CreatePoolWithKnownUncertainty
- Add missing assertion in RandomSampling_InformativenessScores test
- Add meaningful assertion in Rotation_ApplyWithTargets test
- Rename AllRegularizationStrategies to CoreRegularizationStrategies for accuracy
- Fix reflection test to assert if type exists but methods don't
- Fix greedy regex in ChainOfThoughtAgent JSON extraction
- Add BindingFlags for non-public property reflection in Benchmarking
- Add cleanup for default checkpoints directory
- Fix culture-invariant decimal comparison in AutoCheckpointState test

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* fix: address additional pr review comments

- Rename misleading test names in DataVersionControl
- Make MockChatModel thread-safe with Interlocked and ConcurrentBag
- Guard against deleting pre-existing default directories in DataVersioning
- Add delays to prevent timestamp-tie flakes in ordered tests

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* fix: replace generic catch clauses with specific exception types in finetuning

- Replace `catch (Exception ex) when (...)` patterns with specific
  exception type catches (InvalidCastException, FormatException, OverflowException)
- Extract ComputeSequenceMatchLogProbability helper to reduce code duplication
- Fix Equals on collections issue by using EqualityComparer<TOutput>.Default

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* fix: replace generic catch clause with specific exception types in ssl tests

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* fix: guard against single-class divide-by-zero in mock model

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* fix: update lora tests to expect argumentoutofrangeexception

The code correctly throws ArgumentOutOfRangeException (more specific than
ArgumentException) for invalid rank values. Tests now expect the correct
exception type.

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* fix: update vblora test to expect argumentoutofrangeexception

The test Constructor_WithRankExceedingBankSizeA_ThrowsArgumentException was
expecting ArgumentException but the code correctly throws
ArgumentOutOfRangeException which is more specific for range validation.

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

---------

Co-authored-by: franklinic <franklin@ivorycloud.com>
Co-authored-by: Claude Opus 4.5 <noreply@anthropic.com>
ooples added a commit that referenced this pull request Jul 10, 2026
….111.2 -> 0.112.0

Bump AiDotNet.Tensors to 0.112.0, which ships #758 (tape-auto-tracking sparse
engine ops + GPU-resident sparse autograd backward) and #765 (compiled ReduceMax
fills the entire axis-reduction output instead of only output[0] — a stale-tail
NaN leak).

Rewire SparseLinearLayer.Forward to the new tape-tracking path. It previously
copied the input into a Matrix<T> element by element and called the
non-differentiable ISparseEngine.SpMM, then copied the result back by hand —
which DETACHED the autodiff tape, so every trainable parameter received a zero
gradient (TapeGradient_ShouldReachAtLeastOneTrainableParameter). It now computes
output = (W · inputᵀ)ᵀ + bias entirely from tape-tracked Engine ops
(TensorTranspose / ISparseEngine.SparseMatMul / TensorBroadcastAdd), so the
gradient reaches the registered sparse _weights and _biases. The legacy manual
ComputeGradients path (SparseNeuralNetwork.Train) is untouched.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
ooples added a commit that referenced this pull request Jul 10, 2026
….111.2 -> 0.112.0

Bump AiDotNet.Tensors to 0.112.0, which ships #758 (tape-auto-tracking sparse
engine ops + GPU-resident sparse autograd backward) and #765 (compiled ReduceMax
fills the entire axis-reduction output instead of only output[0] — a stale-tail
NaN leak).

Rewire SparseLinearLayer.Forward to the new tape-tracking path. It previously
copied the input into a Matrix<T> element by element and called the
non-differentiable ISparseEngine.SpMM, then copied the result back by hand —
which DETACHED the autodiff tape, so every trainable parameter received a zero
gradient (TapeGradient_ShouldReachAtLeastOneTrainableParameter). It now computes
output = (W · inputᵀ)ᵀ + bias entirely from tape-tracked Engine ops
(TensorTranspose / ISparseEngine.SparseMatMul / TensorBroadcastAdd), so the
gradient reaches the registered sparse _weights and _biases. The legacy manual
ComputeGradients path (SparseNeuralNetwork.Train) is untouched.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
ooples added a commit that referenced this pull request Jul 10, 2026
…mirrors

0.113.0 publishes Tensors PR #763 — this branch's gating dependency. #763 makes
the compiled/persistent backward honor createGraph=true (GradientTape.Compute-
Gradients previously gated the compiled path on !createGraph), so the WGAN-GP
gradient penalty's inner backward now differentiates into the disc weights through
the fused compiled plan instead of silently returning zeros (issue #1844). Against
0.111.2 the fused GPU-resident WGAN-GP path would have silently degraded to plain
WGAN. Also carries #765 (compiled ReduceMax axis fill) and #764 (resident-param
fused-Adam fix).

Now that 0.113.0 is published to NuGet, retire the local mirror classes that stood
in until it shipped (they were 1:1 copies with identical public APIs), and point
every consumer at the Engines.Training.* versions in the package:

- Delete src/Training/{DpSgdFusedStep,WganGpFusedStep,MultiSlotFusedStep}.cs.
- Repoint all call sites (added by #1847's fused-primitive centralization) to
  AiDotNet.Tensors.Engines.Training.*:
    * WganGpFusedStep — CausalGAN, CTGAN, CopulaGAN, TableGAN
    * MultiSlotFusedStep — DiffusionModelBase, CSDI, TabDDPM, TabSyn
    * DpSgdFusedStep — DPCTGAN, MedSynth
  Signatures are identical, so this is a pure namespace swap.

GpuResidentFusedStep stays consumer-side (not part of #763; it composes the loss —
including the createGraph=true gradient penalty — and drives the generic fused plan).

Bumps AiDotNet.Native.OneDNN / OpenBLAS to 0.113.0 in lockstep (both published).
Builds green on net471 / net8.0 / net10.0.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
ooples added a commit that referenced this pull request Jul 11, 2026
…for fused paths)

0.113.0 publishes Tensors PR #763 — this branch's gating dependency. #763's key fix
(4C): the compiled/persistent backward now honors createGraph=true (GradientTape.
ComputeGradients previously gated the compiled path on !createGraph), so the WGAN-GP
gradient penalty's inner backward differentiates into the disc weights through the
fused compiled plan instead of silently returning zeros (issue #1844). Against 0.111.2
the fused GPU-resident WGAN-GP path would have silently degraded to plain WGAN.

The src/Training fused-primitive mirrors (WganGpFusedStep, MultiSlotFusedStep,
DpSgdFusedStep) stay as the consumer-side primitives that #1847/#1848 centralize every
consumer through; they call the public engine API, so the bump alone routes them onto
#763's fixed compiled backward. Also picks up #765 (compiled ReduceMax axis fill) and
#764 (resident-param fused-Adam fix). Native OneDNN/OpenBLAS/CLBlast bumped in lockstep
(all published). Builds green on net8.0.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
ooples added a commit that referenced this pull request Jul 11, 2026
…1843)

* feat(training): gPU-resident fused step for non-TS single-net models

Kicks off the non-time-series GPU-residency sweep. Adds a shared helper for
classes that don't inherit from NeuralNetworkBase or TimeSeriesModelBase but
still want their Train() to route forward + backward + optimizer through the
compiled fused plan (weights / activations / Adam moments resident on-device
across the whole step).

* src/Training/GpuResidentFusedStep.cs — shared helper. TryResolveOptimizerConfig
  maps a runtime IGradientBasedOptimizer to the fused-plan OptimizerType +
  hyperparameters (Adam / AdamW / SGD supported via case-insensitive class-name
  match + reflection over Options.InitialLearningRate / Beta1 / Beta2 / Epsilon
  / WeightDecay). TryStep is the one-shot entry.

Wired the following single-net models through the helper (mirrors
NeuralNetworkBase.TrainWithFusedStep):

* GraphClassificationModel (NN base + GNN + pooling + cross-entropy) —
  routes tape training through the fused plan when float + DirectGpu +
  compilation are live; falls through to the existing eager tape+optimizer
  loop on any failure.
* LinkPredictionModel (NN base + GNN + node embeddings + BCE) — same
  wire-up as GraphClassificationModel.
* FourierNeuralOperator (lift → FourierLayers → project, hardcoded SGD +
  learning rate 0.001) — captures the whole spectral-conv chain in a fused
  SGD plan; falls through to the in-place SGD loop below.

GAN-family generators (CTGAN, DPCTGAN, PATEGAN, ...), CLAPModel (learned
scalar outside layer hierarchy), AutoDiffTabGenerator (per-sample variable
shape defeats plan replay) and other non-conforming shapes are queued for
follow-up commits. The shared helper is the common seam so they all wire
the same way once their per-step structure fits the fused-plan contract
(constant shape + Adam/AdamW/SGD optimizer + ITrainableLayer-carried params).

Builds green net471 / net8.0 / net10.0.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* feat(training): gPU-resident fused step for Finance foundation forecasters

Wires TFC, TOTEM, CSDI, MQCNN through the compiled fused SGD plan (all four use
in-place SGD with lr=0.001 as their eager path). Extends the shared helper with
an inline IsGpuResidentAvailable check so callers don't need to duplicate the
availability gate.

* GpuResidentFusedStep — adds IsGpuResidentAvailable static property (mirrors
  TimeSeriesModelBase.CanTrainOnGpu / NeuralNetworkBase.CanTrainOnGpu). TryStep
  short-circuits when unavailable so callers stay on their eager path.
* TFC — supervised + contrastive branches fused into one closure; the fused
  plan captures both losses in a single backward.
* TOTEM — reconstruction + VQ commitment terms fused; the recompute-loss
  closure re-derives the commitment loss on each replay for graph parity.
* CSDI — denoising-score-matching: the forward closure re-samples timestep +
  noise each step (via ComputeDenoisingPairTape) so replay produces fresh
  training data even though the plan's captured graph shape is fixed.
* MQCNN — multi-quantile pinball loss captured directly via ComputeMultiQuantilePinballLossTape.

All four fall through to the existing in-place SGD path on any fused-plan failure
(unsupported op, non-compilable graph, etc). Builds green net471 / net8.0 / net10.0.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* feat(training): gPU-resident fused generator step for GAN-family synthetic generators + TVAE

Wires the generator side of the dual-network training loop through the fused
compiled plan for 9 SyntheticData generators. Each covers a distinct pattern:

* PATEGAN — per-sample loop; fused plan compiles on first sample, replays
  across the batch with fresh noise per iteration.
* MedSynth — batched noise -> decoder -> discriminator-frozen -> non-saturating
  log-sigmoid loss.
* CausalGAN — batched noise -> generator -> optional causal-structure ->
  output activations -> disc-frozen -> -avgFake.
* OCTGAN — per-sample loop; noise -> gen -> disc-frozen embedding -> SVDD dist².
* TableGAN — noise + real batch -> gen -> composite loss (fake-scores +
  information-loss + optional classification).
* CTGAN — the tricky one: pack (noise, cond, mask) into one persistent input
  tensor so the closure can slice cond and mask back out on replay; captures
  both -avgFake AND conditional cross-entropy on the fused plan.
* DPCTGAN — same pattern as CTGAN but simpler (no mask, no CE).
* CopulaGAN — same pattern as DPCTGAN.
* TVAE — encoder + reparam + decoder + composite ELBO (recon + KL) all fused;
  Reparameterize re-samples inside the closure so training stays stochastic.

Discriminator/critic/student layers are NOT passed to the fused step (their
weights stay frozen on the gen step, matching the eager path semantics).
Falls through to the eager tape+optimizer path on any failure.

The discriminator STEPS of these GANs are not yet wired — they need a separate
pass because their loss involves gradient penalty (WGAN-GP) or per-example DP
clipping (DPCTGAN) which don't compose cleanly with the single-closure fused
step yet.

Builds green net471 / net8.0 / net10.0.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* feat(training): gPU-resident fused step for TabSyn VAE + MisGAN mask/data gens

Two more SyntheticData generators covered — each covers a distinct pattern:

* TabSyn.TrainVAEBatch — per-row VAE ELBO training. Fused plan compiles on
  first row, replays across the batch with refreshed input per iteration.
  Reparameterize re-samples inside the closure so training stays stochastic.
* MisGAN.TrainDataGeneratorStep — dual-noise generator (data noise + mask
  noise) packed into a single persistent input tensor; closure slices them
  back out on replay. Loss = -E[D_x(fakeRow ⊙ fakeMask)].
* MisGAN.TrainMaskGeneratorStep — simpler single-noise version; per-sample
  loop with fused plan replay across the batch.

The DataDiscriminator step (WGAN-style critic + weight clipping) and mask
discriminator step aren't wired — they need separate plans and the
ClipWeights post-step interacts poorly with the fused optimizer's
in-place update. Follow-up.

Builds green net471 / net8.0 / net10.0.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* feat(training): gPU-resident fused step for TimeGAN P1/P2 + MedSynth non-DP disc

* TimeGAN Phase 1 (TrainReconstructionStepBatched) — embedder + recovery
  reconstruction MSE. Straight single-network fused-step conversion.
* TimeGAN Phase 2 (TrainSupervisedStepBatched) — supervisor's next-step MSE
  in the embedded space, with the embedder frozen (its layers not in the
  trainable set for this step). Two-tensor closure (ht, htNext).
* MedSynth non-DP disc step — pack (real, fake) into a single input tensor
  along axis 0; slice back in the loss for BCE-real + BCE-fake. Generator
  runs OUTSIDE the fused plan so its weights stay frozen on the critic step.

TimeGAN Phase 3 (adversarial + supervised joint) and MedSynth DP-SGD paths
still use their eager tape — the WGAN-GP gradient penalty and per-example
DP clipping don't compose with the single-closure fused optimizer step.

Builds green net471 / net8.0 / net10.0.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* feat(training): extra-tensors parameter on fused-step API

Extends the fused compiled-plan training step to accept raw trainable tensors
that aren't naturally carried by an ITrainableLayer (e.g. GOGGLE's soft
adjacency A; CLAP's learned _logTemperature scalar). Solves the SOLID
concerns raised in review of the ITrainableLayer-wrapper alternative:
* ISP — no forcing raw tensors to implement the full ILayer surface with
  no-op Forward / SetTrainingMode / GetParameterGradients members
* LSP — no risk of "layer" wrappers with identity Forward diverging from
  the layer contract's behavioral expectations elsewhere in the codebase

Wire-up:
* CompiledTapeTrainingStep.TryStepWithFusedOptimizer — new optional
  extraTensors param. Threaded into the dedup-aware parameter collection
  (CollectDeduplicatedParametersWithExtras) so extras get moment buffers,
  gradient accumulation, and GPU-residency in exactly the same code path
  that layer-carried params do.
* GpuResidentFusedStep.TryStep — plumbs extras through. A callsite with
  only extras (no layers) is a valid config for models whose whole
  trainable surface is raw tensors.

Dedup is by Tensor<T> reference across BOTH sources, so an extra tensor
that also happens to be layer-carried is registered exactly once — the
same shared/tied-weight protection the layer-only collector provides.

Enables Phase 4E (GOGGLE + CLAP wire-ups).
Builds green net471 / net8.0 / net10.0.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* feat(training): wire GOGGLE + CLAP through fused-step extra-tensors API

Consumes Phase 4A's extraTensors parameter to route models with raw
trainable state (outside the ITrainableLayer contract) through the fused
compiled plan.

* GOGGLE — soft adjacency A (initialised in InitializeModel, projected
  after each optimizer step via ProjectAdjacencyConstraints) is passed
  through extraTensors. The fused optimizer allocates its moment buffer,
  accumulates gradients, and applies the update in-place. Reparameterize
  re-samples inside the forward closure so training stays stochastic;
  ProjectAdjacencyConstraints runs after the fused step to keep the
  adjacency on the valid soft-adjacency manifold.
* CLAP — the learned _logTemperature scalar (Radford 2021 / Wu 2023
  contrastive-alignment temperature) goes through extraTensors. Both
  audio and text encoders are in the layers list; the fused step handles
  all three parameter classes uniformly.

Both fall through to the existing eager tape+optimizer path on any
failure (unsupported optimizer, non-compilable graph, etc).

Builds green net471 / net8.0 / net10.0.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* feat(diffusion): batched per-element timesteps for DDPM-canonical training

Adds industry-standard batched-per-element timestep API to the diffusion stack
(Ho et al. 2020 §3, HuggingFace diffusers reference). Previously
DiffusionModelBase.Train sampled ONE timestep for the whole batch — reducing
the signal-to-noise diversity of the training gradient. Canonical DDPM training
samples a distinct timestep per batch element.

* INoiseScheduler.AddNoiseBatched(cleanBatch, noiseBatch, timesteps) — default
  interface implementation delegates to the scalar AddNoise per element for
  backward compatibility. Concrete schedulers can override with a fused
  batched implementation.
* NoisePredictorBase.PredictNoiseBatched(noisyBatch, timesteps, conditioning) —
  virtual with default slice-then-call-scalar implementation. Subclasses that
  want a fused batched forward override this to keep the training loop on-device.
* DiffusionModelBase.PredictNoiseBatched — model-level counterpart with the
  same default slice-then-call-scalar behavior.
* DiffusionModelBase.Train detects batched vs rank-1 input and routes through
  the batched noise scheduler + batched predictor when input.Rank >= 2. Rank-1
  unbatched inputs stay on the scalar path for backward compatibility.

This is the foundation for the industry-exceeding "batched-per-element +
fused-resident" diffusion training path — see follow-up commits that wire
DiffusionModelBase.Train through the fused compiled plan.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* feat(training): wGAN-GP correctness fix + fused disc + DP-SGD wire-ups (4F/4G/4H)

Consumes Tensors PR ooples/AiDotNet.Tensors#763 (compiled backward with
createGraph=true, DP-SGD helper, multi-slot persistent inputs).

## 4G — WGAN-GP correctness fix + fused disc steps (5 files)

Root cause fixed for AiDotNet issue #1844: every ComputeGradientPenalty
implementation used the inner GradientTape's ComputeGradients WITHOUT
createGraph=true. The inner backward's ops didn't record on the outer
tape, so inputGradients had no GradFn chain back to the discriminator
weights. Effect: the gradient penalty was included in the loss VALUE
but NOT in the disc's gradient — WGAN-GP silently degraded to plain
WGAN with no 1-Lipschitz enforcement.

Fix (5 files: CTGAN, DPCTGAN, CopulaGAN, CausalGAN, TableGAN): add
createGraph: true to the inner ComputeGradients so the outer tape can
differentiate the penalty through the disc weights.

Also wired 4 disc steps through the fused compiled plan
(CausalGAN, CTGAN, CopulaGAN, TableGAN — DPCTGAN's disc is DP so goes
via 4H). Packs (real, fake) into one persistent input along axis 0;
loss closure splits scores back out and computes wasserstein + λ·GP.
The fused path activates once Tensors PR #763 lands (its 4C change
removes the !createGraph gate at GradientTape.cs:727); until then
these wire-ups fall through to the (now-corrected) eager tape path.

## 4H — DP-SGD wire-ups (2 files)

Local AiDotNet.Training.DpSgdStep<T> — drop-in mirror of the same
helper in Tensors PR #763 (Engines.Training.DpSgdStep<T>). Enables
the wire-ups to land in this PR without waiting on the Tensors NuGet
publish. Both implementations enforce the Abadi 2016 §3 Algorithm 1
clip-BEFORE-aggregate order via their structure.

* DPCTGAN.TrainDiscriminatorStepBatchedDP — routes the per-example
  WGAN-GP + DP-SGD critic through DpSgdStep<T>. Objective: Wasserstein
  + λ·GP per example, clipped per-example, aggregated, noised, averaged.
* MedSynth.TrainDiscriminatorStepPerExampleDPSGD — routes the per-
  example non-saturating BCE + DP-SGD critic through the same helper.

Once Tensors PR #763 merges + NuGet publishes, the local mirror can
be swapped for the Tensors version by changing the `using` — the API
surface is identical by design.

## 4F — Batched-per-element diffusion foundation

DiffusionModelBase.Train samples per-batch-element timesteps for
rank ≥ 2 inputs (Ho et al. 2020 canonical pattern; HuggingFace
diffusers reference), routing through NoiseSchedulerBase.AddNoiseBatched
+ DiffusionModelBase.PredictNoiseBatched. Rank-1 unbatched inputs
keep the scalar path for backward compat.

net471 doesn't support default interface implementations, so
AddNoiseBatched lives on NoiseSchedulerBase<T> (as virtual with a
default per-element delegate) rather than on the INoiseScheduler<T>
interface. DiffusionModelBase's fused-training wire-up (via Tensors
PR #763's PersistentInputRegistry) is queued for the NuGet-bump
follow-up commit — the API foundation is already in place.

Builds green net471 / net8.0 / net10.0.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* fix(review): resolve 12 CodeRabbit findings on PR #1843

Addresses every unresolved review comment on the non-TS GPU-residency PR.

## Correctness fixes

* DiffusionModelBase.Train — RecomputeForward closure now dispatches
  isBatched ? PredictNoiseBatched(inp, timesteps) : PredictNoise(inp, timestep),
  so optimizer re-evaluation matches the recorded batched forward objective.
* NoisePredictorBase.PredictNoiseBatched — detect batch-aligned conditioning
  (leading dim == batchSize) and slice per element; preserve shared conditioning
  when the leading dim doesn't match. Fixes shape-mismatch / conditioning-shared
  bugs in classifier-free-guidance-style predictors.
* CSDI.Train — remove the fused-plan branch. ComputeDenoisingPairTape samples
  fresh (timestep, noise) each call, and calling it in both Fwd and Loss produces
  independent samples that don't match. The compiled plan can't refresh the RNG
  per replay either. Path stays on the eager tape until Tensors' PersistentInputRegistry
  (PR ooples/AiDotNet.Tensors#763) lands.
* TFC.Train — ComputeContrastiveLossTape now runs INSIDE ForwardCombined so it
  consumes the current-step persistent input (`inp`), not the outer `input` which
  would freeze at compile time. Closure-captured local; Loss reads it with a
  null-guard covering the Fwd-then-Loss ordering invariant.
* TOTEM.Train — capture the commitment tensor from ForwardNativeForTrainingWithCommitment's
  first call; reuse it in Loss instead of running the quantizer a second time.
  Without this, the compiled path performs an EMA SetCodebookValue update TWICE
  per step and diverges from the eager path.
* NoiseSchedulerBase.AddNoiseBatched — validate full shape parity (rank + every
  dim), not only the leading batch dim. Prevents indexing beyond noiseBatch's span
  when a caller passes [B, smaller...] shapes.

## Correctness cleanup

* CTGAN.TrainGeneratorStepBatched — capture (act, condFromInput, maskFromInput)
  from Fwd's single generator pass; reuse in Loss so the conditional-CE term
  doesn't re-run GeneratorForwardWithResidualBatched + ApplyOutputActivationsBatched
  (was doubling the per-step generator forward cost).
* TabSynGenerator.TrainVAEBatch — capture (mean, logVar) from Fwd's single encoder
  pass; reuse in Loss so the encoder doesn't run twice per row. Also: if the fused
  step fails mid-batch after fusedEngaged=true, drop to the eager path for the
  remaining rows (no row is silently skipped).
* FourierNeuralOperator.TapeTrainStep — remove the duplicate _fourierLayers loop
  in both the fused-path allTrainable collection AND the eager paramList. Fourier
  layers are already registered into Layers at construction, so iterating them
  again double-registered each Fourier parameter and drove the eager SGD to apply
  the update twice per step.

## API cleanup

* GpuResidentFusedStep<T> — public → internal. Aligns with CompiledTapeTrainingStep<T>
  which is already internal; keeps in-assembly training plumbing out of the public
  API surface.

Also fixed an inadvertent null-forgiving operator (`!`) usage that slipped into
the review-fix edits. All new nullable annotations use explicit null-guards with
descriptive InvalidOperationException on invariant violation (per CLAUDE.md's
"never use null-forgiving operators" rule).

INoiseScheduler default-interface issue (net471 incompatibility) was already
fixed in an earlier commit (AddNoiseBatched moved to NoiseSchedulerBase).

Builds green net471 / net8.0 / net10.0.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* feat(training): vectorized fused DP-SGD/WGAN-GP/multi-slot primitives + consumer wire-ups

Adds three fused-plan training primitives (mirrors of AiDotNet.Tensors PR #763)
and wires the DP-SGD consumers through them:

* DpSgdFusedStep<T> — Abadi 2016 §3 Algorithm 1 per-example clip-before-aggregate.
  Runs each per-example forward+backward through a compiled plan (LR=0 so weights
  don't drift between replays), then computes the global L2 norm, clips, noises,
  and aggregates via vectorized IEngine ops (TensorMultiply/ReduceSum for the
  norm, TensorMultiplyScalar/TensorAdd for the accumulator, TensorRandomNormalInto
  for on-device Gaussian noise). Returns aggregated gradients so the caller's
  configured optimizer (Adam/AdamW/SGD/...) applies the update — no hardcoded
  SGD step inside.

* WganGpFusedStep<T> — Gulrajani 2017 WGAN-GP critic step. Composes
  E[D(fake)] − E[D(real)] + λ·(‖∇_x̃ D(x̃)‖₂ − 1)² inside one compiled plan with
  the inner ∇_x̃ D(x̃) recorded via createGraph=true so it differentiates into
  disc weights (issue #1844 fix). OnesLike uses vectorized Engine.TensorFill.

* MultiSlotFusedStep<T> — N-slot persistent input mechanism with plan cache
  keyed by composite shape + parameter identity. Slots refreshed via
  AsSpan().CopyTo(AsWritableSpan()) — no per-element loops on the hot path.

Consumer wire-ups (this PR):

* DPCTGANGenerator.TrainDiscriminatorStepBatchedDP — primary path routes through
  DpSgdFusedStep.TryStep, falls back to the existing ComputePerExampleNoisedGradients
  when the fused path can't engage (non-GPU host / compilation disabled).
* MedSynthGenerator.TrainDiscriminatorStepPerExampleDPSGD — same primary/fallback
  layering.
* Both legacy fallback loops (ComputePerExampleNoisedGradients + MedSynth eager)
  are now themselves vectorized: global-L2-norm via ReduceSum(g·g), clipped
  accumulation via TensorMultiplyScalar+TensorAdd, on-device Gaussian noise via
  TensorRandomNormalInto+TensorAdd — no scalar per-element loops anywhere on the
  DP-SGD path.

Codebase convention adopted:

* Class-scope `private static IEngine Engine => AiDotNetEngine.Current;` and
  `private static readonly INumericOperations<T> Ops = MathHelper.GetNumericOperations<T>();`
  on each fused-step class (matches the ~20 activation/optimizer/etc. bases).
  No IEngine or INumericOperations<T> threaded through method signatures.

Skipped (follow-ups filed):

* WGAN-GP consumer rewire → #1845 (blocked on IGradientBasedOptimizer<T> config
  accessor extension; consumers already correct via GpuResidentFusedStep + local
  createGraph=true ComputeGradientPenalty).
* Diffusion consumer rewire → #1846 (blocked on per-generator forward-refactor
  to move _timestepProjection inside compiled plan while treating raw timestep
  as a persistent slot).

Verification:

* net8.0 + net471 + net10.0 all build clean on both AiDotNet and AiDotNet.Tensors.
* Old DpSgdStep.cs deleted (replaced by DpSgdFusedStep).

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* fix(review): tFC + TOTEM preprocessing rewritten with traceable engine ops

Resolves the two BLOCKING CodeRabbit findings on PR #1843: both fused-plan
training paths were freezing preprocessing into the compiled plan because
their inner ops used host-side .Data.Span loops, so later replays reused the
first batch's normalized values / spectrum / argmin decisions instead of
recomputing per batch.

TFC — ApplyInstanceNormalization (RevIN) + ComputeFrequencyRepresentation:

* ApplyInstanceNormalization now delegates to a new stateless NormalizeWithStats
  that returns (normalized, mean, std) as tensors. Under the hood: ReduceMean +
  ReduceVariance + TensorSqrt + TensorBroadcastSubtract + TensorBroadcastDivide.
  All ops record on the tape and re-execute per replay under the compiled plan.
* DenormalizeForecast likewise delegates to DenormalizeForecastWithStats, which
  takes mean/std as explicit tensor parameters. ForwardNative threads them
  through as locals so the compile-mode replay uses CURRENT-step stats instead
  of frozen trace-time values.
* _revinMean/_revinStd (Vector<T> scalars) → _revinMeanTensor/_revinStdTensor
  (Tensor<T>? nullable) — kept for the abstract override's external callers,
  but the fused path never reads them.
* ComputeFrequencyRepresentation now uses Engine.RFFT for batched real FFT →
  reshape to [B, halfN, 2] pairs → TensorMultiply + ReduceSum(axis=2) for
  magSquared → TensorSqrt + TensorMultiplyScalar(1/n) for the one-sided
  magnitude spectrum → TensorSlice + TensorFlip + TensorConcatenate to mirror
  bins [1..n-halfN] into the tail. Handles even and odd n identically to the
  old scalar impl.
* Fused-plan fast path in TFC.Train restored — now safe because both
  preprocessing methods trace correctly.

TOTEM — VectorQuantize (VQ-VAE) traceable rewrite + post-Step EMA:

* New VectorQuantizeTraceable returns (quantized, commitmentLoss, argmin, head)
  using Engine ops end-to-end: TensorBroadcastSubtract + TensorMultiply +
  ReduceSum for distances, TensorArgMin along the codebookSize axis, per-c
  TensorSliceAxis + TensorIndexSelectDiff + TensorStack for the gather,
  Engine.StopGradient for the straight-through estimator, ReduceSum-based
  commitment loss weighted by β/totalLen.
* EMA moved OUT of the compiled forward into a new UpdateCodebookEMA(head,
  argmin). Called POST-Step by the fused path with the trace-time graph-node
  references — their .Data reflects the LAST replay so the update lands
  exactly once per batch (matches CodeRabbit's "EMA must execute exactly once
  per batch" contract). Under the compiled plan, argmin/head are refreshed
  by every _plan.Step() so post-Step reads see the current batch's values.
* UpdateCodebookEMA expresses the per-codebook scatter as: current codebook
  slice + TensorScatterAdd((1-decay)·(head - gathered), argmin) → new slice,
  then TensorConcatenate across codebook axis into the full [numCodebooks,
  codebookSize, codebookDim] tensor, then Engine.TensorCopy back into the
  _codebooks tensor object to preserve identity (future reads via the same
  reference see the update).
* Legacy VectorQuantize is now a thin adapter around VectorQuantizeTraceable
  for callers that don't need the extras.
* ForwardNativeForTrainingWithCommitment delegates to a new
  ForwardNativeForTrainingWithVQExtras that exposes the argmin/head; the
  original (forecast, commitmentLoss) contract is preserved for callers that
  don't need EMA state.
* Fused-plan fast path in TOTEM.Train restored — now safe because VQ is
  fully traceable AND the EMA runs exactly once per batch in post-Step eager
  code.

No new Tensors primitives were needed — the engine already has RFFT,
ReduceMean/Variance, TensorSqrt, TensorBroadcastSubtract/Divide, TensorFlip,
TensorConcatenate, TensorArgMin, TensorIndexSelectDiff, TensorSliceAxis,
TensorStack, StopGradient, TensorScatterAdd, and TensorCopy, all in the
autodiff/compile registry per OpRegistry.cs.

Verification:

* net8.0 + net471 + net10.0 all build clean.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* feat(training): centralize WGAN-GP + diffusion consumers through fused primitives (#1847)

Closes #1845 and #1846. Routes all four WGAN-GP critics through WganGpFusedStep
(the fused-plan primitive from PR #1843) and adds MultiSlotFusedStep wire-ups
to the three tabular diffusion consumers plus a base-class opt-in hook for
DiffusionModelBase subclasses.

## Optimizer config plumbing

Zero new API surface: IFusedOptimizerSpec.TryGetFusedOptimizerConfig already
exposes (OptimizerType, LR, Beta1, Beta2, Epsilon, WeightDecay, Schedule,
UseBf16Moments). Widened NeuralNetworkBase<T>.TryMapToFusedOptimizerConfig
from private to internal so the sibling generators in the same assembly can
reuse the existing helper.

## WGAN-GP consumers (#1845)

* CTGANGenerator, CopulaGANGenerator, TableGANGenerator, CausalGANGenerator —
  each critic training method now attempts WganGpFusedStep.TryStep FIRST with
  the discriminator's optimizer hyperparameters extracted via
  TryMapToFusedOptimizerConfig, falls back to the existing
  GpuResidentFusedStep path (secondary fused), then the eager tape (final
  fallback). The ε ∈ [0, 1]^B epsilon sampler uses
  Engine.TensorRandomUniformRange to match each critic's local
  ComputeGradientPenalty behavior.
* Non-Adam optimizers (Lion, LBFGS) that don't implement IFusedOptimizerSpec
  cleanly fall through — TryMapToFusedOptimizerConfig returns false and the
  code path skips to GpuResidentFusedStep as before.

## Diffusion consumers (#1846)

* TabDDPMGenerator — refactored to expose a slot-based forward
  (BuildTabDDPMSlots + DenoiserForwardFromTensors +
  ComputeDiffusionLossTapeFromTensors). Per-row TrainBatch loop now attempts
  MultiSlotFusedStep with (numNoisy, actualNoise, catNoisy, catClean,
  rawSinusoidalTimeEmbed) as persistent slots. The learnable
  _timestepProjection stays INSIDE the compiled forward closure so its
  weights participate in the backward pass. Plan is compiled once on the
  first row and replayed via slot-data refresh for subsequent rows.
* TabSynGenerator — TrainDiffusionBatch's per-row loop wired with
  MultiSlotFusedStep on (noisyLatent, actualNoise, projectedTimeEmbed).
  Matches the existing eager path's semantic that _timestepProjection is
  NOT in _diffMLPLayers (kept detached in the eager path too), so the
  projected embedding is precomputed host-side per row and passed as slot
  data.
* Finance/Forecasting/Foundation/CSDI —
  - ApplyInstanceNormalization rewritten with traceable engine ops
    (ReduceMean + ReduceVariance + TensorSqrt + broadcast subtract/divide) —
    same pattern as the TFC RevIN fix. The previous `.Data.Span` per-batch
    loop froze at trace time.
  - New BuildCsdiSlots + DenoiserForwardFromSlots express the DDPM x_t
    formation and packed denoising input via TensorConcatenate + engine
    scalar multiplies. Replaces the `.Data.Span[i] = xt[0, i]` fill that
    baked the trace batch's x_t into the compiled plan.
  - Train() attempts MultiSlotFusedStep first, falls back to the existing
    eager ComputeDenoisingPairTape path when the fused path can't engage.
* Diffusion/DiffusionModelBase —
  - New opt-in `protected virtual bool SupportsFusedDenoising => false;`
    property. Base default is false so no existing subclass changes
    behavior.
  - Train() attempts MultiSlotFusedStep when SupportsFusedDenoising is true
    AND the training optimizer maps cleanly to a fused config. Slots:
    (noisySample, noise). Loss = MSE(pred, noise). QAT shadow restoration
    is preserved on the fused-success path.
  - Subclasses with fully-traceable PredictNoise / PredictNoiseBatched
    (e.g. after auditing to remove `.Data.Span` host loops) can opt in via
    a single-line override; no infrastructure changes needed elsewhere.

## Verification

* net8.0, net471, net10.0 all build clean.
* No API surface changes on IGradientBasedOptimizer<T> — the existing
  IFusedOptimizerSpec interface (already implemented by all fuse-able
  optimizers) provided everything needed.
* All consumers preserve eager fallback path for non-fuse-able optimizers
  and non-GPU hosts.

Co-authored-by: franklinic <franklin@ivorycloud.com>
Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>

* feat(training): wire additional WGAN-GP + diffusion consumers through fused primitives (#1848)

Extends PR #1847 with three additional consumer wire-ups discovered during a
comprehensive audit against the primitives added in PR #1843. Each consumer
attempts the fused path first via the existing IFusedOptimizerSpec /
TryMapToFusedOptimizerConfig plumbing and falls back to eager tape training
when the fused path can't engage.

## WGAN-GP consumers (additional beyond #1845)

* NeuralNetworks/WGANGP.cs — standalone WGAN-GP class. Previously used the
  legacy flat-vector round-trip (three Critic.Predict calls per step,
  GetParameterGradients + host-side vector combine + UpdateCriticWithOptimizer).
  TrainCriticBatchWithGP now attempts WganGpFusedStep.TryStep first, using
  Critic.Layers as the parameter source and a per-batch ε ~ U(0, 1) sampler
  matching each critic's local ComputeGradientPenalty behavior. Falls back to
  the legacy path when the critic's optimizer has no fused-kernel mapping.

## Diffusion consumers (additional beyond #1846)

* NeuralNetworks/SyntheticData/FinDiffGenerator.cs — per-row DDPM training
  (Sattarov et al. 2023). TrainBatch caches a MultiSlotFusedStep across rows
  so the compiled plan is built once and replayed via slot-data refresh for
  subsequent rows. Slots: (packedDenoiserInput, targetNoise). Falls back to
  the existing Train(input, targetNoise) call on miss.

* NeuralNetworks/SyntheticData/AutoDiffTabGenerator.cs — per-row DDPM training
  with a custom TapeStepOver optimizer step. Same MultiSlotFusedStep caching
  pattern as FinDiff. Slots: (denoiserInput, targetNoise). Falls back to the
  existing tape-based TapeStepOver path on miss.

## Explicitly not wired in this PR (documented)

* NeuralNetworks/GenerativeAdversarialNetwork.cs — uses BCE loss with optional
  GP regularization (a separate auxiliary optimizer step), NOT Wasserstein +
  GP. Wiring WganGpFusedStep would change training semantics from BCE to
  Wasserstein. Requires a separate design decision.

* Finance/Forecasting/Foundation/{CCDM,MGTSD,TSDiff,TimeDiff,TimeGrad}.cs and
  Finance/Probabilistic/DiffusionTS.cs — each uses .Data.Span host-side packs
  for the denoising input (same trace-freeze issue TFC/CSDI hit in PR #1843).
  Each needs a TFC/CSDI-scale traceable rewrite BEFORE fused wiring is safe.
  Deferred to individual per-consumer PRs.

* MetaLearning/Algorithms/{MetaDDPMAlgorithm,MetaDMAlgorithm,MetaDiffAlgorithm}.cs
  — hand-rolled Vector&lt;T&gt; params with index arithmetic, NOT the
  ILayer/Engine/tape infrastructure. Would need full rewrite to fit MultiSlotFusedStep.

* DiffusionModelBase subclasses (DDPMModel, DiffWaveModel, LatentDiffusionModelBase)
  — the SupportsFusedDenoising opt-in hook (added in PR #1847) is available,
  but flipping it safely requires per-class audit of PredictNoise → UNet
  traceability. Left as follow-up work; the mechanism is in place.

## Verification

* net8.0, net471, net10.0 all build clean.
* No API surface changes.

Co-authored-by: franklinic <franklin@ivorycloud.com>
Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>

* chore(deps): bump AiDotNet.Tensors 0.111.2 -> 0.113.0 (activates #763 for fused paths)

0.113.0 publishes Tensors PR #763 — this branch's gating dependency. #763's key fix
(4C): the compiled/persistent backward now honors createGraph=true (GradientTape.
ComputeGradients previously gated the compiled path on !createGraph), so the WGAN-GP
gradient penalty's inner backward differentiates into the disc weights through the
fused compiled plan instead of silently returning zeros (issue #1844). Against 0.111.2
the fused GPU-resident WGAN-GP path would have silently degraded to plain WGAN.

The src/Training fused-primitive mirrors (WganGpFusedStep, MultiSlotFusedStep,
DpSgdFusedStep) stay as the consumer-side primitives that #1847/#1848 centralize every
consumer through; they call the public engine API, so the bump alone routes them onto
#763's fixed compiled backward. Also picks up #765 (compiled ReduceMax axis fill) and
#764 (resident-param fused-Adam fix). Native OneDNN/OpenBLAS/CLBlast bumped in lockstep
(all published). Builds green on net8.0.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: franklinic <franklin@ivorycloud.com>
Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>

This branch was successfully deployed

1 active deployment
Preview — edd68a91 Deployed Jan 23, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

feature Feature work item

Projects

None yet

Development

Successfully merging this pull request may close these issues.

test: Add integration tests for Evaluation module [P3]

2 participants