test: add comprehensive integration tests for ModelRegistry module - #774
Conversation
Add 54 tests covering model registration and versioning: - Constructor and directory initialization - RegisterModel with tags and validation - CreateModelVersion for existing models - GetModel, GetLatestModel, and GetModelByStage - TransitionModelStage with archive options - ListModels with filters and tags - ListModelVersions for specific models - SearchModels with various criteria - UpdateModelMetadata and UpdateModelTags - DeleteModelVersion and DeleteModel - CompareModels for version comparison - GetModelLineage for tracking - ArchiveModel for model lifecycle - GetModelStoragePath for file location - ModelCard operations (attach, get, generate, save) - Persistence across registry instances Tests run on both .NET 10.0 and .NET Framework 4.7.1. Closes #664 Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
|
The latest updates on your projects. Learn more about Vercel for GitHub.
|
Summary by CodeRabbit
✏️ Tip: You can customize this high-level summary in your review settings. WalkthroughThis pull request introduces a comprehensive integration test suite for the ModelRegistry module, covering model registration, version management, metadata handling, model retrieval, lifecycle transitions, persistence, and error handling scenarios across 1025 lines of test code. Changes
Estimated code review effort🎯 3 (Moderate) | ⏱️ ~20 minutes Suggested labels
Poem
🚥 Pre-merge checks | ✅ 4 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (4 passed)
✏️ Tip: You can configure your own custom pre-merge checks in the settings. ✨ Finishing touches
🧪 Generate unit tests (beta)
Comment |
|
Fixed 5 critical production-readiness bugs: 1. Mutable internal state returned: GetModel, GetModelByStage, and SearchModels now return clones to prevent external modification of internal state 2. Console.WriteLine in library code: Replaced with LoadErrors property for proper diagnostic exposure without polluting stdout 3. TOCTOU race conditions: Fixed file operations in DeleteModelVersion and GetModelCard to use try-catch pattern instead of File.Exists checks 4. DeleteModelVersion didn't validate modelName: Added ValidateModelName call for consistency with other methods 5. Lineage tracking didn't work: _lineage dictionary is now populated when model versions are created via RegisterModel and CreateModelVersion Also added GetInternalModel helper method for internal mutation operations. Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
* fix: production bugs in recently merged PRs Fixed critical bugs found during production-readiness review: 1. PredictionTypeInference (PR #765): Integer overflow when calculating class label range. maxClass - minClass overflows for extreme values. Fixed by using long for range calculation. 2. GeneticOptimizer (PR #762): IndexOf bug in tournament selection. When population contains duplicate prompts, IndexOf returns first occurrence index, causing wrong fitness selection. Fixed by tracking index directly. 3. NeuralProgramSynthesizer (PR #763): Absolute error comparison fails for large numbers. 1e12 + 0.5 vs 1e12 incorrectly fails with 1e-6 absolute tolerance. Fixed with relative error comparison for large numbers and absolute for small. 4. TrainingMonitor (PR #755): CSV export shows "0" for missing metrics instead of empty string. FirstOrDefault returns default(T) which is 0 for numerics, not null. Fixed by checking if match exists. Added comprehensive tests that expose all bugs and verify fixes. Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * fix: critical FineTuning bugs - log probability and SFT constructor FineTuningBase.cs: - Implemented ComputeLogProbabilityFromPrediction that was returning 0.0 - Added proper handling for probability distributions (cross-entropy) - Added cosine similarity for embeddings - Added scalar comparison for numeric values - Added string similarity using Levenshtein distance - Fixed single-element array bug (cosine similarity is always 1.0) SupervisedFineTuning.cs: - Added single-parameter constructor for Activator.CreateInstance compatibility - This enables reflection-based instantiation used by test frameworks MergedPRBugFixTests.cs: - Added tests for log probability computation - Added test verifying SFT can be instantiated with single-parameter constructor - All 14 tests passing Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * fix: critical Diagnostics bugs - memory leak, thread safety, variance MemoryTracker.cs: - Added MaxHistorySize property with default of 10000 - Enforces limit to prevent unbounded memory growth in long-running apps - Removes oldest snapshots when limit is reached (FIFO) ProfilerSession.cs: - Fixed thread-unsafe System.Random by using RandomHelper.ThreadSafeRandom - System.Random is NOT thread-safe; concurrent access corrupts internal state - Fixed variance calculation: was population (_m2/_count), now sample (_m2/(_count-1)) - Fixed call stack cleanup: handles out-of-order Stop() calls (e.g., due to exceptions) - Now searches stack for timer instead of only checking top, prevents memory leak MergedPRBugFixTests.cs: - Added 5 tests for Diagnostics bug fixes - All 19 tests passing Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * fix: ModelRegistry production bugs (PR #774) Fixed 5 critical production-readiness bugs: 1. Mutable internal state returned: GetModel, GetModelByStage, and SearchModels now return clones to prevent external modification of internal state 2. Console.WriteLine in library code: Replaced with LoadErrors property for proper diagnostic exposure without polluting stdout 3. TOCTOU race conditions: Fixed file operations in DeleteModelVersion and GetModelCard to use try-catch pattern instead of File.Exists checks 4. DeleteModelVersion didn't validate modelName: Added ValidateModelName call for consistency with other methods 5. Lineage tracking didn't work: _lineage dictionary is now populated when model versions are created via RegisterModel and CreateModelVersion Also added GetInternalModel helper method for internal mutation operations. Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * fix: PR #773 Metrics - add null checks, empty tensor handling, and validation Bugs fixed in Metrics module: - PSNR.ComputeBatch: Added null argument validation - STOI.Compute: Added null argument validation - STOI.ComputeNormalizedCorrelation: Fixed bounds check to include both arrays - SI-SDR.Compute: Added null validation and empty tensor handling - SNR.Compute: Added null argument validation - IoU3D.ComputeBoxIoU: Added null validation and coordinate validation (min <= max) - ChamferDistance.ComputeOneWay: Throws for empty target with non-empty source Added 9 new tests to verify fixes. Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * fix: PR #772 Logging - add null checks, validation, and TOCTOU fix Bugs fixed in Logging module: - SummaryWriter.AddScalars: Added null check for tagScalarDict - SummaryWriter.AddHistogram: Added null check for values arrays - SummaryWriter.AddImage: Validate dataformats parameter (must be CHW or HWC) - SummaryWriter.AddImages: Added null check and parameter validation - SummaryWriter.AddPrCurve: Added null checks, length validation - SummaryWriter.LogWeights: Added null check, handle empty weights array - TensorBoardWriter.WriteEmbedding: Added null check, metadata length validation - TensorBoardWriter.EncodePng: Validate pixels array length - TensorBoardWriter.WriteProjectorConfig: Fixed TOCTOU race condition Added 7 new tests to verify fixes. Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * fix(LanguageModels): add null/empty validation for model names and API version Bug fixes for PR #771 LanguageModels: - OpenAIChatModel: Validate modelName not null/empty (was NullReferenceException) - OpenAIChatModel: Validate maxTokens > 0 early with clear error message - AnthropicChatModel: Validate modelName not null/empty (was NullReferenceException) - AzureOpenAIChatModel: Validate apiVersion not null/empty (caused invalid URL) - AzureOpenAIChatModel: Validate maxTokens > 0 early with clear error message Added 12 tests covering these validation scenarios. Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * fix(JitCompiler): add null checks and remove side effect from Validate Bug fixes for PR #770 JitCompiler: - IRGraph.Validate(): Remove side effect that modified TensorShapes (validation should be read-only) - TensorShapeExtensions.GetElementCount(): Add null check - TensorShapeExtensions.ShapeToString(): Add null check - TensorShapeExtensions.GetShapeHashCode(): Add null check - TensorShapeExtensions.GetShape(): Add null tensor check - IRTypeExtensions.FromSystemType(): Add null Type check Added 10 tests covering these validation scenarios. Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * fix(Interpretability): add null argument validation to helper methods Bug fixes for PR #769 Interpretability: - InterpretabilityMetricsHelper.GetUniqueGroups: Add null check for sensitiveFeature - InterpretabilityMetricsHelper.GetGroupIndices: Add null check for sensitiveFeature - InterpretabilityMetricsHelper.GetSubset: Add null checks for vector and indices - InterpretabilityMetricsHelper.ComputePositiveRate: Add null check for predictions - InterpretabilityMetricsHelper.ComputeTruePositiveRate: Add null checks for predictions and actualLabels - InterpretabilityMetricsHelper.ComputeFalsePositiveRate: Add null checks for predictions and actualLabels - InterpretabilityMetricsHelper.ComputePrecision: Add null checks for predictions and actualLabels - InterpretableModelHelper: Add null checks for model, enabledMethods, and input parameters across all async methods (GetGlobalFeatureImportanceAsync, GetLocalFeatureImportanceAsync, GetShapValuesAsync, GetLimeExplanationAsync, GetPartialDependenceAsync, GetCounterfactualAsync, GetModelSpecificInterpretabilityAsync, GenerateTextExplanationAsync, GetFeatureInteractionAsync, ValidateFairnessAsync, GetAnchorExplanationAsync) Added 10 tests covering null argument validation scenarios. Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * fix(InferenceOptimization): add null validation to optimization graph and node methods PR #768 production-readiness fixes: - OptimizationNode.AddInput: add null check for inputNode - OptimizationNode.RemoveInput: add null check for inputNode - OptimizationNode.ReplaceInput: add null checks for oldInput/newInput - OptimizationGraph.FindNodeById: add null check for id - OptimizationGraph.FindNodesByName: add null check for name - IRDataTypeExtensions.FromSystemType: add null check for type - TensorType.IsBroadcastCompatible: add null check for other - GraphOptimizer.Optimize: add null check for graph - GraphOptimizer.AddPass: add null check for pass Added 14 tests covering all 9 bug fixes and valid input scenarios. Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * fix(HyperparameterOptimization): add null validation to trial pruner and distributions PR #767 production-readiness fixes: - TrialPruner.ReportAndCheckPrune(trial): add null check for trial - TrialPruner.ReportAndCheckPrune(trialId): add null check for trialId - TrialPruner.MarkComplete: add null check for trialId - ContinuousDistribution.Sample: add null check for random - IntegerDistribution.Sample: add null check for random - CategoricalDistribution.Sample: add null check for random - HyperparameterOptimizerBase.FindBestTrial: add null check for completedTrials - HyperparameterOptimizerBase.EvaluateTrialSafely: add null checks for trial, objectiveFunction, parameters Added 13 tests covering all 8 bug fixes and valid input scenarios. Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * fix(ExperimentTracking): fix production bugs in experiment tracking Bug fixes: - StartRun: persist experiment timestamp after Touch() to maintain disk consistency - SearchRuns: validate maxResults > 0 to prevent invalid queries - GetRunDirectory: use TryGetValue with descriptive InvalidOperationException - DeleteExperiment/DeleteRun: handle IOException gracefully when directory deletion fails - LogArtifact: validate extracted filename isn't empty for root paths - LogArtifacts: wrap UnauthorizedAccessException with descriptive message - Add null validation to GetExperiment, GetRun, ListRuns, DeleteExperiment, DeleteRun - Add null validation to SerializeToJson, DeserializeFromJson, GetLatestMetric Added 11 tests covering: - Null argument validation - MaxResults validation - Timestamp persistence verification - Thread safety for concurrent metric logging - Graceful error handling for directory operations Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * fix(DistributedTraining): fix production bugs in distributed communication Bug fixes: - CommunicationManager.Broadcast: add root range validation to prevent invalid operations - CommunicationManager.Scatter: add root range validation to prevent invalid operations - CommunicationManager.ReduceScatter: add null validation for data parameter - ParameterAnalyzer.CalculateDistributionStats: throw for null/empty groups (consistent API) - InMemoryCommunicationBackend.PerformReduction: validate all vectors have same length - InMemoryCommunicationBackend.Receive: validate message size BEFORE dequeuing (prevents data loss) - ShardingConfiguration factory methods: add null validation for better error messages Added 10 tests covering: - Root validation in Broadcast/Scatter - Null data validation in ReduceScatter - Null/empty groups in ParameterAnalyzer - Null backend in ShardingConfiguration factory methods - Single-process AllReduce optimization path Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * fix(Reasoning): fix production bugs in reasoning module Bug fixes: - ReasoningChain.AddStep: add null validation for step parameter - DepthFirstSearch.SearchAsync: add null validation for generator, evaluator, config (consistent with other search algorithms) - MonteCarloTreeSearch constructor: validate numSimulations >= 1 and explorationConstant >= 0 - BreadthFirstSearch.CollectAllNodes: use iterative approach instead of recursion to prevent StackOverflow on deep trees Added 9 tests covering: - Null step validation in ReasoningChain - Step number auto-increment - MCTS constructor parameter validation - ThoughtNode path reconstruction - ThoughtNode leaf/root detection - ReasoningConfig default values Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * fix(Serialization): add validation for negative dimensions in JSON converters - VectorJsonConverter: validate that length is non-negative - TensorJsonConverter: validate that shape array is not empty - TensorJsonConverter: validate that all shape dimensions are non-negative - Added 8 tests for serialization validation Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * fix(Serving): add parameter validation to batching strategies and padding Add production bug fixes for PR #758 (Serving module): Batching Strategies: - ContinuousBatchingStrategy: validate maxConcurrency >= 1, minWaitMs >= 0, targetLatencyMs > 0 - TimeoutBatchingStrategy: validate timeoutMs >= 0, maxBatchSize >= 1 - AdaptiveBatchingStrategy: validate minBatchSize >= 1, maxBatchSize >= minBatchSize, maxWaitMs >= 0, targetLatencyMs > 0, latencyToleranceFactor > 0 - SizeBatchingStrategy: validate batchSize >= 1, maxWaitMs >= 0 - BucketBatchingStrategy: validate maxBatchSize >= 1, maxWaitMs >= 0, bucket values > 0 Monitoring: - PerformanceMetrics: validate maxSamples >= 1, maxQueueDepthSamples >= 1 Padding Strategies (MinimalPaddingStrategy, BucketPaddingStrategy, FixedSizePaddingStrategy): - PadBatch: validate no null vectors in input array - UnpadBatch: validate originalLengths are non-negative - FixedSizePaddingStrategy: validate fixedLength > 0 - BucketPaddingStrategy: validate bucketSizes not null/empty Added 25 validation tests across BatchingStrategyTests, PaddingStrategyTests, and PerformanceMetricsTests. Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * fix(Tokenization): add parameter validation to tokenizers and vocabulary Add production bug fixes for PR #757 (Tokenization module): Vocabulary: - AddTokens: validate tokens parameter is not null - Constructor(Dictionary): validate tokenToId parameter is not null Tokenizers: - BpeTokenizer.Train: validate corpus not null, vocabSize >= 1 - WordPieceTokenizer.Train: validate corpus not null, vocabSize >= 1 - WordPieceTokenizer constructor: validate maxInputCharsPerWord >= 1 - CharacterTokenizer.Train: validate corpus not null, minFrequency >= 1 MidiTokenizer: - Constructor: validate ticksPerBeat >= 1, numVelocityBins >= 1 - CreateREMI: validate ticksPerBeat >= 1, numVelocityBins >= 1 - CreateCPWord: validate ticksPerBeat >= 1, numVelocityBins >= 1 - CreateSimpleNote: validate ticksPerBeat >= 1 TokenizationConfig: - ParallelBatchThreshold: validate value >= 1 via property setter Added 17 validation tests across BpeTokenizerTests, CharacterTokenizerTests, WordPieceTokenizerTests, SpecializedTokenizerTests, and VocabularyTests. Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * fix(Tools): add parameter validation and robust type conversion - Add upper bound validation (100) for topK in VectorSearchTool and RAGTool - Add validation that topKAfterRerank cannot exceed topK in RAGTool - Make ToolBase TryGetInt/TryGetDouble/TryGetBool handle type conversion errors gracefully - Add 34 unit tests covering parameter validation and edge cases PR #756 bug fixes - prevent performance issues from excessive topK values and improve robustness when receiving invalid JSON property types. Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * fix(DistributedTraining): fix initialization order and add parameter validation Bugs fixed: 1. Initialization order bug in derived classes (DivideByZeroException): - Base class was calling InitializeSharding() before derived class fields were set - Fix: Remove InitializeSharding() from base constructor, each derived class now calls it explicitly at the end of their constructor 2. ShardingConfiguration missing learningRate validation: - Added validation that learningRate > 0 3. PipelineParallelModel missing microBatchSize validation: - Added validation that microBatchSize >= 1 4. HybridShardedModel missing parallelism size validation: - Added validation that pipelineParallelSize >= 1 - Added validation that tensorParallelSize >= 1 5. Inconsistent learning rate usage: - DDPModel, FSDPModel, PipelineParallelModel, HybridShardedModel were using hardcoded 0.01 instead of Config.LearningRate - Fixed to use Config.LearningRate consistently Affected files: - ShardedModelBase.cs - removed InitializeSharding() call from constructor - DDPModel.cs, FSDPModel.cs, ZeRO1Model.cs, ZeRO2Model.cs - added InitializeSharding() call - TensorParallelModel.cs, PipelineParallelModel.cs, HybridShardedModel.cs - same + validation - ShardingConfiguration.cs - added learningRate validation Tests: Added 26 validation tests in DistributedTrainingValidationTests.cs Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * fix(LoRA): fix matrix indexing bugs and initialization order issues - Add input/output size validation in LoRALayer constructor - Add pruningInterval validation in AdaLoRAAdapter to prevent division by zero - Fix matrix indexing in MergeWeights (returns [inputSize, outputSize]): - LoRAAdapterBase.MergeToDenseOrFullyConnected - StandardLoRAAdapter.MergeToOriginalLayer - QLoRAAdapter.MergeToOriginalLayer - DoRAAdapter.Forward() and MergeToOriginalLayer - Fix DefaultLoRAConfiguration.CreateAdapter to handle different constructor signatures - Fix VeRAAdapter initialization order bug: - Move scaling vector init to CreateLoRALayer (called before ParameterCount) - Add UpdateParametersFromLayers override for VeRA-specific parameter sync - Add 22 validation tests covering all bug fixes Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * fix(Autodiff): add numerical stability validation to tensor operations - Add division by zero validation in TensorOperations.Divide - Add non-positive value validation in TensorOperations.Log - Add negative value validation in TensorOperations.Sqrt - Handle sqrt(0) edge case in backward pass (use 0 instead of infinity) - Fix null axes handling in TensorOperations.Sum OperationParams - Add segmentSize validation in GradientCheckpointing.SequentialCheckpoint - Add 19 validation tests covering all bug fixes Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * fix: address pr review comments and code scanning alerts - ExperimentTracker: add logging to IOException catch block as comment indicated - MonteCarloTreeSearch: add NaN/Infinity guards to explorationConstant validation - TensorJsonConverter: remove validation rejecting empty shapes (scalars are valid) - BucketBatchingStrategy: add validation for empty bucket boundaries array - ToolBase: update exception handling to catch JsonSerializationException - LoRALayer: update XML docs to reflect actual exception types thrown - MergedPRBugFixTests: update size-mismatch test to actually verify behavior - FineTuningBase: fix generic catch clauses and collection equality check - ShardedModelBase: use lazy initialization to avoid virtual calls in constructor Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * fix: correct misleading comment and tighten assertion in sum operation - Update comment in TensorOperations.cs to accurately describe OperationParams behavior - Tighten Sum_NullAxes_OperationParamsHandledCorrectly test assertion to properly verify that Axes key is NOT present when axes parameter is null Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * fix: remove virtual calls from distributed training model constructors - Remove direct InitializeSharding() calls from DDPModel, FSDPModel, ZeRO1Model, and ZeRO2Model constructors - Remove redundant InitializeSharding() call from TensorParallelModel's OnBeforeInitializeSharding() method - All models now rely on lazy initialization via EnsureShardingInitialized() in the base class to avoid virtual calls in constructors Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * fix: address coderabbit review comments for pr #783 - TensorOperations.cs: convert if/else to ternary for operationparams assignment - HybridShardedModel.cs: move pendingconfig.value = null to onbeforeinitializesharding where it is consumed (lazy init compatibility) - FineTuningBase.cs: move string check before ienumerable check since strings implement ienumerable<char> Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * fix: address additional coderabbit review comments - TensorOperations.cs: only emit "Axes" metadata when non-null AND non-empty (empty array means sum-all like null) - FineTuningBase.cs: treat null-null elements as matches in sequence matching fallback Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * fix: guard numeric conversion against mixed runtime types Check both prediction and target are numeric before calling Convert.ToDouble to avoid throwing when TOutput is object and types don't match. Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * chore: remove accidentally committed _playground_publish folder Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * feat: add comprehensive ActiveLearning integration tests Add 62 integration tests covering: - EntropySampling (8 tests) - UncertaintySampling (5 tests) - BALD (4 tests) - RandomSampling (3 tests) - MarginSampling (2 tests) - LeastConfidenceSampling (2 tests) - VariationRatios (2 tests) - DiversitySampling (5 tests with all methods/metrics) - CoreSetSelection (2 tests) - HybridSampling (3 tests with all combination methods) - InformationDensity (2 tests) - DensityWeightedSampling (1 test) - ExpectedModelChange (2 tests) - BatchBALD (2 tests) - QueryByCommittee (2 tests) - Edge cases and mathematical validation (6 tests) Tests include: - Correct batch size validation - Null argument handling - Score range validation - Mathematical properties (entropy of uniform/certain distributions) - Diversity selection across clusters Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * feat: add comprehensive ContinualLearning integration tests Add 53 integration tests covering: - ElasticWeightConsolidation (EWC): 12 tests - SynapticIntelligence (SI): 5 tests - MemoryAwareSynapses (MAS): 4 tests - GradientEpisodicMemory (GEM): 4 tests - LearningWithoutForgetting (LwF): 4 tests - OnlineEWC: 3 tests - ExperienceReplay: 3 tests - PackNet: 3 tests - ProgressiveNeuralNetworks: 2 tests - GenerativeReplay: 2 tests - AveragedGEM (A-GEM): 3 tests - Edge cases and cross-strategy validation: 8 tests Tests verify: - Constructor initialization - BeforeTask/AfterTask lifecycle - ComputeLoss returns non-negative values - ModifyGradients produces valid output - Reset clears stored data - Multiple sequential tasks work correctly - Lambda property can be modified - Null argument handling Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * feat: add comprehensive curriculum learning integration tests - Add 35 integration tests for CurriculumLearning components - Test schedulers: LinearScheduler, SelfPacedScheduler, CompetenceBasedScheduler - Test difficulty estimators: LossBased, ConfidenceBased, TransferBased, ExpertDefined, Ensemble - Test edge cases: empty arrays, zero epochs, reset behavior - Fix bug in CurriculumSchedulerBase.GetIndicesAtPhase that crashed on empty arrays Note: Tests document a bug with generic T? default values - for unconstrained generics, T? with default value is 0.0 for value types, not null. Tests work around this by providing explicit values for optional parameters. Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * test: add comprehensive selfsupervisedlearning integration tests Add 52 integration tests covering the SelfSupervisedLearning module: - NT-Xent Loss tests (4 tests): temperature effects, gradient computation - InfoNCE Loss tests (5 tests): memory bank integration, accuracy metrics - BYOL Loss tests (5 tests): cosine similarity, symmetric loss computation - Linear Projector tests (6 tests): shape validation, gradient backprop - MLP Projector tests (5 tests): batch norm, training mode, backward pass - Symmetric Projector tests (6 tests): predictor head, combined operations - Memory Bank tests (11 tests): FIFO queue, momentum updates, sampling - Momentum Encoder static method tests (3 tests): cosine schedule - Edge cases and error handling tests (6 tests) Uses RandomHelper for secure random number generation and proper Tensor API patterns for cross-framework compatibility. Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * test: add nestedlearning integration tests for associativememory and contextflow Add 30 integration tests covering the NestedLearning module: AssociativeMemory tests (12 tests): - Constructor validation and initialization - Associate/Retrieve with dimension validation - Capacity limit enforcement (FIFO) - Association matrix updates with Hebbian learning - Clear/Reset functionality - Multiple associations and large capacity handling ContextFlow tests (15 tests): - Constructor and matrix initialization - PropagateContext with level validation - ComputeContextGradients backpropagation - UpdateFlow transformation matrix updates - GetContextState and CompressContext operations - Reset clears all context states - Independent states across multiple levels Integration tests (3 tests): - Combined AssociativeMemory + ContextFlow workflow - Large capacity stress testing - Sequential propagation state accumulation Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * test: add mixedprecision integration tests (39 tests) Adds comprehensive integration tests for MixedPrecision module: - LossScaler: scaling, unscaling, overflow detection, dynamic scaling - MixedPrecisionConfig: defaults follow NVIDIA recommendations - MixedPrecisionContext: FP32/FP16 weight management, gradient preparation - Full workflow tests: training iterations, overflow recovery Closes #642 Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * test: add comprehensive integration tests for adversarial robustness module - Add 83 integration tests for AdversarialRobustness module: - FGSM attack tests (8 tests) - PGD attack tests (4 tests) - CW attack tests (3 tests) - AutoAttack tests (2 tests) - Adversarial training defense tests (7 tests) - Randomized smoothing certification tests (8 tests) - Interval bound propagation tests (7 tests) - CROWN verification tests (6 tests) - Safety filter tests (11 tests) - Rule-based content classifier tests (11 tests) - Integration scenarios (6 tests) - Edge cases and error handling (10 tests) - Fix bug in FGSMAttack: add null check for trueLabel parameter to throw ArgumentNullException instead of InvalidOperationException Closes #631 Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * test: fix randomhelper usage in finetuning integration tests Replace new Random(42) with RandomHelper.CreateSeededRandom(42) to follow project security standards for random number generation. Closes #640 Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * test: fix lora integration tests and loralayer exception consistency - Fix LoRALayer to throw ArgumentOutOfRangeException consistently for all invalid rank values (was throwing ArgumentException for rank > min(in, out)) - Update test to expect ArgumentOutOfRangeException - Replace new Random(42) with RandomHelper.CreateSeededRandom(42) Closes #641 Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * test: add knowledgedistillation integration tests Comprehensive integration tests covering: - DistillationLoss: constructor, compute loss/gradient, edge cases - All distillation strategies: Feature, Attention, Contrastive, Probabilistic, Hybrid, Curriculum, Adaptive, Variational, NeuronSelectivity, Relational, SimilarityPreserving, FlowBased, FactorTransfer - DistillationStrategyFactory: all strategy types - DistillationForwardResult and DistillationCheckpointConfig - IntermediateActivations: add/get/count - Numerical stability and edge case testing Total: 85 tests, 0 bugs found Closes #636 Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * test: add comprehensive physicsinformed integration tests Add 94 integration tests covering the PhysicsInformed module: - PhysicsInformedLoss tests (loss computation, gradients, edge cases) - PDE tests (HeatEquation, WaveEquation, PoissonEquation, BurgersEquation, AllenCahnEquation, KdV, AdvectionDiffusion) - PINN tests (PhysicsInformedNeuralNetwork, VariationalPINN, DeepRitzMethod) - Neural Operator tests (FourierNeuralOperator, FourierLayer) - ScientificML tests (HamiltonianNeuralNetwork) - TrainingHistory, PDEDerivatives, PDEResidualGradient tests - Edge cases and numerical stability tests - Integration workflow tests - Serialization tests Closes #637 Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * fix: move tensor type check before text fullname check in datamodalitydetector The Tensor check was happening after the fullName-based Text check, which caused Tensor<T> types to be incorrectly detected as Text modality since the fullName could contain substrings matching other checks. Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * test: add comprehensive augmentation integration tests with 106 tests Coverage includes: - Image augmentations (GaussianNoise, Brightness, Contrast, Cutout, RandomCrop, etc.) - Audio augmentations (AudioNoise, TimeStretch, PitchShift, etc.) - Video augmentations (TemporalFlip, FrameDropout, SpeedChange) - Text augmentations (RandomDeletion, RandomInsertion, SynonymReplacement, etc.) - Object detection augmentations (BoundingBox, Keypoint transformations) - Compose and auto-augment pipelines - DataModalityDetector type detection Tests verify correct behavior for each augmentation category including: - Apply with probability, deterministic behavior, edge cases - Parameter validation, composition, and context management Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * test: add comprehensive federated learning integration tests Adds 64 integration tests for the FederatedLearning module covering: - Aggregation strategies: FedAvg, FedProx, FedBN with weighted averaging - Byzantine-robust aggregators: Krum, MultiKrum, Bulyan, TrimmedMean, WinsorizedMean, GeometricMedian, RFA - Client selection strategies: UniformRandom, WeightedRandom, Stratified, AvailabilityAware, PerformanceAware, Clustered - Privacy mechanisms: GaussianDifferentialPrivacy with clipping and noise - Privacy accounting: BasicComposition and RDP privacy accountants - Cryptography: HKDF key derivation - Server optimizers: FedAdam, FedAdagrad, FedYogi, FedAvgM - Heterogeneity corrections: SCAFFOLD, FedNova, FedDyn - Secure aggregation: SecureAggregationVector, ThresholdSecureAggregationVector - Additional tests: GaussianDifferentialPrivacyVector Tests verify mathematical correctness and proper API behavior. Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * test: add checkpoint management integration tests Adds 38 integration tests for the CheckpointManagement module covering: - CheckpointManager construction and directory creation - Auto-checkpointing configuration (save frequency, keep last, save on improvement) - ShouldAutoSaveCheckpoint logic (frequency-based and improvement-based triggers) - UpdateAutoSaveState for tracking last save step and best metric values - AutoCheckpointState properties and ToString formatting - Thread-safe concurrent configuration updates and state reads - ListCheckpoints, LoadLatestCheckpoint, LoadBestCheckpoint edge cases - CleanupOldCheckpoints and CleanupKeepBest cleanup strategies - MetricOptimizationDirection enum values - Path validation and nested directory support Tests verify proper state tracking for minimization and maximization scenarios. Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * test: add 64 integration tests for Configuration module - AutoMLBudgetOptions: default values, property setting, all presets - RLTrainingOptions: default values, property setting, callbacks - RLCheckpointConfig: default values, property setting - RLEarlyStoppingConfig: default values, generic type support - ExplorationScheduleConfig: default values, all decay types - InferenceOptimizationConfig: default values, validation, all enum values - ResNetConfiguration: variants, block counts, expansion, factory methods - BenchmarkingOptions: default values, federated configs - CurriculumLearningOptions: schedule types, difficulty estimators - Supporting options classes: SelfPacedOptions, CompetenceBasedOptions Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * feat: add dataprocessor integration tests and fix empty matrix bug - Add 24 comprehensive integration tests for DataProcessor module - Tests cover DefaultDataPreprocessor, DataProcessorOptions, SplitData - Tests validate preprocessing pipeline with Matrix, Vector, and Tensor types - Fix DivideByZeroException in FeatureSelectorHelper.CreateFilteredData when handling empty matrices (0 rows) - The fix returns a properly dimensioned empty matrix instead of attempting to call FromColumns on empty column vectors Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * feat: add dataversioncontrol integration tests (62 tests) - Add comprehensive integration tests for DataVersionControl module - Tests cover versioning, hashing, integrity verification, run linking - Tests cover tagging, lineage tracking, snapshots, and persistence - Tests verify thread safety with concurrent version creation - Tests model classes: DatasetVersion, DatasetLineage, DatasetStatistics, etc. Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * test: add 65 integration tests for dataversioning module Tests cover: - Constructor and directory structure creation - CreateDataset with validation, metadata, and duplicate handling - AddVersion with files, directories, and deduplication - GetVersion by ID, version number, and "latest" - ListVersions and ListDatasets with ordering - GetDataPath with directory structure preservation - CompareVersions detecting additions, removals, modifications - DeleteVersion and DeleteDataset with file cleanup - RecordLineage and GetLineage with recursive upstream resolution - Persistence across instance restarts - Model classes (DatasetInfo, DataVersion, DataFileInfo, DataVersionDiff, DataLineage) - Thread safety for concurrent operations - Edge cases (large file count, empty directories, special characters) Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * fix: correct json parsing bug in agent response extraction The ExtractJsonFromResponse method in all agent classes used a non-greedy regex pattern that incorrectly matched nested JSON objects. For example, with input: {"reasoning_steps": [...], "tool_calls": [{"tool_name": "X"}]} The regex would match up to the first "}" (end of tool_calls inner object) instead of the outer closing brace. Fixed by implementing proper brace-balancing algorithm that: - Tracks brace count while respecting string boundaries - Handles escape sequences within strings - Returns the complete outermost JSON object Also added comprehensive integration tests for all agent types: - Agent (ReAct pattern): 15 tests - ChainOfThoughtAgent: 8 tests - PlanAndExecuteAgent: 7 tests - RAGAgent: 12 tests - AgentBase: 3 tests - JSON parsing edge cases: 4 tests - Thread safety: 1 test Total: 50 new tests for the Agents module Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * fix: correct training data dimensions in vision and tabular benchmarking tests The tests were training KNN models on 1-feature data but then running benchmarks that generate test data with different feature dimensions: - CIFAR10/CIFAR100: 3072 features (32x32x3 pixels flattened) - TabularNonIID: FeatureCount features (3 in this test) Fixed by providing training data that matches the benchmark feature dimensions: - CIFAR tests: 2 samples with 3072 features each (normalized pixel values) - TabularNonIID test: 3 samples with 3 features each Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * test: add comprehensive benchmarking integration tests Add 98 integration tests for the Benchmarking module covering: - BenchmarkSuiteRegistry: display names, suite categories, mappings - BenchmarkReport: construction, serialization, aggregation, statistics - BenchmarkMetricValue: creation, edge cases, formatting - BenchmarkExecutionStatus: all status values and transitions - BenchmarkSuite enum: all 23 benchmark suites - BenchmarkMetric enum: all 8 metric types - BenchmarkSuiteKind enum: all 3 suite kinds - Model classes: BenchmarkSuiteReport, BenchmarkDataSelectionSummary Tests verify behavior for: - Enum value coverage for all benchmark-related enums - Display name generation and formatting - Report creation and metric aggregation - Suite category classification (ReasoningSuite vs DatasetSuite) - Edge cases like empty reports and default values Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * test: add comprehensive computervision integration tests (66 tests) Tests cover: - BoundingBox format conversions (XYXY, XYWH, CXCYWH, YOLO) - BoundingBox IoU, Area, Clip, IsValid operations - Detection and DetectionResult classes - DetectionStatistics and BatchDetectionResult - NMS (standard, class-aware, batched) - IoU variants (IoU, GIoU, DIoU, CIoU) - GIoULoss for bounding box regression - SORT tracker with Kalman filtering - Track, TrackingOptions, TrackingResult classes - End-to-end integration scenarios Closes #647 Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * fix: net471 compatibility for string.split and enum.getvalues - Use char array overload for string.Split with StringSplitOptions - Use typeof() overload for Enum.GetValues instead of generic version Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * fix: address pr review comments for integration tests - Fix divide-by-zero guard in ActiveLearning CreatePoolWithKnownUncertainty - Add missing assertion in RandomSampling_InformativenessScores test - Add meaningful assertion in Rotation_ApplyWithTargets test - Rename AllRegularizationStrategies to CoreRegularizationStrategies for accuracy - Fix reflection test to assert if type exists but methods don't - Fix greedy regex in ChainOfThoughtAgent JSON extraction - Add BindingFlags for non-public property reflection in Benchmarking - Add cleanup for default checkpoints directory - Fix culture-invariant decimal comparison in AutoCheckpointState test Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * fix: address additional pr review comments - Rename misleading test names in DataVersionControl - Make MockChatModel thread-safe with Interlocked and ConcurrentBag - Guard against deleting pre-existing default directories in DataVersioning - Add delays to prevent timestamp-tie flakes in ordered tests Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * fix: replace generic catch clauses with specific exception types in finetuning - Replace `catch (Exception ex) when (...)` patterns with specific exception type catches (InvalidCastException, FormatException, OverflowException) - Extract ComputeSequenceMatchLogProbability helper to reduce code duplication - Fix Equals on collections issue by using EqualityComparer<TOutput>.Default Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * fix: replace generic catch clause with specific exception types in ssl tests Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * fix: guard against single-class divide-by-zero in mock model Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * fix: update lora tests to expect argumentoutofrangeexception The code correctly throws ArgumentOutOfRangeException (more specific than ArgumentException) for invalid rank values. Tests now expect the correct exception type. Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * fix: update vblora test to expect argumentoutofrangeexception The test Constructor_WithRankExceedingBankSizeA_ThrowsArgumentException was expecting ArgumentException but the code correctly throws ArgumentOutOfRangeException which is more specific for range validation. Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> --------- Co-authored-by: franklinic <franklin@ivorycloud.com> Co-authored-by: Claude Opus 4.5 <noreply@anthropic.com>



Summary
Test Coverage
Test Results
Key Test Patterns
Closes #664
🤖 Generated with Claude Code