Skip to content

test: add comprehensive integration tests for PromptEngineering module - #762

Merged
ooples merged 1 commit into
masterfrom
test/promptengineering-integration-tests
Jan 24, 2026
Merged

ooples merged 1 commit into
masterfrom
test/promptengineering-integration-tests

Conversation

@ooples

@ooples ooples commented Jan 23, 2026

Copy link
Copy Markdown
Owner

Summary

  • Add 81 comprehensive integration tests for the PromptEngineering module
  • Tests cover Templates, Analysis, FewShot selectors, Compression, Chains, and Tools
  • Remove stale CLBlast project reference from test project

Tests Added

SimplePromptTemplate (14 tests)

  • Constructor validation (null, empty)
  • Variable extraction (single, multiple, duplicate)
  • Format method (substitution, missing variables)
  • Validate method
  • FromTemplate factory method

PromptMetrics (4 tests)

  • Default values
  • Property setters
  • AnalyzedAt timestamp
  • DetectedPatterns collection

PromptIssue & IssueSeverity (3 tests)

  • Default values
  • Property setters
  • Severity enum values

ValidationOptions (3 tests)

  • Default values
  • Strict preset
  • Lenient preset

CompressionResult (7 tests)

  • Default values
  • TokensSaved calculation
  • CompressionRatio calculation
  • Zero original handling
  • IsSuccessful detection
  • CompressedAt timestamp

CompressionOptions (4 tests)

  • Default values
  • Default preset
  • Aggressive preset
  • Conservative preset

FixedExampleSelector (12 tests)

  • Constructor
  • AddExample (validation, null, empty)
  • SelectExamples (order, limits, validation)
  • RemoveExample
  • GetAllExamples

ToolRegistry (16 tests)

  • Constructor
  • RegisterTool (validation, duplicates)
  • GetTool (found, not found, case insensitivity)
  • UnregisterTool
  • HasTool
  • GetAllTools
  • ExecuteTool
  • Clear
  • GenerateToolsDescription

SequentialChain (11 tests)

  • Constructor
  • AddStep (validation, chaining)
  • Run (single step, multiple steps)
  • RunAsync (sync steps, async steps, cancellation)

FewShotExample (2 tests)

  • Default values
  • Property setters

Test plan

  • All 81 tests pass on net10.0 framework
  • Build succeeds with no errors
  • No breaking changes to existing code

Closes #666

🤖 Generated with Claude Code

Add 81 integration tests covering Templates, Analysis, FewShot selectors,
Compression, Chains, and Tools components.

Tests include:
- SimplePromptTemplate: construction, variable extraction, formatting, validation
- PromptMetrics: property values and timestamp behavior
- PromptIssue and IssueSeverity: issue tracking and severity levels
- ValidationOptions: default, strict, and lenient configurations
- CompressionResult: token savings, compression ratio, success detection
- CompressionOptions: default, aggressive, conservative presets
- FixedExampleSelector: example management and selection behavior
- ToolRegistry: registration, execution, case-insensitivity, descriptions
- SequentialChain: step management, sync/async execution, cancellation
- FewShotExample: model properties

Also removes stale CLBlast project reference from test project.

Closes #666

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
@vercel

vercel Bot commented Jan 23, 2026 •

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Review Updated (UTC)
aidotnet-playground-api Ready Ready Preview, Comment Jan 23, 2026 9:46pm

@coderabbitai

coderabbitai Bot commented Jan 23, 2026 •

Copy link
Copy Markdown
Contributor

Summary by CodeRabbit

  • Tests
    • Added comprehensive integration test suite for the PromptEngineering module, including coverage for prompt processing, validation, compression, tool management, and execution workflows.

✏️ Tip: You can customize this high-level summary in your review settings.

Walkthrough

This PR removes a native CLBlast dependency reference from the test project and introduces comprehensive integration tests for the PromptEngineering module, covering template processing, metrics, validation options, compression, few-shot selection, tool orchestration, and sequential chain execution.

Changes

Cohort / File(s) Summary
Test Project Configuration
tests/AiDotNet.Tests/AiDotNetTests.csproj
Removed ProjectReference to AiDotNet.Native.CLBlast, reducing test project native dependencies.
PromptEngineering Integration Tests
tests/AiDotNet.Tests/IntegrationTests/PromptEngineering/PromptEngineeringIntegrationTests.cs
Added 1185+ lines of comprehensive integration tests covering: SimplePromptTemplate (construction, validation, formatting), PromptMetrics (defaults, timestamping, pattern detection), PromptIssue (properties, severity), ValidationOptions (presets), CompressionResult (ratio calculation, success logic), CompressionOptions (presets), FixedExampleSelector (bounds validation, selection logic), ToolRegistry (case-insensitive management, execution), SequentialChain (async/sync execution, cancellation), and FewShotExample with helper MockFunctionTool. Includes extensive edge cases, nulls, empties, and error handling validation.

Estimated code review effort

🎯 4 (Complex) | ⏱️ ~45 minutes

Suggested labels

feature

Poem

🐰 Hops through tests with glee so bright,
PromptEngineering shines in coverage light!
From chains to tools, templates true,
A thousand tests, all fresh and new!
CLBlast removed, dependencies lean,
The finest integration tests we've seen! 🌟

🚥 Pre-merge checks | ✅ 4 | ❌ 1
❌ Failed checks (1 warning)
Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 0.00% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Title check ✅ Passed The PR title accurately describes the main change: adding comprehensive integration tests for the PromptEngineering module.
Description check ✅ Passed The description is comprehensive and directly related to the changeset, detailing the 81 tests added across multiple PromptEngineering components and the CLBlast removal.
Linked Issues check ✅ Passed The PR implements the core objectives from issue #666 by adding integration tests for PromptEngineering module components including templates, metrics, validation, compression, selectors, tools, and chains.
Out of Scope Changes check ✅ Passed All changes are in-scope: 81 integration tests for PromptEngineering module and removal of a stale CLBlast dependency, both directly supporting issue #666 requirements.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing touches
  • 📝 Generate docstrings
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Post copyable unit tests in a comment
  • Commit unit tests in branch test/promptengineering-integration-tests

Comment @coderabbitai help to get the list of available commands and usage tips.

@coderabbitai coderabbitai Bot added the feature Feature work item label Jan 23, 2026
@sonarqubecloud

Copy link
Copy Markdown

@ooples
ooples merged commit 382ac75 into master Jan 24, 2026
41 checks passed
@ooples
ooples deleted the test/promptengineering-integration-tests branch January 24, 2026 02:17
ooples added a commit that referenced this pull request Jan 24, 2026
#762)

Add 81 integration tests covering Templates, Analysis, FewShot selectors,
Compression, Chains, and Tools components.

Tests include:
- SimplePromptTemplate: construction, variable extraction, formatting, validation
- PromptMetrics: property values and timestamp behavior
- PromptIssue and IssueSeverity: issue tracking and severity levels
- ValidationOptions: default, strict, and lenient configurations
- CompressionResult: token savings, compression ratio, success detection
- CompressionOptions: default, aggressive, conservative presets
- FixedExampleSelector: example management and selection behavior
- ToolRegistry: registration, execution, case-insensitivity, descriptions
- SequentialChain: step management, sync/async execution, cancellation
- FewShotExample: model properties

Also removes stale CLBlast project reference from test project.

Closes #666

Co-authored-by: franklinic <franklin@ivorycloud.com>
Co-authored-by: Claude Opus 4.5 <noreply@anthropic.com>
ooples pushed a commit that referenced this pull request Jan 24, 2026
Fixed critical bugs found during production-readiness review:

1. PredictionTypeInference (PR #765): Integer overflow when calculating
   class label range. maxClass - minClass overflows for extreme values.
   Fixed by using long for range calculation.

2. GeneticOptimizer (PR #762): IndexOf bug in tournament selection.
   When population contains duplicate prompts, IndexOf returns first
   occurrence index, causing wrong fitness selection. Fixed by tracking
   index directly.

3. NeuralProgramSynthesizer (PR #763): Absolute error comparison fails
   for large numbers. 1e12 + 0.5 vs 1e12 incorrectly fails with 1e-6
   absolute tolerance. Fixed with relative error comparison for large
   numbers and absolute for small.

4. TrainingMonitor (PR #755): CSV export shows "0" for missing metrics
   instead of empty string. FirstOrDefault returns default(T) which is
   0 for numerics, not null. Fixed by checking if match exists.

Added comprehensive tests that expose all bugs and verify fixes.

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
@ooples ooples mentioned this pull request Jan 24, 2026
3 tasks done
ooples added a commit that referenced this pull request Jan 26, 2026
* test: add comprehensive Serving module integration tests (#670)

Add 62 integration tests covering:
- ContinuousBatcherConfig (defaults, configuration, model presets)
- BatchSchedulerConfig (defaults, model-specific configs)
- BatchScheduler<T> (scheduling, preemption, cancellation, statistics)
- SequenceState<T> (lifecycle, token management, stop conditions)
- GenerationRequest<T> (configuration options)
- GenerationResult<T> (result data)
- BatcherStatistics and SchedulerStatistics
- SequenceStatus, StopReason, SchedulingPolicy enums
- Full integration workflows

Also remove stale CLBlast project reference from test project.

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* test: add comprehensive FineTuning module integration tests (#640) (#753)

Add 49 integration tests covering all 16 fine-tuning methods:
- SFT (Supervised Fine-Tuning)
- DPO (Direct Preference Optimization)
- SimPO (Simple Preference Optimization)
- ORPO (Odds Ratio Preference Optimization)
- IPO (Identity Preference Optimization)
- RDPO (Robust Direct Preference Optimization)
- KTO (Kahneman-Tversky Optimization)
- CPO (Contrastive Preference Optimization)
- RLHF-PPO (Reinforcement Learning Human Feedback)
- GRPO (Group Relative Policy Optimization)
- PRO (Pairwise Ranking Optimization)
- RRHF (Rank Responses Human Feedback)
- RSO (Statistical Rejection Sampling)
- SPIN (Self-Play Fine-Tuning)
- CAI (Constitutional AI)

Tests cover:
- Constructor initialization
- Fine-tuning workflow completion
- Data validation (SFT, Preference, RL, Ranking)
- Serialization/deserialization
- Edge cases and parameter handling
- Cancellation support

Also removes stale CLBlast project reference from test project.

Co-authored-by: franklinic <franklin@ivorycloud.com>
Co-authored-by: Claude Opus 4.5 <noreply@anthropic.com>

* test: add Evaluation integration tests (#655) (#765)

Add 51 comprehensive integration tests for the Evaluation module covering:
- Enums: PredictionType (BinaryClassification, Regression, MultiClass, MultiLabel)
- DefaultModelEvaluator: Construction, options handling
- PredictionStatsOptions: Default values, configuration
- PredictionTypeInference: Inference from vectors, matrices, tensors

Tests validate prediction type inference logic including:
- Binary classification detection (0/1 labels)
- Multi-class detection (integer labels with low unique ratio)
- Regression detection (continuous values, NaN, infinity, high unique ratio)
- Multi-label detection (matrix with multiple positives per row)

Also fixes stale CLBlast project reference in test project.

Co-authored-by: franklinic <franklin@ivorycloud.com>
Co-authored-by: Claude Opus 4.5 <noreply@anthropic.com>

* test: add ProgramSynthesis integration tests (#665) (#763)

Add 50 comprehensive integration tests for the ProgramSynthesis module covering:
- Models: CodePosition, CodeSpan, CodeLocation, CodeIssue, Program<T>
- Enums: ProgramLanguage, CodeTask, CodeIssueSeverity, SqlDialect, etc.
- Execution: CompilationDiagnostic, SqlValue, ProgramExecuteResponse
- Results: CodeGenerationResult

All tests validate model construction, property access, and default values.

Co-authored-by: franklinic <franklin@ivorycloud.com>
Co-authored-by: Claude Opus 4.5 <noreply@anthropic.com>

* test: add comprehensive integration tests for PromptEngineering module (#762)

Add 81 integration tests covering Templates, Analysis, FewShot selectors,
Compression, Chains, and Tools components.

Tests include:
- SimplePromptTemplate: construction, variable extraction, formatting, validation
- PromptMetrics: property values and timestamp behavior
- PromptIssue and IssueSeverity: issue tracking and severity levels
- ValidationOptions: default, strict, and lenient configurations
- CompressionResult: token savings, compression ratio, success detection
- CompressionOptions: default, aggressive, conservative presets
- FixedExampleSelector: example management and selection behavior
- ToolRegistry: registration, execution, case-insensitivity, descriptions
- SequentialChain: step management, sync/async execution, cancellation
- FewShotExample: model properties

Also removes stale CLBlast project reference from test project.

Closes #666

Co-authored-by: franklinic <franklin@ivorycloud.com>
Co-authored-by: Claude Opus 4.5 <noreply@anthropic.com>

* test: add comprehensive integration tests for Prototypes module (#761)

Add 58 integration tests covering PrototypeVector, PrototypeAdamOptimizer,
SimpleLinearRegression, and SimpleNeuralNetwork classes.

Tests include:
- PrototypeVector: construction, indexing, arithmetic operations, factory methods
- PrototypeAdamOptimizer: construction, parameter updates, convergence behavior
- SimpleLinearRegression: training, prediction, MSE/R2 computation
- SimpleNeuralNetwork: forward/backward passes, XOR training scenario
- Cross-component integration scenarios

Also removes stale CLBlast project reference from test project.

Closes #667

Co-authored-by: franklinic <franklin@ivorycloud.com>
Co-authored-by: Claude Opus 4.5 <noreply@anthropic.com>

* test: add comprehensive TrainingMonitoring module integration tests (#673) (#755)

Add 92 integration tests for the TrainingMonitoring module covering:

- TrainingMonitor<T>: session lifecycle, metric logging, progress tracking,
  speed stats, resource usage, issue detection, and data export
- ResourceMonitor: lifecycle, snapshots, history, threshold alerts, events
- ExperimentTracker: experiments, runs, parameters, metrics, artifacts,
  tags, deletion/restoration, persistence, and search
- NotificationManager: services, sending, filtering, buffering, events
- Dashboard implementations (ConsoleDashboard, HtmlDashboard, LiveDashboard):
  scalar/histogram logging, hyperparameters, confusion matrix, reports
- Cross-module integration tests combining multiple components

Also removes stale CLBlast project reference from test project.

Co-authored-by: franklinic <franklin@ivorycloud.com>
Co-authored-by: Claude Opus 4.5 <noreply@anthropic.com>

* Add comprehensive integration tests for AiDotNet.Serving module

Added 42 integration tests covering:
- ServingOptions configuration defaults and custom values
- StartupModel property validation
- PerformanceMetrics recording, percentiles, and statistics
- ContinuousBatchingStrategy concurrency and adaptive behavior
- TimeoutBatchingStrategy timeout-based processing
- SizeBatchingStrategy fixed-size batching
- AdaptiveBatchingStrategy latency-based adaptation
- BucketBatchingStrategy bucket index handling
- Padding strategies (Minimal, Bucket, Fixed) data preservation
- Configuration enum value verification

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* fix: replace thread.sleep with deterministic timestamp in queue time test

Instead of using Thread.Sleep(50) which is flaky under CI load,
directly set GenerationStartedAt to a known offset from CreatedAt.
This makes the test deterministic and asserts the exact expected
duration instead of using >= comparison.

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* fix: use net471-compatible Enum.GetValues in serving tests

Replace Enum.GetValues<T>() (NET 5+ only) with
Enum.GetValues(typeof(T)).Cast<T>() for .NET Framework 4.7.1 compatibility.

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

---------

Co-authored-by: franklinic <franklin@ivorycloud.com>
Co-authored-by: Claude Opus 4.5 <noreply@anthropic.com>
ooples added a commit that referenced this pull request Jan 27, 2026
* fix: production bugs in recently merged PRs

Fixed critical bugs found during production-readiness review:

1. PredictionTypeInference (PR #765): Integer overflow when calculating
   class label range. maxClass - minClass overflows for extreme values.
   Fixed by using long for range calculation.

2. GeneticOptimizer (PR #762): IndexOf bug in tournament selection.
   When population contains duplicate prompts, IndexOf returns first
   occurrence index, causing wrong fitness selection. Fixed by tracking
   index directly.

3. NeuralProgramSynthesizer (PR #763): Absolute error comparison fails
   for large numbers. 1e12 + 0.5 vs 1e12 incorrectly fails with 1e-6
   absolute tolerance. Fixed with relative error comparison for large
   numbers and absolute for small.

4. TrainingMonitor (PR #755): CSV export shows "0" for missing metrics
   instead of empty string. FirstOrDefault returns default(T) which is
   0 for numerics, not null. Fixed by checking if match exists.

Added comprehensive tests that expose all bugs and verify fixes.

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* fix: critical FineTuning bugs - log probability and SFT constructor

FineTuningBase.cs:
- Implemented ComputeLogProbabilityFromPrediction that was returning 0.0
- Added proper handling for probability distributions (cross-entropy)
- Added cosine similarity for embeddings
- Added scalar comparison for numeric values
- Added string similarity using Levenshtein distance
- Fixed single-element array bug (cosine similarity is always 1.0)

SupervisedFineTuning.cs:
- Added single-parameter constructor for Activator.CreateInstance compatibility
- This enables reflection-based instantiation used by test frameworks

MergedPRBugFixTests.cs:
- Added tests for log probability computation
- Added test verifying SFT can be instantiated with single-parameter constructor
- All 14 tests passing

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* fix: critical Diagnostics bugs - memory leak, thread safety, variance

MemoryTracker.cs:
- Added MaxHistorySize property with default of 10000
- Enforces limit to prevent unbounded memory growth in long-running apps
- Removes oldest snapshots when limit is reached (FIFO)

ProfilerSession.cs:
- Fixed thread-unsafe System.Random by using RandomHelper.ThreadSafeRandom
- System.Random is NOT thread-safe; concurrent access corrupts internal state
- Fixed variance calculation: was population (_m2/_count), now sample (_m2/(_count-1))
- Fixed call stack cleanup: handles out-of-order Stop() calls (e.g., due to exceptions)
- Now searches stack for timer instead of only checking top, prevents memory leak

MergedPRBugFixTests.cs:
- Added 5 tests for Diagnostics bug fixes
- All 19 tests passing

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* fix: ModelRegistry production bugs (PR #774)

Fixed 5 critical production-readiness bugs:

1. Mutable internal state returned: GetModel, GetModelByStage, and SearchModels
   now return clones to prevent external modification of internal state

2. Console.WriteLine in library code: Replaced with LoadErrors property for
   proper diagnostic exposure without polluting stdout

3. TOCTOU race conditions: Fixed file operations in DeleteModelVersion and
   GetModelCard to use try-catch pattern instead of File.Exists checks

4. DeleteModelVersion didn't validate modelName: Added ValidateModelName call
   for consistency with other methods

5. Lineage tracking didn't work: _lineage dictionary is now populated when
   model versions are created via RegisterModel and CreateModelVersion

Also added GetInternalModel helper method for internal mutation operations.

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* fix: PR #773 Metrics - add null checks, empty tensor handling, and validation

Bugs fixed in Metrics module:
- PSNR.ComputeBatch: Added null argument validation
- STOI.Compute: Added null argument validation
- STOI.ComputeNormalizedCorrelation: Fixed bounds check to include both arrays
- SI-SDR.Compute: Added null validation and empty tensor handling
- SNR.Compute: Added null argument validation
- IoU3D.ComputeBoxIoU: Added null validation and coordinate validation (min <= max)
- ChamferDistance.ComputeOneWay: Throws for empty target with non-empty source

Added 9 new tests to verify fixes.

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* fix: PR #772 Logging - add null checks, validation, and TOCTOU fix

Bugs fixed in Logging module:
- SummaryWriter.AddScalars: Added null check for tagScalarDict
- SummaryWriter.AddHistogram: Added null check for values arrays
- SummaryWriter.AddImage: Validate dataformats parameter (must be CHW or HWC)
- SummaryWriter.AddImages: Added null check and parameter validation
- SummaryWriter.AddPrCurve: Added null checks, length validation
- SummaryWriter.LogWeights: Added null check, handle empty weights array
- TensorBoardWriter.WriteEmbedding: Added null check, metadata length validation
- TensorBoardWriter.EncodePng: Validate pixels array length
- TensorBoardWriter.WriteProjectorConfig: Fixed TOCTOU race condition

Added 7 new tests to verify fixes.

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* fix(LanguageModels): add null/empty validation for model names and API version

Bug fixes for PR #771 LanguageModels:
- OpenAIChatModel: Validate modelName not null/empty (was NullReferenceException)
- OpenAIChatModel: Validate maxTokens > 0 early with clear error message
- AnthropicChatModel: Validate modelName not null/empty (was NullReferenceException)
- AzureOpenAIChatModel: Validate apiVersion not null/empty (caused invalid URL)
- AzureOpenAIChatModel: Validate maxTokens > 0 early with clear error message

Added 12 tests covering these validation scenarios.

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* fix(JitCompiler): add null checks and remove side effect from Validate

Bug fixes for PR #770 JitCompiler:
- IRGraph.Validate(): Remove side effect that modified TensorShapes
  (validation should be read-only)
- TensorShapeExtensions.GetElementCount(): Add null check
- TensorShapeExtensions.ShapeToString(): Add null check
- TensorShapeExtensions.GetShapeHashCode(): Add null check
- TensorShapeExtensions.GetShape(): Add null tensor check
- IRTypeExtensions.FromSystemType(): Add null Type check

Added 10 tests covering these validation scenarios.

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* fix(Interpretability): add null argument validation to helper methods

Bug fixes for PR #769 Interpretability:
- InterpretabilityMetricsHelper.GetUniqueGroups: Add null check for sensitiveFeature
- InterpretabilityMetricsHelper.GetGroupIndices: Add null check for sensitiveFeature
- InterpretabilityMetricsHelper.GetSubset: Add null checks for vector and indices
- InterpretabilityMetricsHelper.ComputePositiveRate: Add null check for predictions
- InterpretabilityMetricsHelper.ComputeTruePositiveRate: Add null checks for predictions and actualLabels
- InterpretabilityMetricsHelper.ComputeFalsePositiveRate: Add null checks for predictions and actualLabels
- InterpretabilityMetricsHelper.ComputePrecision: Add null checks for predictions and actualLabels
- InterpretableModelHelper: Add null checks for model, enabledMethods, and input parameters
  across all async methods (GetGlobalFeatureImportanceAsync, GetLocalFeatureImportanceAsync,
  GetShapValuesAsync, GetLimeExplanationAsync, GetPartialDependenceAsync,
  GetCounterfactualAsync, GetModelSpecificInterpretabilityAsync,
  GenerateTextExplanationAsync, GetFeatureInteractionAsync, ValidateFairnessAsync,
  GetAnchorExplanationAsync)

Added 10 tests covering null argument validation scenarios.

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* fix(InferenceOptimization): add null validation to optimization graph and node methods

PR #768 production-readiness fixes:
- OptimizationNode.AddInput: add null check for inputNode
- OptimizationNode.RemoveInput: add null check for inputNode
- OptimizationNode.ReplaceInput: add null checks for oldInput/newInput
- OptimizationGraph.FindNodeById: add null check for id
- OptimizationGraph.FindNodesByName: add null check for name
- IRDataTypeExtensions.FromSystemType: add null check for type
- TensorType.IsBroadcastCompatible: add null check for other
- GraphOptimizer.Optimize: add null check for graph
- GraphOptimizer.AddPass: add null check for pass

Added 14 tests covering all 9 bug fixes and valid input scenarios.

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* fix(HyperparameterOptimization): add null validation to trial pruner and distributions

PR #767 production-readiness fixes:
- TrialPruner.ReportAndCheckPrune(trial): add null check for trial
- TrialPruner.ReportAndCheckPrune(trialId): add null check for trialId
- TrialPruner.MarkComplete: add null check for trialId
- ContinuousDistribution.Sample: add null check for random
- IntegerDistribution.Sample: add null check for random
- CategoricalDistribution.Sample: add null check for random
- HyperparameterOptimizerBase.FindBestTrial: add null check for completedTrials
- HyperparameterOptimizerBase.EvaluateTrialSafely: add null checks for trial, objectiveFunction, parameters

Added 13 tests covering all 8 bug fixes and valid input scenarios.

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* fix(ExperimentTracking): fix production bugs in experiment tracking

Bug fixes:
- StartRun: persist experiment timestamp after Touch() to maintain disk consistency
- SearchRuns: validate maxResults > 0 to prevent invalid queries
- GetRunDirectory: use TryGetValue with descriptive InvalidOperationException
- DeleteExperiment/DeleteRun: handle IOException gracefully when directory deletion fails
- LogArtifact: validate extracted filename isn't empty for root paths
- LogArtifacts: wrap UnauthorizedAccessException with descriptive message
- Add null validation to GetExperiment, GetRun, ListRuns, DeleteExperiment, DeleteRun
- Add null validation to SerializeToJson, DeserializeFromJson, GetLatestMetric

Added 11 tests covering:
- Null argument validation
- MaxResults validation
- Timestamp persistence verification
- Thread safety for concurrent metric logging
- Graceful error handling for directory operations

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* fix(DistributedTraining): fix production bugs in distributed communication

Bug fixes:
- CommunicationManager.Broadcast: add root range validation to prevent invalid operations
- CommunicationManager.Scatter: add root range validation to prevent invalid operations
- CommunicationManager.ReduceScatter: add null validation for data parameter
- ParameterAnalyzer.CalculateDistributionStats: throw for null/empty groups (consistent API)
- InMemoryCommunicationBackend.PerformReduction: validate all vectors have same length
- InMemoryCommunicationBackend.Receive: validate message size BEFORE dequeuing (prevents data loss)
- ShardingConfiguration factory methods: add null validation for better error messages

Added 10 tests covering:
- Root validation in Broadcast/Scatter
- Null data validation in ReduceScatter
- Null/empty groups in ParameterAnalyzer
- Null backend in ShardingConfiguration factory methods
- Single-process AllReduce optimization path

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* fix(Reasoning): fix production bugs in reasoning module

Bug fixes:
- ReasoningChain.AddStep: add null validation for step parameter
- DepthFirstSearch.SearchAsync: add null validation for generator, evaluator, config (consistent with other search algorithms)
- MonteCarloTreeSearch constructor: validate numSimulations >= 1 and explorationConstant >= 0
- BreadthFirstSearch.CollectAllNodes: use iterative approach instead of recursion to prevent StackOverflow on deep trees

Added 9 tests covering:
- Null step validation in ReasoningChain
- Step number auto-increment
- MCTS constructor parameter validation
- ThoughtNode path reconstruction
- ThoughtNode leaf/root detection
- ReasoningConfig default values

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* fix(Serialization): add validation for negative dimensions in JSON converters

- VectorJsonConverter: validate that length is non-negative
- TensorJsonConverter: validate that shape array is not empty
- TensorJsonConverter: validate that all shape dimensions are non-negative
- Added 8 tests for serialization validation

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* fix(Serving): add parameter validation to batching strategies and padding

Add production bug fixes for PR #758 (Serving module):

Batching Strategies:
- ContinuousBatchingStrategy: validate maxConcurrency >= 1, minWaitMs >= 0, targetLatencyMs > 0
- TimeoutBatchingStrategy: validate timeoutMs >= 0, maxBatchSize >= 1
- AdaptiveBatchingStrategy: validate minBatchSize >= 1, maxBatchSize >= minBatchSize,
  maxWaitMs >= 0, targetLatencyMs > 0, latencyToleranceFactor > 0
- SizeBatchingStrategy: validate batchSize >= 1, maxWaitMs >= 0
- BucketBatchingStrategy: validate maxBatchSize >= 1, maxWaitMs >= 0, bucket values > 0

Monitoring:
- PerformanceMetrics: validate maxSamples >= 1, maxQueueDepthSamples >= 1

Padding Strategies (MinimalPaddingStrategy, BucketPaddingStrategy, FixedSizePaddingStrategy):
- PadBatch: validate no null vectors in input array
- UnpadBatch: validate originalLengths are non-negative
- FixedSizePaddingStrategy: validate fixedLength > 0
- BucketPaddingStrategy: validate bucketSizes not null/empty

Added 25 validation tests across BatchingStrategyTests, PaddingStrategyTests,
and PerformanceMetricsTests.

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* fix(Tokenization): add parameter validation to tokenizers and vocabulary

Add production bug fixes for PR #757 (Tokenization module):

Vocabulary:
- AddTokens: validate tokens parameter is not null
- Constructor(Dictionary): validate tokenToId parameter is not null

Tokenizers:
- BpeTokenizer.Train: validate corpus not null, vocabSize >= 1
- WordPieceTokenizer.Train: validate corpus not null, vocabSize >= 1
- WordPieceTokenizer constructor: validate maxInputCharsPerWord >= 1
- CharacterTokenizer.Train: validate corpus not null, minFrequency >= 1

MidiTokenizer:
- Constructor: validate ticksPerBeat >= 1, numVelocityBins >= 1
- CreateREMI: validate ticksPerBeat >= 1, numVelocityBins >= 1
- CreateCPWord: validate ticksPerBeat >= 1, numVelocityBins >= 1
- CreateSimpleNote: validate ticksPerBeat >= 1

TokenizationConfig:
- ParallelBatchThreshold: validate value >= 1 via property setter

Added 17 validation tests across BpeTokenizerTests, CharacterTokenizerTests,
WordPieceTokenizerTests, SpecializedTokenizerTests, and VocabularyTests.

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* fix(Tools): add parameter validation and robust type conversion

- Add upper bound validation (100) for topK in VectorSearchTool and RAGTool
- Add validation that topKAfterRerank cannot exceed topK in RAGTool
- Make ToolBase TryGetInt/TryGetDouble/TryGetBool handle type conversion errors gracefully
- Add 34 unit tests covering parameter validation and edge cases

PR #756 bug fixes - prevent performance issues from excessive topK values and
improve robustness when receiving invalid JSON property types.

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* fix(DistributedTraining): fix initialization order and add parameter validation

Bugs fixed:

1. Initialization order bug in derived classes (DivideByZeroException):
   - Base class was calling InitializeSharding() before derived class fields were set
   - Fix: Remove InitializeSharding() from base constructor, each derived class now
     calls it explicitly at the end of their constructor

2. ShardingConfiguration missing learningRate validation:
   - Added validation that learningRate > 0

3. PipelineParallelModel missing microBatchSize validation:
   - Added validation that microBatchSize >= 1

4. HybridShardedModel missing parallelism size validation:
   - Added validation that pipelineParallelSize >= 1
   - Added validation that tensorParallelSize >= 1

5. Inconsistent learning rate usage:
   - DDPModel, FSDPModel, PipelineParallelModel, HybridShardedModel were using
     hardcoded 0.01 instead of Config.LearningRate
   - Fixed to use Config.LearningRate consistently

Affected files:
- ShardedModelBase.cs - removed InitializeSharding() call from constructor
- DDPModel.cs, FSDPModel.cs, ZeRO1Model.cs, ZeRO2Model.cs - added InitializeSharding() call
- TensorParallelModel.cs, PipelineParallelModel.cs, HybridShardedModel.cs - same + validation
- ShardingConfiguration.cs - added learningRate validation

Tests: Added 26 validation tests in DistributedTrainingValidationTests.cs

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* fix(LoRA): fix matrix indexing bugs and initialization order issues

- Add input/output size validation in LoRALayer constructor
- Add pruningInterval validation in AdaLoRAAdapter to prevent division by zero
- Fix matrix indexing in MergeWeights (returns [inputSize, outputSize]):
  - LoRAAdapterBase.MergeToDenseOrFullyConnected
  - StandardLoRAAdapter.MergeToOriginalLayer
  - QLoRAAdapter.MergeToOriginalLayer
  - DoRAAdapter.Forward() and MergeToOriginalLayer
- Fix DefaultLoRAConfiguration.CreateAdapter to handle different constructor signatures
- Fix VeRAAdapter initialization order bug:
  - Move scaling vector init to CreateLoRALayer (called before ParameterCount)
  - Add UpdateParametersFromLayers override for VeRA-specific parameter sync
- Add 22 validation tests covering all bug fixes

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* fix(Autodiff): add numerical stability validation to tensor operations

- Add division by zero validation in TensorOperations.Divide
- Add non-positive value validation in TensorOperations.Log
- Add negative value validation in TensorOperations.Sqrt
- Handle sqrt(0) edge case in backward pass (use 0 instead of infinity)
- Fix null axes handling in TensorOperations.Sum OperationParams
- Add segmentSize validation in GradientCheckpointing.SequentialCheckpoint
- Add 19 validation tests covering all bug fixes

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* fix: address pr review comments and code scanning alerts

- ExperimentTracker: add logging to IOException catch block as comment indicated
- MonteCarloTreeSearch: add NaN/Infinity guards to explorationConstant validation
- TensorJsonConverter: remove validation rejecting empty shapes (scalars are valid)
- BucketBatchingStrategy: add validation for empty bucket boundaries array
- ToolBase: update exception handling to catch JsonSerializationException
- LoRALayer: update XML docs to reflect actual exception types thrown
- MergedPRBugFixTests: update size-mismatch test to actually verify behavior
- FineTuningBase: fix generic catch clauses and collection equality check
- ShardedModelBase: use lazy initialization to avoid virtual calls in constructor

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* fix: correct misleading comment and tighten assertion in sum operation

- Update comment in TensorOperations.cs to accurately describe OperationParams behavior
- Tighten Sum_NullAxes_OperationParamsHandledCorrectly test assertion to properly
  verify that Axes key is NOT present when axes parameter is null

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* fix: remove virtual calls from distributed training model constructors

- Remove direct InitializeSharding() calls from DDPModel, FSDPModel,
  ZeRO1Model, and ZeRO2Model constructors
- Remove redundant InitializeSharding() call from TensorParallelModel's
  OnBeforeInitializeSharding() method
- All models now rely on lazy initialization via EnsureShardingInitialized()
  in the base class to avoid virtual calls in constructors

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* fix: address coderabbit review comments for pr #783

- TensorOperations.cs: convert if/else to ternary for operationparams assignment
- HybridShardedModel.cs: move pendingconfig.value = null to onbeforeinitializesharding where it is consumed (lazy init compatibility)
- FineTuningBase.cs: move string check before ienumerable check since strings implement ienumerable<char>

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* fix: address additional coderabbit review comments

- TensorOperations.cs: only emit "Axes" metadata when non-null AND non-empty (empty array means sum-all like null)
- FineTuningBase.cs: treat null-null elements as matches in sequence matching fallback

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* fix: guard numeric conversion against mixed runtime types

Check both prediction and target are numeric before calling Convert.ToDouble
to avoid throwing when TOutput is object and types don't match.

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* chore: remove accidentally committed _playground_publish folder

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* feat: add comprehensive ActiveLearning integration tests

Add 62 integration tests covering:
- EntropySampling (8 tests)
- UncertaintySampling (5 tests)
- BALD (4 tests)
- RandomSampling (3 tests)
- MarginSampling (2 tests)
- LeastConfidenceSampling (2 tests)
- VariationRatios (2 tests)
- DiversitySampling (5 tests with all methods/metrics)
- CoreSetSelection (2 tests)
- HybridSampling (3 tests with all combination methods)
- InformationDensity (2 tests)
- DensityWeightedSampling (1 test)
- ExpectedModelChange (2 tests)
- BatchBALD (2 tests)
- QueryByCommittee (2 tests)
- Edge cases and mathematical validation (6 tests)

Tests include:
- Correct batch size validation
- Null argument handling
- Score range validation
- Mathematical properties (entropy of uniform/certain distributions)
- Diversity selection across clusters

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* feat: add comprehensive ContinualLearning integration tests

Add 53 integration tests covering:
- ElasticWeightConsolidation (EWC): 12 tests
- SynapticIntelligence (SI): 5 tests
- MemoryAwareSynapses (MAS): 4 tests
- GradientEpisodicMemory (GEM): 4 tests
- LearningWithoutForgetting (LwF): 4 tests
- OnlineEWC: 3 tests
- ExperienceReplay: 3 tests
- PackNet: 3 tests
- ProgressiveNeuralNetworks: 2 tests
- GenerativeReplay: 2 tests
- AveragedGEM (A-GEM): 3 tests
- Edge cases and cross-strategy validation: 8 tests

Tests verify:
- Constructor initialization
- BeforeTask/AfterTask lifecycle
- ComputeLoss returns non-negative values
- ModifyGradients produces valid output
- Reset clears stored data
- Multiple sequential tasks work correctly
- Lambda property can be modified
- Null argument handling

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* feat: add comprehensive curriculum learning integration tests

- Add 35 integration tests for CurriculumLearning components
- Test schedulers: LinearScheduler, SelfPacedScheduler, CompetenceBasedScheduler
- Test difficulty estimators: LossBased, ConfidenceBased, TransferBased, ExpertDefined, Ensemble
- Test edge cases: empty arrays, zero epochs, reset behavior
- Fix bug in CurriculumSchedulerBase.GetIndicesAtPhase that crashed on empty arrays

Note: Tests document a bug with generic T? default values - for unconstrained generics,
T? with default value is 0.0 for value types, not null. Tests work around this by
providing explicit values for optional parameters.

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* test: add comprehensive selfsupervisedlearning integration tests

Add 52 integration tests covering the SelfSupervisedLearning module:
- NT-Xent Loss tests (4 tests): temperature effects, gradient computation
- InfoNCE Loss tests (5 tests): memory bank integration, accuracy metrics
- BYOL Loss tests (5 tests): cosine similarity, symmetric loss computation
- Linear Projector tests (6 tests): shape validation, gradient backprop
- MLP Projector tests (5 tests): batch norm, training mode, backward pass
- Symmetric Projector tests (6 tests): predictor head, combined operations
- Memory Bank tests (11 tests): FIFO queue, momentum updates, sampling
- Momentum Encoder static method tests (3 tests): cosine schedule
- Edge cases and error handling tests (6 tests)

Uses RandomHelper for secure random number generation and proper
Tensor API patterns for cross-framework compatibility.

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* test: add nestedlearning integration tests for associativememory and contextflow

Add 30 integration tests covering the NestedLearning module:

AssociativeMemory tests (12 tests):
- Constructor validation and initialization
- Associate/Retrieve with dimension validation
- Capacity limit enforcement (FIFO)
- Association matrix updates with Hebbian learning
- Clear/Reset functionality
- Multiple associations and large capacity handling

ContextFlow tests (15 tests):
- Constructor and matrix initialization
- PropagateContext with level validation
- ComputeContextGradients backpropagation
- UpdateFlow transformation matrix updates
- GetContextState and CompressContext operations
- Reset clears all context states
- Independent states across multiple levels

Integration tests (3 tests):
- Combined AssociativeMemory + ContextFlow workflow
- Large capacity stress testing
- Sequential propagation state accumulation

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* test: add mixedprecision integration tests (39 tests)

Adds comprehensive integration tests for MixedPrecision module:
- LossScaler: scaling, unscaling, overflow detection, dynamic scaling
- MixedPrecisionConfig: defaults follow NVIDIA recommendations
- MixedPrecisionContext: FP32/FP16 weight management, gradient preparation
- Full workflow tests: training iterations, overflow recovery

Closes #642

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* test: add comprehensive integration tests for adversarial robustness module

- Add 83 integration tests for AdversarialRobustness module:
  - FGSM attack tests (8 tests)
  - PGD attack tests (4 tests)
  - CW attack tests (3 tests)
  - AutoAttack tests (2 tests)
  - Adversarial training defense tests (7 tests)
  - Randomized smoothing certification tests (8 tests)
  - Interval bound propagation tests (7 tests)
  - CROWN verification tests (6 tests)
  - Safety filter tests (11 tests)
  - Rule-based content classifier tests (11 tests)
  - Integration scenarios (6 tests)
  - Edge cases and error handling (10 tests)

- Fix bug in FGSMAttack: add null check for trueLabel parameter
  to throw ArgumentNullException instead of InvalidOperationException

Closes #631

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* test: fix randomhelper usage in finetuning integration tests

Replace new Random(42) with RandomHelper.CreateSeededRandom(42) to follow
project security standards for random number generation.

Closes #640

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* test: fix lora integration tests and loralayer exception consistency

- Fix LoRALayer to throw ArgumentOutOfRangeException consistently for all
  invalid rank values (was throwing ArgumentException for rank > min(in, out))
- Update test to expect ArgumentOutOfRangeException
- Replace new Random(42) with RandomHelper.CreateSeededRandom(42)

Closes #641

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* test: add knowledgedistillation integration tests

Comprehensive integration tests covering:
- DistillationLoss: constructor, compute loss/gradient, edge cases
- All distillation strategies: Feature, Attention, Contrastive,
  Probabilistic, Hybrid, Curriculum, Adaptive, Variational,
  NeuronSelectivity, Relational, SimilarityPreserving, FlowBased,
  FactorTransfer
- DistillationStrategyFactory: all strategy types
- DistillationForwardResult and DistillationCheckpointConfig
- IntermediateActivations: add/get/count
- Numerical stability and edge case testing

Total: 85 tests, 0 bugs found

Closes #636

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* test: add comprehensive physicsinformed integration tests

Add 94 integration tests covering the PhysicsInformed module:
- PhysicsInformedLoss tests (loss computation, gradients, edge cases)
- PDE tests (HeatEquation, WaveEquation, PoissonEquation, BurgersEquation,
  AllenCahnEquation, KdV, AdvectionDiffusion)
- PINN tests (PhysicsInformedNeuralNetwork, VariationalPINN, DeepRitzMethod)
- Neural Operator tests (FourierNeuralOperator, FourierLayer)
- ScientificML tests (HamiltonianNeuralNetwork)
- TrainingHistory, PDEDerivatives, PDEResidualGradient tests
- Edge cases and numerical stability tests
- Integration workflow tests
- Serialization tests

Closes #637

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* fix: move tensor type check before text fullname check in datamodalitydetector

The Tensor check was happening after the fullName-based Text check, which
caused Tensor<T> types to be incorrectly detected as Text modality since
the fullName could contain substrings matching other checks.

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* test: add comprehensive augmentation integration tests with 106 tests

Coverage includes:
- Image augmentations (GaussianNoise, Brightness, Contrast, Cutout, RandomCrop, etc.)
- Audio augmentations (AudioNoise, TimeStretch, PitchShift, etc.)
- Video augmentations (TemporalFlip, FrameDropout, SpeedChange)
- Text augmentations (RandomDeletion, RandomInsertion, SynonymReplacement, etc.)
- Object detection augmentations (BoundingBox, Keypoint transformations)
- Compose and auto-augment pipelines
- DataModalityDetector type detection

Tests verify correct behavior for each augmentation category including:
- Apply with probability, deterministic behavior, edge cases
- Parameter validation, composition, and context management

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* test: add comprehensive federated learning integration tests

Adds 64 integration tests for the FederatedLearning module covering:

- Aggregation strategies: FedAvg, FedProx, FedBN with weighted averaging
- Byzantine-robust aggregators: Krum, MultiKrum, Bulyan, TrimmedMean,
  WinsorizedMean, GeometricMedian, RFA
- Client selection strategies: UniformRandom, WeightedRandom, Stratified,
  AvailabilityAware, PerformanceAware, Clustered
- Privacy mechanisms: GaussianDifferentialPrivacy with clipping and noise
- Privacy accounting: BasicComposition and RDP privacy accountants
- Cryptography: HKDF key derivation
- Server optimizers: FedAdam, FedAdagrad, FedYogi, FedAvgM
- Heterogeneity corrections: SCAFFOLD, FedNova, FedDyn
- Secure aggregation: SecureAggregationVector, ThresholdSecureAggregationVector
- Additional tests: GaussianDifferentialPrivacyVector

Tests verify mathematical correctness and proper API behavior.

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* test: add checkpoint management integration tests

Adds 38 integration tests for the CheckpointManagement module covering:

- CheckpointManager construction and directory creation
- Auto-checkpointing configuration (save frequency, keep last, save on improvement)
- ShouldAutoSaveCheckpoint logic (frequency-based and improvement-based triggers)
- UpdateAutoSaveState for tracking last save step and best metric values
- AutoCheckpointState properties and ToString formatting
- Thread-safe concurrent configuration updates and state reads
- ListCheckpoints, LoadLatestCheckpoint, LoadBestCheckpoint edge cases
- CleanupOldCheckpoints and CleanupKeepBest cleanup strategies
- MetricOptimizationDirection enum values
- Path validation and nested directory support

Tests verify proper state tracking for minimization and maximization scenarios.

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* test: add 64 integration tests for Configuration module

- AutoMLBudgetOptions: default values, property setting, all presets
- RLTrainingOptions: default values, property setting, callbacks
- RLCheckpointConfig: default values, property setting
- RLEarlyStoppingConfig: default values, generic type support
- ExplorationScheduleConfig: default values, all decay types
- InferenceOptimizationConfig: default values, validation, all enum values
- ResNetConfiguration: variants, block counts, expansion, factory methods
- BenchmarkingOptions: default values, federated configs
- CurriculumLearningOptions: schedule types, difficulty estimators
- Supporting options classes: SelfPacedOptions, CompetenceBasedOptions

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* feat: add dataprocessor integration tests and fix empty matrix bug

- Add 24 comprehensive integration tests for DataProcessor module
- Tests cover DefaultDataPreprocessor, DataProcessorOptions, SplitData
- Tests validate preprocessing pipeline with Matrix, Vector, and Tensor types
- Fix DivideByZeroException in FeatureSelectorHelper.CreateFilteredData when
  handling empty matrices (0 rows)
- The fix returns a properly dimensioned empty matrix instead of attempting
  to call FromColumns on empty column vectors

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* feat: add dataversioncontrol integration tests (62 tests)

- Add comprehensive integration tests for DataVersionControl module
- Tests cover versioning, hashing, integrity verification, run linking
- Tests cover tagging, lineage tracking, snapshots, and persistence
- Tests verify thread safety with concurrent version creation
- Tests model classes: DatasetVersion, DatasetLineage, DatasetStatistics, etc.

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* test: add 65 integration tests for dataversioning module

Tests cover:
- Constructor and directory structure creation
- CreateDataset with validation, metadata, and duplicate handling
- AddVersion with files, directories, and deduplication
- GetVersion by ID, version number, and "latest"
- ListVersions and ListDatasets with ordering
- GetDataPath with directory structure preservation
- CompareVersions detecting additions, removals, modifications
- DeleteVersion and DeleteDataset with file cleanup
- RecordLineage and GetLineage with recursive upstream resolution
- Persistence across instance restarts
- Model classes (DatasetInfo, DataVersion, DataFileInfo, DataVersionDiff, DataLineage)
- Thread safety for concurrent operations
- Edge cases (large file count, empty directories, special characters)

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* fix: correct json parsing bug in agent response extraction

The ExtractJsonFromResponse method in all agent classes used a non-greedy
regex pattern that incorrectly matched nested JSON objects. For example,
with input:
  {"reasoning_steps": [...], "tool_calls": [{"tool_name": "X"}]}
The regex would match up to the first "}" (end of tool_calls inner object)
instead of the outer closing brace.

Fixed by implementing proper brace-balancing algorithm that:
- Tracks brace count while respecting string boundaries
- Handles escape sequences within strings
- Returns the complete outermost JSON object

Also added comprehensive integration tests for all agent types:
- Agent (ReAct pattern): 15 tests
- ChainOfThoughtAgent: 8 tests
- PlanAndExecuteAgent: 7 tests
- RAGAgent: 12 tests
- AgentBase: 3 tests
- JSON parsing edge cases: 4 tests
- Thread safety: 1 test

Total: 50 new tests for the Agents module

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* fix: correct training data dimensions in vision and tabular benchmarking tests

The tests were training KNN models on 1-feature data but then running
benchmarks that generate test data with different feature dimensions:
- CIFAR10/CIFAR100: 3072 features (32x32x3 pixels flattened)
- TabularNonIID: FeatureCount features (3 in this test)

Fixed by providing training data that matches the benchmark feature dimensions:
- CIFAR tests: 2 samples with 3072 features each (normalized pixel values)
- TabularNonIID test: 3 samples with 3 features each

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* test: add comprehensive benchmarking integration tests

Add 98 integration tests for the Benchmarking module covering:
- BenchmarkSuiteRegistry: display names, suite categories, mappings
- BenchmarkReport: construction, serialization, aggregation, statistics
- BenchmarkMetricValue: creation, edge cases, formatting
- BenchmarkExecutionStatus: all status values and transitions
- BenchmarkSuite enum: all 23 benchmark suites
- BenchmarkMetric enum: all 8 metric types
- BenchmarkSuiteKind enum: all 3 suite kinds
- Model classes: BenchmarkSuiteReport, BenchmarkDataSelectionSummary

Tests verify behavior for:
- Enum value coverage for all benchmark-related enums
- Display name generation and formatting
- Report creation and metric aggregation
- Suite category classification (ReasoningSuite vs DatasetSuite)
- Edge cases like empty reports and default values

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* test: add comprehensive computervision integration tests (66 tests)

Tests cover:
- BoundingBox format conversions (XYXY, XYWH, CXCYWH, YOLO)
- BoundingBox IoU, Area, Clip, IsValid operations
- Detection and DetectionResult classes
- DetectionStatistics and BatchDetectionResult
- NMS (standard, class-aware, batched)
- IoU variants (IoU, GIoU, DIoU, CIoU)
- GIoULoss for bounding box regression
- SORT tracker with Kalman filtering
- Track, TrackingOptions, TrackingResult classes
- End-to-end integration scenarios

Closes #647

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* fix: net471 compatibility for string.split and enum.getvalues

- Use char array overload for string.Split with StringSplitOptions
- Use typeof() overload for Enum.GetValues instead of generic version

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* fix: address pr review comments for integration tests

- Fix divide-by-zero guard in ActiveLearning CreatePoolWithKnownUncertainty
- Add missing assertion in RandomSampling_InformativenessScores test
- Add meaningful assertion in Rotation_ApplyWithTargets test
- Rename AllRegularizationStrategies to CoreRegularizationStrategies for accuracy
- Fix reflection test to assert if type exists but methods don't
- Fix greedy regex in ChainOfThoughtAgent JSON extraction
- Add BindingFlags for non-public property reflection in Benchmarking
- Add cleanup for default checkpoints directory
- Fix culture-invariant decimal comparison in AutoCheckpointState test

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* fix: address additional pr review comments

- Rename misleading test names in DataVersionControl
- Make MockChatModel thread-safe with Interlocked and ConcurrentBag
- Guard against deleting pre-existing default directories in DataVersioning
- Add delays to prevent timestamp-tie flakes in ordered tests

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* fix: replace generic catch clauses with specific exception types in finetuning

- Replace `catch (Exception ex) when (...)` patterns with specific
  exception type catches (InvalidCastException, FormatException, OverflowException)
- Extract ComputeSequenceMatchLogProbability helper to reduce code duplication
- Fix Equals on collections issue by using EqualityComparer<TOutput>.Default

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* fix: replace generic catch clause with specific exception types in ssl tests

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* fix: guard against single-class divide-by-zero in mock model

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* fix: update lora tests to expect argumentoutofrangeexception

The code correctly throws ArgumentOutOfRangeException (more specific than
ArgumentException) for invalid rank values. Tests now expect the correct
exception type.

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* fix: update vblora test to expect argumentoutofrangeexception

The test Constructor_WithRankExceedingBankSizeA_ThrowsArgumentException was
expecting ArgumentException but the code correctly throws
ArgumentOutOfRangeException which is more specific for range validation.

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

---------

Co-authored-by: franklinic <franklin@ivorycloud.com>
Co-authored-by: Claude Opus 4.5 <noreply@anthropic.com>

This branch was successfully deployed

1 active deployment
Preview — 579297d8 Deployed Jan 23, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

feature Feature work item

Projects

None yet

Development

Successfully merging this pull request may close these issues.

test: Add integration tests for PromptEngineering module [P3]

2 participants