feat: add ILayeredModel<T> interface for layer-aware model access (#462) - #847
Conversation
Implements Phase 1 of issue #462: the ILayeredModel<T> interface that provides per-layer metadata access across the AiDotNet stack, enabling pipeline parallelism, quantization, pruning, LoRA, and checkpointing tools to make layer-aware decisions instead of operating on flat vectors. New types: - ILayeredModel<T>: interface with Layers, LayerCount, GetLayerInfo, GetAllLayerInfo, ValidatePartitionPoint - LayerInfo<T>: metadata class with parameter offsets, shapes, category, cost estimates (FLOPs, activation memory) - LayerCategory: 15-value enum for automated per-layer decisions Changes: - INeuralNetwork<T> now extends ILayeredModel<T> - LayerBase<T> gains GetLayerCategory(), EstimateFlops(), EstimateActivationMemory() virtual methods - NeuralNetworkBase<T> implements all ILayeredModel<T> members Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
…nd meta-learning Phase 2 subsystem integration for issue #462: PipelineParallelModel: - Layer-aware partitioning using ILayeredModel metadata - FLOPs-balanced stage assignment respecting layer boundaries - Falls back to naive parameter splitting for non-neural models TrainingMemoryManager: - New ComputeCheckpointIndices(ILayeredModel) overload - Uses LayerCategory and EstimatedActivationMemory for selective checkpointing decisions - Layers with >1MB activation memory auto-checkpointed DefaultLoRAConfiguration: - ShouldApplyLoRA(LayerCategory) for category-based adapter decisions - ApplyLoRAToModel(ILayeredModel) for automatic adapter injection - Dense, Conv, Attention, Recurrent, Embedding, FeedForward, Graph, Residual categories eligible for LoRA MetaLearnerBase: - ApplyPerLayerGradients using LayerInfo parameter offsets - Enables per-layer learning rates for fine-grained meta-learning Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Phase 3 of issue #462: ModelExporterBase: - GetInputShape now auto-infers from first layer via ILayeredModel - New GetOutputShape auto-infers from last layer via ILayeredModel - New GetLayerSummary provides layer metadata for export metadata ExportConfiguration: - Add OutputShape property for explicit output shape specification PruningConfig: - Add CategorySparsityTargets for per-LayerCategory sparsity levels - Enables attention layers at lower sparsity than dense layers Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
|
The latest updates on your projects. Learn more about Vercel for GitHub.
|
|
Note Reviews pausedIt looks like this branch is under active development. To avoid overwhelming you with review comments due to an influx of new commits, CodeRabbit has automatically paused this review. You can configure this behavior by changing the Use the following commands to manage reviews:
Use the checkboxes below for quick actions:
WalkthroughAdds a layer-aware model abstraction (ILayeredModel, LayerInfo, LayerCategory) and integrates layer-aware logic across exporters, partitioners, quantization, pruning, LoRA, meta-learning, checkpointing, distillation, model metadata, distributed sharding, and tests. Some duplicate method definitions and placeholder implementations appear in the diff. Changes
Sequence Diagram(s)sequenceDiagram
participant Caller as Pipeline Parallelizer
participant Pipeline as PipelineParallelModel
participant Model as ILayeredModel<T>
participant Sharder as InitializeLayerAwareSharding
Caller->>Pipeline: InitializeSharding(model, stages, ...)
Pipeline->>Model: Query layering (ILayeredModel?)
alt layered with layers
Pipeline->>Sharder: InitializeLayerAwareSharding(model, stages)
Sharder->>Model: GetAllLayerInfo()
Model-->>Sharder: List<LayerInfo<T>>
Sharder->>Sharder: Read ParameterOffset/ParameterCount/EstimatedFlops
Sharder->>Sharder: Compute per-stage FLOP targets & assign layers
Sharder-->>Pipeline: Return per-stage shard indices/sizes
else non-layered
Pipeline->>Pipeline: Fallback param-count chunking
end
Pipeline-->>Caller: Shards configured
sequenceDiagram
participant Caller as LoRA Applier
participant LoRA as DefaultLoRAConfiguration
participant Model as ILayeredModel<T>
participant Layer as ILayer<T>
Caller->>LoRA: ApplyLoRAToModel(layeredModel)
LoRA->>Model: GetAllLayerInfo()
Model-->>LoRA: List<LayerInfo<T>>
loop per layer
LoRA->>LoRA: ShouldApplyLoRA(category)?
alt apply
LoRA->>Layer: ApplyLoRA(originalLayer)
Layer-->>LoRA: LoRA-wrapped layer
else skip
LoRA-->>LoRA: keep original
end
end
LoRA-->>Caller: Layers returned (applied/unchanged)
Estimated code review effort🎯 5 (Critical) | ⏱️ ~120 minutes Possibly related PRs
Suggested labels
🚥 Pre-merge checks | ✅ 5 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (5 passed)
✏️ Tip: You can configure your own custom pre-merge checks in the settings. ✨ Finishing touches🧪 Generate unit tests (beta)
Comment |
There was a problem hiding this comment.
Pull request overview
Adds a new ILayeredModel<T> abstraction to expose layer-level metadata (category, shapes, parameter offsets, cost estimates) and wires it into existing subsystems (pipeline sharding, checkpointing, LoRA, meta-learning, export, pruning).
Changes:
- Introduces
ILayeredModel<T>,LayerCategory, andLayerInfo<T>for layer-aware tooling. - Implements layer metadata access in
NeuralNetworkBase<T>and adds default categorization/cost estimation hooks toLayerBase<T>. - Updates pipeline parallel sharding, activation checkpointing, LoRA eligibility, meta-learning gradient application, and exporter shape inference to use layer metadata.
Reviewed changes
Copilot reviewed 15 out of 15 changed files in this pull request and generated 7 comments.
Show a summary per file
| File | Description |
|---|---|
| tests/AiDotNet.Tests/Helpers/MockNeuralNetwork.cs | Adds minimal ILayeredModel<double> stub for tests. |
| tests/AiDotNet.Tests/Helpers/ContinualLearningTestHelper.cs | Adds ILayeredModel<T> implementation to test helper model. |
| src/Training/Memory/TrainingMemoryManager.cs | Adds layer-aware checkpoint index selection using LayerInfo. |
| src/NeuralNetworks/NeuralNetworkBase.cs | Implements ILayeredModel<T> metadata + partition point validation. |
| src/NeuralNetworks/Layers/LayerBase.cs | Adds default layer categorization and FLOPs/activation memory estimates. |
| src/MetaLearning/MetaLearnerBase.cs | Adds per-layer learning-rate gradient application using offsets. |
| src/LoRA/DefaultLoRAConfiguration.cs | Adds category-based LoRA eligibility + model-wide application helper. |
| src/Interfaces/LayerInfo.cs | Introduces LayerInfo<T> metadata container. |
| src/Interfaces/LayerCategory.cs | Introduces LayerCategory enum for automated decisions. |
| src/Interfaces/IPruningStrategy.cs | Extends pruning config with per-category sparsity targets. |
| src/Interfaces/INeuralNetwork.cs | Makes all neural networks implement ILayeredModel<T>. |
| src/Interfaces/ILayeredModel.cs | Adds new interface for layer access, metadata, and partition validation. |
| src/DistributedTraining/PipelineParallelModel.cs | Adds FLOPs-balanced, layer-boundary-aware parameter sharding. |
| src/Deployment/Export/ModelExporterBase.cs | Infers input/output shapes and exports layer summaries via ILayeredModel<T>. |
| src/Deployment/Export/ExportConfiguration.cs | Adds OutputShape to support export shape inference/override. |
💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.
There was a problem hiding this comment.
Actionable comments posted: 7
🤖 Fix all issues with AI agents
In `@src/Interfaces/ILayeredModel.cs`:
- Around line 31-70: ILayeredModel<T> is missing the ExtractSubModel(start,end)
method and should extend IParameterizable<T> so layer parameter offsets are
guaranteed to refer to the model's flat parameter vector; add
"IParameterizable<T>" to the interface inheritance, add a method signature
"ILayeredModel<T> ExtractSubModel(int startLayerIndex, int endLayerIndex)"
(inclusive/exclusive semantics per objectives) and ensure
LayerInfo<T>.ParameterOffset semantics are documented/consistent with the
IParameterizable<T>.ParameterVector; update GetLayerInfo and GetAllLayerInfo
behaviors to return offsets that map into the parameter vector exposed by
IParameterizable<T>.
In `@src/Interfaces/INeuralNetwork.cs`:
- Around line 31-33: You added ILayeredModel<T> to the public interface
INeuralNetwork<T>, which is a breaking change for external implementers; revert
the breaking change by either: 1) removing the ILayeredModel<T> inheritance from
INeuralNetwork<T> and introducing a new interface (e.g.,
ILayeredNeuralNetwork<T>) that extends INeuralNetwork<T> plus ILayeredModel<T>
for consumers who need layered features, or 2) if the change must land now,
document the breaking change clearly in the PR and bump the package major
version; locate INeuralNetwork<T> in the diff and apply one of these two fixes
so external implementations won’t be forced to add new members.
In `@src/LoRA/DefaultLoRAConfiguration.cs`:
- Around line 397-415: Add a null guard at the start of ApplyLoRAToModel to
validate the public parameter layeredModel and throw an ArgumentNullException if
it's null; specifically, check layeredModel before calling
layeredModel.GetAllLayerInfo(), and keep the rest of the method (including calls
to GetAllLayerInfo and ApplyLoRA) unchanged so behavior is preserved for
non-null inputs.
In `@src/NeuralNetworks/Layers/LayerBase.cs`:
- Around line 1039-1062: EstimateActivationMemory currently hardcodes 4 bytes
per element; change it to compute the element size from the layer's generic type
parameter T (instead of assuming float) by calling Marshal.SizeOf<T>() (or
Marshal.SizeOf(typeof(T)) for broader framework compatibility), then multiply
the computed elementSize by the total element count from GetOutputShape() to
return the correct byte count; ensure you import System.Runtime.InteropServices
and add a comment to verify target frameworks support the chosen Marshal API.
In `@src/Training/Memory/TrainingMemoryManager.cs`:
- Around line 165-173: ComputeCheckpointIndices currently assumes layeredModel
is non-null and will throw later; add an explicit null guard at the start of the
method (before _checkpointIndices.Clear()) that validates the ILayeredModel<T>
layeredModel parameter and throws an ArgumentNullException (or similar) with the
parameter name if null, so callers fail fast; update ComputeCheckpointIndices in
TrainingMemoryManager to perform this check and only proceed to call
layeredModel.GetAllLayerInfo() when layeredModel is non-null.
- Around line 198-201: The hardcoded 1_048_576 threshold in
TrainingMemoryManager (the if checking info.EstimatedActivationMemory) should be
replaced with a named constant or configurable setting: introduce a descriptive
constant (e.g., CheckpointActivationMemoryThreshold) or read from a
TrainingMemoryOptions/Config value and use that instead of the literal, update
the if to compare against that symbol, and adjust the comment to reference the
constant/config so the threshold is explicit and adjustable at runtime or via
config.
In `@tests/AiDotNet.Tests/Helpers/ContinualLearningTestHelper.cs`:
- Around line 108-158: ValidatePartitionPoint currently only checks index
bounds; update it to ensure the layer at afterLayerIndex and the next layer have
compatible shapes by comparing the left layer's OutputShape (from GetLayerInfo
or directly via _layers[afterLayerIndex].GetOutputShape()) with the right
layer's InputShape (from _layers[afterLayerIndex + 1].GetInputShape()),
returning false if shapes are null, lengths differ, or any corresponding
dimension differs; keep the existing bounds check and use LayerInfo/
GetLayerInfo only if helpful for clarity, and return true only when the shapes
match.
…ware pruning, and KD activation matching - Add ExtractSubModel to ILayeredModel<T> and SubModel<T> class for pipeline parallelism - Add CategoryBitWidths, GetBitWidthForLayer, ForMixedPrecision to QuantizationConfiguration - Add ApplyLayerAwarePruning to StructuredPruningStrategy with per-category sparsity - Add CollectIntermediateActivations and ComputeIntermediateActivationLoss to KD trainer - Add LayerCount, LayerCategorySummary, TotalTrainableParameters, TotalEstimatedFlops to AiModelResult - Update test mocks with ExtractSubModel implementations Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
…titioning - SubModel<T> now implements ILayeredModel<T> for composable sub-model extraction - Add GetLayerInfo, GetAllLayerInfo, ValidatePartitionPoint, ExtractSubModel to SubModel - Update EdgeOptimizer.CalculateAdaptivePartitionPoint to use ILayeredModel FLOPs - Add CalculateLoadBalancedPartitionPoint using per-layer FLOP estimates Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
There was a problem hiding this comment.
Pull request overview
Copilot reviewed 22 out of 22 changed files in this pull request and generated 7 comments.
💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.
There was a problem hiding this comment.
Actionable comments posted: 11
🤖 Fix all issues with AI agents
In `@src/Deployment/Edge/EdgeOptimizer.cs`:
- Around line 259-304: The CalculateLoadBalancedPartitionPoint method can return
an invalid mid-point if no boundary passes layeredModel.ValidatePartitionPoint;
modify it to track whether any valid partition was found (e.g., a boolean
foundValid), update foundValid when a partition is accepted inside the loop, and
after the loop check foundValid—if false, either throw a clear exception (e.g.,
InvalidOperationException) or apply a safe fallback heuristic such as scanning
for the first valid partition via ValidatePartitionPoint or returning 0/last
layer explicitly documented; ensure you update/return bestPartition only when a
valid partition was actually found.
In `@src/Deployment/Optimization/Quantization/QuantizationConfiguration.cs`:
- Around line 550-580: In ForMixedPrecision validate inputs at the top: ensure
sensitiveBitWidth and aggressiveBitWidth are > 0 and within a sane max (e.g., <=
32), ensure groupSize > 0, and enforce sensitiveBitWidth >= aggressiveBitWidth;
if any check fails throw ArgumentOutOfRangeException (with clear param name and
message). Update QuantizationConfiguration.ForMixedPrecision to perform these
guards before creating the new QuantizationConfiguration so invalid values
cannot silently produce an unusable config.
- Around line 505-518: GetBitWidthForCategory currently returns values from
CategoryBitWidths without validation, allowing zero/negative bit widths to
propagate; update GetBitWidthForCategory to check the looked-up override
(CategoryBitWidths.TryGetValue) and if bitWidth <= 0 either throw an
ArgumentOutOfRangeException with a clear message referencing the invalid
category and value or clamp to a minimum positive value (decide which strategy
project prefers), otherwise return the valid override; keep the fallback to
EffectiveBitWidth when no override exists and consider also validating
EffectiveBitWidth elsewhere if desired.
In `@src/Interfaces/SubModel.cs`:
- Around line 67-97: The public SubModel constructor must validate that the
incoming collections and index span are consistent: check layers and layerInfos
are non-null (already done) and additionally verify layerInfos.Count ==
layers.Count, that startIndex and endIndex are within a valid range (startIndex
>= 0, endIndex >= startIndex) and that (endIndex - startIndex + 1) ==
layers.Count; throw ArgumentException or ArgumentOutOfRangeException with clear
parameter names when any check fails so GetLayerInfo / ExtractSubModel will not
encounter mismatched metadata at runtime.
In `@src/KnowledgeDistillation/IntermediateActivations.cs`:
- Around line 90-94: LayerNames currently returns the live view
_activations.Keys which can throw if the dictionary mutates during enumeration;
change LayerNames to return a defensive snapshot (e.g. create and return a new
list/array from _activations.Keys) to match the defensive-copy semantics used by
AllActivations and avoid concurrent-enumeration issues when callers iterate
layer names.
In `@src/KnowledgeDistillation/KnowledgeDistillationTrainerBase.cs`:
- Around line 848-852: The code assumes a batched tensor shape and divides by
batch size unguarded; update the flattening logic around current, outputShape,
batchSize and features to handle null/empty/1D shapes and avoid division by
zero: if outputShape is null or outputShape.Length == 0 set batchSize = 1 and
features = current.Length; if outputShape.Length == 1 set batchSize = 1 and
features = outputShape[0] (or current.Length) ; otherwise set batchSize =
outputShape[0], and if batchSize <= 0 set batchSize = 1 before computing
features = current.Length / batchSize; ensure all branches use integer-safe math
and keep variable names current, outputShape, batchSize, features so callers
remain unchanged.
- Around line 907-935: The current loop in KnowledgeDistillationTrainerBase.cs
that computes per-element MSE between studentMatrix and teacherMatrix (using
rows = Min(...), cols = Min(...)) silently truncates mismatched activation
shapes; change this to validate exact shape equality for each matched layer pair
(compare studentMatrix.Rows == teacherMatrix.Rows and studentMatrix.Columns ==
teacherMatrix.Columns) and throw a clear exception (or return a failure) when
shapes differ so the distillation setup fails fast; also ensure you handle the
case where matchedLayers remains zero by throwing or returning an error instead
of dividing by zero before averaging totalLoss.
In `@src/Models/Results/AiModelResult.cs`:
- Around line 545-562: Change the public mutable Dictionary property
LayerCategorySummary to expose a read-only surface: replace the public
Dictionary<LayerCategory,int>? LayerCategorySummary with a public
IReadOnlyDictionary<LayerCategory,int>? LayerCategorySummary (or a
ReadOnlyDictionary wrapper) while keeping an internal/private mutable backing
field that the class can set; ensure the setter remains non-public (internal or
private) and when populating the backing Dictionary convert/wrap it to an
IReadOnlyDictionary for the public getter so external callers cannot mutate the
summary.
- Around line 1223-1251: The Deserialize path doesn't restore layer metadata
(LayerCount, LayerCategorySummary, TotalTrainableParameters,
TotalEstimatedFlops) after the Model is rehydrated; extract the layer-metadata
population logic currently in the constructor into a private helper (e.g.,
PopulateLayerMetadataFromModel or UpdateLayerMetadata) that accepts the Model
(and uses ILayeredModel<T>, GetAllLayerInfo, LayerCount, etc.), then call that
helper both from the constructor and at the end of Deserialize() after Model is
set (or copy values from the deserialized instance into these properties) so
loaded AiModelResult objects have the same layer metadata as newly constructed
ones.
In `@src/NeuralNetworks/NeuralNetworkBase.cs`:
- Around line 4054-4062: The explicit ILayeredModel<T>.Layers implementation
returns the internal mutable _layers collection, allowing callers to cast and
mutate it; wrap or copy _layers to an immutable/read-only wrapper before
returning (e.g., return a ReadOnlyCollection or AsReadOnly view or a defensive
copy) so callers cannot mutate the underlying list and bypass cache
invalidation; update the ILayeredModel<T>.Layers getter (currently referencing
_layers) to return the read-only wrapper and ensure any caching/invalidations
still use the original _layers field.
In `@tests/AiDotNet.Tests/Helpers/ContinualLearningTestHelper.cs`:
- Around line 183-196: The hardcoded EstimatedActivationMemory = 40 in the
submodel LayerInfo<T> entries makes tests invalid; replace the constant with a
computed value based on the layer's output shape (or reuse the existing
GetLayerInfo helper) when building subInfos: compute activation bytes from
layer.GetOutputShape() (consider element count * sizeof(float) or the same logic
used by GetLayerInfo) and assign that result to EstimatedActivationMemory
instead of 40 so LayerInfo<T>.EstimatedActivationMemory reflects the real
activation memory for the layer.
- Add input validation to ForMixedPrecision (bit widths, group size) - Return defensive snapshot from IntermediateActivations.LayerNames - Guard against non-batched/empty output shapes in KD activation collection - Expose LayerCategorySummary as IReadOnlyDictionary - Recompute layer metadata after Deserialize to avoid stale data - Replace ContainsKey+indexer with TryGetValue pattern - Fix Array.Copy bug in MetaLearnerBase - Cache GetAllLayerInfo to avoid O(n) per GetLayerInfo call - Fix hardcoded bytes-per-element in LayerBase - Fix ToArray() repeated allocation in PipelineParallelModel - Fix integer overflow in stage parameter end calculation - Add null guards for layeredModel parameters - Make HighActivationMemoryThresholdBytes configurable - Fix ValidatePartitionPoint with shape compatibility check - Fix KD shape mismatch handling and averaging threshold - Add SubModel constructor validation - Add load-balanced partition fallback - Prevent Layers list mutation via AsReadOnly - Compute activation memory from output shape Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
There was a problem hiding this comment.
Actionable comments posted: 10
Caution
Some comments are outside the diff and can’t be posted inline due to platform limitations.
⚠️ Outside diff range comments (1)
src/Deployment/Edge/EdgeOptimizer.cs (1)
312-370:⚠️ Potential issue | 🔴 CriticalBlocking: PartitionModel cannot succeed because layer extraction always throws.
ExtractEdgeLayersandExtractCloudLayersunconditionally throw, soPartitionModelis unusable even when a model implementsILayeredModel<T>. Implement extraction for layered models (e.g., viaExtractSubModelwith end-exclusive handling) and keepNotSupportedExceptiononly for unsupported models.🛠️ Suggested fix (extract sub-models for layered models)
private IFullModel<T, TInput, TOutput>? ExtractEdgeLayers(IFullModel<T, TInput, TOutput> model, int start, int end) { - throw new NotSupportedException( - "Model partitioning requires layer-wise structure information not provided by IFullModel. " + - "To enable model partitioning, use one of these approaches:\n" + - "1. Export your model to ONNX format using OnnxModelExporter, then use ONNX graph manipulation tools to split the computational graph at specific operator nodes.\n" + - "2. Implement a custom IPartitionable interface on your model type that exposes layer extraction methods.\n" + - "3. Use framework-specific partitioning (TensorFlow SavedModel, PyTorch TorchScript) before importing to AiDotNet.\n" + - "\n" + - "Example ONNX approach:\n" + - " var exporter = new OnnxModelExporter<T, TInput, TOutput>(exportConfig);\n" + - " var onnxBytes = exporter.Export(model);\n" + - " // Use ONNX graph manipulation library to split graph\n" + - " // Load edge and cloud portions separately via ONNX Runtime\n" + - "\n" + - "Arbitrary parameter splitting is not implemented as it would create invalid models with incomplete layers."); + if (model is ILayeredModel<T> layeredModel) + { + if (end <= start) return null; + var sub = layeredModel.ExtractSubModel(start, end - 1); // end is exclusive + return (IFullModel<T, TInput, TOutput>)sub; + } + + throw new NotSupportedException( + "Model partitioning requires layer-wise structure information not provided by IFullModel."); } @@ private IFullModel<T, TInput, TOutput>? ExtractCloudLayers(IFullModel<T, TInput, TOutput> model, int startFrom) { - throw new NotSupportedException( - "Model partitioning requires layer-wise structure information not provided by IFullModel. " + - "See EdgeOptimizer.ExtractEdgeLayers documentation for production-ready partitioning approaches."); + if (model is ILayeredModel<T> layeredModel) + { + var sub = layeredModel.ExtractSubModel(startFrom, layeredModel.LayerCount - 1); + return (IFullModel<T, TInput, TOutput>)sub; + } + + throw new NotSupportedException( + "Model partitioning requires layer-wise structure information not provided by IFullModel."); }As per coding guidelines: “Production Readiness (CRITICAL - Flag as BLOCKING): Stubs/Placeholders … methods with
throw new NotImplementedException(), placeholder return values, or incomplete features…”.
🤖 Fix all issues with AI agents
In `@src/Deployment/Edge/EdgeOptimizer.cs`:
- Around line 233-257: The current fallback returns hardcoded partition indices
(7/5/3) based on parameterCount which can exceed actual layer boundaries;
replace this by failing fast for non-layered adaptive partitioning or compute a
valid layer index from real layer metadata: inside the method that calls
CalculateLoadBalancedPartitionPoint, keep the ILayeredModel<T> branch to compute
a partition via CalculateLoadBalancedPartitionPoint(layeredModel) and remove the
parameterCount-based hardcoded returns; for non-layered models either throw an
InvalidOperationException (or a specific DomainException) indicating adaptive
partitioning requires ILayeredModel, or implement a safe fallback that maps
parameterCount to a clamped layer index using layeredModel.LayerCount (e.g.,
compute fraction -> Math.Clamp(index, 0, layeredModel.LayerCount-1)); reference
ILayeredModel<T>, layeredModel.LayerCount, CalculateLoadBalancedPartitionPoint,
and model.GetParameters() when making the change.
In `@src/Deployment/Optimization/Quantization/QuantizationConfiguration.cs`:
- Around line 537-542: The name-based override currently ignores non-positive
BitWidth values instead of rejecting them; update the logic in the method
handling layer bit-width resolution (the block using
CustomLayerParams.TryGetValue and layerParams.BitWidth) to validate and throw an
ArgumentException (or similar) when BitWidth.HasValue is true but BitWidth.Value
<= 0, matching the validation behavior in ForMixedPrecision; ensure the
exception message names the offending layer (layerInfo.Name) and the invalid
value so users know their CustomLayerParams entry was invalid rather than
silently ignored.
- Around line 583-589: The Mapping for layer precisions in
QuantizationConfiguration.cs currently forces
AiDotNet.Interfaces.LayerCategory.Normalization to 16 regardless of
sensitiveBitWidth; update the XML documentation for the
QuantizationConfiguration (or the property/method that builds this mapping) to
explicitly state that Normalization layers are overridden to 16-bit due to
precision sensitivity and will not follow the sensitiveBitWidth parameter,
referencing LayerCategory.Normalization and sensitiveBitWidth so readers
understand the intentional exception.
In `@src/DistributedTraining/PipelineParallelModel.cs`:
- Around line 186-219: When choosing stage boundaries in the partition loop
(variables currentLayer, stageFlops, layersInStage, stageStartLayer,
allLayerInfo, _numStages, layerCount), honor the model's
ILayeredModel.ValidatePartitionPoint: before accepting a candidate boundary
(i.e., before breaking when stageFlops >= targetFlopsPerStage), call
ValidatePartitionPoint(candidateIndex) and only break if it returns true; if it
returns false, continue consuming layers to find the next valid boundary, and if
you reach the end without any valid boundary for a stage, fail fast (throw or
return an error) rather than silently split across an invalid partition point.
In `@src/Interfaces/SubModel.cs`:
- Around line 103-110: The aggregation of parameter counts uses int accumulators
(totalParams and localOffset) and can overflow for large models; change the
accumulation to use a long accumulator (e.g., long totalParamsLong) or use
checked arithmetic when summing layerInfos[i].ParameterCount, then validate
bounds before assigning to ParameterCount (cast or clamp with explicit overflow
handling). Update all similar loops (including the block referencing
localOffset) to use the new long accumulator or checked operations and ensure
any final assignment to ParameterCount performs a safe conversion with a clear
overflow error or clamp.
In `@src/KnowledgeDistillation/KnowledgeDistillationTrainerBase.cs`:
- Around line 910-914: When the code in KnowledgeDistillationTrainerBase that
compares studentMatrix and teacherMatrix dimensions (the block with the
Rows/Columns checks) skips a layer due to mismatched shapes, add a diagnostic
log entry instead of silently continuing: log the layer identifier or index and
both shapes (studentMatrix.Rows x studentMatrix.Columns and teacherMatrix.Rows x
teacherMatrix.Columns) and a clear message indicating a configuration mismatch.
Use the class's existing logging facility (e.g., _logger or Logger) and include
enough context (layer name/index and tensor names) so debugging which layer was
skipped is straightforward.
- Around line 849-865: The code computes features as current.Length / batchSize
which can silently truncate; add a validation in the block that checks
outputShape, e.g. verify current.Length % batchSize == 0 before computing
features, and if not throw an informative exception (or log and continue)
mentioning current.Length and batchSize; then compute int features =
current.Length / batchSize and proceed to build Matrix<T> as before (references:
current, outputShape, batchSize, features, Matrix<T>, ToVector()) so any
mis-shaped tensor fails fast instead of producing truncated activations.
In `@src/MetaLearning/MetaLearnerBase.cs`:
- Around line 781-798: The loop that applies per-layer learning rates currently
iterates allLayerInfo.Where(l => l.ParameterCount > 0) and updates parameters
even for frozen layers; change the filter to skip non-trainable layers by
checking IsTrainable and validate ParameterOffset/ParameterCount against
parameters.Length and gradients.Length before updating. Specifically, update the
iteration over allLayerInfo (used with perLayerLearningRates,
defaultLearningRate) to only include info where info.IsTrainable is true and
ensure idx = info.ParameterOffset + i is within bounds for parameters and
gradients before assigning updated[idx]; keep using NumOps.FromDouble,
NumOps.Multiply and NumOps.Subtract as shown.
In `@src/NeuralNetworks/NeuralNetworkBase.cs`:
- Around line 4085-4133: Add a cache invalidation helper that clears
_cachedLayerInfo and resets _cachedLayerCount (e.g., private void
InvalidateLayerInfoCache() { _cachedLayerInfo = null; _cachedLayerCount = -1; })
in the NeuralNetworkBase class, and call InvalidateLayerInfoCache() from all
mutation/deserialization paths: InsertLayerIntoCollection, AddLayerToCollection,
RemoveLayerFromCollection, ClearLayers, and at the end of Deserialize so
GetAllLayerInfo() cannot return stale metadata when layers are replaced or
reloaded with the same count.
In `@src/Training/Memory/TrainingMemoryManager.cs`:
- Around line 184-206: The code uses Config.CheckpointEveryNLayers in a modulo
(i % Config.CheckpointEveryNLayers) which can throw for zero/negative values and
Config.HighActivationMemoryThresholdBytes can be negative causing unintended
behavior; in the initializer or at the start of the method that contains the
checkpoint logic (look for TrainingMemoryManager and the block using
shouldCheckpoint, i, info, and Config.*) validate Config.CheckpointEveryNLayers
is > 0 and throw a clear ArgumentException/InvalidOperation if not, and also
validate Config.HighActivationMemoryThresholdBytes is >= 0 (throw if negative)
so the subsequent checks (interval-based modulo, attention/residual flags, and
info.EstimatedActivationMemory > Config.HighActivationMemoryThresholdBytes) are
safe and deterministic.
…and cache invalidation - EdgeOptimizer: fail fast for non-ILayeredModel adaptive partitioning instead of returning hardcoded indices that could exceed layer boundaries - QuantizationConfiguration: throw on invalid BitWidth in CustomLayerParams instead of silently ignoring; document Normalization 16-bit override - PipelineParallelModel: honor ValidatePartitionPoint before accepting stage boundaries to prevent splitting across incompatible layers - SubModel: use long accumulators for parameter counts with overflow checks to prevent silent truncation on large models - KDTrainerBase: log diagnostic for shape mismatches; validate tensor length is divisible by batch size before flattening activations - MetaLearnerBase: filter by IsTrainable and validate ParameterOffset bounds before applying per-layer learning rates to frozen layers - NeuralNetworkBase: add InvalidateLayerInfoCache helper called from all mutation paths and Deserialize to prevent stale cached metadata - TrainingMemoryManager: validate CheckpointEveryNLayers > 0 and HighActivationMemoryThresholdBytes >= 0 to prevent divide-by-zero and unintended checkpoint behavior Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
There was a problem hiding this comment.
Pull request overview
Copilot reviewed 23 out of 23 changed files in this pull request and generated no new comments.
💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.
There was a problem hiding this comment.
Actionable comments posted: 6
🤖 Fix all issues with AI agents
In `@src/Deployment/Optimization/Quantization/QuantizationConfiguration.cs`:
- Around line 511-522: GetBitWidthForCategory currently silently falls back to
EffectiveBitWidth when a configured CategoryBitWidths entry is <= 0; make this
consistent with GetBitWidthForLayer by failing fast: if CategoryBitWidths
contains the category but the retrieved bitWidth is <= 0, throw an
ArgumentException that includes the category name/enum value and the invalid
bitWidth; keep the existing behavior of returning EffectiveBitWidth only when
the category key is absent (CategoryBitWidths is null or TryGetValue returns
false). This aligns validation with GetBitWidthForLayer and CustomLayerParams.
In `@src/DistributedTraining/PipelineParallelModel.cs`:
- Around line 186-234: The partition loop can silently produce invalid stage
boundaries if no valid partition point is found; update the code in the loop
that computes stage boundaries (symbols: currentLayer, candidateBoundary,
layeredModel.ValidatePartitionPoint, stageStartLayer, _numStages) to attempt a
backward search for a valid partition when candidateBoundary == -1 (iterate from
currentLayer-1 down to the stage's start layer and call
layeredModel.ValidatePartitionPoint for each potential partitionAfter), and if
none are valid throw a clear InvalidOperationException (or log and fail fast)
including stage index and layer range; also remove the redundant assignment
currentLayer = candidateBoundary or only assign currentLayer = candidateBoundary
when candidateBoundary != currentLayer.
- Around line 253-262: The computation of stageParamEndLong and subsequent
clamping to fullParameters.Length can silently truncate the shard when the
computed end > int.MaxValue or > fullParameters.Length; update the logic around
stageParamEndLong, stageParamEnd and the assignment to ShardStartIndex/ShardSize
to validate the computed range against fullParameters.Length and int.MaxValue
and emit a clear diagnostic (throw or processLogger/error log) if the requested
end exceeds fullParameters.Length or would require truncation, or otherwise fail
fast; reference the symbols stageParamStart, stageParamEndLong, stageParamEnd,
fullParameters.Length, ShardStartIndex and ShardSize when adding the validation
and diagnostic so the code does not silently truncate shards.
In `@src/MetaLearning/MetaLearnerBase.cs`:
- Around line 783-789: The loop that checks each LayerInfo currently silently
continues when info.ParameterOffset or info.ParameterCount would make the slice
go out of bounds (the condition using info.ParameterOffset, info.ParameterCount,
parameters.Length, gradients.Length), which can hide configuration errors;
instead, replace the silent continue with a clear warning (or error) log that
includes the layer identifier (e.g., the LayerInfo instance or its name/index)
and the offending offsets/counts, then skip the layer if you still want to
proceed; update the code around the boundary check in MetaLearnerBase (the block
that reads info.ParameterOffset / info.ParameterCount) to call your existing
logger (or throw) with those values so mis-sized LayerInfo metadata is visible
during debugging.
In `@src/NeuralNetworks/NeuralNetworkBase.cs`:
- Around line 2849-2852: The Deserialize() method currently calls both
InvalidateParameterCountCache() and InvalidateLayerInfoCache(), but
InvalidateParameterCountCache() already calls InvalidateLayerInfoCache(), so
remove the redundant InvalidateLayerInfoCache() invocation from Deserialize()
and leave only a single call to InvalidateParameterCountCache() to avoid
duplicate work.
- Around line 4098-4146: GetAllLayerInfo builds a mutable List<LayerInfo<T>>
into _cachedLayerInfo and returns it, allowing callers to cast and mutate the
cached list; change the caching/return to an immutable/read-only snapshot so
callers cannot modify the cached instance: after populating the List (result)
wrap it with a read-only wrapper (e.g. call AsReadOnly() or otherwise produce an
IReadOnlyList<LayerInfo<T>> snapshot) and assign that to _cachedLayerInfo, then
return _cachedLayerInfo; keep references to GetAllLayerInfo, _cachedLayerInfo,
_cachedLayerCount, LayerInfo<T>, and _layers unchanged so only the
assignment/return is updated.
…ilayeredmodel - QuantizationConfiguration: throw ArgumentException for invalid CategoryBitWidths instead of silently falling back to EffectiveBitWidth - PipelineParallelModel: add backward search for valid partition points and throw on out-of-range shard parameters instead of silent clamping - MetaLearnerBase: add Debug.WriteLine for skipped out-of-bounds LayerInfo - NeuralNetworkBase: remove redundant InvalidateLayerInfoCache call in Deserialize and use AsReadOnly() for cached layer info Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
- PipelineParallelModel: remove redundant candidateBoundary assignment, add debug log when forcing partition boundary without valid point - QuantizationConfiguration: refactor nested if to early return for clarity - AiModelResult: use ternary for categorySummary TryGetValue pattern (2 places) - LayerBase: catch ArgumentException instead of generic catch in SizeOf Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
There was a problem hiding this comment.
Pull request overview
Copilot reviewed 23 out of 23 changed files in this pull request and generated 7 comments.
💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.
There was a problem hiding this comment.
Actionable comments posted: 4
🤖 Fix all issues with AI agents
In `@src/DistributedTraining/PipelineParallelModel.cs`:
- Around line 171-263: The forward search may consume layers and then the
backward search can fail, leaving currentLayer advanced and potentially creating
an invalid partition; before committing stageStartLayer[stage + 1] =
currentLayer ensure the boundary is validated: require
layeredModel.ValidatePartitionPoint(currentLayer - 1) (or, if you prefer, throw
an InvalidOperationException) when candidateBoundary < 0 and foundBackward ==
false and layersInStage > 0; alternatively, instead of silently proceeding,
explicitly throw with a clear message referencing stage, stageStartLayer[stage],
and currentLayer so the caller knows the partitioning failed.
In `@src/MetaLearning/MetaLearnerBase.cs`:
- Around line 773-780: The code copies model.GetParameters() into updated but
lacks a guard that gradients.Length matches parameters.Length; add an upfront
validation at the start of the method (before creating updated or iterating
per-layer info) that compares gradients.Length to parameters.Length and fails
fast (throw an ArgumentException/InvalidOperationException with a clear message)
if they differ, so functions like model.GetParameters(),
layeredModel.GetAllLayerInfo(), and the updated Vector<T> creation never produce
a partially-applied result.
In `@src/NeuralNetworks/Layers/LayerBase.cs`:
- Around line 1068-1088: GetBytesPerElement silently defaults to 4 bytes when
Marshal.SizeOf<T>() throws, which can mask incorrect memory estimates for custom
numeric types; update GetBytesPerElement to emit a diagnostic (e.g., Debug/Trace
or the project’s logging facility) when catching ArgumentException and include
typeof(T) and the exception message, or alternatively rethrow a more descriptive
exception to fail fast instead of silently returning 4; ensure the change
references GetBytesPerElement, the Marshal.SizeOf<T>() call, the
ArgumentException catch, and the current fallback return 4 so reviewers can find
and validate the fix.
In `@src/NeuralNetworks/NeuralNetworkBase.cs`:
- Around line 4203-4258: The method ExtractSubModel computes a local baseOffset
(variable baseOffset) by summing ParameterCount for layers before startLayer but
never uses it, leaving dead code; remove the unused baseOffset declaration and
the for-loop that accumulates it (the lines creating and updating baseOffset)
from ExtractSubModel so the method only computes localOffset and builds
subInfos/subLayers without the unused global-offset logic.
- Add gradient/parameter length mismatch guard in MetaLearnerBase - Remove unused baseOffset variable in NeuralNetworkBase.ExtractSubModel - Remove unnecessary .ToList() allocation in IntermediateActivations.LayerNames - Graceful fallback in EdgeOptimizer for non-ILayeredModel (debug log + default partition) - Clear stale layer-metadata in AiModelResult.Deserialize for non-layered models - Add debug trace for implicit else branch in PipelineParallelModel partition loop - Add debug trace for Marshal.SizeOf fallback in LayerBase.GetBytesPerElement Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
…) (#847) * feat: add ILayeredModel<T> interface for layer-aware model access Implements Phase 1 of issue #462: the ILayeredModel<T> interface that provides per-layer metadata access across the AiDotNet stack, enabling pipeline parallelism, quantization, pruning, LoRA, and checkpointing tools to make layer-aware decisions instead of operating on flat vectors. New types: - ILayeredModel<T>: interface with Layers, LayerCount, GetLayerInfo, GetAllLayerInfo, ValidatePartitionPoint - LayerInfo<T>: metadata class with parameter offsets, shapes, category, cost estimates (FLOPs, activation memory) - LayerCategory: 15-value enum for automated per-layer decisions Changes: - INeuralNetwork<T> now extends ILayeredModel<T> - LayerBase<T> gains GetLayerCategory(), EstimateFlops(), EstimateActivationMemory() virtual methods - NeuralNetworkBase<T> implements all ILayeredModel<T> members Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * feat: integrate ILayeredModel across pipeline, checkpointing, LoRA, and meta-learning Phase 2 subsystem integration for issue #462: PipelineParallelModel: - Layer-aware partitioning using ILayeredModel metadata - FLOPs-balanced stage assignment respecting layer boundaries - Falls back to naive parameter splitting for non-neural models TrainingMemoryManager: - New ComputeCheckpointIndices(ILayeredModel) overload - Uses LayerCategory and EstimatedActivationMemory for selective checkpointing decisions - Layers with >1MB activation memory auto-checkpointed DefaultLoRAConfiguration: - ShouldApplyLoRA(LayerCategory) for category-based adapter decisions - ApplyLoRAToModel(ILayeredModel) for automatic adapter injection - Dense, Conv, Attention, Recurrent, Embedding, FeedForward, Graph, Residual categories eligible for LoRA MetaLearnerBase: - ApplyPerLayerGradients using LayerInfo parameter offsets - Enables per-layer learning rates for fine-grained meta-learning Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * feat: add ILayeredModel support to model export, pruning config Phase 3 of issue #462: ModelExporterBase: - GetInputShape now auto-infers from first layer via ILayeredModel - New GetOutputShape auto-infers from last layer via ILayeredModel - New GetLayerSummary provides layer metadata for export metadata ExportConfiguration: - Add OutputShape property for explicit output shape specification PruningConfig: - Add CategorySparsityTargets for per-LayerCategory sparsity levels - Enables attention layers at lower sparsity than dense layers Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix: add ILayeredModel implementation to test mock neural networks Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * feat: add sub-model extraction, mixed-precision quantization, layer-aware pruning, and KD activation matching - Add ExtractSubModel to ILayeredModel<T> and SubModel<T> class for pipeline parallelism - Add CategoryBitWidths, GetBitWidthForLayer, ForMixedPrecision to QuantizationConfiguration - Add ApplyLayerAwarePruning to StructuredPruningStrategy with per-category sparsity - Add CollectIntermediateActivations and ComputeIntermediateActivationLoss to KD trainer - Add LayerCount, LayerCategorySummary, TotalTrainableParameters, TotalEstimatedFlops to AiModelResult - Update test mocks with ExtractSubModel implementations Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * feat: make SubModel implement ILayeredModel and add load-balanced partitioning - SubModel<T> now implements ILayeredModel<T> for composable sub-model extraction - Add GetLayerInfo, GetAllLayerInfo, ValidatePartitionPoint, ExtractSubModel to SubModel - Update EdgeOptimizer.CalculateAdaptivePartitionPoint to use ILayeredModel FLOPs - Add CalculateLoadBalancedPartitionPoint using per-layer FLOP estimates Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix: address pr review comments for accessibility and code quality - Add input validation to ForMixedPrecision (bit widths, group size) - Return defensive snapshot from IntermediateActivations.LayerNames - Guard against non-batched/empty output shapes in KD activation collection - Expose LayerCategorySummary as IReadOnlyDictionary - Recompute layer metadata after Deserialize to avoid stale data - Replace ContainsKey+indexer with TryGetValue pattern - Fix Array.Copy bug in MetaLearnerBase - Cache GetAllLayerInfo to avoid O(n) per GetLayerInfo call - Fix hardcoded bytes-per-element in LayerBase - Fix ToArray() repeated allocation in PipelineParallelModel - Fix integer overflow in stage parameter end calculation - Add null guards for layeredModel parameters - Make HighActivationMemoryThresholdBytes configurable - Fix ValidatePartitionPoint with shape compatibility check - Fix KD shape mismatch handling and averaging threshold - Add SubModel constructor validation - Add load-balanced partition fallback - Prevent Layers list mutation via AsReadOnly - Compute activation memory from output shape Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix: harden layer-aware subsystems with validation, overflow safety, and cache invalidation - EdgeOptimizer: fail fast for non-ILayeredModel adaptive partitioning instead of returning hardcoded indices that could exceed layer boundaries - QuantizationConfiguration: throw on invalid BitWidth in CustomLayerParams instead of silently ignoring; document Normalization 16-bit override - PipelineParallelModel: honor ValidatePartitionPoint before accepting stage boundaries to prevent splitting across incompatible layers - SubModel: use long accumulators for parameter counts with overflow checks to prevent silent truncation on large models - KDTrainerBase: log diagnostic for shape mismatches; validate tensor length is divisible by batch size before flattening activations - MetaLearnerBase: filter by IsTrainable and validate ParameterOffset bounds before applying per-layer learning rates to frozen layers - NeuralNetworkBase: add InvalidateLayerInfoCache helper called from all mutation paths and Deserialize to prevent stale cached metadata - TrainingMemoryManager: validate CheckpointEveryNLayers > 0 and HighActivationMemoryThresholdBytes >= 0 to prevent divide-by-zero and unintended checkpoint behavior Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix: fail-fast validation, debug logging, and cache immutability for ilayeredmodel - QuantizationConfiguration: throw ArgumentException for invalid CategoryBitWidths instead of silently falling back to EffectiveBitWidth - PipelineParallelModel: add backward search for valid partition points and throw on out-of-range shard parameters instead of silent clamping - MetaLearnerBase: add Debug.WriteLine for skipped out-of-bounds LayerInfo - NeuralNetworkBase: remove redundant InvalidateLayerInfoCache call in Deserialize and use AsReadOnly() for cached layer info Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix: address pr review comments for accessibility and code quality - PipelineParallelModel: remove redundant candidateBoundary assignment, add debug log when forcing partition boundary without valid point - QuantizationConfiguration: refactor nested if to early return for clarity - AiModelResult: use ternary for categorySummary TryGetValue pattern (2 places) - LayerBase: catch ArgumentException instead of generic catch in SizeOf Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix: address remaining pr review comments for ilayeredmodel interface - Add gradient/parameter length mismatch guard in MetaLearnerBase - Remove unused baseOffset variable in NeuralNetworkBase.ExtractSubModel - Remove unnecessary .ToList() allocation in IntermediateActivations.LayerNames - Graceful fallback in EdgeOptimizer for non-ILayeredModel (debug log + default partition) - Clear stale layer-metadata in AiModelResult.Deserialize for non-layered models - Add debug trace for implicit else branch in PipelineParallelModel partition loop - Add debug trace for Marshal.SizeOf fallback in LayerBase.GetBytesPerElement Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
…uling, activation checkpointing (#845) * feat: add pipeline parallelism with load balancing and scheduling (#463) - Add IPipelinePartitionStrategy and UniformPartitionStrategy (default) - Add LoadBalancedPartitionStrategy using DP min-max partitioning - Add IPipelineSchedule with GPipeSchedule and 1F1B schedule - Add ActivationCheckpointConfig with configurable recompute strategies - Integrate optimizations into PipelineParallelModel - 1F1B schedule reduces pipeline bubble from ~50% to ~12-15% - Activation checkpointing reduces memory from O(L) to O(sqrt(L)) Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix: integrate pipeline parallelism options into aimodelbuilder facade ConfigureDistributedTraining() now accepts optional pipeline-specific parameters (schedule, partition strategy, checkpoint config, micro-batch size) that are passed through to PipelineParallelModel when the user selects DistributedStrategy.PipelineParallel. All parameters are optional with backward-compatible defaults. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * feat: add zero bubble and interleaved pipeline schedules with backward decomposition Add 5 new pipeline schedule implementations based on 2024-2025 research: - ZB-H1: splits backward into B+W, ~1/3 bubble of 1F1B (same memory) - ZB-H2: aggressive scheduling for zero bubble (higher memory) - ZB-V: 2 virtual stages per rank, zero bubble with 1F1B memory - Interleaved 1F1B: V virtual stages per rank, depth-first ordering - Looped BFS: V virtual stages per rank, breadth-first ordering Expand IPipelineSchedule with VirtualStagesPerRank and BackwardInput/ BackwardWeight operation types. Update PipelineParallelModel to handle split backward passes with cached input gradients. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix: implement backward decomposition, virtual stages, and micro-batching - Add IPipelineDecomposableModel<T> for true B/W split - Emulated B/W split fallback for non-decomposable models - Virtual stage partitioning with non-contiguous chunks - Proper micro-batch slicing via vector conversion - Activation checkpoint recomputation from nearest checkpoint - Virtual-stage-aware communication routing with unique tags Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix: add source generator exclusions, validation, and tag safety for pipeline parallelism - Add missing AiDotNet.Generators compile exclusion and ProjectReference to csproj (fixes CS0579 duplicate assembly attributes build error) - Add property setter validation in ActivationCheckpointConfig (CheckpointEveryNLayers >= 1, MaxActivationsInMemory >= 0) - Reorder GPipeSchedule validation to check numStages/numMicroBatches before stageId - Add _isAutoDetect flag and boundary validation to LoadBalancedPartitionStrategy - Add tag constants (ActivationTagBase, GradientTagBase, PredictTagBase) to prevent communication collisions in PipelineParallelModel - Add partition validation, checkpointing fail-fast, and internal property visibility Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix: clean up all schedule implementations and pipeline model code quality issues - Reorder validation in all 7 schedule classes: check numStages/numMicroBatches before stageId - Fix integer overflow in EstimateBubbleFraction across all schedules (use long arithmetic) - Remove unused variables: forwardCount/backwardCount (Interleaved1F1B), isFirstLoop/isLastLoop (LoopedBFS), totalVirtualStages/totalWarmupForwards (ZeroBubbleV) - Remove redundant operations: / 1 (Interleaved1F1B), - 0 (ZeroBubbleH1) - Replace generic catch clauses with specific InvalidOperationException in PipelineParallelModel - Combine nested if statements in GetStageInput - Remove unused globalVirtualStageId variable - Use ternary operator for cost estimation in LoadBalancedPartitionStrategy Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix: split configure methods, fix virtual-stage routing, and fail-fast micro-batch slicing - Split ConfigureDistributedTraining (7 params) into ConfigureDistributedTraining (3 params) + ConfigurePipelineParallelism (4 params) to avoid breaking the interface - Fix virtual-stage routing: non-first virtual stages now use forward output from the previous virtual stage instead of falling back to raw micro-batch input - Fail fast on micro-batch slicing failures instead of silently duplicating data to all micro-batches (which produces incorrect gradient averages) - Apply partition strategy for multi-stage (V>1) schedules instead of ignoring it - Limit checkpoint recompute search to current micro-batch boundaries Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix: add schedule bounds validation, checkpoint guards, and cost doc accuracy - Validate schedule output bounds (MicroBatchIndex, VirtualStageIndex) before executing ops to guard against externally injectable schedules - Integrate CheckpointFirstLayer config into ShouldCheckpointActivation - Add defensive guard for CheckpointEveryNLayers modulo-by-zero - Fix cost estimator doc to correctly explain paramCount^1.5 derivation Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix: address remaining pr review comments for pipeline parallelism - AiModelBuilder: validate microBatchCount > 0, rename from microBatchSize - IAiModelBuilder: rename microBatchSize to microBatchCount for clarity - PipelineParallelModel: add multi-stage partition bounds validation, pre-init guard on EstimatedBubbleFraction, gradient length check, comment on unreachable recompute code, debug log for forced partition - ZeroBubbleH2Schedule: make warmup count stage-dependent using stageId - ZeroBubbleVSchedule: fix BackwardWeight IsCooldown to use computed flag instead of hardcoded true - ActivationCheckpointConfig: use property names in exceptions - AiDotNet.csproj: gate EmitCompilerGeneratedFiles to Debug config only Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix: resolve ci build failure from duplicate generator references and tfm mismatch - Remove duplicate source generator sections from merge with master - Add SetTargetFramework=netstandard2.0 to generator ProjectReference to prevent MSBuild from building it for net10.0/net471 Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix: use long arithmetic in bubble fraction and widen tag ranges - Use long variables in EstimateBubbleFraction across all 6 schedule classes to prevent integer overflow in numerator arithmetic - Increase communication tag ranges from 100K to 1M between bases to prevent collisions with many micro-batches and virtual stages Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix: resolve sonarcloud build failure from vbcscompiler file lock Disable shared compilation (-p:UseSharedCompilation=false) in the SonarCloud analysis build step to prevent VBCSCompiler from holding file locks on AiDotNet.Generators.dll during parallel project builds. Also use long arithmetic in bubble fraction calculations and widen communication tag ranges from 100K to 1M to prevent collisions. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix: disable shared compilation in both build jobs to prevent file locks Apply -p:UseSharedCompilation=false to both Build (Windows) and SonarCloud Analysis build steps. VBCSCompiler holds file locks on AiDotNet.Generators.dll when building the solution, causing CS2012 errors when multiple projects compile the generator concurrently. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * feat: add ILayeredModel<T> interface for layer-aware model access (#462) (#847) * feat: add ILayeredModel<T> interface for layer-aware model access Implements Phase 1 of issue #462: the ILayeredModel<T> interface that provides per-layer metadata access across the AiDotNet stack, enabling pipeline parallelism, quantization, pruning, LoRA, and checkpointing tools to make layer-aware decisions instead of operating on flat vectors. New types: - ILayeredModel<T>: interface with Layers, LayerCount, GetLayerInfo, GetAllLayerInfo, ValidatePartitionPoint - LayerInfo<T>: metadata class with parameter offsets, shapes, category, cost estimates (FLOPs, activation memory) - LayerCategory: 15-value enum for automated per-layer decisions Changes: - INeuralNetwork<T> now extends ILayeredModel<T> - LayerBase<T> gains GetLayerCategory(), EstimateFlops(), EstimateActivationMemory() virtual methods - NeuralNetworkBase<T> implements all ILayeredModel<T> members Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * feat: integrate ILayeredModel across pipeline, checkpointing, LoRA, and meta-learning Phase 2 subsystem integration for issue #462: PipelineParallelModel: - Layer-aware partitioning using ILayeredModel metadata - FLOPs-balanced stage assignment respecting layer boundaries - Falls back to naive parameter splitting for non-neural models TrainingMemoryManager: - New ComputeCheckpointIndices(ILayeredModel) overload - Uses LayerCategory and EstimatedActivationMemory for selective checkpointing decisions - Layers with >1MB activation memory auto-checkpointed DefaultLoRAConfiguration: - ShouldApplyLoRA(LayerCategory) for category-based adapter decisions - ApplyLoRAToModel(ILayeredModel) for automatic adapter injection - Dense, Conv, Attention, Recurrent, Embedding, FeedForward, Graph, Residual categories eligible for LoRA MetaLearnerBase: - ApplyPerLayerGradients using LayerInfo parameter offsets - Enables per-layer learning rates for fine-grained meta-learning Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * feat: add ILayeredModel support to model export, pruning config Phase 3 of issue #462: ModelExporterBase: - GetInputShape now auto-infers from first layer via ILayeredModel - New GetOutputShape auto-infers from last layer via ILayeredModel - New GetLayerSummary provides layer metadata for export metadata ExportConfiguration: - Add OutputShape property for explicit output shape specification PruningConfig: - Add CategorySparsityTargets for per-LayerCategory sparsity levels - Enables attention layers at lower sparsity than dense layers Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix: add ILayeredModel implementation to test mock neural networks Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * feat: add sub-model extraction, mixed-precision quantization, layer-aware pruning, and KD activation matching - Add ExtractSubModel to ILayeredModel<T> and SubModel<T> class for pipeline parallelism - Add CategoryBitWidths, GetBitWidthForLayer, ForMixedPrecision to QuantizationConfiguration - Add ApplyLayerAwarePruning to StructuredPruningStrategy with per-category sparsity - Add CollectIntermediateActivations and ComputeIntermediateActivationLoss to KD trainer - Add LayerCount, LayerCategorySummary, TotalTrainableParameters, TotalEstimatedFlops to AiModelResult - Update test mocks with ExtractSubModel implementations Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * feat: make SubModel implement ILayeredModel and add load-balanced partitioning - SubModel<T> now implements ILayeredModel<T> for composable sub-model extraction - Add GetLayerInfo, GetAllLayerInfo, ValidatePartitionPoint, ExtractSubModel to SubModel - Update EdgeOptimizer.CalculateAdaptivePartitionPoint to use ILayeredModel FLOPs - Add CalculateLoadBalancedPartitionPoint using per-layer FLOP estimates Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix: address pr review comments for accessibility and code quality - Add input validation to ForMixedPrecision (bit widths, group size) - Return defensive snapshot from IntermediateActivations.LayerNames - Guard against non-batched/empty output shapes in KD activation collection - Expose LayerCategorySummary as IReadOnlyDictionary - Recompute layer metadata after Deserialize to avoid stale data - Replace ContainsKey+indexer with TryGetValue pattern - Fix Array.Copy bug in MetaLearnerBase - Cache GetAllLayerInfo to avoid O(n) per GetLayerInfo call - Fix hardcoded bytes-per-element in LayerBase - Fix ToArray() repeated allocation in PipelineParallelModel - Fix integer overflow in stage parameter end calculation - Add null guards for layeredModel parameters - Make HighActivationMemoryThresholdBytes configurable - Fix ValidatePartitionPoint with shape compatibility check - Fix KD shape mismatch handling and averaging threshold - Add SubModel constructor validation - Add load-balanced partition fallback - Prevent Layers list mutation via AsReadOnly - Compute activation memory from output shape Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix: harden layer-aware subsystems with validation, overflow safety, and cache invalidation - EdgeOptimizer: fail fast for non-ILayeredModel adaptive partitioning instead of returning hardcoded indices that could exceed layer boundaries - QuantizationConfiguration: throw on invalid BitWidth in CustomLayerParams instead of silently ignoring; document Normalization 16-bit override - PipelineParallelModel: honor ValidatePartitionPoint before accepting stage boundaries to prevent splitting across incompatible layers - SubModel: use long accumulators for parameter counts with overflow checks to prevent silent truncation on large models - KDTrainerBase: log diagnostic for shape mismatches; validate tensor length is divisible by batch size before flattening activations - MetaLearnerBase: filter by IsTrainable and validate ParameterOffset bounds before applying per-layer learning rates to frozen layers - NeuralNetworkBase: add InvalidateLayerInfoCache helper called from all mutation paths and Deserialize to prevent stale cached metadata - TrainingMemoryManager: validate CheckpointEveryNLayers > 0 and HighActivationMemoryThresholdBytes >= 0 to prevent divide-by-zero and unintended checkpoint behavior Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix: fail-fast validation, debug logging, and cache immutability for ilayeredmodel - QuantizationConfiguration: throw ArgumentException for invalid CategoryBitWidths instead of silently falling back to EffectiveBitWidth - PipelineParallelModel: add backward search for valid partition points and throw on out-of-range shard parameters instead of silent clamping - MetaLearnerBase: add Debug.WriteLine for skipped out-of-bounds LayerInfo - NeuralNetworkBase: remove redundant InvalidateLayerInfoCache call in Deserialize and use AsReadOnly() for cached layer info Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix: address pr review comments for accessibility and code quality - PipelineParallelModel: remove redundant candidateBoundary assignment, add debug log when forcing partition boundary without valid point - QuantizationConfiguration: refactor nested if to early return for clarity - AiModelResult: use ternary for categorySummary TryGetValue pattern (2 places) - LayerBase: catch ArgumentException instead of generic catch in SizeOf Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix: address remaining pr review comments for ilayeredmodel interface - Add gradient/parameter length mismatch guard in MetaLearnerBase - Remove unused baseOffset variable in NeuralNetworkBase.ExtractSubModel - Remove unnecessary .ToList() allocation in IntermediateActivations.LayerNames - Graceful fallback in EdgeOptimizer for non-ILayeredModel (debug log + default partition) - Clear stale layer-metadata in AiModelResult.Deserialize for non-layered models - Add debug trace for implicit else branch in PipelineParallelModel partition loop - Add debug trace for Marshal.SizeOf fallback in LayerBase.GetBytesPerElement Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com> * feat: comprehensive diffusion model infrastructure & model library (#843) * feat: comprehensive diffusion model infrastructure & model library (Issue #262) Add 60+ diffusion models, 13 noise schedulers, 2 noise predictors, conditioning modules, and comprehensive tests covering text-to-image, video, 3D, audio, super-resolution, image editing, and control/adapter architectures. Models: SD 1.5/2/3, FLUX.1, DALL-E 2/3, Kandinsky, Imagen, DeepFloyd IF, SDXL Turbo, PixArt-Sigma/Delta, RAPHAEL, eDiff-I, Hunyuan-DiT, Kolors, AuraFlow, Lumina-T2X, OmniGen, Stable Cascade, Playground v2.5, LCM, ControlNet-XS/Union, InstantID, PhotoMaker, IP-Adapter FaceID, UniControlNet, InstructPix2Pix, Prompt-to-Prompt, DiffEdit, LEDITS++, MagicBrush, Imagic, SDEdit, PaintByExample, Blended Diffusion, SD Upscaler, Real-ESRGAN, StableSR, DiffBIR, SUPIR, Upscale-A-Video, Sora, ModelScope T2V, Latte, Open-Sora, Runway Gen, Make-A-Video, Kling, Veo, Mochi 1, HunyuanVideo, LTX-Video, Wan, SyncDreamer, Wonder3D, One-2-3-45, Instant3D, DreamGaussian, LGM, TripoSR, Meshy, Stable Audio, Bark, VoiceCraft, SoundStorm, Udio. Schedulers: DDPM, Euler, Euler Ancestral, DPM++ Multistep, DPM++ SDE, DPM++ Singlestep, DEIS, Heun, LMS, LCM, FlowMatching, UniPC, Consistency. Noise Predictors: MMDiT (joint attention), U-ViT (vision transformer). Conditioning: ChatGLM3 text conditioner for Chinese language models. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * refactor: upgrade diffusion shell models to golden pattern & reorganize folders (Issue #262) - Reorganized 92+ models from flat Models/ dir into domain subdirectories (TextToImage, ImageEditing, Video, Audio, ThreeD, Control, SuperResolution, FastGeneration) - Moved schedulers from NeuralNetworks/Diffusion/Schedulers to Diffusion/Schedulers - Moved DiffusionModelBase and DDPMModel from NeuralNetworks/Diffusion to Diffusion/ - Added NeuralNetworkArchitecture<T>? to all base classes and model constructors - Upgraded ~40 shell models to production golden pattern with: - Rich XML docs (architecture list, For Beginners, examples, paper references) - [MemberNotNull] + InitializeLayers with Architecture support - #region organization, parameter validation, 8-10+ metadata properties - Updated all noise predictors, VAE classes, and test files for new namespaces Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * refactor: add NeuralNetworkArchitecture<T> support to all remaining diffusion models (Issue #262) Add NeuralNetworkArchitecture<T>? constructor parameter and architecture forwarding to UNet noise predictors across all 53 remaining diffusion model files: - TextToImage (19): DallE2/3, DeepFloydIF, Flux1, SDXL, SD2/3, PixArt variants, etc. - ImageEditing (8): BlendedDiffusion, DiffEdit, Imagic, LEDITSPP, MagicBrush, etc. - Control (5): ControlNet, ControlNetXS, IPAdapter, InstantID, T2IAdapter - Audio (6): AudioLDM, AudioLDM2, MusicGen, JEN1, DiffWave, Riffusion - ThreeD (6): DreamFusion, Magic3D, MVDream, PointE, ShapE, Zero123 - Video (5): AnimateDiff, CogVideo, LuminaT2X, StableVideoDiffusion, VideoCrafter - FastGeneration (4): AuraFlow, ConsistencyModel, LatentConsistency, SDTurbo Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * refactor: apply golden pattern to 19 pre-existing diffusion models (Issue #262) Add #region organization, [MemberNotNull]+InitializeLayers, <remarks> on constants, and SetParameters validation to all pre-existing models: Audio: AudioLDM, AudioLDM2, MusicGen, Riffusion, DiffWave TextToImage: SDXL, DallE3, PixArt ThreeD: PointE, ShapE, Zero123, DreamFusion, MVDream Video: AnimateDiff, VideoCrafter Control: ControlNet, IPAdapter FastGeneration: ConsistencyModel Core: DDPMModel Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix: address pr review comments for diffusion model infrastructure - Fix LCMScheduler timesteps: computed timestep list was never applied. Added SetTimestepArray protected method to NoiseSchedulerBase so derived schedulers can override the timestep schedule. - Fix DiTNoisePredictor MLP1 assignment: use pattern-match (as) instead of ternary with type-check cast for clearer layer handling. - Document ControlNetXS encoder channel multiplier design choice: intentionally uses 3 levels [1,2,4] vs main U-Net's [1,2,4,4]. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> --------- Co-authored-by: franklinic <franklin@ivorycloud.com> Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com> * feat: implement private set intersection protocols for entity alignment (#538) Add complete PSI subsystem for vertical federated learning entity alignment: - 4 PSI protocols: Diffie-Hellman, Oblivious Transfer (cuckoo hashing), Circuit-Based (garbled circuits), Bloom Filter (probabilistic) - Multi-party PSI with star topology for 3+ parties - Unbalanced PSI optimized for asymmetric set sizes - 5 fuzzy matchers: exact, edit distance (Levenshtein), phonetic (Soundex), n-gram overlap, Jaccard similarity - EntityAligner orchestrator with overlap sufficiency checks - Full net10.0 and net471 compatibility Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * feat: implement production-grade vertical federated learning (#542) Add complete VFL system with split neural networks, entity alignment, secure gradient exchange, label differential privacy, and GDPR unlearning. - 4 enums: VflAggregationMode, MissingFeatureStrategy, SplitPointStrategy, VflUnlearningMethod - 4 options: SplitModelOptions, MissingFeatureOptions, VflUnlearningOptions, VerticalFederatedLearningOptions - 4 interfaces: IVerticalParty, IVerticalFederatedTrainer, ISplitModel, ILabelProtector - 10 implementations: VerticalFederatedTrainer, VerticalPartyClient, VerticalPartyLabelHolder, SplitNeuralNetwork, VerticalDataPartitioner, SecureGradientExchange, MissingFeatureHandler, LabelDifferentialPrivacy, VerticalFederatedUnlearner, VerticalFederatedBenchmark - 1 result types file: VflResults (VflAlignmentSummary, VflEpochResult, VflTrainingResult) Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * feat: implement MPC beyond secure aggregation (#540) Add complete multi-party computation framework with arithmetic/boolean secret sharing, oblivious transfer, garbled circuits, secure comparison, and clipping. - 2 enums: MpcProtocol, MpcSecurityModel - 1 options: MpcOptions - 4 interfaces: ISecureComputationProtocol, IObliviousTransfer, IGarbledCircuit, ISecretSharingScheme - 9 implementations: ArithmeticSecretSharing (Beaver triples), BooleanSecretSharing (AND triples), BaseObliviousTransfer, ExtendedObliviousTransfer (IKNP), GarbledCircuitGenerator (free XOR, half-gates, point-and-permute), GarbledCircuitEvaluator, SecureComparisonProtocol, SecureClippingProtocol (norm + value clipping), HybridMpcProtocol (SS + GC orchestrator) Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * feat: implement verifiable FL with zero-knowledge proofs (#539) Add complete ZK verification framework with commitment schemes, range proofs, loss threshold proofs, computation integrity, and verifiable aggregation. - 2 enums: ZkProofSystem, VerificationLevel - 2 options: VerificationOptions, CommitmentOptions - 3 interfaces: IVerifiableComputation, IGradientCommitment, IZkProofSystem - 8 implementations: HashCommitmentScheme (SHA-256), PedersenCommitment (homomorphic), GradientNormRangeProof (Bulletproofs-style), GradientBoundednessProof (element-wise), LossThresholdProof (ZKP-FedEval), ComputationIntegrityProof (hash-chain), ModelUpdateVerifier (server-side engine), VerifiableAggregationStrategy (decorator) Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * feat: add trusted execution environment integration for federated learning (#537) Implements TEE-based secure aggregation with support for Intel SGX, Intel TDX, AMD SEV-SNP, ARM CCA, and a simulated provider for testing. Includes remote attestation verification, enclave lifecycle management, and AES-256 authenticated encryption with cross-framework support (AES-GCM on .NET Core+, AES-CBC-HMAC on .NET Framework). 16 new files: 4 options/enums, 3 interfaces, 1 model, 7 providers, 1 AES helper. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * feat: add federated graph learning for subgraph-level GNN training (#541) Implements graph FL with cross-client edge discovery (PSI-based), pseudo-node expansion (ZeroFill, FeatureAverage, GeneratorBased), graph-aware aggregation (edge/label/degree composite weighting), METIS/community/stream graph partitioning, LDP neighborhood privacy, prototype-based FGL, and a learned pseudo-node generator. 18 new files: 6 options/enums, 4 interfaces, 8 implementations. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * feat: add federated unlearning system with 4 methods (#849) Implements GDPR right-to-be-forgotten for federated learning with: - ExactRetraining: gold-standard re-aggregation excluding target client - GradientAscent: fast approximate unlearning via reversed training - InfluenceFunction: Newton step via LiSSA inverse Hessian approximation - DiffusiveNoise: structured noise injection targeting memorized samples Includes UnlearningCertificate with membership inference scoring, model divergence tracking, and audit trail hashing. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * feat: add fairness constraints, contribution evaluation, and incentives (#850) Implements client contribution evaluation with 3 methods: - ShapleyValue: exact marginal contribution (exponential, small federations) - DataShapley: Monte Carlo approximation with early convergence - Prototypical: alignment/diversity/magnitude scoring (constant cost) Adds group fairness constraints (demographic parity, equalized odds, equal opportunity, minimax) with automatic weight adjustment. Includes contribution-based incentive mechanism with trust scoring using exponential moving average and consistency bonus. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * feat: add advanced communication compression for federated learning (#851) Implements 4 advanced compression methods beyond basic TopK/Quantization: - PowerSGD: low-rank gradient approximation with warm-start factor reuse - GradientSketch: Count Sketch compression with median estimation and top-k - ErrorFeedback: residual accumulation wrapper with 1-bit SGD support - Adaptive: bandwidth/importance/staleness-aware per-client compression Extends FederatedCompressionOptions with AdvancedCompressionOptions sub-property. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * feat: add federated concept drift detection and adaptation (#852) Implements drift detection for federated learning with: - StatisticalDriftDetector: Page-Hinkley, ADWIN, DDM tests on client metrics - ModelDriftDetector: gradient direction and weight divergence analysis - DriftAdaptiveAggregator: wraps any strategy with drift-aware weight adjustment Includes DriftReport with per-client scores, drift type classification (sudden/gradual/recurring), and recommended actions (monitor, reduce weight, selective retrain, temporary exclude). Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * test: add unit tests for federated unlearning and compression Add 92 unit tests covering issues #849-852: - FederatedUnlearningTests: exact retraining, gradient ascent, influence functions, diffusive noise - FederatedFairnessAndContributionTests: Shapley, DataShapley, prototypical evaluators, group fairness, incentives - AdvancedCompressionTests: PowerSGD, gradient sketch, error feedback, adaptive compression - FederatedDriftDetectionTests: statistical drift detection, model-based detection, drift adaptive aggregator Also adds contribution score fields to RoundMetadata and FederatedLearningMetadata for #850 acceptance criteria. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix: resolve 5 bugs in federated learning mpc, tee, and vfl subsystems - ArithmeticSecretSharing: beaver triples now match actual party count instead of always using constructor's numberOfParties - GarbledCircuitGenerator: fix key derivation index mismatch between garbler (used outputWire) and evaluator (used gate array index) - GarbledCircuitGenerator: implement proper point-and-permute table permutation using color bits for correct gate evaluation - TeeSecureAggregation: auto-generate session key when EncryptForSubmission or GenerateSessionKey is called before BeginRound - SplitNeuralNetwork: fix ConcatenateEmbeddings to handle 1D embeddings by detecting rank and setting batchSize=1 for 1D tensors Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * test: add 599 comprehensive integration tests for 6 federated learning subsystems Add integration tests covering all 6 advanced federated learning features: - TEE (#537): 40+ tests for SimulatedTeeProvider, TeeSecureAggregation, TeeAesHelper, attestation, data sealing/unsealing, and all 5 provider types - PSI (#538): 35+ tests for DiffieHellmanPsi, BloomFilterPsi, MultiPartyPsi, FuzzyPsi, EntityAligner, and intersection computation - ZK Proofs (#539): 25+ tests for HashCommitmentScheme, PedersenCommitment, GradientNormRangeProof, GradientBoundednessProof, LossThresholdProof, ComputationIntegrityProof, and ModelUpdateVerifier - MPC (#540): 35+ tests for ArithmeticSecretSharing, BooleanSecretSharing, BaseObliviousTransfer, ExtendedObliviousTransfer, GarbledCircuit, SecureComparisonProtocol, SecureClippingProtocol, and HybridMpcProtocol - Graph FL (#541): 25+ tests for FedGnnAggregationStrategy, SubgraphExpander, FederatedGraphPartitioner, GraphNeighborhoodPrivacy, and SubgraphFederatedTrainer - VFL (#542): 40+ tests for VerticalFederatedTrainer, SplitNeuralNetwork, SecureGradientExchange, MissingFeatureHandler, LabelDifferentialPrivacy, VerticalPartyClient, VerticalPartyLabelHolder, and VerticalFederatedBenchmark All 599 tests pass. Tests discovered and helped fix 5 source code bugs. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix: address pr review comments for accessibility and code quality - ActivationCheckpointConfig: default RecomputeStrategy to None since Selective/Full are not yet implemented (avoids NotImplementedException) - LoadBalancedPartitionStrategy: enforce first boundary == 0 to prevent orphaned parameters, add overflow checks for long-to-int casts, remove useless local variable assignment - Pipeline schedules (1F1B, Interleaved, ZB-H1/H2/V, LoopedBFS): use double arithmetic directly in EstimateBubbleFraction to avoid precision loss from long-to-double implicit conversions - PipelineParallelModel: combine nested if statements, preserve inner exceptions in catch clauses for better debugging Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix: add missing lazy initialization calls and looped bfs forward output retention PipelineParallelModel had several methods that accessed _numStages and other derived state before EnsureShardingInitialized() was triggered: - Train() used _numStages for GetSchedule before GatherFullParameters - GetModelMetadata() accessed EstimatedBubbleFraction without init - EstimatedBubbleFraction property lacked init call - Serialize() used _virtualStagesPerRank without init - Deserialize() called InitializeSharding directly, skipping OnBeforeInitializeSharding which sets _numStages Also fixes LoopedBFS forward output retention: when processing multiple virtual stages, vStage 0's forward outputs must be retained during backward since vStage 1 uses them as input. Adds 45 integration tests covering all 7 pipeline schedules, micro-batch slicing, activation checkpointing, load-balanced partitioning, virtual stage dependencies, and serialization. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix: correct virtual-stage routing, shard update, and checkpointing eviction Virtual-stage dataflow (V>1 schedules): - Every virtual stage now communicates with adjacent physical ranks, not just the first/last virtual stage on the rank. For rank i holding global chunks {i, i+P, i+2P, ...}, each chunk sends forward to rank (i+1)%P and receives from rank (i-1+P)%P, matching the global pipeline flow: chunk 0 -> 1 -> 2 -> ... -> V*P-1. - Tags use global virtual stage indices for unique identification. - Single-rank V>1 uses local forwarding without communication. Shard update for non-contiguous partitions: - Added UpdateLocalShardFromFullParameters that correctly gathers non-contiguous virtual stage chunks for V>1, instead of using base class UpdateLocalShardFromFull which assumes contiguous shard. Checkpointing eviction policy: - Renamed FreeNonCheckpointedActivations to FreeConsumedActivations. - After backward consumes an activation, frees it unconditionally from forwardInputs (not just for non-checkpointed ops). - With RecomputeStrategy.None, also frees checkpointed copies since they cannot be used for recomputation. Also fixes ContributionBasedIncentive operator spacing. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix: add node reuse flag to prevent generator dll file lock Add -nodeReuse:false to both build commands in the CI workflow to prevent MSBuild nodes from holding file handles on the source generator DLL while another build process tries to write to it. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * ci: trigger workflow runs after commit message cleanup Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * style: remove extra blank line in pipeline parallel model imports Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix: build source generator separately to prevent dll file lock Build AiDotNet.Generators.csproj first, then shutdown build servers to release file handles before building the full solution. The generator is both a solution project and an analyzer reference, causing the compiler to lock the DLL while MSBuild tries to rebuild. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix: address pr review comments for pipeline parallelism and diffusion - fix error message in activationcheckpointconfig (no division-by-zero) - rename _pipelinemicrobatchsize to _pipelinemicrobatchcount for consistency - change rectified flow beta_end from 1.0 to 0.02 (standard value) - remove dead log-sigma code in lmsdiscretescheduler - implement proper two-pass heun method with predictor/corrector state - fix random usage in consistencymodelscheduler to use randomhelper - fix consistencymodel get/setparameters to include vae parameters - make ipadaptermodel fields readonly, inline init into constructor - make audioldmmodel fields readonly, inline init into constructor - add using directive for layercategory in quantizationconfiguration - pre-compute valid partition points in edgeoptimizer - remove redundant settargetframework from generator project reference Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix: use double locals in estimatebubblefraction to prevent int precision loss Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix: address 9 new pr review comments for pipeline parallelism - EdgeOptimizer: implement strategy-aware partition (EarlyLayers=25%, LateLayers=75%, Adaptive=FLOP-balanced) instead of all-same fallback - DualTextConditioner: re-encode via CLIP in GetPooledEmbedding to avoid sending T5 embeddings to CLIP pooling (dimension mismatch) - LCMScheduler: guard originalStride with Math.Max(1, ...) to prevent zero stride when originalInferenceSteps > TrainTimesteps - LMSDiscreteScheduler: trim _derivativeHistory to last _order entries to prevent unbounded memory growth over many steps - sonarcloud.yml: fix -nodeReuse:false MSBuild syntax to -p:nodeReuse=false for proper dotnet CLI forwarding (2 occurrences) - QuantizationConfiguration: return 16 for skip layers in Float16 mode instead of hardcoded 32 Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix: align ef core package versions to 10.0.3 in aidotnet.serving Microsoft.EntityFrameworkCore.Sqlite 10.0.3 requires Microsoft.EntityFrameworkCore.Relational >= 10.0.3 transitively, but the project had Relational at 10.0.2 causing NU1605 downgrade errors. Updated all EF Core packages to 10.0.3 for consistency. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix: address 12 new pr review comments for security and correctness - ILabelProtector: fix inverted epsilon/delta privacy budget docs (larger epsilon = more privacy spent, not smaller) - VerifiableAggregationStrategy: skip clients missing from weights dict to prevent mismatched models/weights in inner aggregator - UnbalancedPsi: guard tagBytes with Math.Max(1, ...) to prevent empty tags when securityParameter < 8 - CircuitBasedPsi: same guard for outputBytes in ComputePrf - SecureClippingProtocol: document that reconstructing scale factor reveals min(1, C/||g||) not raw norm; note secure sqrt/divide would be needed for full privacy - ExtendedObliviousTransfer: derive pad0 from _baseKeys0 and pad1 from _baseKeys1 so receiver can only decrypt chosen message - HeunDiscreteScheduler: validate timestep matches stored predictor state in corrector pass, throw if mismatched - DualTextConditioner: use CLIP unconditional embedding for pooling fallback instead of re-encoding T5 embeddings through CLIP - LossThresholdProof: remove unused loss variable in GenerateProof - SecureCrossClientEdgeDiscovery: use numerically stable sigmoid 1/(1+exp(-eps)) to prevent overflow for large epsilon values Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix: address code scanning alerts — overflow guards, method extraction, validation - Interleaved1F1BSchedule: use long arithmetic for int*int multiplications (numStages * virtualStagesPerRank, numMicroBatches * virtualStagesPerRank) - PipelineParallelModel: use checked() for tag computation and operation keys; combine nested if into single condition in ExecuteForward - ZeroBubbleH1/H2/VSchedule: extract EmitWarmup, EmitSteadyState, EmitCooldown helper methods to reduce GetSchedule block complexity - LoadBalancedPartitionStrategy: extract ComputePrefixSums, SolveMinMaxDP, BacktrackSplitPoints, ConvertToPartitions from OptimalPartition; use ternary in AssignOneLayerPerStage; use ArrayPolyfill.Fill for .NET Framework compat - GraphNeighborhoodPrivacy: validate embeddingDim and check tensor divisibility - SecureClippingProtocol: document plaintext norm reconstruction limitation - UnbalancedPsi: validate securityParameter >= 8 in ComputeTag Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * feat: add facade integration for all federated learning v2 subsystems Extends ConfigureFederatedLearning with 8 new injectable interface parameters covering PSI, MPC, TEE, ZK proofs, unlearning, drift detection, contribution evaluation, and fairness constraints. Adds FederatedLearningMode enum (Horizontal/Vertical) and 6 new options properties to FederatedLearningOptions for vertical learning, PSI, MPC, verification, and advanced compression. Changes all 21 options property defaults from null to new() to follow industry-standard defaults convention. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * feat: add federated learning research gaps — distillation, PEFT, decentralized, continual learning, backdoor defense Implements 5 missing FL capabilities identified from 2024-2026 research audit: - Federated Knowledge Distillation (FedMD, FedDF, FedGEN) for model-heterogeneous FL - Federated PEFT/LoRA adapters (FedEx-LoRA, HeLoRA, prompt tuning) for LLM fine-tuning - Decentralized P2P FL (gossip protocol, ring all-reduce) for serverless deployments - Federated Continual Learning (EWC, orthogonal projection) for catastrophic forgetting prevention - Backdoor Defense (Neural Cleanse, Direction Alignment Inspector) for poisoning detection All wired into FederatedLearningOptions facade with options classes. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix: address 7 copilot review comments for correctness and naming - LossThresholdProof: reject NaN/Infinity loss values in proof generation and verification - TeeProviderBase: enforce 64-byte max on attestation report data - SimulatedTeeProvider: fix misleading "stable" comment to "ephemeral" - DualTextConditioner: use input when CLIP-compatible instead of ignoring it - SecureCrossClientEdgeDiscovery: validate privacy epsilon is positive - SecureClippingProtocol: remove unused SecureCompare, use direct norm check - PipelineParallelModel: rename microBatchSize to microBatchCount for clarity Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * feat: add arbitrary tensor rank support for all neural network layers Enable all layers to accept any tensor rank by flattening leading dimensions into a batch dimension, processing as fixed-rank, then restoring original shape. This follows the canonical pattern already in ConvolutionalLayer across 18 files spanning pooling, 3D convolution, graph, normalization, and sequence layers. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix: address PR #845 review comments for correctness and clarity - Add comment explaining betaEnd=0.02 for Rectified Flow (DDPM standard) - Fix misleading comment in LCMScheduler about stride=0 guard - Extract GetFullPrecisionBitWidth() helper in QuantizationConfiguration - Upgrade EdgeOptimizer partition fallback from Debug.WriteLine to Console warning - Throw ArgumentException in DualTextConditioner when embedding dims don't match instead of silently returning unconditional embedding - Fix potential integer overflow in EstimateBubbleFraction: use 1.0 instead of 1 in double arithmetic (OneForwardOneBackward, Interleaved1F1B, LoopedBFS) Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix: rename microbatchsize to microbatchcount in pipeline parallelism tests The PipelineParallelModel constructor parameter was renamed from microBatchSize to microBatchCount but tests were not updated to match, causing CS1739 build errors in CI. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix: convert pipeline parallelism to generic types and fix code scanning alerts Replace hardcoded double types with generic T using INumericOperations<T> and replace double[]/double[][] with Vector<T>/Matrix<T>: - IPipelineSchedule -> IPipelineSchedule<T> with T EstimateBubbleFraction - All 7 schedule classes made generic with NumOps arithmetic - LoadBalancedPartitionStrategy: Func<int,double> -> Func<int,T>, double[] -> Vector<T>, double[][] -> Matrix<T> - PipelineParallelModel/AiModelBuilder: updated to IPipelineSchedule<T> Fix code scanning alerts: - ContributionBasedIncentive: cap floor budget when n is large - ExtendedObliviousTransfer: validate securityParameter > 0 - EdgeOptimizer: replace Console.WriteLine with Debug.WriteLine - TeeProviderBase: replace magic number 17 with named constant - SecureCrossClientEdgeDiscovery: remove dead negative-epsilon branch Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix: add missing generic type arguments to pipeline schedule test usages All pipeline schedule classes (GPipeSchedule, OneForwardOneBackwardSchedule, ZeroBubbleH1Schedule, ZeroBubbleH2Schedule, ZeroBubbleVSchedule, Interleaved1F1BSchedule, LoopedBFSSchedule) are generic and require <T> type arguments. Added <double> to all instantiations and typeof() references in both DistributedTrainingValidationTests and PipelineParallelismIntegrationTests. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com> Co-authored-by: franklinic <franklin@ivorycloud.com>
Summary
ILayeredModel<T>interface providing per-layer access to neural network architecture and parameters, enabling layer-aware operations across the stackLayerCategoryenum (15 categories) andLayerInfo<T>metadata record for automated per-layer decisionsLayerBase<T>:GetLayerCategory(),EstimateFlops(),EstimateActivationMemory()ILayeredModel<T>inNeuralNetworkBase<T>(all 80+ neural networks get layer access automatically)Changed Files
New files
src/Interfaces/ILayeredModel.cs- Core interface withLayers,LayerCount,GetLayerInfo(),GetAllLayerInfo(),ValidatePartitionPoint()src/Interfaces/LayerCategory.cs- 15-value enum: Dense, Convolution, Attention, Normalization, Activation, Pooling, Embedding, Recurrent, Regularization, Residual, FeedForward, Graph, Structural, Input, Othersrc/Interfaces/LayerInfo.cs- Metadata record: Index, Name, Category, Layer, ParameterOffset, ParameterCount, InputShape, OutputShape, IsTrainable, EstimatedFlops, EstimatedActivationMemoryCore implementation
src/Interfaces/INeuralNetwork.cs- AddedILayeredModel<T>to inheritancesrc/NeuralNetworks/Layers/LayerBase.cs- Added 3 virtual methods for layer classification and cost estimationsrc/NeuralNetworks/NeuralNetworkBase.cs- FullILayeredModel<T>implementation with parameter offset computationSubsystem integrations
src/DistributedTraining/PipelineParallelModel.cs- FLOPs-balanced layer-boundary partitioningsrc/Training/Memory/TrainingMemoryManager.cs- Category-based activation checkpointingsrc/LoRA/DefaultLoRAConfiguration.cs- Automatic LoRA layer eligibility by categorysrc/MetaLearning/MetaLearnerBase.cs- Per-layer gradient application with parameter offsetssrc/Deployment/Export/ModelExporterBase.cs- Auto-infer input/output shapes from layerssrc/Deployment/Export/ExportConfiguration.cs- Added OutputShape propertysrc/Interfaces/IPruningStrategy.cs- Per-category sparsity targets in PruningConfigTests
tests/AiDotNet.Tests/Helpers/MockNeuralNetwork.cs- Added ILayeredModel implementationtests/AiDotNet.Tests/Helpers/ContinualLearningTestHelper.cs- Added ILayeredModel implementationTest plan
Closes #462
🤖 Generated with Claude Code
Summary by CodeRabbit