Skip to content

perf: replace scalar for-loops with IEngine hardware-accelerated ops in 147 files - #913

Merged
ooples merged 60 commits into
masterfrom
perf/iengine-tensor-ops-910
Mar 1, 2026
Merged

ooples merged 60 commits into
masterfrom
perf/iengine-tensor-ops-910

Conversation

@ooples

@ooples ooples commented Feb 28, 2026 •

Copy link
Copy Markdown
Owner

Summary

  • Replaces scalar NumOps.Subtract/NumOps.Multiply/NumOps.Add for-loops with hardware-accelerated IEngine vector and tensor operations across 147 files
  • Net reduction of ~1,218 lines of code (2,395 deletions, 1,177 insertions) by replacing verbose element-wise loops with single Engine method calls
  • Adds Engine property to DeepReinforcementLearningAgentBase and ContinualLearningStrategyBase for subclass access

Categories of files converted:

  • Video models (30+ files): frame interpolation, motion estimation, segmentation, generation, recognition, inpainting
  • Finance models: LSTNet, TSMixer, neural forecasting, transformers, graph neural networks
  • Document models (29 files): OCR, layout analysis, page segmentation, pixel-to-sequence
  • Audio models (3 files): VITS, Tacotron2, Wav2Vec2
  • SSL methods (9 files): SimCLR, MoCo, MoCoV3, BYOL, SimSiam, BarlowTwins, MAE, TeacherStudentSSL, SSLFineTuningPipeline
  • RL agents (2 files): DQN, MuZero
  • Regression (2 files): RegressionBase, NonLinearRegressionBase
  • Neural networks (4 files): Gpt4Vision, Blip2, AudioVisual networks
  • Tabular models (10+ files): layers, base classes, utilities
  • SSM layers (10+ files): fill operations, reduce-sum, concatenation
  • Other: NeRF, TimeSeries, AiModelBuilder, diffusion layers

Patterns replaced:

  1. SGD parameter updates: for(i) params[i] = NumOps.Subtract(params[i], NumOps.Multiply(lr, grads[i])) → Engine.Subtract(params, Engine.Multiply(grads, lr))
  2. EMA updates: for(i) target[i] = NumOps.Add(NumOps.Multiply(m, target[i]), NumOps.Multiply(1-m, online[i])) → Engine.Add(Engine.Multiply(target, m), Engine.Multiply(online, 1-m))
  3. Tensor element-wise ops: scalar loops → Engine.TensorAdd, TensorSubtract, TensorMultiplyScalar, TensorFill
  4. Vector subtract/add: manual loops → Engine.Subtract(vec1, vec2), Engine.Add(vec1, vec2)
  5. Reduce-sum: manual accumulation loops → Engine.ReduceSum

Test plan

  • Build succeeds with dotnet build src/AiDotNet.csproj --framework net10.0 -c Release (0 errors, 0 warnings)
  • Run existing unit tests to verify behavioral equivalence
  • Verify GPU acceleration path works when available

🤖 Generated with Claude Code

Closes #910

Summary by CodeRabbit

  • Chores

    • Bumped core tensor and optional native acceleration packages to v0.8.0.
  • Refactor

    • Replaced many element-wise loops with centralized engine-backed tensor/vector operations and added shared vector math helpers (norms, normalization, distances, cosine similarity) for consistency and performance.
  • New Features

    • Many layers/models now expose standardized training/export interfaces and engine integrations for improved interoperability and exportability.

ooples and others added 24 commits February 28, 2026 13:19
Upgrade AiDotNet.Tensors, AiDotNet.Native.OpenBLAS,
AiDotNet.Native.CLBlast, and AiDotNet.Native.OneDNN from 0.7.0
to 0.8.0 to get new IEngine tensor-level operations needed for
replacing scalar for-loops across the codebase (issue #910).

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Replace 3 scalar for-loops in VideoNeuralNetworkBase with
hardware-accelerated IEngine tensor operations:
- NormalizeFrames: scalar divide loop → TensorDivideScalar
- DenormalizeFrames: scalar multiply+clamp → TensorMultiplyScalar+TensorClamp
- ConcatenateFeatures: manual copy loops → TensorConcatenate

Part of #910.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Replace scalar element-wise loops with hardware-accelerated IEngine
tensor operations in RAFT and GMFlow optical flow models:
- AddTensors: Engine.TensorAdd
- ConcatenateChannels: Engine.TensorConcatenate
- ApplySigmoid: Engine.Sigmoid
- Loss gradient subtraction: Engine.TensorSubtract
- GRU gate arithmetic: Engine.TensorMultiply/TensorSubtract

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
- TimeEmbeddingLayer: replace scalar matmul forward/backward loops with Engine.TensorMatMul and Engine.TensorBroadcastAdd
- TimeEmbeddingLayer: replace scalar gradient accumulation with Engine.ReduceSum and Engine.TensorTranspose
- CrossAttentionLayer: replace scalar weight update loop with Engine.TensorMultiplyScalar and Engine.TensorSubtract
- ConvolutionalLayer: replace scalar UpdateParameters loops with Engine.TensorSubtract/TensorMultiplyScalar
- DeformableConvolutionalLayer: replace 6 scalar UpdateParameters loops with Engine tensor ops
- DeformableConvolutionalLayer: replace scalar sigmoid derivative loop with Engine.TensorMultiply/TensorSubtract
- DeformableConvolutionalLayer: replace scalar gradient sum loop with Engine.TensorAdd

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
RIFE: AddTensors, ConcatenateChannels, ScaleFlow, CombineFlowGradients, loss gradient
FILM: AddTensors, ConcatenateChannels, ApplySigmoid, ScaleFlow, loss gradient
DRVI/ABME: timestep interpolation loops

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
…ormers, and graph

Replace element-wise scalar for-loops with hardware-accelerated IEngine
tensor operations in Finance module files:
- TCN, WaveNet, NHiTSFinance, NBEATSFinance, DeepFactor, DeepState
- TimesFM, Autoformer, TimesNet, Crossformer, TFT
- DCRNN diffusion convolution

Patterns replaced:
- AddTensors scalar loops -> Engine.TensorAdd
- SubtractTensors scalar loops -> Engine.TensorSubtract
- MultiplyTensors scalar loops -> Engine.TensorMultiply
- Scalar multiply + add patterns -> Engine.TensorMultiplyScalar + Engine.TensorAdd
- Broadcasting add patterns -> Engine.TensorBroadcastAdd

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Replace 34 instances of `for (int i = 0; i < ones.Length; i++) ones[i] = NumOps.One`
with `ones.Fill(NumOps.One)` across 28 SSM layer files. The Fill() method uses
hardware-accelerated memory operations instead of scalar element-by-element assignment.

Also fix null-forgiving operators in ABCLayer, BASEDLayer, and DeltaFormerLayer
UpdateParameters methods by adding proper null checks for all gradient fields.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
- SwinTransformerBlockLayer: replace AddTensors scalar loop with Engine.TensorAdd
- SpyNetLayer: replace AddTensors and AccumulatePyramidGradient scalar loops with Engine.TensorAdd
- UNetDiscriminator: replace AddTensors scalar loop with Engine.TensorAdd
- ResidualDenseBlock: replace AddResidual, ScaleGradient, AddTensors loops with Engine tensor ops
- RRDBLayer: replace AddResidual, ScaleGradient, AddTensors loops with Engine tensor ops
- RRDBNetGenerator: replace AddTensors scalar loop with Engine.TensorAdd

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
SAM2: AddTensors, ConcatenateChannels, ApplySigmoid, loss gradient
XMem: ConcatenateChannels, QueryMemory accumulation and scaling

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
StableVideoDiffusion: AddTensors, ConcatenateChannels, loss gradient,
  ApplyGuidance, InterpolateLatents, AddNoise, AddNoiseAtLevel, DenoisingStep
OpenSora: AddTensors, ApplySigmoid, ApplyGuidance, noisyInput mixing,
  InitializeLatentsFromImage, DenoisingStep

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
InteractingLayer and IntersampleAttention now inherit from LayerBase<T>,
gaining Engine, NumOps, Random from the base class. Replaced scalar
for-loops with Engine.TensorMatMul, Engine.Softmax, Engine.ReLU,
Engine.TensorAdd, Engine.TensorSubtract, Engine.TensorMultiplyScalar,
Engine.TensorFill, and Tensor.Transpose operations.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
…understanding

VideoMAE: ApplySoftmax -> Engine.Softmax
VideoCLIP: AddTensors, loss gradient

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
ProPainter: AddTensors, ConcatenateChannelsDim1, loss gradient
E2FGVI: AddTensors, ConcatenateChannels, ScaleTensor, ApplySigmoid

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
…ixer

Replace element-wise scalar loops with hardware-accelerated IEngine
tensor operations:
- LSTNet: highway connection add loop -> Engine.TensorAdd
- TSMixer: RevIN normalize loop -> Engine.TensorSubtractScalar +
  Engine.TensorDivideScalar
- TSMixer: RevIN denormalize loop -> Engine.TensorMultiplyScalar +
  Engine.TensorAddScalar

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
…giving operators

- HighwayLayer: replace scalar double-loop sum in ComputeAuxiliaryLoss with Engine.ReduceSum
- GroupedQueryAttentionLayer: replace single null check with comprehensive null checks
  for all 5 gradient fields and remove all null-forgiving operators

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
… layers

Remove all null-forgiving operators (!) from UpdateParameters methods across
32 SSM layer files. Replace single-field null checks with comprehensive
validation of all gradient fields before use. This follows the project rule
of never using the null-forgiving operator.

Also fix null-forgiving operators in Backward methods for HedgehogLayer,
MEGALayer, MegalodonLayer, RetNetLayer, and TransNormerLLMLayer gradient
accumulation code.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
…r ops

Convert AttentiveTransformer, BatchEnsembleLayer, FeatureTransformer,
GatedFeatureLearningUnit, and SoftTree to inherit from LayerBase<T> and
replace scalar for-loops with hardware-accelerated IEngine operations
including TensorMultiply, TensorSubtract, TensorAdd, TensorMultiplyScalar,
TensorFill, and Sigmoid.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
…inear encoding with iengine ops

Convert ObliviousDecisionTree and PiecewiseLinearEncoding to inherit from
LayerBase<T> and replace scalar gradient zeroing and parameter update loops
with Engine.TensorFill, Engine.TensorSubtract, and Engine.TensorMultiplyScalar.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
… iengine ops

Add Engine access via AiDotNetEngine.Current to GANDALFBase, TabTransformerBase,
TabMBase, TabNetBase, AutoIntBase, CLSToken, ColumnEmbedding, and
ContrastivePretraining. Replace scalar parameter update loops with
Engine.TensorSubtract/TensorMultiplyScalar, gradient zero-fill with
Engine.TensorFill, sigmoid with Engine.Sigmoid, element-wise multiply with
Engine.TensorMultiply, and tensor creation with Tensor<T>.CreateDefault.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Replace scalar for-loops with hardware-accelerated IEngine tensor
operations in SAINTBase, NODEBase, TabDPTBase, TabPFNBase, and
FeatureTokenizer. Key replacements include AddTensors helpers,
tree output aggregation, TransformerBlock residual connections,
MatMul, UpdateParameters, ResetState gradient zeroing, and
categorical embedding additions.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Replace element-wise for loops using NumOps.Subtract/Multiply with
Engine.Subtract/Engine.Multiply vector operations for hardware-accelerated
parameter updates across Document, Audio, NeuralNetworks, Regression,
TimeSeries, and NeuralRadianceFields modules.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
…files

Replace element-wise for loops using NumOps.Subtract/Multiply with
Engine.Subtract/Engine.Multiply vector operations across all remaining
Document models (VisionLanguage, PixelToSequence, OCR, LayoutAware,
GraphBased).

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
…l files

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
…d base classes

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Copilot AI review requested due to automatic review settings February 28, 2026 20:02
@vercel

vercel Bot commented Feb 28, 2026 •

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated (UTC)
aidotnet-playground-api Ready Ready Preview, Comment Mar 1, 2026 4:33am

@coderabbitai

coderabbitai Bot commented Feb 28, 2026 •

Copy link
Copy Markdown
Contributor

Note

Reviews paused

It looks like this branch is under active development. To avoid overwhelming you with review comments due to an influx of new commits, CodeRabbit has automatically paused this review. You can configure this behavior by changing the reviews.auto_review.auto_pause_after_reviewed_commits setting.

Use the following commands to manage reviews:

  • @coderabbitai resume to resume automatic reviews.
  • @coderabbitai review to trigger a single review.

Use the checkboxes below for quick actions:

  • ▶️ Resume reviews
  • 🔍 Trigger review

Walkthrough

Widespread refactor replacing element-wise numeric loops with AiDotNetEngine tensor ops and new VectorHelper utilities, adding Engine accessors, migrating many tabular components to LayerBase, and bumping AiDotNet.Tensors and native packages to 0.8.0. Public APIs largely preserved.

Changes

Cohort / File(s) Summary
Dependency Updates
src/AiDotNet.csproj
Bumped AiDotNet.Tensors, AiDotNet.Native.OpenBLAS, AiDotNet.Native.CLBlast, AiDotNet.Native.OneDNN 0.7.0 → 0.8.0.
Engine accessors added
src/**/AutoML/*, src/**/ContinualLearning/*, src/**/ReinforcementLearning/*, src/**/Tabular/*, src/**/TextSafetyModuleBase.cs, ...
Added protected/protected static IEngine Engine => AiDotNetEngine.Current in many bases and modules to route tensor ops through AiDotNetEngine.Current.
VectorHelper: new vector ops
src/Helpers/VectorHelper.cs, callers in src/**/*
Added L2Norm/Normalize/NormalizeInPlace/CosineSimilarity/DotProduct/EuclideanDistance/ManhattanDistance and replaced numerous local dot/norm/cosine implementations with VectorHelper calls.
Vectorized parameter updates (ubiquitous)
src/AiModelBuilder.cs, src/*/UpdateParameters, src/*/ApplyGradients (many files)
Replaced per-element SGD/EMA/EMA-like updates with vector ops: Engine.Subtract(params, Engine.Multiply(grads, lr)) or Engine.Add/Multiply patterns across many UpdateParameters/ApplyGradients paths.
Tensor op replacements (Finance, Video, Diffusion, Vision, etc.)
src/Finance/Forecasting/**, src/Video/**, src/Video/Generation/**, src/Diffusion/**, src/Video/Inpainting/**
Replaced manual loops/Transforms with Engine.TensorAdd/Subtract/Multiply/MultiplyScalar/Concatenate/MatMul/Sigmoid/Softmax and similar engine helpers.
Tabular layers → LayerBase migrations
src/NeuralNetworks/Tabular/{AttentiveTransformer,BatchEnsembleLayer,FeatureTransformer,GatedFeatureLearningUnit,InteractingLayer,IntersampleAttention,ObliviousDecisionTree,PiecewiseLinearEncoding,SoftTree,...}
Many tabular components now inherit LayerBase<T> and expose overrides (SupportsTraining, Forward/Backward, Get/SetParameters, UpdateParameters, ResetState, ExportComputationGraph); internals converted to NumOps/Engine patterns—large public-surface changes requiring API/behavior review.
High-density internal rewrites
src/NeuralNetworks/Layers/{DeformableConvolutionalLayer,TimeEmbeddingLayer,ConvolutionalLayer,Layer*}, src/NeuralNetworks/Layers/SSM/*
Significant changes: matmul, reduce, gradient accumulation, parameter updates moved to Engine ops. These files are dense and flagged for deep review (see notes).
Helper consolidations & removed local helpers
many src/** (dot/norm/cosine/add/subtract helpers removed)
Removed many private helpers (DotProduct, Normalize, CosineSimilarity, EuclideanDistance, AddTensors, etc.) in favor of VectorHelper and Engine methods; impacts consistency and numeric behavior across callers.
Minor initialization and fill optimizations
src/NeuralNetworks/Layers/SSM/*, various layers
Replaced manual loops filling ones/zeros with ones.Fill(NumOps.One) / Engine.TensorFill and strengthened null-checks on gradients.
Estimated hotspots needing attention
See detailed files below
Files with high review priority: TimeEmbeddingLayer.cs, DeformableConvolutionalLayer.cs, ConvolutionalLayer.cs, InteractingLayer.cs, FeatureTransformer.cs, BatchEnsembleLayer.cs, tabular model migrations, and VectorHelper.cs (new, broad impact).

Sequence Diagram(s)

sequenceDiagram
  participant Model as Model (UpdateParameters / Forward)
  participant Engine as AiDotNetEngine (IEngine)
  participant Helper as VectorHelper
  Model->>Engine: Engine.Multiply(gradients, lr) / Tensor ops
  Model->>Engine: Engine.Subtract(parameters, scaledGrads)
  Model->>Helper: VectorHelper.{L2Norm,Normalize,CosineSimilarity,...}
  Engine-->>Model: Tensor/Vector result
  Helper-->>Model: scalar/distance/similarity
Loading

Estimated code review effort

🎯 4 (Complex) | ⏱️ ~60 minutes

Possibly related PRs

Suggested labels

feature

"Loops turned to vectors, engines hum,
Helpers unified, the math now runs,
New LayerBase paths, big files to comb,
Flag dense layers and export hooks as BLOCKING,
Review VectorHelper and numeric-edge cases, then ship."

🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 61.61% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed Title clearly summarizes main change: systematic replacement of scalar loops with IEngine hardware-accelerated operations across 147 files.
Linked Issues check ✅ Passed PR fully addresses issue #910 Phase 1 objectives: replaced scalar loops with Engine tensor ops (TensorAdd, TensorSubtract, TensorMultiply, TensorMultiplyScalar) across video, finance, document, audio, SSL, RL, and tabular models. Added Engine properties to base classes. Scope, patterns, and priority modules align with issue requirements.
Out of Scope Changes check ✅ Passed All changes are in-scope replacements per #910 Phase 1. Changes consistently replace scalar loops with IEngine operations, add Engine properties to base classes, and introduce VectorHelper for vector operations. No unrelated refactoring or feature additions detected outside Engine-based optimization scope.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Post copyable unit tests in a comment
  • Commit unit tests in branch perf/iengine-tensor-ops-910

Comment @coderabbitai help to get the list of available commands and usage tips.

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR modernizes the AiDotNet model implementations by replacing many hand-written element-wise scalar loops with IEngine tensor/vector operations (SIMD/GPU-accelerated), and bumps the AiDotNet.Tensors/native acceleration package versions to align with the new API usage.

Changes:

  • Replaced numerous scalar NumOps.* loops with Engine.* / Engine.Tensor* operations (e.g., add/subtract/multiply, concatenate, clamp, softmax, reduce-sum).
  • Introduced/standardized access to the global engine in additional base classes (e.g., RL + continual learning) to support subclass conversions.
  • Updated NuGet references to AiDotNet.Tensors and native backends from 0.7.0 → 0.8.0.

Reviewed changes

Copilot reviewed 147 out of 147 changed files in this pull request and generated 9 comments.

Show a summary per file
File Description
src/Video/VideoNeuralNetworkBase.cs Uses engine scalar-divide/multiply+clamp for (de)normalization; replaces manual feature concatenation with TensorConcatenate.
src/Video/Understanding/VideoCLIP.cs Replaces per-element transform add/subtract with TensorAdd/TensorSubtract.
src/Video/Segmentation/XMem.cs Replaces manual accumulation/scale and channel concatenation with engine ops.
src/Video/Segmentation/SAM2.cs Replaces loss gradient, concatenation, add, sigmoid with engine ops.
src/Video/Motion/RAFT.cs Replaces loss gradient, GRU update math, concatenation, add, sigmoid with engine ops.
src/Video/Motion/GMFlow.cs Replaces loss gradient/diff/concat/add with engine ops.
src/Video/Inpainting/ProPainter.cs Replaces loss gradient, concatenation, add with engine ops.
src/Video/Inpainting/E2FGVI.cs Replaces scale/add/concat/sigmoid with engine ops.
src/Video/Generation/StableVideoDiffusion.cs Refactors diffusion algebra to engine tensor ops; replaces add helpers.
src/Video/Generation/OpenSora.cs Refactors diffusion algebra and guidance to engine tensor ops; replaces add/sigmoid helpers.
src/Video/FrameInterpolation/RIFE.cs Replaces loss gradient, concatenation, adds, flow scaling with engine ops.
src/Video/FrameInterpolation/FILM.cs Replaces loss gradient, concatenation, sigmoid, add, flow scaling with engine ops.
src/Video/FrameInterpolation/DRVI.cs Replaces scalar blending loop with tensor multiply/add.
src/Video/FrameInterpolation/ABME.cs Replaces scalar blending loop with tensor multiply/add.
src/Video/ActionRecognition/VideoMAE.cs Replaces manual softmax implementation with Engine.Softmax.
src/TimeSeries/TimeSeriesModelBase.cs Uses engine vector ops for SGD-style parameter updates.
src/SelfSupervisedLearning/TeacherStudentSSL.cs Uses engine vector ops for EMA + SGD parameter updates.
src/SelfSupervisedLearning/SimSiam.cs Uses engine vector ops for SGD parameter updates.
src/SelfSupervisedLearning/SimCLR.cs Uses engine vector ops for SGD parameter updates.
src/SelfSupervisedLearning/SSLFineTuningPipeline.cs Adds engine access and uses it for SGD parameter updates.
src/SelfSupervisedLearning/MoCoV3.cs Uses engine vector ops for SGD parameter updates.
src/SelfSupervisedLearning/MoCo.cs Uses engine vector ops for SGD parameter updates.
src/SelfSupervisedLearning/MAE.cs Uses engine vector ops for SGD parameter updates.
src/SelfSupervisedLearning/BarlowTwins.cs Uses engine vector ops for SGD parameter updates.
src/SelfSupervisedLearning/BYOL.cs Uses engine ops for SGD + EMA parameter updates.
src/ReinforcementLearning/Agents/MuZeroAgent.cs Uses engine vector ops for gradient application.
src/ReinforcementLearning/Agents/DeepReinforcementLearningAgentBase.cs Adds protected Engine accessor for subclasses.
src/ReinforcementLearning/Agents/DQNAgent.cs Uses engine vector ops for gradient application.
src/Regression/RegressionBase.cs Uses engine vector ops for gradient application.
src/Regression/NonLinearRegressionBase.cs Uses engine vector ops for gradient application.
src/NeuralRadianceFields/Models/NeRF.cs Uses engine vector ops for gradient descent update.
src/NeuralNetworks/Tabular/TabTransformerBase.cs Adds engine accessor; uses engine tensor ops for embedding updates.
src/NeuralNetworks/Tabular/TabPFNBase.cs Uses engine tensor ops for residuals/matmul/fills.
src/NeuralNetworks/Tabular/TabNetBase.cs Adds engine accessor; replaces ones/zeros/mask loops with engine ops.
src/NeuralNetworks/Tabular/TabMBase.cs Adds engine accessor; refactors embedding updates.
src/NeuralNetworks/Tabular/TabDPTBase.cs Uses engine tensor ops for residuals/matmul/fills.
src/NeuralNetworks/Tabular/SoftTree.cs Converts to LayerBase<T> and uses engine ops for fills/updates; adds parameter export plumbing.
src/NeuralNetworks/Tabular/SAINTBase.cs Adds engine accessor; replaces add helper with TensorAdd.
src/NeuralNetworks/Tabular/PiecewiseLinearEncoding.cs Converts to LayerBase<T> and uses engine ops for fills/updates; adds parameter export plumbing.
src/NeuralNetworks/Tabular/ObliviousDecisionTree.cs Converts to LayerBase<T> and uses engine ops for fills/updates; adds parameter export plumbing.
src/NeuralNetworks/Tabular/NODEBase.cs Adds engine accessor; replaces ensemble accumulation/scaling loops with engine ops.
src/NeuralNetworks/Tabular/IntersampleAttention.cs Converts to LayerBase<T>; replaces transpose/attention math with engine ops.
src/NeuralNetworks/Tabular/GatedFeatureLearningUnit.cs Converts to LayerBase<T>; replaces sigmoid/multiply/add loops with engine ops.
src/NeuralNetworks/Tabular/GANDALFBase.cs Adds engine accessor; replaces sigmoid/multiply/add/update loops with engine ops.
src/NeuralNetworks/Tabular/FeatureTransformer.cs Converts to LayerBase<T>; refactors GLU/residual scaling to engine ops.
src/NeuralNetworks/Tabular/FeatureTokenizer.cs Uses engine ops for parameter updates.
src/NeuralNetworks/Tabular/ContrastivePretraining.cs Adds engine accessor; uses engine ops for parameter updates and fill helpers.
src/NeuralNetworks/Tabular/ColumnEmbedding.cs Adds engine accessor; uses engine ops for updates + fill.
src/NeuralNetworks/Tabular/CLSToken.cs Adds engine accessor; uses engine ops for updates + fill.
src/NeuralNetworks/Tabular/BatchEnsembleLayer.cs Converts to LayerBase<T>; replaces some scalar ops with shared NumOps and adds parameter plumbing.
src/NeuralNetworks/Tabular/AutoIntBase.cs Adds engine accessor; replaces gradient reset loops and some updates with engine ops.
src/NeuralNetworks/Tabular/AttentiveTransformer.cs Converts to LayerBase<T>; replaces element-wise ops with engine ops.
src/NeuralNetworks/Layers/UNetDiscriminator.cs Replaces AddTensors loop with Engine.TensorAdd.
src/NeuralNetworks/Layers/TimeEmbeddingLayer.cs Refactors FC layers and gradient computations to use matmul/broadcast add/reduce-sum engine ops.
src/NeuralNetworks/Layers/SwinTransformerBlockLayer.cs Replaces AddTensors loop with Engine.TensorAdd.
src/NeuralNetworks/Layers/SpyNetLayer.cs Replaces add helper with TensorAdd and changes gradient accumulation to use engine add + copy-back.
src/NeuralNetworks/Layers/SSM/* Replaces repeated “create ones” loops with Fill(NumOps.One) and tightens null-guard checks in parameter updates.
src/NeuralNetworks/Layers/ResidualDenseBlock.cs Refactors residual add/scale helpers to engine ops.
src/NeuralNetworks/Layers/RRDBNetGenerator.cs Replaces AddTensors loop with Engine.TensorAdd.
src/NeuralNetworks/Layers/RRDBLayer.cs Refactors residual add/scale helpers to engine ops.
src/NeuralNetworks/Layers/HighwayLayer.cs Replaces nested summation loops with Engine.ReduceSum across all axes.
src/NeuralNetworks/Layers/GroupedQueryAttentionLayer.cs Tightens null-guard checks in parameter updates; uses engine ops consistently.
src/NeuralNetworks/Layers/DeformableConvolutionalLayer.cs Uses engine ops for sigmoid derivative and gradient summation; refactors parameter updates to engine ops.
src/NeuralNetworks/Layers/CrossAttentionLayer.cs Refactors weight update to use engine ops then copies into target tensor.
src/NeuralNetworks/Layers/ConvolutionalLayer.cs Refactors CPU parameter update to engine tensor ops.
src/NeuralNetworks/Gpt4VisionNeuralNetwork.cs Uses engine vector ops for gradient descent update.
src/NeuralNetworks/Blip2NeuralNetwork.cs Uses engine vector ops for gradient descent update.
src/NeuralNetworks/AudioVisualEventLocalizationNetwork.cs Uses engine vector ops for gradient descent update.
src/NeuralNetworks/AudioVisualCorrespondenceNetwork.cs Uses engine vector ops for gradient descent update.
src/Finance/Graph/DCRNN.cs Replaces scalar combine/accumulate/scale loops with engine ops.
src/Finance/Forecasting/Transformers/TimesNet.cs Replaces residual add loop with Engine.TensorAdd.
src/Finance/Forecasting/Transformers/TSMixer.cs Uses engine ops for RevIN normalize/denormalize steps.
src/Finance/Forecasting/Transformers/TFT.cs Refactors gated connection to engine ops.
src/Finance/Forecasting/Transformers/Crossformer.cs Replaces residual add loop with Engine.TensorAdd.
src/Finance/Forecasting/Transformers/Autoformer.cs Replaces add/subtract helpers with engine ops.
src/Finance/Forecasting/Neural/WaveNet.cs Replaces add/multiply helpers with engine ops.
src/Finance/Forecasting/Neural/TCN.cs Replaces add helper with engine ops.
src/Finance/Forecasting/Neural/NHiTSFinance.cs Replaces add/subtract helpers with engine ops.
src/Finance/Forecasting/Neural/NBEATSFinance.cs Replaces add/subtract helpers with engine ops.
src/Finance/Forecasting/Neural/LSTNet.cs Uses engine ops for highway residual add.
src/Finance/Forecasting/Neural/DeepState.cs Replaces mismatched-shape handling with TensorBroadcastAdd.
src/Finance/Forecasting/Neural/DeepFactor.cs Replaces manual “maxLen” broadcasting behavior with TensorBroadcastAdd.
src/Finance/Forecasting/Foundation/TimesFM.cs Replaces add helper with engine ops.
src/ContinualLearning/Strategies/ContinualLearningStrategyBase.cs Adds protected static Engine accessor for subclasses.
src/Audio/TextToSpeech/VITSModel.cs Uses engine vector ops for gradient descent update.
src/Audio/TextToSpeech/Tacotron2Model.cs Uses engine vector ops for gradient descent update.
src/Audio/SpeechRecognition/Wav2Vec2Model.cs Uses engine vector ops for gradient descent update.
src/AiModelBuilder.cs Refactors distillation SGD update to engine vector ops.
src/AiDotNet.csproj Bumps AiDotNet.Tensors and native backend packages to 0.8.0.

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment thread src/Video/Segmentation/XMem.cs
Comment thread src/Finance/Forecasting/Neural/WaveNet.cs
Comment thread src/Finance/Forecasting/Neural/WaveNet.cs
Comment thread src/Finance/Forecasting/Neural/TCN.cs
Comment thread src/Finance/Forecasting/Transformers/Crossformer.cs
Comment thread src/Finance/Forecasting/Neural/NHiTSFinance.cs
Comment thread src/Finance/Forecasting/Neural/NHiTSFinance.cs
Comment thread src/Finance/Forecasting/Transformers/TimesNet.cs
Comment thread src/Finance/Forecasting/Foundation/TimesFM.cs
ooples and others added 2 commits February 28, 2026 19:35
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Replace hand-rolled scalar LayerNorm implementations with hardware-accelerated
Engine.LayerNorm calls. Converts T[] gamma/beta arrays to Tensor<T> where needed.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
ooples and others added 3 commits February 28, 2026 20:39
…ith iengine ops

Convert 11 AddTensors, 1 Concatenate, and 1 Softmax scalar loops to
hardware-accelerated IEngine tensor operations found during thorough audit.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Add protected IEngine Engine => AiDotNetEngine.Current to:
DiffusionModelBase, NoisePredictorBase, ObjectDetectorBase, OCRBase,
BackboneBase, VAEModelBase, FTTransformerBase, TabRBase, MambularBase,
TabDPTBase, TabPFNBase. This cascades to all derived classes via
inheritance (LatentDiffusionModelBase, VideoDiffusionModelBase, etc).

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
…lasses

Update concrete classes to use inherited Engine property instead of
direct AiDotNetEngine.Current access. Static methods and classes
without base class inheritance retain AiDotNetEngine.Current.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Copilot AI review requested due to automatic review settings March 1, 2026 02:02

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot wasn't able to review this pull request because it exceeds the maximum number of files (300). Try reducing the number of changed files and requesting a review from Copilot again.

ooples and others added 5 commits February 28, 2026 21:13
…ed using

- Fix Engine => Engine infinite recursion in LinearFeatureMapper and CLIPScore
- Fix cosine similarity clamping from [0,1] to [-1,1] with Math.Clamp
- Remove unused using directive in VectorHelper

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Restore tolerance for mismatched tensor shapes that was lost when
switching from scalar loops to Engine ops. Falls back to scalar
overlap-add when shapes differ, uses fast Engine path when equal.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
- Replace hardcoded learning rates with configurable options properties
- Restore adaptive sigmoid gating in TFT instead of fixed 0.5 averaging
- Add BN parameter updates in AttentiveTransformer
- Fix SetParameters order in NonLinearRegressionBase to match GetParameters
- Add gradient length validation in MuZeroAgent, DiffusionAutoML, MAE
- Remove dead ApplySigmoid from RAFT and orphaned docs from StubEmbeddingModel
- Remove redundant using directives and deduplicate ComputeAlignment

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
- Narrow Engine visibility to private in 5 concrete classes
- Add NumTrees validation in NODEBase to prevent divide-by-zero
- Add constructor validation in ObliviousDecisionTree
- Add empty-vector guard in PTMAPAlgorithm
- Use configured loss derivative in SAM2 instead of hardcoded subtraction
- Add tensor rank validation in VideoNeuralNetworkBase
- Remove unnecessary Vector<T> allocations in BYOL
- Snapshot engine in AiModelBuilder to avoid mid-operation race
- Eliminate intermediate tensor allocations in SpyNetLayer and
  CrossAttentionLayer hot paths
- Inline euclidean distance in LSCPDetector to avoid Vector allocations

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Implement proper ExportComputationGraph for 7 tabular layer classes
(AttentiveTransformer, BatchEnsembleLayer, FeatureTransformer,
InteractingLayer, IntersampleAttention, ObliviousDecisionTree,
PiecewiseLinearEncoding) and proper SoftTree backward pass with
split decision backpropagation through sigmoid derivatives.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Math.Clamp was introduced in .NET Core 2.0 and is not available in
.NET Framework 4.7.1. Use Math.Max/Math.Min combination instead.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Copilot AI review requested due to automatic review settings March 1, 2026 03:06

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot wasn't able to review this pull request because it exceeds the maximum number of files (300). Try reducing the number of changed files and requesting a review from Copilot again.

This branch was successfully deployed

1 active deployment
Preview — 6186e94c Deployed Mar 1, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Replace scalar for-loops with existing IEngine hardware-accelerated tensor operations across codebase

2 participants