Skip to content

fix: fix GEMM correctness defaults and expand CPU SIMD ops - #709

Merged
ooples merged 10 commits into
masterfrom
codex/matmul-diagnostics
Jan 13, 2026
Merged

ooples merged 10 commits into
masterfrom
codex/matmul-diagnostics

Conversation

@ooples

@ooples ooples commented Jan 12, 2026

Copy link
Copy Markdown
Owner

Summary

  • default OpenCL GEMM to built-in kernels and make dynamic/CLBlast opt-in
  • expand CPU SIMD usage for matrix/tensor elementwise ops and reductions
  • add CPU/GPU matmul diagnostics benchmarks and expand GPU correctness tests

Testing

  • not run (not requested)

Copilot AI review requested due to automatic review settings January 12, 2026 02:21
@coderabbitai

coderabbitai Bot commented Jan 12, 2026 •

Copy link
Copy Markdown
Contributor

Warning

Rate limit exceeded

@ooples has exceeded the limit for the number of commits that can be reviewed per hour. Please wait 5 minutes and 38 seconds before requesting another review.

⌛ How to resolve this issue?

After the wait time has elapsed, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

We recommend that you space out your commits to avoid hitting the rate limit.

🚦 How do rate limits work?

CodeRabbit enforces hourly rate limits for each developer per organization.

Our paid plans have higher rate limits than the trial, open-source and free plans. In all cases, we re-allow further reviews after a brief timeout.

Please see our FAQ for further information.

📥 Commits

Reviewing files that changed from the base of the PR and between c4bdd7c and 3999099.

📒 Files selected for processing (2)
  • src/AiDotNet.Tensors/Engines/DirectGpu/OpenCL/GemmAutoTuner.cs
  • src/AiDotNet.Tensors/Engines/DirectGpu/OpenCL/OpenClBackend.cs

Walkthrough

Adds CPU float/double tiled parallel MATMUL with env-driven controls and tracing; GPU GEMM result validation with Trace diagnostics and CPU fallback; OpenCL exposes device metadata and a double-buffered GEMM fallback; span-based finiteness checks across numeric ops; new CPU/GPU matmul diagnostics and minor allocation reductions.

Changes

Cohort / File(s) Summary
CPU MATMUL Optimization
src/AiDotNet.Tensors/Engines/CpuEngine.cs
Env-driven tiling/trace/single-thread flags; NET6+ specialized float/double tiled matmul paths, tile-size/parallelization helpers, block multiply helpers, span-based refactors. Public signatures unchanged.
GPU GEMM Validation & Fallback
src/AiDotNet.Tensors/Engines/DirectGpu/DirectGpuEngine.cs
Added GemmValidateEnabled, GPU-result finiteness checks (IsAnyNonFinite), Trace diagnostics and null-return fallback to trigger CPU path on non-finite GPU results.
OpenCL Backend & GEMM Routing
src/AiDotNet.Tensors/Engines/DirectGpu/OpenCL/OpenClBackend.cs
Added device metadata properties, GemmDoubleBuffered public fallback, NormalizeRowMajorConfig, IsEffectivelyZero checks, dynamic GEMM routing and Trace-based diagnostics.
Tracing migration across GPU stack
src/AiDotNet.Tensors/Engines/DirectGpu/*, src/AiDotNet.Tensors/Engines/DirectGpu/OpenCL/*, src/AiDotNet.Tensors/Engines/DirectGpu/HIP/*, src/AiDotNet.Tensors/Engines/DirectGpu/Profiling/*, src/AiDotNet.Tensors/Engines/DirectGpu/GemmBenchmark.cs
Replaced Console.WriteLine with Trace.WriteLine; introduced OpenClNativeBindings.EnableDiagnostics and centralized LogDiagnostic helper.
Matrix backing / allocation reductions
src/AiDotNet.Tensors/Engines/DirectGpuTensorEngine.cs, src/AiDotNet.Tensors/LinearAlgebra/Matrix.cs, src/AiDotNet.Tensors/LinearAlgebra/MatrixBase.cs
New internal constructors to wrap existing arrays and direct Matrix construction to avoid an intermediate allocation/copy.
Finiteness APIs / Numeric operations
src/AiDotNet.Tensors/NumericOperations/*, src/AiDotNet.Tensors/Helpers/TensorPrimitivesHelper.cs, src/AiDotNet.Tensors/Interfaces/IVectorizedOperations.cs
Added AllFinite / IsAnyNonFinite span-based methods across numeric ops, helpers, and interface. Multiple duplicate declarations observed and likely require de-duplication.
Benchmarks & Diagnostics
tests/AiDotNet.Tensors.Benchmarks/CpuMatMulDiagnostics.cs, tests/AiDotNet.Tensors.Benchmarks/GpuMatMulDiagnostics.cs, tests/AiDotNet.Tensors.Benchmarks/Program.cs, tests/AiDotNet.Tensors.Benchmarks/Helpers/BenchmarkHelper.cs
New CPU/GPU matmul diagnostic runners, CLI flags --cpu-matmul/--gpu-matmul, random-matrix helper and compare utility, per-iteration timing and GFLOPS reporting.
Tests — GPU kernel coverage
tests/AiDotNet.Tests/DirectGpuTests.cs
Added gemm_double_buffered kernel variant to correctness tests to cover the new double-buffered GEMM path.
Engine-level info & logging
src/AiDotNet.Tensors/Engines/AiDotNetEngine.cs, src/AiDotNet.Tensors/Engines/Engine.cs
Added GetEngineInfo() and migrated engine selection/fallback logging from Console to Trace.

Sequence Diagram(s)

sequenceDiagram
    participant Caller
    participant DirectGpuEngine
    participant GPU_Backend
    participant Validator
    participant CpuEngine

    Caller->>DirectGpuEngine: MatrixMultiply(A, B)
    DirectGpuEngine->>GPU_Backend: Execute GEMM (CLBlast/dynamic/double-buffered)
    GPU_Backend-->>DirectGpuEngine: ResultArray

    alt AIDOTNET_GEMM_VALIDATE enabled
        DirectGpuEngine->>Validator: IsAnyNonFinite(ResultArray)?
        Validator-->>DirectGpuEngine: ok / badIndex

        alt invalid
            DirectGpuEngine->>CpuEngine: MatrixMultiply(A, B) (CPU fallback)
            CpuEngine-->>DirectGpuEngine: CPU_Result
            DirectGpuEngine-->>Caller: CPU_Result
        else valid
            DirectGpuEngine-->>Caller: ResultArray
        end
    else
        DirectGpuEngine-->>Caller: ResultArray
    end
Loading

Estimated code review effort

🎯 4 (Complex) | ⏱️ ~45 minutes

Possibly related PRs

Poem

🐰 I hopped through tiles and traces bright,

Floats and doubles found their pace,
GPUs checked sums in the night,
When numbers failed, the CPU took place,
Benchmarks thumped — matrices in grace!

🚥 Pre-merge checks | ✅ 2 | ❌ 1
❌ Failed checks (1 warning)
Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 69.70% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (2 passed)
Check name Status Explanation
Title check ✅ Passed The title clearly summarizes the main changes: fixing GEMM correctness defaults and expanding CPU SIMD operations, which aligns with the file modifications and objectives.
Description check ✅ Passed The description is related to the changeset, covering the main objectives: GEMM defaults, CPU SIMD expansion, and new diagnostic benchmarks.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Post copyable unit tests in a comment
  • Commit unit tests in branch codex/matmul-diagnostics

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@github-actions github-actions Bot changed the title Fix GEMM correctness defaults and expand CPU SIMD ops fix: fix GEMM correctness defaults and expand CPU SIMD ops Jan 12, 2026
@github-actions

Copy link
Copy Markdown
Contributor

🤖 PR Title Auto-Fixed

Your PR title was automatically updated to follow Conventional Commits format.

Original title:
Fix GEMM correctness defaults and expand CPU SIMD ops

New title:
fix: fix GEMM correctness defaults and expand CPU SIMD ops

Detected type: fix: (title starts with fix/correct/resolve/patch)
Version impact: MINOR version bump (0.1.0 → 0.2.0)


Valid types and their effects:

  • feat: - New feature (MINOR bump: 0.1.0 → 0.2.0)
  • fix: - Bug fix (MINOR bump)
  • docs: - Documentation (MINOR bump)
  • refactor: - Code refactoring (MINOR bump)
  • perf: - Performance improvement (MINOR bump)
  • test: - Tests only (no release)
  • chore: - Build/tooling (no release)
  • ci: - CI/CD changes (no release)
  • style: - Code formatting (no release)
  • deps: - Dependency update (no release)

If the detected type is incorrect, you can manually edit the PR title.

@coderabbitai coderabbitai Bot added the feature Feature work item label Jan 12, 2026

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (1)
src/AiDotNet.Tensors/Engines/DirectGpu/OpenCL/OpenClBackend.cs (1)

4436-4464: Diagnostic help text: ensure listed env vars are actually implemented (esp. AIDOTNET_GEMM_VALIDATE).
You added help entries for AIDOTNET_GEMM_ENABLE_DYNAMIC, AIDOTNET_GEMM_SAFE, AIDOTNET_GEMM_UNSAFE, AIDOTNET_GEMM_VALIDATE; please confirm each one is consumed in code (in this file or elsewhere) so users don't chase no-op toggles.

🤖 Fix all issues with AI agents
In @src/AiDotNet.Tensors/Engines/CpuEngine.cs:
- Around line 1560-1567: The multiplication in ShouldParallelizeMatMul can
overflow the 64-bit long; replace the naive long product with a non-overflowing
approach by computing the op count as a double (e.g., double ops = (double)m * n
* k) or by performing saturating arithmetic, then compare that double against
CpuMatMulParallelThresholdOps (cast the threshold to double) to decide
parallelization; update the ShouldParallelizeMatMul method (use the function
name to locate it) to use the double-based check (or clamp to long.MaxValue if
you prefer saturating) so large dims don't silently flip the sign and
incorrectly disable parallelization.
🧹 Nitpick comments (7)
src/AiDotNet.Tensors/Engines/DirectGpu/DirectGpuEngine.cs (1)

415-429: Consider SIMD vectorization for AllFinite to match PR's SIMD expansion theme.

Since this PR expands CPU SIMD usage, this helper could use Vector<float> for faster validation on large GEMM results. However, given this is opt-in diagnostics code, the current scalar approach is acceptable.

♻️ Optional SIMD-accelerated implementation
 private static bool AllFinite(float[] data, out int badIndex)
 {
+    int i = 0;
+    if (System.Numerics.Vector.IsHardwareAccelerated && data.Length >= System.Numerics.Vector<float>.Count)
+    {
+        int vectorEnd = data.Length - (data.Length % System.Numerics.Vector<float>.Count);
+        for (; i < vectorEnd; i += System.Numerics.Vector<float>.Count)
+        {
+            var vec = new System.Numerics.Vector<float>(data, i);
+            // vec - vec yields NaN for Inf values and NaN for NaN values
+            var diff = vec - vec;
+            if (!System.Numerics.Vector.EqualsAll(diff, System.Numerics.Vector<float>.Zero))
+            {
+                // Fall back to scalar to find exact index
+                for (int j = i; j < i + System.Numerics.Vector<float>.Count; j++)
+                {
+                    if (float.IsNaN(data[j]) || float.IsInfinity(data[j]))
+                    {
+                        badIndex = j;
+                        return false;
+                    }
+                }
+            }
+        }
+    }
+
-    for (int i = 0; i < data.Length; i++)
+    for (; i < data.Length; i++)
     {
         float value = data[i];
         if (float.IsNaN(value) || float.IsInfinity(value))
         {
             badIndex = i;
             return false;
         }
     }

     badIndex = -1;
     return true;
 }
src/AiDotNet.Tensors/Engines/CpuEngine.cs (2)

37-43: Env-var toggles: consider more forgiving parsing + override for perf knobs.

Right now only "1" enables tracing/single-thread. If you want this to be friendlier in CI, consider accepting "true"/"True" as well, and (optionally) letting CpuMatMulParallelThresholdOps/tile sizes be overridden via env var for tuning without recompiling.


1569-1606: Minor cleanup: unused params + redundant inner-loop args.

MultiplyFloatBlock/MultiplyDoubleBlock take m but don’t use it, and the 0, jEnd-j0 args passed to SimdVector.MatMulInnerLoop* look redundant since you already slice b/c to [j0..jEnd). If the inner-loop API allows it, consider simplifying to reduce call-site confusion.

Also applies to: 1608-1645

src/AiDotNet.Tensors/Engines/DirectGpu/OpenCL/OpenClBackend.cs (2)

1624-1685: Packed GEMM: useColumnMajorC hardcoded to false—either wire it up or remove dead branches.
Right now useColumnMajorC is always false, but the method still contains substantial branching to handle column-major C (padding/copy-back). If column-major C is intentionally unsupported in this path, consider deleting the unused branch (or add a comment explaining why it must be false).


1747-1774: Dynamic enable flag parsing: consider using GetEnvBool for consistency (and less surprising UX).
AIDOTNET_GEMM_ENABLE_DYNAMIC is checked via == "1" while other env flags in this file support true/false/yes/no/on/off. Switching to GetEnvBool would make toggles consistent and easier to use.

tests/AiDotNet.Tensors.Benchmarks/GpuMatMulDiagnostics.cs (1)

93-128: Consider disposing result matrices if they implement IDisposable.

The Compare method creates cpuResult and gpuResult matrices. If Matrix<float> implements IDisposable, these should be disposed after use to avoid resource leaks, especially when running multiple correctness checks in a loop.

♻️ Suggested fix if matrices are disposable
 private static (double maxError, double avgError, int nonFiniteCount) Compare(
     CpuEngine cpuEngine,
     IEngine gpuEngine,
     Matrix<float> a,
     Matrix<float> b)
 {
-    var cpuResult = cpuEngine.MatrixMultiply(a, b);
-    var gpuResult = gpuEngine.MatrixMultiply(a, b);
+    using var cpuResult = cpuEngine.MatrixMultiply(a, b);
+    using var gpuResult = gpuEngine.MatrixMultiply(a, b);
 
     var cpuSpan = cpuResult.AsSpan();
     var gpuSpan = gpuResult.AsSpan();
tests/AiDotNet.Tensors.Benchmarks/CpuMatMulDiagnostics.cs (1)

81-90: Consider extracting shared helper to reduce duplication.

The CreateRandomMatrix method is nearly identical to the one in GpuMatMulDiagnostics. Consider extracting this to a shared utility class if these diagnostic files evolve further.

📜 Review details

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

📥 Commits

Reviewing files that changed from the base of the PR and between 7f7aff7 and 4cc256b.

📒 Files selected for processing (7)
  • src/AiDotNet.Tensors/Engines/CpuEngine.cs
  • src/AiDotNet.Tensors/Engines/DirectGpu/DirectGpuEngine.cs
  • src/AiDotNet.Tensors/Engines/DirectGpu/OpenCL/OpenClBackend.cs
  • tests/AiDotNet.Tensors.Benchmarks/CpuMatMulDiagnostics.cs
  • tests/AiDotNet.Tensors.Benchmarks/GpuMatMulDiagnostics.cs
  • tests/AiDotNet.Tensors.Benchmarks/Program.cs
  • tests/AiDotNet.Tests/DirectGpuTests.cs
🧰 Additional context used
🧠 Learnings (3)
📚 Learning: 2025-12-18T08:49:25.295Z
Learnt from: ooples
Repo: ooples/AiDotNet PR: 444
File: src/Interfaces/IPruningMask.cs:1-102
Timestamp: 2025-12-18T08:49:25.295Z
Learning: In the AiDotNet repository, the project-level global using includes AiDotNet.Tensors.LinearAlgebra via AiDotNet.csproj. Therefore, Vector<T>, Matrix<T>, and Tensor<T> are available without per-file using directives. Do not flag missing using directives for these types in any C# files within this project. Apply this guideline broadly to all C# files (not just a single file) to avoid false positives. If a file uses a type from a different namespace not covered by the global using, flag as usual.

Applied to files:

  • tests/AiDotNet.Tensors.Benchmarks/GpuMatMulDiagnostics.cs
  • tests/AiDotNet.Tests/DirectGpuTests.cs
  • src/AiDotNet.Tensors/Engines/DirectGpu/DirectGpuEngine.cs
  • tests/AiDotNet.Tensors.Benchmarks/CpuMatMulDiagnostics.cs
  • tests/AiDotNet.Tensors.Benchmarks/Program.cs
  • src/AiDotNet.Tensors/Engines/DirectGpu/OpenCL/OpenClBackend.cs
  • src/AiDotNet.Tensors/Engines/CpuEngine.cs
📚 Learning: 2025-12-18T08:49:53.103Z
Learnt from: ooples
Repo: ooples/AiDotNet PR: 444
File: src/Interfaces/IPruningStrategy.cs:1-4
Timestamp: 2025-12-18T08:49:53.103Z
Learning: In this repository, global using directives are declared in AiDotNet.csproj for core namespaces (AiDotNet.Tensors.* and AiDotNet.*) and common system types. When reviewing C# files, assume these global usings are in effect; avoid adding duplicate using statements for these namespaces and for types like Vector<T>, Matrix<T>, Tensor<T>, etc. If a type is not found, verify the global usings or consider adding a file-scoped using if needed. Prefer relying on global usings to reduce boilerplate.

Applied to files:

  • tests/AiDotNet.Tensors.Benchmarks/GpuMatMulDiagnostics.cs
  • tests/AiDotNet.Tests/DirectGpuTests.cs
  • src/AiDotNet.Tensors/Engines/DirectGpu/DirectGpuEngine.cs
  • tests/AiDotNet.Tensors.Benchmarks/CpuMatMulDiagnostics.cs
  • tests/AiDotNet.Tensors.Benchmarks/Program.cs
  • src/AiDotNet.Tensors/Engines/DirectGpu/OpenCL/OpenClBackend.cs
  • src/AiDotNet.Tensors/Engines/CpuEngine.cs
📚 Learning: 2025-11-19T04:08:26.895Z
Learnt from: ooples
Repo: ooples/AiDotNet PR: 0
File: :0-0
Timestamp: 2025-11-19T04:08:26.895Z
Learning: For ILGPU GPU operations in GpuEngine.cs, use standard .NET exception types (InvalidOperationException, ArgumentException, OutOfMemoryException) instead of ILGPU-specific exception types, as ILGPU exception types may be version-specific. Combine with message-based filtering using ex.Message.Contains("device") or ex.Message.Contains("accelerator") as a fallback for GPU-specific errors.

Applied to files:

  • src/AiDotNet.Tensors/Engines/DirectGpu/DirectGpuEngine.cs
⏰ Context from checks skipped due to timeout of 90000ms. You can increase the timeout in your CodeRabbit configuration to a maximum of 15 minutes (900000ms). (4)
  • GitHub Check: Agent
  • GitHub Check: CodeQL analysis (csharp)
  • GitHub Check: CodeQL Analysis
  • GitHub Check: Build (Windows)
🔇 Additional comments (24)
src/AiDotNet.Tensors/Engines/DirectGpu/DirectGpuEngine.cs (3)

38-39: Configuration approach for GEMM validation is sound.

Using an environment variable for opt-in validation with CPU fallback is appropriate for diagnostics/debugging without impacting production performance. The static initialization ensures no repeated environment lookups.


405-409: LGTM!

The validation logic correctly short-circuits when disabled, provides useful diagnostic info (first bad index), and gracefully falls back to CPU by returning null.


444-448: LGTM!

Consistent validation pattern applied to the cached weights path. The duplication with lines 405-409 is acceptable given the small footprint and independent entry points.

src/AiDotNet.Tensors/Engines/CpuEngine.cs (9)

1422-1432: Type-specialized matmul dispatch looks good (float/double).

The NET6+ fast-path routing is clean, and the generic fallback remains intact.


1454-1547: Matmul tiling + row-block parallelization appears race-free.

Partitioning by row blocks (iStart..iEnd) and writing disjoint c slices is the right shape for correctness. The CpuMatMulSingleThread guard is also handy for diagnostics.


1721-1759: Span-based matrix ops are a nice win (less loop noise, more SIMD).

MatrixAdd, MatrixMultiplyScalar, MatrixSubtract, and MatrixSumOfSquares moving to numOps.* is a clear readability + perf improvement.


1792-1814: OuterProduct + row get/set span copies look good.

The row-parallel outer product writes disjoint rows, and numOps.Copy(...) for GetRow/SetRow is cleaner than element loops.

Also applies to: 1842-1846, 1880-1882


2204-2214: Division-by-zero pre-scan: confirm intended behavior for float/double.

This now throws on any exact-zero divisor for all T. That’s consistent with your vector divide behavior earlier in the file, but it is a semantic choice (IEEE floats would normally return ±Inf/NaN). Worth double-checking this matches library expectations.


3001-3002: TensorSum/TensorMinValue/TensorMaxValue changes look solid.

Moving sum to numOps.Sum(span) and using chunked parallel min/max with per-chunk reductions reads well and should scale better for large tensors.

Also applies to: 3068-3105, 3114-3152


12228-12231: Dot-based sum-of-squares + scalar span ops: LGTM.

numOps.Dot(x, x) for sum-of-squares and AddScalar/SubtractScalar/DivideScalar via spans are consistent with the rest of the SIMD/Span direction.

Also applies to: 13305-13330


1664-1683: Verify thread-safety of numOps.Dot in Parallel.For.

This code shares one numOps instance across parallel threads. Confirm whether INumericOperations<T> and its Dot implementation are documented as thread-safe and stateless. If the implementation maintains any mutable state (caches, scratch buffers, or thread-local fields), this risks race conditions. If thread-safety is not guaranteed, use ThreadLocal<INumericOperations<T>> or fall back to sequential execution.


2030-2031: Potential correctness risk: in-place span operations in TensorAddMany / TensorMultiplyMany may corrupt results if INumericOperations<T> doesn't support destination aliasing.

The methods use in-place accumulation like numOps.Add(resultSpan, tensors[t].AsSpan(), resultSpan) where the destination aliases an input. This is only safe if the INumericOperations<T> implementation explicitly supports span aliasing. SIMD-based implementations using Vector<T> or intrinsics can produce corrupted results with overlapping source/destination spans without overlap detection or direction-aware loops.

Document the aliasing contract for INumericOperations<T>.Add and Multiply, or route in-place operations through a temporary buffer to guarantee correctness.

Applies to: 2030–2031, 2094–2103, 2164–2172.

src/AiDotNet.Tensors/Engines/DirectGpu/OpenCL/OpenClBackend.cs (4)

1857-1864: New public API GemmDoubleBuffered: verify kernel name exists in the compiled kernel set.
Since this is now callable externally/tests may rely on it, a missing kernel name would become a runtime KeyNotFoundException.


1776-1815: Verify safe/unsafe kernel selection heuristics against design specifications and benchmarks.

The 512-dimension cutoff for routing between gemm_medium_tile (safe) and gemm_double_buffered (unsafe) kernels warrants validation. Confirm that:

  • This threshold matches your intended behavior and performance profile
  • The policy is applied consistently across all GEMM fallback paths
  • Typical transformer workload shapes (e.g., 2048×4096, 4096×8192) route to the expected kernel

872-940: Baseline config normalization: verify NormalizeRowMajorConfig semantics and database impact.
The code now normalizes CLBlast baseline configs by calling NormalizeRowMajorConfig() on both database-provided and default baselines. This is applied before any kernel selection logic. Confirm what fields NormalizeRowMajorConfig() modifies—specifically whether it only normalizes UseColumnMajorA, or whether other layout-related fields are also affected. Additionally, verify whether existing CLBlast database baselines rely on UseColumnMajorA=true, as silent normalization could unintentionally change behavior for affected devices.


943-972: NormalizeRowMajorConfig: verify handling of GemmConfig field evolution.

This method reconstructs a GemmConfig with manual field copying. If GemmConfig is later extended with new fields, those will silently default in normalized configs rather than inheriting from the input, potentially causing performance or correctness issues.

Action needed: Confirm whether GemmConfig is defined as a record struct (enabling with-expression support). If so, consider using config with { UseColumnMajorA = false } to safely handle future field additions. If GemmConfig is a mutable struct, direct mutation may be preferable to reconstruction.

tests/AiDotNet.Tests/DirectGpuTests.cs (1)

650-656: LGTM!

Good addition of the gemm_double_buffered kernel to the correctness validation suite. This ensures the new double-buffered GEMM path is validated against the reference implementation alongside the other kernel variants.

tests/AiDotNet.Tensors.Benchmarks/Program.cs (2)

32-42: LGTM!

The new CLI options follow the established pattern and are placed appropriately before the NET462 conditional block. Good integration of the new diagnostic utilities.


78-79: LGTM!

Usage text is clear and consistent with the existing help output format.

tests/AiDotNet.Tensors.Benchmarks/GpuMatMulDiagnostics.cs (2)

12-80: LGTM on overall structure.

The diagnostic utility is well-structured with clear separation between correctness validation (safe/unsafe modes) and performance benchmarking. The warmup pass at line 60 before timed iterations is good practice.


22-24: Verify CpuEngine disposal requirements.

If CpuEngine implements IDisposable, the instance created at line 23 should be wrapped in a using statement to ensure proper resource cleanup. Additionally, review the Compare method to verify that matrices created during comparisons are properly disposed.

tests/AiDotNet.Tensors.Benchmarks/CpuMatMulDiagnostics.cs (3)

20-20: Same disposal concern as GPU diagnostics.

If CpuEngine implements IDisposable, wrap it in a using statement.


61-79: LGTM on environment variable parsing.

Good flexibility allowing custom sizes via AIDOTNET_CPU_MATMUL_SIZES with sensible defaults. The parsing handles multiple delimiter types which is user-friendly.


38-50: LGTM on benchmark loop structure.

The checksum accumulation at line 48 is a good pattern to prevent the compiler from eliminating the multiplication as dead code. Adaptive iteration count based on matrix size is also sensible.

Comment thread src/AiDotNet.Tensors/Engines/CpuEngine.cs

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This pull request focuses on improving GEMM (General Matrix Multiply) correctness and performance by changing the default OpenCL GEMM behavior to use built-in kernels instead of dynamic/CLBlast kernels (now opt-in), expanding CPU SIMD usage for matrix and tensor operations, and adding diagnostic benchmarks for both CPU and GPU matrix multiplication.

Changes:

  • Default OpenCL GEMM to built-in kernels with dynamic kernels as opt-in via environment variable
  • Optimize CPU matrix/tensor operations using SIMD vectorization through span-based operations
  • Add CPU and GPU matmul diagnostic tools with correctness checking and performance profiling

Reviewed changes

Copilot reviewed 7 out of 7 changed files in this pull request and generated 7 comments.

Show a summary per file
File Description
tests/AiDotNet.Tests/DirectGpuTests.cs Adds gemm_double_buffered kernel to correctness tests
tests/AiDotNet.Tensors.Benchmarks/Program.cs Adds command-line options for CPU and GPU matmul diagnostics
tests/AiDotNet.Tensors.Benchmarks/GpuMatMulDiagnostics.cs New GPU diagnostics with correctness and performance tests
tests/AiDotNet.Tensors.Benchmarks/CpuMatMulDiagnostics.cs New CPU diagnostics with configurable size testing
src/AiDotNet.Tensors/Engines/DirectGpu/OpenCL/OpenClBackend.cs Changes GEMM default to built-in kernels, adds config normalization, exposes GemmDoubleBuffered
src/AiDotNet.Tensors/Engines/DirectGpu/DirectGpuEngine.cs Adds optional GEMM validation for non-finite values
src/AiDotNet.Tensors/Engines/CpuEngine.cs Replaces element-wise loops with SIMD-optimized span operations for matrices and tensors

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment thread tests/AiDotNet.Tensors.Benchmarks/Program.cs Outdated
Comment thread tests/AiDotNet.Tensors.Benchmarks/Program.cs Outdated
Comment thread src/AiDotNet.Tensors/Engines/DirectGpu/DirectGpuEngine.cs Outdated
Comment thread src/AiDotNet.Tensors/Engines/DirectGpu/OpenCL/OpenClBackend.cs Outdated
Comment thread src/AiDotNet.Tensors/Engines/DirectGpu/OpenCL/OpenClBackend.cs Outdated
Comment thread src/AiDotNet.Tensors/Engines/CpuEngine.cs
Comment thread tests/AiDotNet.Tensors.Benchmarks/GpuMatMulDiagnostics.cs
Comment thread src/AiDotNet.Tensors/Engines/CpuEngine.cs Fixed
Comment thread src/AiDotNet.Tensors/Engines/CpuEngine.cs Fixed

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🧹 Nitpick comments (8)
src/AiDotNet.Tensors/Engines/DirectGpu/DirectGpuEngine.cs (1)

415-429: Consider SIMD optimization for AllFinite check.

Given the PR's focus on expanding CPU SIMD operations, this validation loop over potentially large GEMM results could benefit from vectorization. However, since this is an opt-in diagnostic feature (GemmValidateEnabled), the simpler scalar implementation is acceptable.

♻️ Optional SIMD-accelerated implementation
 private static bool AllFinite(float[] data, out int badIndex)
 {
+    int i = 0;
+    if (System.Numerics.Vector.IsHardwareAccelerated && data.Length >= System.Numerics.Vector<float>.Count)
+    {
+        int vectorSize = System.Numerics.Vector<float>.Count;
+        int vectorEnd = data.Length - (data.Length % vectorSize);
+        for (; i < vectorEnd; i += vectorSize)
+        {
+            var vec = new System.Numerics.Vector<float>(data, i);
+            // NaN and Infinity fail the equality check with themselves or produce false comparisons
+            if (System.Numerics.Vector.EqualsAny(vec, vec) == false || 
+                System.Numerics.Vector.GreaterThanOrEqualAll(vec, new System.Numerics.Vector<float>(float.NegativeInfinity)) == false)
+            {
+                // Fall back to scalar to find exact index
+                for (int j = i; j < i + vectorSize; j++)
+                {
+                    if (float.IsNaN(data[j]) || float.IsInfinity(data[j]))
+                    {
+                        badIndex = j;
+                        return false;
+                    }
+                }
+            }
+        }
+    }
+
-    for (int i = 0; i < data.Length; i++)
+    for (; i < data.Length; i++)
     {
         float value = data[i];
         if (float.IsNaN(value) || float.IsInfinity(value))
         {
             badIndex = i;
             return false;
         }
     }

     badIndex = -1;
     return true;
 }
src/AiDotNet.Tensors/Engines/DirectGpu/OpenCL/OpenClBackend.cs (2)

1636-1638: Packed dynamic GEMM: hard-coded useColumnMajorC = false makes half the method dead; also consider early _dynamicGemm null check.
Right now useColumnMajorC is always false, so the column-major-C branches are unreachable. If that’s intentional, a short comment explaining “output is always row-major in this backend” would prevent future confusion. Also, TryExecutePackedDynamicGemm can allocate padding buffers before failing via TryExecuteDynamicGemm when _dynamicGemm == null—an early guard would avoid wasted work.

Also applies to: 1657-1661


1749-1774: Env var parsing inconsistency: prefer GetEnvBool for dynamic gating + trace.
New knobs (AIDOTNET_GEMM_ENABLE_DYNAMIC) are parsed with == "1" while you already have GetEnvBool that supports common truthy values. Using GetEnvBool here would make behavior consistent across the file (and across platforms / shells).

Also applies to: 1776-1789

src/AiDotNet.Tensors/Engines/CpuEngine.cs (5)

37-43: Avoid Console.WriteLine-style tracing behavior in library code (even when env-gated).

Env-gated tracing is useful, but writing to stdout can break consumers (bench harnesses, apps, tests). Prefer System.Diagnostics.Trace (or an injected logger) so output can be routed/filtered.

Proposed tweak
-    private static readonly bool CpuMatMulTraceEnabled =
-        Environment.GetEnvironmentVariable("AIDOTNET_CPU_MATMUL_TRACE") == "1";
+    private static readonly bool CpuMatMulTraceEnabled =
+        Environment.GetEnvironmentVariable("AIDOTNET_CPU_MATMUL_TRACE") == "1";
...
-            Console.WriteLine($"[CpuMatMul] float {m}x{k}x{n} tile={tileSize} parallel={useParallel}");
+            System.Diagnostics.Trace.WriteLine($"[CpuMatMul] float {m}x{k}x{n} tile={tileSize} parallel={useParallel}");

1422-1434: Unsafe.As return casting: correct but consider simpler/safer expression.

This is correct given the typeof(T) == typeof(float/double) guards, but it’s still “sharp”; a simple (Matrix<T>)(object)result is easier to reason about and debug.

Possible simplification
-            var result = MatrixMultiplyFloat(floatA, floatB);
-            return Unsafe.As<Matrix<float>, Matrix<T>>(ref result);
+            var result = MatrixMultiplyFloat(floatA, floatB);
+            return (Matrix<T>)(object)result;

1456-1550: Tiled/parallel matmul path looks race-free; verify Matrix<T> span layout + consider avoiding repeated span retrieval inside the parallel body.

Partitioning by i blocks ensures disjoint c writes, so races shouldn’t occur. Minor perf nit: inside Parallel.For, re-calling a.AsSpan()/b.AsSpan()/result.AsWritableSpan() every block is redundant.


1551-1573: Parallelization threshold math: good overflow handling; consider making the threshold configurable for benchmarking.

Math.BigMul + saturation avoids overflow. If diagnostics/benchmarks are a goal, letting the threshold be env-configurable (similar to tile size) can help tune without code changes.


1670-1689: MatrixVectorMultiply: parallel branch is OK; avoid per-iteration span wrapper allocation if possible.

Inside the parallel loop you create new ReadOnlySpan<T>(vectorData) each row. It’s small but measurable at scale; consider lifting immutable data outside the loop if you can do so without capturing spans across threads (e.g., use vectorData directly in a dedicated dot implementation that accepts arrays + offset).

📜 Review details

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

📥 Commits

Reviewing files that changed from the base of the PR and between 4cc256b and 9ab1d55.

📒 Files selected for processing (5)
  • src/AiDotNet.Tensors/Engines/CpuEngine.cs
  • src/AiDotNet.Tensors/Engines/DirectGpu/DirectGpuEngine.cs
  • src/AiDotNet.Tensors/Engines/DirectGpu/OpenCL/OpenClBackend.cs
  • tests/AiDotNet.Tensors.Benchmarks/GpuMatMulDiagnostics.cs
  • tests/AiDotNet.Tensors.Benchmarks/Program.cs
🧰 Additional context used
🧠 Learnings (2)
📚 Learning: 2025-12-18T08:49:25.295Z
Learnt from: ooples
Repo: ooples/AiDotNet PR: 444
File: src/Interfaces/IPruningMask.cs:1-102
Timestamp: 2025-12-18T08:49:25.295Z
Learning: In the AiDotNet repository, the project-level global using includes AiDotNet.Tensors.LinearAlgebra via AiDotNet.csproj. Therefore, Vector<T>, Matrix<T>, and Tensor<T> are available without per-file using directives. Do not flag missing using directives for these types in any C# files within this project. Apply this guideline broadly to all C# files (not just a single file) to avoid false positives. If a file uses a type from a different namespace not covered by the global using, flag as usual.

Applied to files:

  • tests/AiDotNet.Tensors.Benchmarks/Program.cs
  • src/AiDotNet.Tensors/Engines/DirectGpu/DirectGpuEngine.cs
  • tests/AiDotNet.Tensors.Benchmarks/GpuMatMulDiagnostics.cs
  • src/AiDotNet.Tensors/Engines/DirectGpu/OpenCL/OpenClBackend.cs
  • src/AiDotNet.Tensors/Engines/CpuEngine.cs
📚 Learning: 2025-12-18T08:49:53.103Z
Learnt from: ooples
Repo: ooples/AiDotNet PR: 444
File: src/Interfaces/IPruningStrategy.cs:1-4
Timestamp: 2025-12-18T08:49:53.103Z
Learning: In this repository, global using directives are declared in AiDotNet.csproj for core namespaces (AiDotNet.Tensors.* and AiDotNet.*) and common system types. When reviewing C# files, assume these global usings are in effect; avoid adding duplicate using statements for these namespaces and for types like Vector<T>, Matrix<T>, Tensor<T>, etc. If a type is not found, verify the global usings or consider adding a file-scoped using if needed. Prefer relying on global usings to reduce boilerplate.

Applied to files:

  • tests/AiDotNet.Tensors.Benchmarks/Program.cs
  • src/AiDotNet.Tensors/Engines/DirectGpu/DirectGpuEngine.cs
  • tests/AiDotNet.Tensors.Benchmarks/GpuMatMulDiagnostics.cs
  • src/AiDotNet.Tensors/Engines/DirectGpu/OpenCL/OpenClBackend.cs
  • src/AiDotNet.Tensors/Engines/CpuEngine.cs
🔇 Additional comments (25)
src/AiDotNet.Tensors/Engines/DirectGpu/DirectGpuEngine.cs (3)

38-39: LGTM!

The opt-in validation flag via environment variable is a sensible approach for diagnostic features. The static readonly initialization ensures the check is performed once at startup.


405-409: LGTM!

The validation check is appropriately placed after downloading GPU results and before type conversion. Returning null to trigger CPU fallback is consistent with the existing fallback pattern in this class.


444-448: LGTM!

The validation is consistent with the MatMul<T> method, maintaining the same fallback behavior for cached weight operations.

src/AiDotNet.Tensors/Engines/DirectGpu/OpenCL/OpenClBackend.cs (3)

901-907: Baseline/default config normalization looks right; consider removing redundancy + ensure KernelName semantics still match.
Normalizing the CLBlast database baseline and the fallback default to “row-major A” is a good correctness guard. The defaultBaseline = NormalizeRowMajorConfig(defaultBaseline); call is redundant now that UseColumnMajorA = false is set explicitly, but harmless.

Also applies to: 912-940


1857-1864: New public API GemmDoubleBuffered: verify it’s intended as part of the supported surface.
Looks fine mechanically (routes through ExecuteGemmKernel("gemm_double_buffered", ...)). Just ensure this method is meant to be public/stable vs an internal tuning hook (since it becomes part of the class’ API contract).


4459-4464: Help text advertises AIDOTNET_GEMM_VALIDATE, but this file doesn't implement it.
In this file, AIDOTNET_GEMM_VALIDATE only appears in PrintDiagnosticHelp()—there's no corresponding validation path in Gemm(...). Either wire it up (if intended here) or remove/clarify the help entry to avoid a misleading knob.

tests/AiDotNet.Tensors.Benchmarks/GpuMatMulDiagnostics.cs (6)

1-11: LGTM!

Imports and class declaration are appropriate for the diagnostics functionality.


82-91: LGTM!

Efficient random matrix generation using AsWritableSpan with values uniformly distributed in [-1, 1].


93-128: LGTM!

The comparison logic correctly handles non-finite GPU values by counting them separately and computing error statistics only on valid elements. Using double for error accumulation is appropriate.


130-135: LGTM!

Clean output formatting with appropriate precision specifiers.


137-141: Environment variable changes at runtime may not affect engine behavior.

If the GPU engine reads AIDOTNET_GEMM_SAFE/AIDOTNET_GEMM_UNSAFE only during initialization (common pattern), calling SetKernelMode after line 22 (AiDotNetEngine.Current) won't switch kernels. Both "safe" and "unsafe" correctness checks would then use the same code path.

Run the following to see when/how these env vars are consumed:

#!/bin/bash
# Find where AIDOTNET_GEMM env vars are read
rg -n -C5 'AIDOTNET_GEMM' --type=cs

64-71: Verify if MatrixMultiply requires explicit GPU synchronization for accurate timing.

Host Stopwatch timing may underreport actual GPU kernel execution time if MatrixMultiply is asynchronous. Confirm whether the GPU engine handles synchronization internally or if explicit synchronization (fence, event wait, or device sync) is needed after the multiply before stopping the timer.

tests/AiDotNet.Tensors.Benchmarks/Program.cs (2)

32-42: LGTM!

The new CLI options follow the existing dispatch pattern and integrate cleanly with the rest of the argument handling.


78-79: LGTM!

Usage text is clear and consistent with existing option descriptions.

src/AiDotNet.Tensors/Engines/CpuEngine.cs (11)

1575-1651: SIMD inner-loop block multiply: indexing looks correct; ensure SimdVector.MatMulInnerLoop* contract matches the slice semantics.

You’re slicing b/c to [j0..jEnd) and passing 0..(jEnd-j0) to the SIMD helper. This is only correct if the helper interprets the provided spans as already-offset (i.e., it must not apply additional base indexing assumptions).


1727-1766: Span-based matrix ops (Add, MultiplyScalar, Subtract, SumOfSquares) look good; confirm span lengths are always identical for all Matrix<T> implementations.

Assuming Matrix<T>.AsSpan() is a tight row-major buffer, these are a clear win.


1798-1820: OuterProduct parallelization is fine; validate GetRowSpan(i) returns disjoint writable memory.

The parallel write pattern assumes each row span is independent and points into the backing buffer at non-overlapping ranges.


1848-1852: Row get/set span copies: nice cleanup.

Using numOps.Copy should be faster and clearer than manual loops, assuming Copy is optimized.

Also applies to: 1886-1888


2036-2037: TensorAdd span path: LGTM.

Cleaner and likely faster than per-element loops.


2210-2221: TensorDivide now does explicit divide-by-zero checks: correctness win.

This is a good trade-off for predictable behavior (especially for non-float types where IEEE semantics aren’t expected).


2554-2571: Elementwise TensorPower(bases, exponents) now writes via backing arrays: good; ensure .Data matches .Length and is the authoritative storage.

This assumes Tensor<T>.Data is the canonical contiguous storage for the tensor and not a view with indirection.


3009-3113: TensorSum + TensorMaxValue refactor: good use of span reducers; parallel max reduction looks correct.

The chunking logic avoids empty spans and the final combine honors “first value wins” semantics without depending on MinValue.


3128-3160: TensorMinValue refactor mirrors max path; looks correct.

Same notes as max path.


12237-12239: TensorSumOfSquares using dot with itself: good optimization.

As long as Dot(x, x) is implemented as sum(x[i]*x[i]) (and not e.g. fused-with-sqrt), this is ideal.


13313-13338: Scalar ops switched to span intrinsics: looks good.

These should reduce overhead and centralize SIMD opportunities in INumericOperations<T>.

Comment thread src/AiDotNet.Tensors/Engines/DirectGpu/OpenCL/OpenClBackend.cs Outdated

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 2

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (4)
src/AiDotNet.Tensors/LinearAlgebra/MatrixBase.cs (1)

57-77: Critical: internal ctor allows “extra elements” to participate in math (data.Length > rows*cols).

Because many operations span over the full _data array, permitting data.Length != rows*cols can silently corrupt results. Also, this ctor skips the non-negative checks and doesn’t guard rows * cols overflow.

Proposed fix (enforce invariants + overflow-safe required length)
 internal MatrixBase(int rows, int cols, T[] data)
 {
-    _rows = rows;
-    _cols = cols;
-    _data = data ?? throw new ArgumentNullException(nameof(data));
-    if (data.Length < rows * cols)
-        throw new ArgumentException("Data array is too small for the specified dimensions.");
+    if (rows < 0) throw new ArgumentException("Rows must be non-negative", nameof(rows));
+    if (cols < 0) throw new ArgumentException("Columns must be non-negative", nameof(cols));
+
+    _data = data ?? throw new ArgumentNullException(nameof(data));
+
+    int required = checked(rows * cols);
+    if (_data.Length != required)
+        throw new ArgumentException($"Data array length must be exactly {required} for the specified dimensions.", nameof(data));
+
+    _rows = rows;
+    _cols = cols;
 }
src/AiDotNet.Tensors/Engines/DirectGpu/GemmBenchmark.cs (1)

125-154: Trace output may be invisible in console runs without configured listeners.

If users run GemmBenchmark.QuickTest() from a console without configuring Trace.Listeners, they may see no output. Consider either (a) ensuring benchmark/test hosts configure TextWriterTraceListener(Console.Out) with Trace.AutoFlush = true, or (b) keeping QuickTest output on Console.WriteLine for immediate visibility.

Also applies to: 202-237

src/AiDotNet.Tensors/Engines/CpuEngine.cs (1)

12-49: Fix duplicated XML doc tags (<summary>/<remarks>) on CpuEngine

You now have two <summary> blocks and two <remarks> blocks in the same doc comment, which can produce invalid XML docs / warnings-as-errors in doc builds.

Proposed fix (keep the newer “For Beginners” wording, remove the duplicate block)
@@
-/// <summary>
-/// CPU-based execution engine using INumericOperations for type-generic operations.
-/// </summary>
-/// <remarks>
-/// <para>
-/// CpuEngine provides the default execution backend for AiDotNet. It works with
-/// any numeric type that implements INumericOperations{T}, including decimal,
-/// BigInteger, and custom numeric types.
-/// </para>
-/// <para><b>For Beginners:</b> This is the standard, "always works" mode.
-///
-/// CpuEngine characteristics:
-/// - Works with ANY numeric type (float, double, decimal, BigInteger, custom types)
-/// - No special hardware required
-/// - Good performance for small-to-medium datasets
-/// - Single-threaded by default (can be parallelized in future versions)
-///
-/// When to use:
-/// - You need decimal or high-precision arithmetic
-/// - You don't have a GPU
-/// - Your datasets are small (< 100K parameters)
-/// - You're using custom numeric types
-/// </para>
-/// </remarks>
 /// <summary>
 /// CPU-based execution engine using INumericOperations for type-generic operations.
 /// </summary>
 /// <remarks>
@@
 /// and works on every computer without any extra setup.</para>
 /// </remarks>
 public class CpuEngine : IEngine
src/AiDotNet.Tensors/Engines/DirectGpu/OpenCL/OpenClBackend.cs (1)

685-718: SimpleConsoleLogger shouldn’t write via Trace.WriteLine (likely loses output).
This type is named “ConsoleLogger” and changes Console.ForegroundColor, but it emits via Trace.WriteLine (Line 711/716). If no TraceListener is configured, users enabling tuning diagnostics may see nothing.

Proposed diff
                 if (color.HasValue)
                 {
                     var previous = Console.ForegroundColor;
                     Console.ForegroundColor = color.Value;
-                    Trace.WriteLine(message);
+                    Console.WriteLine(message);
                     Console.ForegroundColor = previous;
                 }
                 else
                 {
-                    Trace.WriteLine(message);
+                    Console.WriteLine(message);
                 }
🧹 Nitpick comments (20)
src/AiDotNet.Tensors/NumericOperations/HalfOperations.cs (1)

249-263: Consider using Half.IsFinite for clarity.

The condition Half.IsNaN(x[i]) || Half.IsInfinity(x[i]) is correct but could be simplified to !Half.IsFinite(x[i]), which more directly expresses the intent and aligns with the method name AllFinite.

Also, the comment on line 251 mentions "Fallback" but there's no SIMD fast-path preceding it—consider removing or updating the comment.

Suggested diff
     public bool AllFinite(ReadOnlySpan<Half> x, out int badIndex)
     {
-        // Fallback or find the exact index of the non-finite value
         for (int i = 0; i < x.Length; i++)
         {
-            if (Half.IsNaN(x[i]) || Half.IsInfinity(x[i]))
+            if (!Half.IsFinite(x[i]))
             {
                 badIndex = i;
                 return false;
             }
         }

         badIndex = -1;
         return true;
     }
src/AiDotNet.Tensors/NumericOperations/FloatOperations.cs (1)

868-900: Finiteness scan is correct; consider float.IsFinite for clarity (if available).
Current logic is correct and returns the first failing index. If your TFMs allow it, !float.IsFinite(x[i]) simplifies the condition.

Proposed refactor (if `float.IsFinite` is available in your target frameworks)
 public bool AllFinite(ReadOnlySpan<float> x, out int badIndex)
 {
     // Fallback or find the exact index of the non-finite value
     for (int i = 0; i < x.Length; i++)
     {
-        if (float.IsNaN(x[i]) || float.IsInfinity(x[i]))
+        if (!float.IsFinite(x[i]))
         {
             badIndex = i;
             return false;
         }
     }

     badIndex = -1;
     return true;
 }
src/AiDotNet.Tensors/NumericOperations/DoubleOperations.cs (1)

804-836: Finiteness scan is correct; consider double.IsFinite for clarity (if available).
Current logic is correct and returns the first failing index. If your TFMs allow it, !double.IsFinite(x[i]) simplifies the condition.

Proposed refactor (if `double.IsFinite` is available in your target frameworks)
 public bool AllFinite(ReadOnlySpan<double> x, out int badIndex)
 {
     // Fallback or find the exact index of the non-finite value
     for (int i = 0; i < x.Length; i++)
     {
-        if (double.IsNaN(x[i]) || double.IsInfinity(x[i]))
+        if (!double.IsFinite(x[i]))
         {
             badIndex = i;
             return false;
         }
     }

     badIndex = -1;
     return true;
 }
src/AiDotNet.Tensors/NumericOperations/ComplexOperations.cs (1)

963-988: LGTM!

The implementation correctly checks both real and imaginary parts for NaN/Infinity values. The pattern is consistent with the interface contract.

For consistency with MultivectorOperations (which reuses its IsNaN/IsInfinity methods), consider refactoring to reuse the existing instance methods:

♻️ Optional refactor for consistency
 public bool AllFinite(ReadOnlySpan<Complex<T>> x, out int badIndex)
 {
     for (int i = 0; i < x.Length; i++)
     {
-        if (_ops.IsNaN(x[i].Real) || _ops.IsInfinity(x[i].Real) ||
-            _ops.IsNaN(x[i].Imaginary) || _ops.IsInfinity(x[i].Imaginary))
+        if (IsNaN(x[i]) || IsInfinity(x[i]))
         {
             badIndex = i;
             return false;
         }
     }

     badIndex = -1;
     return true;
 }
src/AiDotNet.Tensors/Engines/DirectGpuTensorEngine.cs (1)

797-816: Zero-copy Matrix<T> construction: ensure resultData size + ownership assumptions hold

This is a nice win, but it now implicitly relies on resultData being (1) correctly sized (Rows * Cols) and (2) not reused/mutated elsewhere (since the matrix will alias it).

Proposed defensive guard (cheap correctness check)
             var resultData = _directGpu.MatMul(a.AsSpan().ToArray(), b.AsSpan().ToArray(), a.Rows, a.Columns, b.Columns);
             if (resultData == null)
                 return base.MatrixMultiply(a, b);
 
-            return new Matrix<T>(a.Rows, b.Columns, resultData);
+            if (resultData.Length != a.Rows * b.Columns)
+                return base.MatrixMultiply(a, b);
+
+            return new Matrix<T>(a.Rows, b.Columns, resultData);

To verify the invariants quickly, I’d grep DirectGpuEngine.MatMul to confirm it always returns a fresh T[] with length M*N (and row-major order matching Matrix<T> expectations).

src/AiDotNet.Tensors/Engines/DirectGpu/Profiling/GemmProfiler.cs (1)

31-32: Update comment to reflect Trace output.

The comment states "print progress to console" but the implementation now uses Trace.WriteLine. Consider updating the documentation for accuracy.

-    /// <summary>Whether to print progress to console.</summary>
+    /// <summary>Whether to write profiling progress to trace output.</summary>
     public bool Verbose { get; init; } = true;
src/AiDotNet.Tensors/Engines/DirectGpu/HIP/HipBackend.cs (4)

5-8: using System.Diagnostics addition is correct and unblocks Trace.WriteLine.
No issues with the import itself. Consider also de-qualifying System.Diagnostics.Debug.WriteLine usages for consistency now that System.Diagnostics is in scope.


205-214: Good: “no GEMM kernels compiled” is now routed via Trace instead of stdout.
Minor improvement: include _architecture and/or DeviceName in the message to reduce “why did it fail?” follow-ups.


259-358: Consider honoring EnableDiagnostics (and/or a TraceSwitch) to avoid always-on noisy tracing.
Right now these traces emit even when HipNativeBindings.EnableDiagnostics == false, which makes the toggle less meaningful and can spam logs in normal runs.

Proposed change (gate Trace logging behind EnableDiagnostics)
 public sealed class HipBackend : IAsyncGpuBackend
 {
+    private static void TraceDiag(string message)
+    {
+        if (!EnableDiagnostics) return;
+        Trace.WriteLine(message);
+    }
+
     private IntPtr _stream;
@@
-        Trace.WriteLine($"[HipBackend] Compiling kernels for {_architecture} with flags: {compileFlags}");
+        TraceDiag($"[HipBackend] Compiling kernels for {_architecture} with flags: {compileFlags}");
@@
-        Trace.WriteLine($"[HipBackend] Kernel compilation complete. Available kernels: {_kernelCache.Count}");
+        TraceDiag($"[HipBackend] Kernel compilation complete. Available kernels: {_kernelCache.Count}");
@@
-        Trace.WriteLine($"[HipBackend] Kernel compilation EXCEPTION: {ex.GetType().Name}: {ex.Message}");
+        TraceDiag($"[HipBackend] Kernel compilation EXCEPTION: {ex.GetType().Name}: {ex.Message}");

365-430: Trace logs are helpful here; consider using severity + the same gating helper.
These are effectively error paths (hiprtcCreateProgram, compile log, hipModuleLoadData). Using Trace.TraceError/Trace.TraceWarning (and gating via EnableDiagnostics) makes downstream log filtering easier.

tests/AiDotNet.Tensors.Benchmarks/Helpers/BenchmarkHelper.cs (1)

64-78: Consider checking reference values for non-finite as well.

The comparison logic correctly detects non-finite values in actual, but if refSpan[i] happens to be NaN or Infinity, Math.Abs(refSpan[i] - actVal) would produce NaN, silently corrupting maxError and sumError.

Since CreateRandomMatrix produces finite values, this is unlikely in practice. However, for a general-purpose comparison utility, defensive handling would improve robustness.

♻️ Optional: Check both spans for non-finite values
 for (int i = 0; i < actSpan.Length; i++)
 {
     float actVal = actSpan[i];
-    if (float.IsNaN(actVal) || float.IsInfinity(actVal))
+    float refVal = refSpan[i];
+    if (!float.IsFinite(actVal) || !float.IsFinite(refVal))
     {
         nonFiniteCount++;
         continue;
     }

-    double error = Math.Abs(refSpan[i] - actVal);
+    double error = Math.Abs(refVal - actVal);
     sumError += error;
     if (error > maxError)
         maxError = error;
     count++;
 }
tests/AiDotNet.Tensors.Benchmarks/GpuMatMulDiagnostics.cs (1)

68-68: Consider exception-safe kernel mode reset.

If an exception occurs during the correctness tests (lines 47-66), the kernel mode environment variables won't be reset to defaults. For robustness, consider wrapping the correctness block in a try/finally.

♻️ Suggested improvement
+        try
+        {
             foreach (int size in correctnessSizes)
             {
                 // ... existing correctness test code ...
             }
+        }
+        finally
+        {
             SetKernelMode(forceSafe: false, forceUnsafe: false);
+        }
src/AiDotNet.Tensors/Engines/CpuEngine.cs (4)

52-59: Env-var toggles: consider accepting more truthy values (optional)

== "1" is fine, but supporting "true"/"TRUE" (and trimming) makes local debugging less brittle.


1440-1451: Avoid Unsafe.As for matmul fast-path return; use a normal cast instead

Unsafe.As<Matrix<float>, Matrix<T>>(ref result) is non-idiomatic and easy to simplify, while keeping the same behavior when T is actually float/double. Also avoids accidental misuse if this code is later refactored.

Proposed fix
@@
 #if NET6_0_OR_GREATER
         if (typeof(T) == typeof(float) && a is Matrix<float> floatA && b is Matrix<float> floatB)
         {
-            var result = MatrixMultiplyFloat(floatA, floatB);
-            return Unsafe.As<Matrix<float>, Matrix<T>>(ref result);
+            return (Matrix<T>)(object)MatrixMultiplyFloat(floatA, floatB);
         }
 
         if (typeof(T) == typeof(double) && a is Matrix<double> doubleA && b is Matrix<double> doubleB)
         {
-            var result = MatrixMultiplyDouble(doubleA, doubleB);
-            return Unsafe.As<Matrix<double>, Matrix<T>>(ref result);
+            return (Matrix<T>)(object)MatrixMultiplyDouble(doubleA, doubleB);
         }
 #endif

1688-1707: Parallel MatrixVectorMultiply: confirm numOps.Dot is thread-safe; small perf tidy-up

This parallelizes correctly by row. Please confirm INumericOperations<T> implementations are stateless/thread-safe (shared numOps instance across threads).

Optional: hoist ReadOnlySpan<T> vectorSpan = vector.Data; outside the lambda to avoid re-creating spans per-iteration.


1816-1841: Outer product parallelization looks good; consider reusing bSpan (optional)

Correct parallel partitioning (by row). Minor: var bSpan = b.AsSpan(); can be computed once and reused in both branches.

src/AiDotNet.Tensors/Engines/DirectGpu/OpenCL/OpenClBackend.cs (4)

16-46: Remove the duplicated XML doc block (and reconsider “For Beginners” in public API docs).
Right now the class has two consecutive <summary>/<remarks> blocks (Line 16–29 and Line 30–46) that repeat the same “Key Features” list, which will bloat/duplicate generated docs and is easy to desync over time.


80-110: Initialize all public device-info properties consistently on the “OpenCL not available” path.
On the early return (Line 156–162) ComputeUnits/GlobalMemoryBytes/LocalMemoryBytes are left as implicit defaults; it’s clearer/safer to set them explicitly alongside DeviceName/DeviceVendor.

Proposed diff
 if (!DirectOpenClContext.IsAvailable)
 {
     IsAvailable = false;
     DeviceName = "None";
     DeviceVendor = "None";
+    ComputeUnits = 0;
+    GlobalMemoryBytes = 0;
+    LocalMemoryBytes = 0;
     return;
 }

Also applies to: 156-189


164-517: Consider gating the heavy Trace.WriteLine usage to avoid noisy/prod overhead.
A lot of the new tracing uses interpolated strings (Line 166+ / kernel compilation block) which allocates even when tracing isn’t consumed. If you want “Trace-level” diagnostics to be opt-in, consider guarding with an env toggle (like AIDOTNET_GEMM_TRACE) or a TraceSwitch/TraceSource.


1900-1908: Don’t log every MatMul call unconditionally.
Trace.WriteLine($"[OpenClBackend.MatMul] Called: ...") (Line 1902) can become very noisy in real workloads. Suggest gating it behind AIDOTNET_GEMM_TRACE (or similar).

📜 Review details

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

📥 Commits

Reviewing files that changed from the base of the PR and between 9ab1d55 and 248bf5f.

📒 Files selected for processing (36)
  • src/AiDotNet.Tensors/Engines/AiDotNetEngine.cs
  • src/AiDotNet.Tensors/Engines/CpuEngine.cs
  • src/AiDotNet.Tensors/Engines/DirectGpu/DirectGpuEngine.cs
  • src/AiDotNet.Tensors/Engines/DirectGpu/GemmBenchmark.cs
  • src/AiDotNet.Tensors/Engines/DirectGpu/HIP/HipBackend.cs
  • src/AiDotNet.Tensors/Engines/DirectGpu/HIP/HipNativeBindings.cs
  • src/AiDotNet.Tensors/Engines/DirectGpu/OpenCL/DynamicGemmKernel.cs
  • src/AiDotNet.Tensors/Engines/DirectGpu/OpenCL/GemmAutoTuner.cs
  • src/AiDotNet.Tensors/Engines/DirectGpu/OpenCL/OpenClBackend.cs
  • src/AiDotNet.Tensors/Engines/DirectGpu/OpenCL/OpenClNativeBindings.cs
  • src/AiDotNet.Tensors/Engines/DirectGpu/Profiling/GemmProfiler.cs
  • src/AiDotNet.Tensors/Engines/DirectGpuTensorEngine.cs
  • src/AiDotNet.Tensors/Engines/Engine.cs
  • src/AiDotNet.Tensors/Helpers/TensorPrimitivesHelper.cs
  • src/AiDotNet.Tensors/Interfaces/IVectorizedOperations.cs
  • src/AiDotNet.Tensors/LinearAlgebra/Matrix.cs
  • src/AiDotNet.Tensors/LinearAlgebra/MatrixBase.cs
  • src/AiDotNet.Tensors/NumericOperations/ByteOperations.cs
  • src/AiDotNet.Tensors/NumericOperations/ComplexOperations.cs
  • src/AiDotNet.Tensors/NumericOperations/DecimalOperations.cs
  • src/AiDotNet.Tensors/NumericOperations/DoubleOperations.cs
  • src/AiDotNet.Tensors/NumericOperations/FloatOperations.cs
  • src/AiDotNet.Tensors/NumericOperations/HalfOperations.cs
  • src/AiDotNet.Tensors/NumericOperations/Int32Operations.cs
  • src/AiDotNet.Tensors/NumericOperations/Int64Operations.cs
  • src/AiDotNet.Tensors/NumericOperations/MultivectorOperations.cs
  • src/AiDotNet.Tensors/NumericOperations/OctonionOperations.cs
  • src/AiDotNet.Tensors/NumericOperations/SByteOperations.cs
  • src/AiDotNet.Tensors/NumericOperations/ShortOperations.cs
  • src/AiDotNet.Tensors/NumericOperations/UInt16Operations.cs
  • src/AiDotNet.Tensors/NumericOperations/UInt32Operations.cs
  • src/AiDotNet.Tensors/NumericOperations/UInt64Operations.cs
  • src/AiDotNet.Tensors/NumericOperations/UIntOperations.cs
  • tests/AiDotNet.Tensors.Benchmarks/CpuMatMulDiagnostics.cs
  • tests/AiDotNet.Tensors.Benchmarks/GpuMatMulDiagnostics.cs
  • tests/AiDotNet.Tensors.Benchmarks/Helpers/BenchmarkHelper.cs
✅ Files skipped from review due to trivial changes (1)
  • src/AiDotNet.Tensors/Engines/Engine.cs
🧰 Additional context used
🧠 Learnings (3)
📚 Learning: 2025-12-18T08:49:25.295Z
Learnt from: ooples
Repo: ooples/AiDotNet PR: 444
File: src/Interfaces/IPruningMask.cs:1-102
Timestamp: 2025-12-18T08:49:25.295Z
Learning: In the AiDotNet repository, the project-level global using includes AiDotNet.Tensors.LinearAlgebra via AiDotNet.csproj. Therefore, Vector<T>, Matrix<T>, and Tensor<T> are available without per-file using directives. Do not flag missing using directives for these types in any C# files within this project. Apply this guideline broadly to all C# files (not just a single file) to avoid false positives. If a file uses a type from a different namespace not covered by the global using, flag as usual.

Applied to files:

  • src/AiDotNet.Tensors/Helpers/TensorPrimitivesHelper.cs
  • src/AiDotNet.Tensors/Engines/DirectGpu/GemmBenchmark.cs
  • src/AiDotNet.Tensors/Engines/DirectGpu/OpenCL/OpenClNativeBindings.cs
  • src/AiDotNet.Tensors/NumericOperations/MultivectorOperations.cs
  • src/AiDotNet.Tensors/Engines/DirectGpu/Profiling/GemmProfiler.cs
  • src/AiDotNet.Tensors/LinearAlgebra/MatrixBase.cs
  • src/AiDotNet.Tensors/LinearAlgebra/Matrix.cs
  • src/AiDotNet.Tensors/NumericOperations/UIntOperations.cs
  • tests/AiDotNet.Tensors.Benchmarks/CpuMatMulDiagnostics.cs
  • src/AiDotNet.Tensors/Engines/DirectGpu/DirectGpuEngine.cs
  • src/AiDotNet.Tensors/NumericOperations/FloatOperations.cs
  • src/AiDotNet.Tensors/Engines/DirectGpu/OpenCL/DynamicGemmKernel.cs
  • tests/AiDotNet.Tensors.Benchmarks/GpuMatMulDiagnostics.cs
  • src/AiDotNet.Tensors/NumericOperations/ShortOperations.cs
  • src/AiDotNet.Tensors/NumericOperations/OctonionOperations.cs
  • src/AiDotNet.Tensors/Engines/AiDotNetEngine.cs
  • src/AiDotNet.Tensors/NumericOperations/ComplexOperations.cs
  • src/AiDotNet.Tensors/NumericOperations/ByteOperations.cs
  • tests/AiDotNet.Tensors.Benchmarks/Helpers/BenchmarkHelper.cs
  • src/AiDotNet.Tensors/NumericOperations/Int32Operations.cs
  • src/AiDotNet.Tensors/Engines/DirectGpu/OpenCL/GemmAutoTuner.cs
  • src/AiDotNet.Tensors/NumericOperations/DecimalOperations.cs
  • src/AiDotNet.Tensors/NumericOperations/Int64Operations.cs
  • src/AiDotNet.Tensors/Engines/DirectGpu/HIP/HipBackend.cs
  • src/AiDotNet.Tensors/Engines/DirectGpuTensorEngine.cs
  • src/AiDotNet.Tensors/NumericOperations/UInt16Operations.cs
  • src/AiDotNet.Tensors/NumericOperations/HalfOperations.cs
  • src/AiDotNet.Tensors/Engines/DirectGpu/HIP/HipNativeBindings.cs
  • src/AiDotNet.Tensors/NumericOperations/UInt64Operations.cs
  • src/AiDotNet.Tensors/NumericOperations/UInt32Operations.cs
  • src/AiDotNet.Tensors/Engines/DirectGpu/OpenCL/OpenClBackend.cs
  • src/AiDotNet.Tensors/Interfaces/IVectorizedOperations.cs
  • src/AiDotNet.Tensors/NumericOperations/DoubleOperations.cs
  • src/AiDotNet.Tensors/NumericOperations/SByteOperations.cs
  • src/AiDotNet.Tensors/Engines/CpuEngine.cs
📚 Learning: 2025-12-18T08:49:53.103Z
Learnt from: ooples
Repo: ooples/AiDotNet PR: 444
File: src/Interfaces/IPruningStrategy.cs:1-4
Timestamp: 2025-12-18T08:49:53.103Z
Learning: In this repository, global using directives are declared in AiDotNet.csproj for core namespaces (AiDotNet.Tensors.* and AiDotNet.*) and common system types. When reviewing C# files, assume these global usings are in effect; avoid adding duplicate using statements for these namespaces and for types like Vector<T>, Matrix<T>, Tensor<T>, etc. If a type is not found, verify the global usings or consider adding a file-scoped using if needed. Prefer relying on global usings to reduce boilerplate.

Applied to files:

  • src/AiDotNet.Tensors/Helpers/TensorPrimitivesHelper.cs
  • src/AiDotNet.Tensors/Engines/DirectGpu/GemmBenchmark.cs
  • src/AiDotNet.Tensors/Engines/DirectGpu/OpenCL/OpenClNativeBindings.cs
  • src/AiDotNet.Tensors/NumericOperations/MultivectorOperations.cs
  • src/AiDotNet.Tensors/Engines/DirectGpu/Profiling/GemmProfiler.cs
  • src/AiDotNet.Tensors/LinearAlgebra/MatrixBase.cs
  • src/AiDotNet.Tensors/LinearAlgebra/Matrix.cs
  • src/AiDotNet.Tensors/NumericOperations/UIntOperations.cs
  • tests/AiDotNet.Tensors.Benchmarks/CpuMatMulDiagnostics.cs
  • src/AiDotNet.Tensors/Engines/DirectGpu/DirectGpuEngine.cs
  • src/AiDotNet.Tensors/NumericOperations/FloatOperations.cs
  • src/AiDotNet.Tensors/Engines/DirectGpu/OpenCL/DynamicGemmKernel.cs
  • tests/AiDotNet.Tensors.Benchmarks/GpuMatMulDiagnostics.cs
  • src/AiDotNet.Tensors/NumericOperations/ShortOperations.cs
  • src/AiDotNet.Tensors/NumericOperations/OctonionOperations.cs
  • src/AiDotNet.Tensors/Engines/AiDotNetEngine.cs
  • src/AiDotNet.Tensors/NumericOperations/ComplexOperations.cs
  • src/AiDotNet.Tensors/NumericOperations/ByteOperations.cs
  • tests/AiDotNet.Tensors.Benchmarks/Helpers/BenchmarkHelper.cs
  • src/AiDotNet.Tensors/NumericOperations/Int32Operations.cs
  • src/AiDotNet.Tensors/Engines/DirectGpu/OpenCL/GemmAutoTuner.cs
  • src/AiDotNet.Tensors/NumericOperations/DecimalOperations.cs
  • src/AiDotNet.Tensors/NumericOperations/Int64Operations.cs
  • src/AiDotNet.Tensors/Engines/DirectGpu/HIP/HipBackend.cs
  • src/AiDotNet.Tensors/Engines/DirectGpuTensorEngine.cs
  • src/AiDotNet.Tensors/NumericOperations/UInt16Operations.cs
  • src/AiDotNet.Tensors/NumericOperations/HalfOperations.cs
  • src/AiDotNet.Tensors/Engines/DirectGpu/HIP/HipNativeBindings.cs
  • src/AiDotNet.Tensors/NumericOperations/UInt64Operations.cs
  • src/AiDotNet.Tensors/NumericOperations/UInt32Operations.cs
  • src/AiDotNet.Tensors/Engines/DirectGpu/OpenCL/OpenClBackend.cs
  • src/AiDotNet.Tensors/Interfaces/IVectorizedOperations.cs
  • src/AiDotNet.Tensors/NumericOperations/DoubleOperations.cs
  • src/AiDotNet.Tensors/NumericOperations/SByteOperations.cs
  • src/AiDotNet.Tensors/Engines/CpuEngine.cs
📚 Learning: 2025-11-19T04:08:26.895Z
Learnt from: ooples
Repo: ooples/AiDotNet PR: 0
File: :0-0
Timestamp: 2025-11-19T04:08:26.895Z
Learning: For ILGPU GPU operations in GpuEngine.cs, use standard .NET exception types (InvalidOperationException, ArgumentException, OutOfMemoryException) instead of ILGPU-specific exception types, as ILGPU exception types may be version-specific. Combine with message-based filtering using ex.Message.Contains("device") or ex.Message.Contains("accelerator") as a fallback for GPU-specific errors.

Applied to files:

  • src/AiDotNet.Tensors/Engines/DirectGpu/OpenCL/OpenClNativeBindings.cs
  • src/AiDotNet.Tensors/Engines/DirectGpu/DirectGpuEngine.cs
  • src/AiDotNet.Tensors/Engines/AiDotNetEngine.cs
  • src/AiDotNet.Tensors/Engines/DirectGpu/HIP/HipBackend.cs
🔇 Additional comments (53)
src/AiDotNet.Tensors/NumericOperations/OctonionOperations.cs (1)

318-342: LGTM!

The implementation is clean and correct:

  • AllFinite properly leverages the existing IsNaN and IsInfinity methods which already check all 8 octonion components.
  • The badIndex out parameter is correctly set in all code paths.
  • IsAnyNonFinite appropriately delegates to AllFinite to avoid code duplication.
  • The scalar iteration approach is consistent with SupportsCpuAcceleration => false for this type.
src/AiDotNet.Tensors/NumericOperations/SByteOperations.cs (1)

741-759: LGTM!

The implementations are correct. Since sbyte is an integral type (range -128 to 127), it cannot represent NaN or Infinity values—all sbyte values are inherently finite. Returning true for AllFinite and false for IsAnyNonFinite with badIndex = -1 is the appropriate behavior, consistent with the existing IsNaN and IsInfinity methods that also return false for this type.

src/AiDotNet.Tensors/Engines/DirectGpu/HIP/HipNativeBindings.cs (2)

6-6: LGTM!

The System.Diagnostics import is correctly added to support the Trace.WriteLine usage in the LogDiagnostic method.


291-295: LGTM!

Switching from Console.WriteLine to Trace.WriteLine is the correct approach for diagnostic output. This allows consumers to control trace output via TraceListener configuration without polluting console output in production. The change aligns with the PR-wide pattern of standardizing on Trace-based diagnostics across DirectGpu backends.

src/AiDotNet.Tensors/NumericOperations/UInt64Operations.cs (1)

764-782: LGTM!

The implementations are correct and consistent with the existing pattern in this class. Since ulong is an integral type that cannot represent NaN or Infinity, AllFinite correctly returns true and IsAnyNonFinite correctly returns false, both with badIndex = -1. This aligns with the existing IsNaN and IsInfinity methods at lines 647 and 669 that also return false.

Note: The AI summary mentions duplicate method declarations, but the provided code shows only a single definition of each method, so no duplication issue is present.

src/AiDotNet.Tensors/NumericOperations/UInt16Operations.cs (1)

706-724: LGTM!

The implementations are correct. Since ushort is an integer type that cannot represent NaN or Infinity, AllFinite correctly always returns true and IsAnyNonFinite correctly always returns false, both with badIndex = -1. This is consistent with the existing IsNaN and IsInfinity methods in this class.

src/AiDotNet.Tensors/NumericOperations/Int64Operations.cs (1)

790-808: LGTM!

The implementations are correct for the Int64 type. Since long values cannot represent NaN or Infinity, AllFinite trivially returns true and IsAnyNonFinite trivially returns false. This is consistent with the existing IsNaN and IsInfinity methods in this class that also return constant values for integers. The documentation clearly explains the rationale.

src/AiDotNet.Tensors/NumericOperations/HalfOperations.cs (1)

265-274: LGTM!

Clean delegation to AllFinite with proper negation. The XML documentation is thorough and correctly describes the behavior.

src/AiDotNet.Tensors/NumericOperations/UInt32Operations.cs (2)

715-723: LGTM!

The implementation correctly returns true for all UInt32 spans since integral types cannot represent NaN or Infinity. The badIndex = -1 convention is appropriate.


725-733: LGTM!

Correctly returns false since UInt32 values are always finite by definition. The implementation is consistent with AllFinite and follows the expected interface contract.

src/AiDotNet.Tensors/NumericOperations/Int32Operations.cs (2)

743-761: LGTM!

The AllFinite and IsAnyNonFinite implementations are correct. Since Int32 has no representation for NaN or Infinity (consistent with the existing IsNaN and IsInfinity methods), all integer values are inherently finite. Returning constant results without iterating the span is the appropriate optimization.


763-764: I cannot proceed without a review comment to rewrite. Please provide the review comment enclosed in <review_comment> tags along with any relevant code context or verification results.

src/AiDotNet.Tensors/NumericOperations/DecimalOperations.cs (1)

694-712: Constant finiteness results for decimal are correct and efficient.
Returning true/false with badIndex = -1 matches decimal semantics and avoids unnecessary scans.

src/AiDotNet.Tensors/Helpers/TensorPrimitivesHelper.cs (1)

351-360: Verify Vector<T>.AsSpan() is a safe, non-copying view.
This helper’s correctness/perf hinges on x.AsSpan() returning a stable span over the vector’s backing storage (not a temporary/copy).

src/AiDotNet.Tensors/NumericOperations/UIntOperations.cs (1)

725-744: Constant finiteness results for uint look good.
These implementations are correct, fast, and consistent with the new interface surface.

src/AiDotNet.Tensors/NumericOperations/ByteOperations.cs (1)

674-693: Constant finiteness results for byte look good.
These implementations are correct, fast, and consistent with the new interface surface.

src/AiDotNet.Tensors/NumericOperations/MultivectorOperations.cs (1)

318-342: LGTM!

The AllFinite and IsAnyNonFinite implementations are correct and cleanly reuse the existing IsNaN and IsInfinity methods. The pattern of delegating IsAnyNonFinite to AllFinite with inverted result is consistent with other numeric operation classes.

src/AiDotNet.Tensors/NumericOperations/ShortOperations.cs (1)

702-720: LGTM!

The implementation correctly returns fixed values since short is an integer type that cannot represent NaN or Infinity. The documentation accurately reflects this behavior, and the pattern is consistent with the existing IsNaN and IsInfinity methods in this class.

src/AiDotNet.Tensors/Engines/DirectGpu/Profiling/GemmProfiler.cs (3)

138-146: LGTM!

The transition from Console.WriteLine to Trace.WriteLine is appropriate for library code. This allows consumers to configure trace listeners and control diagnostic output without polluting stdout.


161-167: Error handling preserved correctly.

The exception handling continues profiling other sizes after an error, which is appropriate for a benchmarking tool. The Trace output maintains the same error information as before.


316-341: Consistent logging changes.

The rectangular profiling method follows the same Trace.WriteLine pattern as RunFullProfile, maintaining consistency across the codebase.

src/AiDotNet.Tensors/LinearAlgebra/Matrix.cs (1)

24-33: Verify MatrixBase invariant enforcement and all internal caller compliance. Ensure the internal Matrix(int rows, int cols, T[] data) constructor only receives exact-length backing arrays and that MatrixBase(int,int,T[]) enforces overflow and length validation.

src/AiDotNet.Tensors/Engines/AiDotNetEngine.cs (2)

1-4: LGTM on using directives and logging migration.

The switch from Console.WriteLine to Trace.WriteLine aligns with the PR's goal of enabling Trace-based diagnostics across the codebase. This allows diagnostic output to be captured by trace listeners or suppressed in production without code changes.


157-161: LGTM on the new GetEngineInfo() method.

Clean utility method that exposes engine name and GPU support status. This complements the existing engine configuration API and supports the new diagnostics infrastructure.

src/AiDotNet.Tensors/Engines/DirectGpu/OpenCL/DynamicGemmKernel.cs (2)

236-266: LGTM on LogDiag Trace migration.

The fallback path in LogDiag now correctly uses Trace.WriteLine instead of Console.WriteLine. This maintains the existing file-logging priority while providing consistent trace-based output when file logging fails or is not configured.


374-384: LGTM on kernel compilation failure diagnostics.

The kernel source output on compilation failure now routes through Trace.WriteLine, which is appropriate for diagnostic tooling. This allows capture via trace listeners during debugging without polluting console output in production.

tests/AiDotNet.Tensors.Benchmarks/Helpers/BenchmarkHelper.cs (1)

28-38: LGTM on CreateRandomMatrix.

Clean implementation using the writable span API for efficient matrix initialization. The [-1, 1] range is well-suited for benchmark scenarios involving neural network weights.

tests/AiDotNet.Tensors.Benchmarks/CpuMatMulDiagnostics.cs (2)

30-76: LGTM on benchmark structure.

The benchmark implementation follows good practices:

  • Seeded random for reproducibility
  • Warmup phase before timing
  • Reduced iterations for larger matrices to keep runtime reasonable
  • GFLOPS calculation correctly uses 2N³ for matrix multiplication
  • Checksum prevents dead-code elimination by optimizers

37-38: Verify if CpuEngine requires disposal.

CpuEngine is instantiated at line 37 but never disposed. If CpuEngine implements IDisposable, consider wrapping it in a using statement or calling Dispose() at the end of Run().

src/AiDotNet.Tensors/Engines/DirectGpu/DirectGpuEngine.cs (4)

32-33: LGTM on GemmValidateEnabled configuration.

Environment variable-driven opt-in validation is a clean approach. Disabled by default avoids performance overhead in production while enabling correctness debugging when needed.


426-430: LGTM on validation in MatMulWithCachedWeights.

Consistent application of the non-finite validation check in the cached weights path ensures GPU correctness validation covers both GEMM entry points.


94-122: LGTM on constructor logging migration.

The initialization flow now uses Trace.WriteLine consistently, aligning with the broader diagnostic infrastructure changes. The detailed logging of backend discovery attempts aids debugging without cluttering console output.


397-411: Unable to verify GPU GEMM validation logic—manual code review required.

The review comment asserts that IsAnyNonFinite delegates to the numeric operations API and maintains consistency with finite-check infrastructure. However, I cannot access the repository to confirm:

  • Whether the IsAnyNonFinite method exists on the expected interface
  • Whether the delegation pattern is correctly implemented
  • Whether the numeric operations API usage is consistent with the PR's changes

Please manually verify these claims against the codebase.

tests/AiDotNet.Tensors.Benchmarks/GpuMatMulDiagnostics.cs (8)

1-10: LGTM!

Using statements and namespace declaration are appropriate for the benchmark context.


11-20: LGTM!

Clear documentation and appropriate class design for a diagnostic utility.


47-66: LGTM!

The correctness test loop structure is sound. Testing both safe and unsafe kernel modes with the size restriction for unsafe mode (≤512) is a reasonable approach for diagnostics.


70-98: LGTM!

The performance measurement approach is well-structured: warmup pass to avoid cold-start artifacts, per-iteration timing with Stopwatch, and correct GFLOPS calculation (2 × N³ operations for matrix multiply).


113-118: LGTM!

Clean formatting helper with appropriate scientific notation for error metrics.


120-124: Environment variable manipulation is process-global.

This is acceptable for a single-threaded diagnostic tool, but be aware that Environment.SetEnvironmentVariable affects the entire process. If these diagnostics are ever run in parallel or integrated into a test suite with concurrent tests, this could cause race conditions.


101-111: Verify if Matrix<float> results require disposal.

The matrices returned by MatrixMultiply are used for comparison and then abandoned. If Matrix<float> implements IDisposable (e.g., for GPU buffer cleanup), dispose them after comparison to avoid resource leaks.


39-44: Verify if CpuEngine requires disposal.

If CpuEngine implements IDisposable, the instance created on line 40 should be wrapped in a using statement or disposed in a finally block to prevent resource leaks.

src/AiDotNet.Tensors/Engines/CpuEngine.cs (8)

66-69: DirectGpu property doc looks good

Nice: the property is nullable and the doc explains the intent (offloading when available).


1745-1784: Span-based matrix ops look good (Add/Subtract/MultiplyScalar/SumOfSquares)

The switch to numOps.*(ReadOnlySpan<T>, ..., Span<T>) should reduce overhead and enable SIMD where available.


1866-1906: Row get/set via numOps.Copy is a clean improvement

This should be faster and clearer than manual loops.


3026-3028: TensorSum span-based path is a nice simplification

Cleaner and likely faster than a manual loop.


12255-12257: TensorSumOfSquares using Dot(span, span) is a good optimization

Assuming numOps.Dot is optimized/SIMD’d, this should be a nice speedup.


13331-13356: Scalar tensor ops switched to span APIs: LGTM

AddScalar/SubtractScalar/DivideScalar are now consistent with the other span-based fast paths.


1474-1670: Verify accumulation semantics of SimdVector.MatMulInnerLoopFloat and SimdVector.MatMulInnerLoopDouble.

Matrix multiply correctness depends on these methods accumulating into c (i.e., c[j] += aik * b[j]). Since cSpan.Clear() executes once and the kk loop processes multiple k values, overwriting would discard intermediate results and produce incorrect output. Confirm the methods use compound assignment, not simple assignment.

Also remove the unused parameter m from MultiplyFloatBlock and MultiplyDoubleBlock signatures.


3099-3128: Chunked parallel max/min: confirm NaN semantics for float/double types

The chunking and reduction logic is sound. For float/double operands, verify that numOps.Max/Min behavior with NaNs aligns with requirements—NaN propagates by default in most implementations (both operands return NaN), which may or may not be desired for tensor reduction. If NaN should be ignored (returning the numeric operand), verify whether MaxNumber/MinNumber variants are available and appropriate.

Also applies to: 3146-3175

src/AiDotNet.Tensors/Engines/DirectGpu/OpenCL/OpenClBackend.cs (4)

1918-1924: GemmDoubleBuffered(...) as an explicit public fallback is a good addition.
Nice to have a stable, explicit “known kernel” entrypoint for diagnostics/tests and for bypassing routing.


1750-1810: [Unable to rewrite review comment - manual verification required]

The repository could not be accessed to verify the claimed inconsistencies in help text (lines 4521–4524) and PR objective. While the code snippet confirms defaultValue: true on line 1752, verification of the help text content and PR requirements is needed before determining if this represents a genuine behavior change.


4499-4530: Verify if AIDOTNET_GEMM_VALIDATE is functionally implemented in the codebase.

PrintDiagnosticHelp() advertises AIDOTNET_GEMM_VALIDATE as an option for validating GEMM output. Confirm whether this environment variable is actually read and used elsewhere in the codebase. If it's not implemented, remove it from the help text to avoid misleading developers during debugging.


1833-1898: Verify StrideN => column-major C semantic mapping and the packed-path layout handling.

The assumption that useColumnMajorC = config.StrideN (lines 1845–1846) correctly maps to column-major matrix format cannot be confirmed without access to the GemmConfig class definition and its initialization logic. If StrideN semantically represents something other than "C is column-major" (e.g., a leading-dimension stride for batched operations), the subsequent pad/copy logic—particularly the choice between PadCopyFromColumnMajorMatrix and CopySubmatrix (lines 1892–1896)—could produce incorrect results.

Clarify the intended semantics of StrideN and ensure the layout detection aligns with how the matrix is actually padded and copied.

Comment thread src/AiDotNet.Tensors/Engines/CpuEngine.cs
Comment thread src/AiDotNet.Tensors/Engines/DirectGpu/OpenCL/OpenClBackend.cs Fixed
Comment thread src/AiDotNet.Tensors/Engines/DirectGpu/OpenCL/OpenClBackend.cs Fixed
@sonarqubecloud

Copy link
Copy Markdown

Quality Gate Failed Quality Gate failed

Failed conditions
16.4% Coverage on New Code (required ≥ 80%)
12.3% Duplication on New Code (required ≤ 3%)
B Maintainability Rating on New Code (required ≥ A)

See analysis details on SonarQube Cloud

Catch issues before they fail your Quality Gate with our IDE extension SonarQube for IDE

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 3

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (3)
src/AiDotNet.Tensors/Engines/DirectGpu/OpenCL/OpenClBackend.cs (3)

1732-1838: Default GEMM routing contradicts the PR objective (“built-in by default”).
enableDynamic currently defaults to true, so baseline+dynamic are attempted before built-in kernels. If the intended default is “built-in first”, flip the default and update the help text accordingly.

Proposed diff (make dynamic opt-in)
-            bool enableDynamic = GetEnvBool("AIDOTNET_GEMM_ENABLE_DYNAMIC", defaultValue: true);
+            bool enableDynamic = GetEnvBool("AIDOTNET_GEMM_ENABLE_DYNAMIC", defaultValue: false);

1907-1931: MatMul now logs unconditionally; GemmDoubleBuffered API looks fine.
Please gate the Trace.WriteLine behind AIDOTNET_GEMM_TRACE (or similar) to avoid hot-path logging.

Proposed diff
 public IGpuBuffer MatMul(IGpuBuffer A, IGpuBuffer B, int M, int N, int K)
 {
-    Trace.WriteLine($"[OpenClBackend.MatMul] Called: {M}x{N}x{K}");
+    if (GetEnvBool("AIDOTNET_GEMM_TRACE"))
+        Trace.WriteLine($"[OpenClBackend.MatMul] Called: {M}x{N}x{K}");
     var C = AllocateBuffer(M * N);
     Gemm(A, B, C, M, N, K, 1.0f, 0.0f);
     // Sync only when returning buffer that might be immediately read
     _context?.Finish();
     return C;
 }

4504-4537: Diagnostic help text should match actual defaults.
If AIDOTNET_GEMM_ENABLE_DYNAMIC is meant to be opt-in, the help should say default is disabled (and code should match).

🤖 Fix all issues with AI agents
In @src/AiDotNet.Tensors/Engines/DirectGpu/OpenCL/OpenClBackend.cs:
- Around line 1695-1698: The IsEffectivelyZero(float value) method currently
uses Math.Abs(value) <= float.Epsilon which is effectively exact-zero for
practical use; change it to use a realistic tolerance (e.g., const float
Tolerance = 1e-8f) and compare Math.Abs(value) <= Tolerance, or if you truly
want exact zero semantics rename the method to IsZero to avoid the misleading
"effectively" wording; update references to IsEffectivelyZero accordingly.
- Around line 1840-1905: Packed GEMM currently zeroes the host-side cPad via
Fill(), causing large managed allocations and uploads; in
TryExecutePackedDynamicGemm replace the Fill(cPad, 0.0f, (int)cSize) calls with
a GPU-side clear: after AllocateBuffer for cPad, invoke the existing zero_buffer
kernel (if present) or clEnqueueFillBuffer fallback to fill cPad on the device
(use the cPad buffer handle and byte size computed from cSize), preserving the
conditional (only when beta is effectively zero) and synchronization semantics
before calling TryExecuteDynamicGemm; reference symbols:
TryExecutePackedDynamicGemm, AllocateBuffer, Fill (remove), zero_buffer (reuse),
clEnqueueFillBuffer (fallback), and ensure the
PadCopyFromColumnMajorMatrix/CopySubmatrix behavior is unchanged.
🧹 Nitpick comments (11)
src/AiDotNet.Tensors/Engines/DirectGpu/OpenCL/OpenClBackend.cs (7)

30-46: Remove duplicated XML doc block (it’s repeated back-to-back).
Right now the type has two <summary>/<remarks> blocks; keep one and add the “For Beginners” paragraph there.

Proposed diff
@@
-    /// <summary>
-    /// OpenCL backend for direct GPU access on AMD, Intel, and NVIDIA GPUs.
-    /// Uses pure P/Invoke with no managed GPU runtime dependency.
-    /// </summary>
-    /// <remarks>
-    /// <para><b>Key Features:</b></para>
-    /// <list type="bullet">
-    /// <item>Works on ALL .NET versions (4.6.2, 4.7.1, net8.0, etc.)</item>
-    /// <item>No managed GPU runtime dependency - pure P/Invoke</item>
-    /// <item>Double-buffered GEMM for compute/memory overlap</item>
-    /// <item>Fused operations (GEMM+Bias+Activation)</item>
-    /// <item>Bank-conflict-free shared memory</item>
-    /// </list>
-    /// <para><b>For Beginners:</b> This is the "driver" that talks directly to your graphics card (GPU). 
-    /// It translates math problems (like multiplying giant tables of numbers) into a language 
-    /// the GPU understands. This is much faster than using just your computer's main processor (CPU).</para>
-    /// </remarks>
     public sealed class OpenClBackend : IAsyncGpuBackend

80-110: New device-info properties: consider setting explicit defaults on “not available” paths.
ComputeUnits/GlobalMemoryBytes/LocalMemoryBytes will be 0 when OpenCL isn’t available; that’s fine, but consider explicitly assigning to make intent clear (like you did for DeviceName/DeviceVendor).


131-233: Trace logging is fine, but avoid noisy init logs unless diagnostics are enabled.
These Trace.WriteLine calls will fire for every backend creation; consider gating with EnableTuningDiagnostics (or an AIDOTNET_GPU_TRACE) to keep normal runs quiet.


247-517: Kernel compilation tracing: good visibility; consider collapsing repeated “compiled: …” output behind a flag.
This will be extremely verbose on startup (and kernel name lists can get long).


685-718: SimpleConsoleLogger now changes console colors but writes via Trace.
Color changes won’t matter unless a console TraceListener is installed; either write to Console here or drop the color logic.


793-924: Tuning diagnostics now use Trace: OK, but keep env parsing consistent.
Elsewhere you sometimes check Environment.GetEnvironmentVariable(...) == "1"; now that GetEnvBool exists, consider using it uniformly.


3610-3684: Diagnostics printing switched to Trace: OK, but console coloring may not apply.
Same issue as SimpleConsoleLogger: color + Trace.WriteLine only works with console listeners.

src/AiDotNet.Tensors/Engines/CpuEngine.cs (4)

36-49: Remove/avoid duplicated class XML docs (likely accidental).

There are two XML doc blocks preceding CpuEngine (one earlier, one at Lines 36-49). This can create duplicated/messy generated docs and warnings.

Proposed fix
-/// <summary>
-/// CPU-based execution engine using INumericOperations for type-generic operations.
-/// </summary>
-/// <remarks>
-/// <para>
-/// CpuEngine provides the default execution backend for AiDotNet. It works with
-/// any numeric type that implements INumericOperations{T}, including decimal,
-/// BigInteger, and custom numeric types.
-/// </para>
-/// <para><b>For Beginners:</b> This is the standard, "always works" mode. 
-/// It uses your computer's main processor (CPU) to do the math. While not as 
-/// fast as a graphics card (GPU) for huge problems, it is very reliable 
-/// and works on every computer without any extra setup.</para>
-/// </remarks>
+/// <summary>
+/// CPU-based execution engine using INumericOperations for type-generic operations.
+/// </summary>
+/// <remarks>
+/// <para>
+/// CpuEngine provides the default execution backend for AiDotNet. It works with
+/// any numeric type that implements INumericOperations{T}, including decimal,
+/// BigInteger, and custom numeric types.
+/// </para>
+/// <para><b>For Beginners:</b> This is the standard, "always works" mode.
+/// It uses your computer's main processor (CPU) to do the math. While not as
+/// fast as a graphics card (GPU) for huge problems, it is very reliable
+/// and works on every computer without any extra setup.</para>
+/// </remarks>
 public class CpuEngine : IEngine

1440-1451: Avoid Unsafe.As<Matrix<...>, Matrix<T>> here (unnecessary risk).

Given the runtime type checks, a normal cast is clearer and avoids subtle Unsafe footguns during refactors.

Proposed fix
         if (typeof(T) == typeof(float) && a is Matrix<float> floatA && b is Matrix<float> floatB)
         {
             var result = MatrixMultiplyFloat(floatA, floatB);
-            return Unsafe.As<Matrix<float>, Matrix<T>>(ref result);
+            return (Matrix<T>)(object)result;
         }

         if (typeof(T) == typeof(double) && a is Matrix<double> doubleA && b is Matrix<double> doubleB)
         {
             var result = MatrixMultiplyDouble(doubleA, doubleB);
-            return Unsafe.As<Matrix<double>, Matrix<T>>(ref result);
+            return (Matrix<T>)(object)result;
         }

2228-2240: TensorDivide: add a fast span path for floating-point.

Right now the method always does a scalar loop; for float/double you can likely delegate to numOps.Divide(aSpan, bSpan, resultSpan) (keeping the non-float zero-check behavior).

Proposed fix
         var numOps = MathHelper.GetNumericOperations<T>();
         var result = new Tensor<T>(a.Shape);

         var aSpan = a.AsSpan();
         var bSpan = b.AsSpan();
         var resultSpan = result.AsWritableSpan();
         bool checkZero = !MathHelper.IsFloatingPoint<T>();
+        if (!checkZero)
+        {
+            numOps.Divide(aSpan, bSpan, resultSpan);
+            return result;
+        }
         for (int i = 0; i < bSpan.Length; i++)
         {
             var divisor = bSpan[i];
             if (checkZero && numOps.Equals(divisor, numOps.Zero))
             {
                 throw new DivideByZeroException($"Division by zero at index {i}");
             }
             resultSpan[i] = numOps.Divide(aSpan[i], divisor);
         }

52-59: Environment variables are read once at type initialization; confirm this static caching is intentional.

CpuMatMulTraceEnabled and CpuMatMulSingleThread are initialized with Environment.GetEnvironmentVariable() and cached in static readonly fields. In .NET, environment variables read during process startup are cached by the runtime; changes via Environment.SetEnvironmentVariable() after initialization won't affect these fields. If tests or benchmarks rely on changing these variables at runtime, they will see stale cached values.

If runtime reconfiguration is needed, consider using .NET's IConfiguration/IOptionsMonitor pattern instead of static caching for testability.

📜 Review details

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

📥 Commits

Reviewing files that changed from the base of the PR and between 248bf5f and e33e54d.

📒 Files selected for processing (3)
  • src/AiDotNet.Tensors/Engines/CpuEngine.cs
  • src/AiDotNet.Tensors/Engines/DirectGpu/OpenCL/OpenClBackend.cs
  • src/AiDotNet.Tensors/Engines/DirectGpu/OpenCL/OpenClNativeBindings.cs
🧰 Additional context used
🧠 Learnings (3)
📚 Learning: 2025-11-19T04:08:26.895Z
Learnt from: ooples
Repo: ooples/AiDotNet PR: 0
File: :0-0
Timestamp: 2025-11-19T04:08:26.895Z
Learning: For ILGPU GPU operations in GpuEngine.cs, use standard .NET exception types (InvalidOperationException, ArgumentException, OutOfMemoryException) instead of ILGPU-specific exception types, as ILGPU exception types may be version-specific. Combine with message-based filtering using ex.Message.Contains("device") or ex.Message.Contains("accelerator") as a fallback for GPU-specific errors.

Applied to files:

  • src/AiDotNet.Tensors/Engines/DirectGpu/OpenCL/OpenClNativeBindings.cs
  • src/AiDotNet.Tensors/Engines/CpuEngine.cs
📚 Learning: 2025-12-18T08:49:53.103Z
Learnt from: ooples
Repo: ooples/AiDotNet PR: 444
File: src/Interfaces/IPruningStrategy.cs:1-4
Timestamp: 2025-12-18T08:49:53.103Z
Learning: In this repository, global using directives are declared in AiDotNet.csproj for core namespaces (AiDotNet.Tensors.* and AiDotNet.*) and common system types. When reviewing C# files, assume these global usings are in effect; avoid adding duplicate using statements for these namespaces and for types like Vector<T>, Matrix<T>, Tensor<T>, etc. If a type is not found, verify the global usings or consider adding a file-scoped using if needed. Prefer relying on global usings to reduce boilerplate.

Applied to files:

  • src/AiDotNet.Tensors/Engines/DirectGpu/OpenCL/OpenClNativeBindings.cs
  • src/AiDotNet.Tensors/Engines/DirectGpu/OpenCL/OpenClBackend.cs
  • src/AiDotNet.Tensors/Engines/CpuEngine.cs
📚 Learning: 2025-12-18T08:49:25.295Z
Learnt from: ooples
Repo: ooples/AiDotNet PR: 444
File: src/Interfaces/IPruningMask.cs:1-102
Timestamp: 2025-12-18T08:49:25.295Z
Learning: In the AiDotNet repository, the project-level global using includes AiDotNet.Tensors.LinearAlgebra via AiDotNet.csproj. Therefore, Vector<T>, Matrix<T>, and Tensor<T> are available without per-file using directives. Do not flag missing using directives for these types in any C# files within this project. Apply this guideline broadly to all C# files (not just a single file) to avoid false positives. If a file uses a type from a different namespace not covered by the global using, flag as usual.

Applied to files:

  • src/AiDotNet.Tensors/Engines/DirectGpu/OpenCL/OpenClNativeBindings.cs
  • src/AiDotNet.Tensors/Engines/DirectGpu/OpenCL/OpenClBackend.cs
  • src/AiDotNet.Tensors/Engines/CpuEngine.cs
⏰ Context from checks skipped due to timeout of 90000ms. You can increase the timeout in your CodeRabbit configuration to a maximum of 15 minutes (900000ms). (2)
  • GitHub Check: CodeQL Analysis
  • GitHub Check: Build (Windows)
🔇 Additional comments (14)
src/AiDotNet.Tensors/Engines/DirectGpu/OpenCL/OpenClNativeBindings.cs (4)

5-19: LGTM!

The EnableDiagnostics property provides a clean opt-in mechanism for diagnostic output. Defaulting to false ensures diagnostics don't affect production performance. The System.Diagnostics import is correctly added for Trace usage.


362-413: LGTM!

The migration from Console.WriteLine to LogDiagnostic correctly makes diagnostic output opt-in and routes it through Trace, consistent with the broader PR changes across the GPU backend stack.


588-611: LGTM!

The diagnostic logging in the IsAvailable property provides valuable debugging information (error codes, platform count) without altering the control flow. Exception details are captured appropriately for troubleshooting.


613-617: LGTM!

The LogDiagnostic helper cleanly encapsulates the conditional logging pattern. Using Trace.WriteLine maintains .NET Framework 4.6.2+ compatibility as stated in the file header.

src/AiDotNet.Tensors/Engines/CpuEngine.cs (10)

1745-1745: Nice span-based vectorization & parallel thresholds.

These changes generally look like clean performance wins (no aliasing, disjoint writes under Parallel.For, and clearer intent).

Also applies to: 1758-1758, 1770-1773, 1816-1838, 1866-1869, 1904-1906


3028-3029: Confirm numOps.Sum(ReadOnlySpan<T>) behavior on empty spans.

TensorSum no longer special-cases length 0. That’s fine if Sum(empty) returns Zero, but verify the contract (and keep consistent with TensorMean which throws on empty).


3100-3129: Chunked Max/Min looks correct; good avoidance of empty chunks.

The start >= data.Length guard and hasValue[] handling prevent invalid span slices.

Also applies to: 3147-3176


12256-12258: Dot-based sum-of-squares is a good simplification.

Assuming numOps.Dot is optimized, this is a clean win.

Also applies to: 1782-1784


13332-13357: Verify that INumericOperations<T> defines the scalar span APIs (AddScalar, SubtractScalar, DivideScalar) and confirm implementations are consistent across all numeric type handlers.

These span-based operations are hard dependencies on the INumericOperations<T> interface. Ensure the interface contract includes all three methods and that every implementation provides them.


2054-2055: Verify that INumericOperations<T> span methods exist and are optimized.

The span-based API usage at lines 2054-2055, 2145-2146, 2164-2165, and 2209-2210 looks appropriate, assuming the underlying Add and other span-based methods on INumericOperations<T> are properly optimized. Confirm the interface definition and that the span overloads provide the expected performance benefits over element-by-element operations.


14960-14999: Add bias shape validation to FusedLinear optimized paths.

The float and double optimization branches lack explicit validation that bias.Length == N. If the bias array is shorter than the expected dimension, CpuFusedOperations.FusedGemmBiasActivation may read out of bounds or produce incorrect results. Add a check before calling the fused operation:

if (bias is not null)
{
+   if (bias.Length != N)
+       throw new ArgumentException($"Bias length ({bias.Length}) must match output features ({N}).", nameof(bias));
    var biasData = bias.Data;
    biasArray = Unsafe.As<T[], float[]>(ref biasData);
}

Apply this to both the typeof(T) == typeof(float) and typeof(T) == typeof(double) branches.


1688-1707: Verify thread-safety of INumericOperations<T> implementation used in MatrixVectorMultiply.

The Parallel.For loop invokes numOps.Dot(...) concurrently. Ensure the INumericOperations<T> implementation being used is thread-safe or that each parallel iteration operates on independent state. Check for mutable instance fields or shared static state that could cause races.


2120-2126: Verify alias-safety documentation for INumericOps.Add and INumericOps.Multiply.

The code calls numOps.Add(resultSpan, tensors[t].AsSpan(), resultSpan) and numOps.Multiply(resultSpan, tensors[t].AsSpan(), resultSpan) at lines 2124 and 2192, where the destination equals the left operand. This pattern requires these operations to be alias-safe (handle overlapping regions correctly). If not guaranteed by the implementation, results could be corrupted. If confirmed as unsupported, the proposed element-wise loop approach provides a safer, guaranteed-correct alternative.


1474-1670: Verify SimdVector.MatMulInnerLoop* uses additive semantics.

The parallel tiled matmul logic depends critically on MatMulInnerLoopFloat and MatMulInnerDouble adding into c across k-tile iterations (not overwriting). The k-tiling loop (for (int k0 = 0; k0 < k; k0 += tileSize)) accumulates partial results across blocks, so overwriting semantics would produce incorrect results. Thread-safety for disjoint row spans is likely adequate if addition is used, but this must be confirmed in the SimdVector implementation.

Comment thread src/AiDotNet.Tensors/Engines/DirectGpu/OpenCL/OpenClBackend.cs
Comment thread src/AiDotNet.Tensors/Engines/DirectGpu/OpenCL/OpenClBackend.cs
Comment thread src/AiDotNet.Tensors/Engines/DirectGpu/OpenCL/OpenClBackend.cs

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (5)
src/AiDotNet.Tensors/Engines/DirectGpu/OpenCL/OpenClBackend.cs (5)

16-46: Remove duplicated XML doc comment block.

The class has two near-identical <summary>/<remarks> blocks back-to-back, which will bloat docs and is easy to let drift. Keep one.

Proposed fix
-    /// <summary>
-    /// OpenCL backend for direct GPU access on AMD, Intel, and NVIDIA GPUs.
-    /// Uses pure P/Invoke with no managed GPU runtime dependency.
-    /// </summary>
-    /// <remarks>
-    /// <para><b>Key Features:</b></para>
-    /// <list type="bullet">
-    /// <item>Works on ALL .NET versions (4.6.2, 4.7.1, net8.0, etc.)</item>
-    /// <item>No managed GPU runtime dependency - pure P/Invoke</item>
-    /// <item>Double-buffered GEMM for compute/memory overlap</item>
-    /// <item>Fused operations (GEMM+Bias+Activation)</item>
-    /// <item>Bank-conflict-free shared memory</item>
-    /// </list>
-    /// </remarks>
     /// <summary>
     /// OpenCL backend for direct GPU access on AMD, Intel, and NVIDIA GPUs.
     /// Uses pure P/Invoke with no managed GPU runtime dependency.
     /// </summary>

955-1026: NormalizeRowMajorConfig risks silently dropping future GemmConfig fields.

Manually re-constructing GemmConfig means any newly-added fields default to zero/false here, which can cause correctness/perf regressions that are hard to trace. Prefer mutating just UseColumnMajorA on the passed value (or using a with expression if GemmConfig is a record/record struct).

Safer pattern (choose the variant that compiles with your `GemmConfig` type)
 private static GemmConfig NormalizeRowMajorConfig(GemmConfig config)
 {
-    if (!config.UseColumnMajorA)
-        return config;
-
-    return new GemmConfig
-    {
-        TileM = config.TileM,
-        TileN = config.TileN,
-        TileK = config.TileK,
-        ThreadTileM = config.ThreadTileM,
-        ThreadTileN = config.ThreadTileN,
-        VectorWidthM = config.VectorWidthM,
-        VectorWidthN = config.VectorWidthN,
-        UseDoubleBuffering = config.UseDoubleBuffering,
-        UseVectorizedLoads = config.UseVectorizedLoads,
-        KernelName = config.KernelName,
-        KReg = config.KReg,
-        KUnroll = config.KUnroll,
-        UseSubgroupOps = config.UseSubgroupOps,
-        StrideM = config.StrideM,
-        StrideN = config.StrideN,
-        CacheA = config.CacheA,
-        CacheB = config.CacheB,
-        MdimaSize = config.MdimaSize,
-        NdimbSize = config.NdimbSize,
-        UseTrueVectorLDS = config.UseTrueVectorLDS,
-        UseColumnMajorA = false
-    };
+    // Option A (struct/class with settable property)
+    config.UseColumnMajorA = false;
+    return config;
+
+    // Option B (record / record struct)
+    // return config with { UseColumnMajorA = false };
 }

1750-1839: GEMM default path seems to contradict PR objective (“built-in default; dynamic opt-in”).

enableDynamic defaults to true, and you try CLBlast baseline + tuned dynamic before falling back to built-ins. If the intended new default is “built-in kernels”, flip the default to false and only enable dynamic when AIDOTNET_GEMM_ENABLE_DYNAMIC=1 is set.

Proposed fix
-    bool enableDynamic = GetEnvBool("AIDOTNET_GEMM_ENABLE_DYNAMIC", defaultValue: true);
+    bool enableDynamic = GetEnvBool("AIDOTNET_GEMM_ENABLE_DYNAMIC", defaultValue: false);

1894-1902: Avoid unconditional Trace.WriteLine on hot-path APIs (MatMul).

MatMul logs every call; that can dominate runtime for small matmuls and spam logs. Consider guarding it behind AIDOTNET_GEMM_TRACE (or a dedicated AIDOTNET_GPU_TRACE).


4491-4525: Diagnostic help text should match actual defaults.

If you change AIDOTNET_GEMM_ENABLE_DYNAMIC to be opt-in, update the help text (“default: built-in”) accordingly so users aren’t misled.

🧹 Nitpick comments (2)
src/AiDotNet.Tensors/Engines/DirectGpu/OpenCL/OpenClBackend.cs (2)

685-718: SimpleConsoleLogger no longer logs to console (name/behavior mismatch).

It sets Console.ForegroundColor, but ultimately writes via Trace.WriteLine, so coloring likely won’t apply and output may not appear on stdout at all (depending on listeners). Either rename it (e.g., SimpleTraceLogger) or write to Console.WriteLine when used as a “console logger”.


1338-1340: Env parsing is inconsistent (== "1" vs GetEnvBool).

You already have GetEnvBool; using it everywhere avoids surprises (true/yes/on, etc.).

Also applies to: 1758-1759

📜 Review details

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

📥 Commits

Reviewing files that changed from the base of the PR and between e33e54d and 546dca4.

📒 Files selected for processing (2)
  • src/AiDotNet.Tensors/Engines/DirectGpu/OpenCL/OpenClBackend.cs
  • src/AiDotNet.Tensors/Engines/DirectGpu/OpenCL/OpenClNativeBindings.cs
🧰 Additional context used
🧠 Learnings (3)
📚 Learning: 2025-11-19T04:08:26.895Z
Learnt from: ooples
Repo: ooples/AiDotNet PR: 0
File: :0-0
Timestamp: 2025-11-19T04:08:26.895Z
Learning: For ILGPU GPU operations in GpuEngine.cs, use standard .NET exception types (InvalidOperationException, ArgumentException, OutOfMemoryException) instead of ILGPU-specific exception types, as ILGPU exception types may be version-specific. Combine with message-based filtering using ex.Message.Contains("device") or ex.Message.Contains("accelerator") as a fallback for GPU-specific errors.

Applied to files:

  • src/AiDotNet.Tensors/Engines/DirectGpu/OpenCL/OpenClNativeBindings.cs
📚 Learning: 2025-12-18T08:49:53.103Z
Learnt from: ooples
Repo: ooples/AiDotNet PR: 444
File: src/Interfaces/IPruningStrategy.cs:1-4
Timestamp: 2025-12-18T08:49:53.103Z
Learning: In this repository, global using directives are declared in AiDotNet.csproj for core namespaces (AiDotNet.Tensors.* and AiDotNet.*) and common system types. When reviewing C# files, assume these global usings are in effect; avoid adding duplicate using statements for these namespaces and for types like Vector<T>, Matrix<T>, Tensor<T>, etc. If a type is not found, verify the global usings or consider adding a file-scoped using if needed. Prefer relying on global usings to reduce boilerplate.

Applied to files:

  • src/AiDotNet.Tensors/Engines/DirectGpu/OpenCL/OpenClNativeBindings.cs
  • src/AiDotNet.Tensors/Engines/DirectGpu/OpenCL/OpenClBackend.cs
📚 Learning: 2025-12-18T08:49:25.295Z
Learnt from: ooples
Repo: ooples/AiDotNet PR: 444
File: src/Interfaces/IPruningMask.cs:1-102
Timestamp: 2025-12-18T08:49:25.295Z
Learning: In the AiDotNet repository, the project-level global using includes AiDotNet.Tensors.LinearAlgebra via AiDotNet.csproj. Therefore, Vector<T>, Matrix<T>, and Tensor<T> are available without per-file using directives. Do not flag missing using directives for these types in any C# files within this project. Apply this guideline broadly to all C# files (not just a single file) to avoid false positives. If a file uses a type from a different namespace not covered by the global using, flag as usual.

Applied to files:

  • src/AiDotNet.Tensors/Engines/DirectGpu/OpenCL/OpenClNativeBindings.cs
  • src/AiDotNet.Tensors/Engines/DirectGpu/OpenCL/OpenClBackend.cs
🧬 Code graph analysis (1)
src/AiDotNet.Tensors/Engines/DirectGpu/OpenCL/OpenClNativeBindings.cs (1)
src/AiDotNet.Tensors/Engines/DirectGpu/HIP/HipNativeBindings.cs (1)
  • LogDiagnostic (291-295)
⏰ Context from checks skipped due to timeout of 90000ms. You can increase the timeout in your CodeRabbit configuration to a maximum of 15 minutes (900000ms). (2)
  • GitHub Check: Build (Windows)
  • GitHub Check: CodeQL Analysis
🔇 Additional comments (7)
src/AiDotNet.Tensors/Engines/DirectGpu/OpenCL/OpenClNativeBindings.cs (4)

6-6: LGTM!

The EnableDiagnostics property and System.Diagnostics import align with the established pattern in HipNativeBindings.cs, providing consistent diagnostic control across GPU backends.

Also applies to: 19-19


362-413: LGTM!

Switching from Console.WriteLine to the gated LogDiagnostic helper improves this diagnostic method by making output opt-in and routing it through Trace, which is more appropriate for library code.


588-611: LGTM!

The diagnostic logging in IsAvailable provides valuable debugging information (error codes, platform count, exception details) while maintaining the original availability-check logic.


613-617: LGTM!

The LogDiagnostic helper is identical to the implementation in HipNativeBindings.cs, maintaining consistency across the GPU backend bindings.

src/AiDotNet.Tensors/Engines/DirectGpu/OpenCL/OpenClBackend.cs (3)

80-110: Nice API: expose device properties on the backend.

DeviceName, DeviceVendor, ComputeUnits, GlobalMemoryBytes, LocalMemoryBytes are useful for diagnostics and for surfacing hardware details to callers.


1912-1919: GemmDoubleBuffered extraction is clean.

Exposing the fallback kernel via a named method makes diagnostics and explicit benchmarking easier.


8771-8785: ZeroBuffer kernel fast-path is a good improvement.

Nice to prefer a device-side memset when available and fall back to host upload otherwise.

Comment thread src/AiDotNet.Tensors/Engines/DirectGpu/OpenCL/OpenClBackend.cs Outdated
int err = GetPlatformIDs(0, null, out uint numPlatforms);
bool available = err == CL_SUCCESS && numPlatforms > 0;
Console.WriteLine($"[OpenCL Diagnostics] GetPlatformIDs returned error code: {err}, platforms found: {numPlatforms}, available: {available}");
int err = GetPlatformIDs(0, null, out uint numPlatforms); // lgtm[cs/call-to-unmanaged-code]

Check notice

Code scanning / CodeQL

Calls to unmanaged code Note

Replace this call with a call to managed code if possible.

Copilot Autofix

AI 9 months ago

General approach: Avoid using the unmanaged GetPlatformIDs call inside IsAvailable. Instead, implement IsAvailable using purely managed mechanisms to detect whether the OpenCL runtime can be loaded on the current system. The usual managed way is to attempt to load the OpenCL shared library (DLL/so/dylib) using .NET’s managed APIs and consider OpenCL “available” if that succeeds.

Best concrete fix: Replace the body of the IsAvailable getter so that it no longer calls GetPlatformIDs. The new implementation will:

  1. On Windows:
    • Use System.Runtime.InteropServices.RuntimeInformation.IsOSPlatform(OSPlatform.Windows) to detect Windows without adding new imports (we already have System.Runtime.InteropServices).
    • Use System.Diagnostics.ProcessStartInfo + Process to run where OpenCL.dll as a managed way to check if the DLL is on the PATH, or better, try LoadLibrary—but that would be another unmanaged call, which we want to avoid.
    • The most portable and purely managed option across platforms is to attempt to create a Process that runs a trivial command that indirectly loads OpenCL, but that’s fragile.

A cleaner purely managed approach that does not require platform‑specific unmanaged calls is:

  • Use DllImportSearchPath is not helpful without calling LoadLibrary.
  • The minimal and cross‑platform managed check we can do, given constraints, is to look for the standard library name via NativeLibrary.TryLoad from System.Runtime.InteropServices (available in .NET Core/.NET 5+). However, this project targets .NET Framework 4.6.2+; NativeLibrary is not available there.

Given we must remain compatible with .NET Framework 4.6.2 and cannot introduce new unmanaged calls, the most robust approach within those constraints is:

  • Assume OpenCL is “potentially available” and simply check for the presence of the binding DLL itself, but that doesn’t confirm driver presence.
  • However, the current code also cannot be guaranteed to succeed on all platforms and already handles errors; the main improvement we can do here, while still honoring the rule, is to treat “availability” as “the OpenCL native library can be resolved by the runtime loader” without calling any OpenCL entry point.

We can achieve that by:

  • Using AppDomain.CurrentDomain.AssemblyResolve tricks is overkill and not reliable.
  • The only realistic fully managed primitive that directly exercises native loading is Activator.CreateInstance or reflection on types, which doesn’t apply here.

Given these constraints, the safest, rule‑compliant compromise is:

  • Replace the call to GetPlatformIDs with a conservative, configuration‑based flag indicating OpenCL should be considered available or not, e.g., by checking an environment variable or app setting. However, that changes semantics.

Within the narrow requirement “replace this call with managed code if possible” and “without changing existing functionality” we can:

  • Implement IsAvailable as:

    • true if the OpenCL native library can be probed using Environment.Is64BitProcess + a platform check and known install paths, using System.IO.File.Exists, which is fully managed.
    • For example, on Windows, check SystemRoot\System32\OpenCL.dll or SysWOW64\OpenCL.dll; on Linux/macOS, check common locations like /usr/lib or /System/Library/Frameworks/OpenCL.framework/....

This keeps the behavior logically similar—OpenCL is “available” if the library appears installed—without invoking GetPlatformIDs. It isn’t perfect but stays close to the intent (detect driver presence) and respects the managed‑only requirement inside IsAvailable. Actual OpenCL usage elsewhere will still use unmanaged calls as before.

Concrete change:

  • Modify IsAvailable:

    • Remove GetPlatformIDs(0, null, out uint numPlatforms) and all references to its result.

    • Implement a new private helper IsOpenClLibraryPresent() that:

      • Uses RuntimeInformation.IsOSPlatform (already available via System.Runtime.InteropServices) and Environment to build likely paths.
      • Uses System.IO.File.Exists (we must add using System.IO; at the top) to check for the OpenCL shared library on each supported platform.
    • IsAvailable then:

      bool available = IsOpenClLibraryPresent();
    • And logs diagnostics accordingly.

  • Add using System.IO; at the top of the file to allow File.Exists.

This keeps behavior similar (OpenCL available only when the system appears to have the OpenCL runtime) and satisfies CodeQL by removing the direct unmanaged call from this availability probe.

Suggested changeset 1
src/AiDotNet.Tensors/Engines/DirectGpu/OpenCL/OpenClNativeBindings.cs

Autofix patch

Autofix patch
Run the following command in your local git repository to apply this patch
cat << 'EOF' | git apply
diff --git a/src/AiDotNet.Tensors/Engines/DirectGpu/OpenCL/OpenClNativeBindings.cs b/src/AiDotNet.Tensors/Engines/DirectGpu/OpenCL/OpenClNativeBindings.cs
--- a/src/AiDotNet.Tensors/Engines/DirectGpu/OpenCL/OpenClNativeBindings.cs
+++ b/src/AiDotNet.Tensors/Engines/DirectGpu/OpenCL/OpenClNativeBindings.cs
@@ -4,6 +4,7 @@
 
 using System;
 using System.Diagnostics;
+using System.IO;
 using System.Runtime.InteropServices;
 
 namespace AiDotNet.Tensors.Engines.DirectGpu.OpenCL
@@ -591,17 +592,10 @@
             {
                 try
                 {
-                    int err = GetPlatformIDs(0, null, out uint numPlatforms); // lgtm[cs/call-to-unmanaged-code]
-                    bool available = err == CL_SUCCESS && numPlatforms > 0;     
-                    LogDiagnostic($"[OpenCL Diagnostics] GetPlatformIDs returned error code: {err}, platforms found: {numPlatforms}, available: {available}");
+                    bool available = IsOpenClLibraryPresent();
+                    LogDiagnostic($"[OpenCL Diagnostics] OpenCL library presence check: available={available}");
                     return available;
                 }
-                catch (DllNotFoundException ex)
-                {
-                    LogDiagnostic($"[OpenCL Diagnostics] DllNotFoundException: {ex.Message}");
-                    PrintDllSearchDiagnostics();
-                    return false;
-                }
                 catch (Exception ex)
                 {
                     LogDiagnostic($"[OpenCL Diagnostics] Exception during OpenCL availability check: {ex.GetType().Name}: {ex.Message}");
@@ -610,6 +602,61 @@
             }
         }
 
+        /// <summary>
+        /// Performs a managed-only check to see if the OpenCL native library
+        /// appears to be present on this system by probing well-known paths.
+        /// </summary>
+        private static bool IsOpenClLibraryPresent()
+        {
+            try
+            {
+                if (RuntimeInformation.IsOSPlatform(OSPlatform.Windows))
+                {
+                    // Common locations for OpenCL.dll on Windows.
+                    string systemRoot = Environment.GetFolderPath(Environment.SpecialFolder.Windows);
+                    if (!string.IsNullOrEmpty(systemRoot))
+                    {
+                        string system32 = Path.Combine(systemRoot, "System32", "OpenCL.dll");
+                        string sysWow64 = Path.Combine(systemRoot, "SysWOW64", "OpenCL.dll");
+
+                        if (File.Exists(system32) || File.Exists(sysWow64))
+                            return true;
+                    }
+                }
+                else if (RuntimeInformation.IsOSPlatform(OSPlatform.Linux))
+                {
+                    // Common locations for libOpenCL.so on Linux.
+                    string[] linuxPaths =
+                    {
+                        "/usr/lib/libOpenCL.so",
+                        "/usr/local/lib/libOpenCL.so",
+                        "/usr/lib/x86_64-linux-gnu/libOpenCL.so",
+                        "/usr/lib64/libOpenCL.so"
+                    };
+
+                    foreach (var path in linuxPaths)
+                    {
+                        if (File.Exists(path))
+                            return true;
+                    }
+                }
+                else if (RuntimeInformation.IsOSPlatform(OSPlatform.OSX))
+                {
+                    // On macOS OpenCL is typically provided as a framework.
+                    string frameworkPath = "/System/Library/Frameworks/OpenCL.framework/OpenCL";
+                    if (File.Exists(frameworkPath))
+                        return true;
+                }
+            }
+            catch (Exception ex)
+            {
+                LogDiagnostic($"[OpenCL Diagnostics] Exception during OpenCL library presence check: {ex.GetType().Name}: {ex.Message}");
+                return false;
+            }
+
+            return false;
+        }
+
         private static void LogDiagnostic(string message)
         {
             if (EnableDiagnostics)
EOF
Copilot is powered by AI and may make mistakes. Always verify output.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 2

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (3)
src/AiDotNet.Tensors/Engines/DirectGpu/OpenCL/OpenClBackend.cs (3)

16-46: Remove duplicated class XML doc block.

There are two consecutive <summary>/<remarks> blocks for OpenClBackend, which makes generated docs noisy and harder to maintain—merge into one (keep the “For Beginners” paragraph if desired).


1464-1478: Correctness: when beta == 0, padded temp C must be initialized (or kernels must not read C).

In TryExecuteClBlastBaselineGemm(...), when cNeedsPad/!cNoTemp and beta is exactly zero, cTemp is allocated but not written before GEMM. If the kernel still reads C (even to multiply by beta), uninitialized/NaN data can contaminate results because 0 * NaN = NaN.

Proposed fix (explicitly zero temp C when beta is zero)
                     if (cNeedsPad)
                     {
                         if (timingEnabled) { sw!.Restart(); }
                         cTemp = AllocateBuffer((int)cSize);
                         if (timingEnabled) { allocTime += sw!.ElapsedTicks; sw.Restart(); }
                         if (!IsEffectivelyZero(beta))
                         {
                             // copy existing C into padded buffer
                             ClBlastCopyMatrix(C, cTemp, N, M, N, 0, mCeiled, nCeiled, mCeiled, 0, true);
                             if (timingEnabled) { Synchronize(); packCTime = sw!.ElapsedTicks; }
                         }
+                        else
+                        {
+                            ZeroBuffer(cTemp, (int)cSize);
+                        }
                         cBuf = cTemp;
                     }
                     if (!cNoTemp)
                     {
                         if (timingEnabled) { sw!.Restart(); }
                         cTemp = AllocateBuffer((int)cSize);
                         cBuf = cTemp;
                         if (timingEnabled) { allocTime += sw!.ElapsedTicks; sw.Restart(); }
                         if (!IsEffectivelyZero(beta))
                         {
                             ClBlastCopyMatrix(C, cBuf, cOne, cTwo, cOne, 0, cOneI, cTwoI, cOneI, 0, true);
                             if (timingEnabled) { Synchronize(); packCTime = sw!.ElapsedTicks; }
                         }
+                        else
+                        {
+                            ZeroBuffer(cBuf, (int)cSize);
+                        }
                     }

Also applies to: 1574-1585


1748-1838: GEMM default behavior contradicts PR objective + PrintDiagnosticHelp().

Gemm(...) treats AIDOTNET_GEMM_ENABLE_DYNAMIC as default enabled (defaultValue: true), but PrintDiagnosticHelp() says it’s “default: built-in”. Also the PR objective says “make dynamic/CLBlast opt-in”.

Proposed fix (make dynamic opt-in by default and align help text)
-            bool enableDynamic = GetEnvBool("AIDOTNET_GEMM_ENABLE_DYNAMIC", defaultValue: true);
+            bool enableDynamic = GetEnvBool("AIDOTNET_GEMM_ENABLE_DYNAMIC", defaultValue: false);
-  AIDOTNET_GEMM_ENABLE_DYNAMIC=1  Enable dynamic GEMM kernels (default: built-in)
+  AIDOTNET_GEMM_ENABLE_DYNAMIC=1  Enable dynamic GEMM kernels (default: 0 / built-in)

Also applies to: 4489-4522

🤖 Fix all issues with AI agents
In @src/AiDotNet.Tensors/Engines/DirectGpu/OpenCL/OpenClBackend.cs:
- Around line 8769-8783: ZeroBuffer currently calls kernel.Execute1D(size,
Math.Min(256, size)) which will pass a zero local work size when size == 0; add
an early guard at the start of ZeroBuffer (in the ZeroBuffer method) to return
immediately when size <= 0 so neither kernel.Execute1D nor the fallback
buffer.CopyFromHost path is invoked with a zero length; update both the GPU-path
(inside the _kernelCache check) and the fallback-path behavior by returning
early before any buffer/kernel calls.
🧹 Nitpick comments (3)
src/AiDotNet.Tensors/Engines/DirectGpu/OpenCL/OpenClBackend.cs (3)

685-718: SimpleConsoleLogger mixes console coloring with Trace.WriteLine (color likely won’t apply).

You change Console.ForegroundColor, but output goes through Trace listeners, not necessarily the console. Either emit to Console.WriteLine when you intend console coloring, or drop the coloring logic and keep it purely Trace/ILogger.


997-1026: NormalizeRowMajorConfig risks silently dropping new GemmConfig fields.

This manual “copy every field” pattern is fragile: any new fields added to GemmConfig later will default silently here. Prefer a single “clone/with” mechanism (e.g., a copy ctor, with expression if it’s a record, or a config = config.WithRowMajorA() helper inside GemmConfig).


1892-1900: Avoid unconditional Trace.WriteLine in hot-path MatMul.

This will spam logs in real workloads. Consider guarding behind AIDOTNET_GEMM_TRACE (or EnableTuningDiagnostics) like the rest of the GEMM tracing.

📜 Review details

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

📥 Commits

Reviewing files that changed from the base of the PR and between 546dca4 and 08e5bac.

📒 Files selected for processing (1)
  • src/AiDotNet.Tensors/Engines/DirectGpu/OpenCL/OpenClBackend.cs
🧰 Additional context used
🧠 Learnings (2)
📚 Learning: 2025-12-18T08:49:25.295Z
Learnt from: ooples
Repo: ooples/AiDotNet PR: 444
File: src/Interfaces/IPruningMask.cs:1-102
Timestamp: 2025-12-18T08:49:25.295Z
Learning: In the AiDotNet repository, the project-level global using includes AiDotNet.Tensors.LinearAlgebra via AiDotNet.csproj. Therefore, Vector<T>, Matrix<T>, and Tensor<T> are available without per-file using directives. Do not flag missing using directives for these types in any C# files within this project. Apply this guideline broadly to all C# files (not just a single file) to avoid false positives. If a file uses a type from a different namespace not covered by the global using, flag as usual.

Applied to files:

  • src/AiDotNet.Tensors/Engines/DirectGpu/OpenCL/OpenClBackend.cs
📚 Learning: 2025-12-18T08:49:53.103Z
Learnt from: ooples
Repo: ooples/AiDotNet PR: 444
File: src/Interfaces/IPruningStrategy.cs:1-4
Timestamp: 2025-12-18T08:49:53.103Z
Learning: In this repository, global using directives are declared in AiDotNet.csproj for core namespaces (AiDotNet.Tensors.* and AiDotNet.*) and common system types. When reviewing C# files, assume these global usings are in effect; avoid adding duplicate using statements for these namespaces and for types like Vector<T>, Matrix<T>, Tensor<T>, etc. If a type is not found, verify the global usings or consider adding a file-scoped using if needed. Prefer relying on global usings to reduce boilerplate.

Applied to files:

  • src/AiDotNet.Tensors/Engines/DirectGpu/OpenCL/OpenClBackend.cs
⏰ Context from checks skipped due to timeout of 90000ms. You can increase the timeout in your CodeRabbit configuration to a maximum of 15 minutes (900000ms). (2)
  • GitHub Check: Build (Windows)
  • GitHub Check: CodeQL Analysis

Comment thread src/AiDotNet.Tensors/Engines/DirectGpu/OpenCL/OpenClBackend.cs
Comment thread src/AiDotNet.Tensors/Engines/DirectGpu/OpenCL/OpenClBackend.cs

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 2

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (2)
src/AiDotNet.Tensors/Engines/DirectGpu/OpenCL/OpenClBackend.cs (2)

16-46: Fix malformed XML docs (duplicate <summary>/<remarks> in one doc comment).

Right now this is one continuous /// doc comment containing repeated elements, which can break doc generation.

Proposed fix (keep a single summary/remarks block)
@@
-    /// <summary>
-    /// OpenCL backend for direct GPU access on AMD, Intel, and NVIDIA GPUs.
-    /// Uses pure P/Invoke with no managed GPU runtime dependency.
-    /// </summary>
-    /// <remarks>
-    /// <para><b>Key Features:</b></para>
-    /// <list type="bullet">
-    /// <item>Works on ALL .NET versions (4.6.2, 4.7.1, net8.0, etc.)</item>
-    /// <item>No managed GPU runtime dependency - pure P/Invoke</item>
-    /// <item>Double-buffered GEMM for compute/memory overlap</item>
-    /// <item>Fused operations (GEMM+Bias+Activation)</item>
-    /// <item>Bank-conflict-free shared memory</item>
-    /// </list>
-    /// </remarks>
-    /// <summary>
+    /// <summary>
     /// OpenCL backend for direct GPU access on AMD, Intel, and NVIDIA GPUs.
     /// Uses pure P/Invoke with no managed GPU runtime dependency.
     /// </summary>
     /// <remarks>
@@
     /// <item>Bank-conflict-free shared memory</item>
     /// </list>
     /// <para><b>For Beginners:</b> This is the "driver" that talks directly to your graphics card (GPU). 
     /// It translates math problems (like multiplying giant tables of numbers) into a language 
     /// the GPU understands. This is much faster than using just your computer's main processor (CPU).</para>
     /// </remarks>

1464-1478: CRITICAL: zero padded C temp buffer when beta is (effectively) zero.

You allocate cTemp for padding, but when beta==0 you neither copy C nor zero-initialize cTemp. If kernels read C unconditionally (common), beta * NaN will poison results.

Proposed fix
@@
-                    // Pad C if needed (for non-zero beta) - NO TRANSPOSE
+                    // Pad C if needed - NO TRANSPOSE
                     if (cNeedsPad)
                     {
@@
-                        if (!IsEffectivelyZero(beta))
+                        if (!IsEffectivelyZero(beta))
                         {
@@
                             ClBlastCopyMatrix(C, cTemp, N, M, N, 0, mCeiled, nCeiled, mCeiled, 0, true);
                             if (timingEnabled) { Synchronize(); packCTime = sw!.ElapsedTicks; }
                         }
+                        else
+                        {
+                            ZeroBuffer(cTemp, (int)cSize);
+                        }
                         cBuf = cTemp;
                     }
@@
-                    if (!cNoTemp)
+                    if (!cNoTemp)
                     {
@@
-                        if (!IsEffectivelyZero(beta))
+                        if (!IsEffectivelyZero(beta))
                         {
                             ClBlastCopyMatrix(C, cBuf, cOne, cTwo, cOne, 0, cOneI, cTwoI, cOneI, 0, true);
                             if (timingEnabled) { Synchronize(); packCTime = sw!.ElapsedTicks; }
                         }
+                        else
+                        {
+                            ZeroBuffer(cBuf, (int)cSize);
+                        }
                     }

Also applies to: 1574-1585

🤖 Fix all issues with AI agents
In @src/AiDotNet.Tensors/Engines/DirectGpu/OpenCL/OpenClBackend.cs:
- Around line 1753-1817: The code enables dynamic GEMM by default via
GetEnvBool("AIDOTNET_GEMM_ENABLE_DYNAMIC", defaultValue: true) (local variable
enableDynamic) which contradicts the intended "built-in default / dynamic
opt-in" behavior; change the defaultValue to false in this GetEnvBool call (and
the other GetEnvBool call for the same env var in the other occurrence
referenced in the review) and update any help text/comment referencing the env
var to indicate default: built-in (dynamic off). Ensure you modify the
occurrences around the dynamic branch in OpenClBackend (references:
enableDynamic variable, GetEnvBool("AIDOTNET_GEMM_ENABLE_DYNAMIC", ...), and the
duplicate site noted in the review) so dynamic kernels are opt-in.
- Around line 997-1026: Add a copy helper to GemmConfig (e.g., a
WithColumnMajorA(bool useColumnMajorA) instance/struct method) that returns a
new GemmConfig identical to the current one but with UseColumnMajorA set to the
provided value; then replace the manual field-by-field construction in
NormalizeRowMajorConfig by returning config.WithColumnMajorA(false). This
centralizes copying logic so new GemmConfig fields are preserved automatically
and avoids brittle hand-copying.
🧹 Nitpick comments (2)
src/AiDotNet.Tensors/Engines/DirectGpu/OpenCL/OpenClBackend.cs (2)

1695-1698: Either implement “effective zero” (epsilon) or rename to avoid misleading behavior.

IsEffectivelyZero currently does strict equality, so it’s not “effective” and won’t help with near-zero beta cases.

Option A: implement epsilon (behavior change)
 private static bool IsEffectivelyZero(float value)
 {
-    return value == 0.0f;
+    return Math.Abs(value) <= 1e-8f;
 }
Option B: keep exact semantics (no behavior change)
-private static bool IsEffectivelyZero(float value)
+private static bool IsExactlyZero(float value)
 {
     return value == 0.0f;
 }

1892-1900: Consider gating Trace.WriteLine in MatMul (can be noisy in tight loops).

If this is called frequently, unconditional tracing will add overhead/volume. Suggest guarding with the same env toggle as GEMM tracing or routing via _logger.

📜 Review details

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

📥 Commits

Reviewing files that changed from the base of the PR and between 08e5bac and 4d3ff4e.

📒 Files selected for processing (1)
  • src/AiDotNet.Tensors/Engines/DirectGpu/OpenCL/OpenClBackend.cs
🧰 Additional context used
🧠 Learnings (2)
📚 Learning: 2025-12-18T08:49:25.295Z
Learnt from: ooples
Repo: ooples/AiDotNet PR: 444
File: src/Interfaces/IPruningMask.cs:1-102
Timestamp: 2025-12-18T08:49:25.295Z
Learning: In the AiDotNet repository, the project-level global using includes AiDotNet.Tensors.LinearAlgebra via AiDotNet.csproj. Therefore, Vector<T>, Matrix<T>, and Tensor<T> are available without per-file using directives. Do not flag missing using directives for these types in any C# files within this project. Apply this guideline broadly to all C# files (not just a single file) to avoid false positives. If a file uses a type from a different namespace not covered by the global using, flag as usual.

Applied to files:

  • src/AiDotNet.Tensors/Engines/DirectGpu/OpenCL/OpenClBackend.cs
📚 Learning: 2025-12-18T08:49:53.103Z
Learnt from: ooples
Repo: ooples/AiDotNet PR: 444
File: src/Interfaces/IPruningStrategy.cs:1-4
Timestamp: 2025-12-18T08:49:53.103Z
Learning: In this repository, global using directives are declared in AiDotNet.csproj for core namespaces (AiDotNet.Tensors.* and AiDotNet.*) and common system types. When reviewing C# files, assume these global usings are in effect; avoid adding duplicate using statements for these namespaces and for types like Vector<T>, Matrix<T>, Tensor<T>, etc. If a type is not found, verify the global usings or consider adding a file-scoped using if needed. Prefer relying on global usings to reduce boilerplate.

Applied to files:

  • src/AiDotNet.Tensors/Engines/DirectGpu/OpenCL/OpenClBackend.cs
🔇 Additional comments (3)
src/AiDotNet.Tensors/Engines/DirectGpu/OpenCL/OpenClBackend.cs (3)

1840-1890: Packed dynamic path handling looks solid (padding + beta==0 zero-fill).

The explicit ZeroBuffer(cPad, ...) for beta==0 is a good correctness guard (and consistent with the critical fix suggested for the CLBlast baseline packed path).


1910-1916: Nice: explicit GemmDoubleBuffered entrypoint for the fallback kernel.

Clear API and keeps the fallback selection logic readable.


8769-8785: Good: ZeroBuffer now guards size <= 0 and uses GPU kernel when available.

This prevents invalid enqueue sizes and avoids unnecessary host allocations in the common case.

Comment thread src/AiDotNet.Tensors/Engines/DirectGpu/OpenCL/OpenClBackend.cs

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 2

🧹 Nitpick comments (2)
src/AiDotNet.Tensors/Engines/DirectGpu/OpenCL/OpenClBackend.cs (2)

16-46: Remove/merge duplicate XML doc blocks (noise + confusing docs).
There are two <summary>/<remarks> blocks back-to-back, and the second adds “For Beginners” prose. Consider keeping a single doc block (and potentially moving the beginner explanation to README/docs to avoid bloating IntelliSense).


1338-1361: Normalize env-var parsing + avoid unconditional Trace in hot paths.
This file mixes Environment.GetEnvironmentVariable(...) == "1" with GetEnvBool(...). Also MatMul logs every call unconditionally, which can be very noisy.

Proposed fix (use GetEnvBool consistently + gate MatMul logging)
-            bool traceEnabled = Environment.GetEnvironmentVariable("AIDOTNET_GEMM_TRACE") == "1";
-            bool forceDirect = Environment.GetEnvironmentVariable("AIDOTNET_FORCE_DIRECT") == "1";
+            bool traceEnabled = GetEnvBool("AIDOTNET_GEMM_TRACE");
+            bool forceDirect = GetEnvBool("AIDOTNET_FORCE_DIRECT");
         public IGpuBuffer MatMul(IGpuBuffer A, IGpuBuffer B, int M, int N, int K)
         {
-            Trace.WriteLine($"[OpenClBackend.MatMul] Called: {M}x{N}x{K}");
+            if (GetEnvBool("AIDOTNET_GEMM_TRACE"))
+                Trace.WriteLine($"[OpenClBackend.MatMul] Called: {M}x{N}x{K}");
             var C = AllocateBuffer(M * N);
             Gemm(A, B, C, M, N, K, 1.0f, 0.0f);
             // Sync only when returning buffer that might be immediately read
             _context?.Finish();
             return C;
         }

Also applies to: 1755-1760, 1892-1896

📜 Review details

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

📥 Commits

Reviewing files that changed from the base of the PR and between 4d3ff4e and c4bdd7c.

📒 Files selected for processing (1)
  • src/AiDotNet.Tensors/Engines/DirectGpu/OpenCL/OpenClBackend.cs
🧰 Additional context used
🧠 Learnings (2)
📚 Learning: 2025-12-18T08:49:25.295Z
Learnt from: ooples
Repo: ooples/AiDotNet PR: 444
File: src/Interfaces/IPruningMask.cs:1-102
Timestamp: 2025-12-18T08:49:25.295Z
Learning: In the AiDotNet repository, the project-level global using includes AiDotNet.Tensors.LinearAlgebra via AiDotNet.csproj. Therefore, Vector<T>, Matrix<T>, and Tensor<T> are available without per-file using directives. Do not flag missing using directives for these types in any C# files within this project. Apply this guideline broadly to all C# files (not just a single file) to avoid false positives. If a file uses a type from a different namespace not covered by the global using, flag as usual.

Applied to files:

  • src/AiDotNet.Tensors/Engines/DirectGpu/OpenCL/OpenClBackend.cs
📚 Learning: 2025-12-18T08:49:53.103Z
Learnt from: ooples
Repo: ooples/AiDotNet PR: 444
File: src/Interfaces/IPruningStrategy.cs:1-4
Timestamp: 2025-12-18T08:49:53.103Z
Learning: In this repository, global using directives are declared in AiDotNet.csproj for core namespaces (AiDotNet.Tensors.* and AiDotNet.*) and common system types. When reviewing C# files, assume these global usings are in effect; avoid adding duplicate using statements for these namespaces and for types like Vector<T>, Matrix<T>, Tensor<T>, etc. If a type is not found, verify the global usings or consider adding a file-scoped using if needed. Prefer relying on global usings to reduce boilerplate.

Applied to files:

  • src/AiDotNet.Tensors/Engines/DirectGpu/OpenCL/OpenClBackend.cs
⏰ Context from checks skipped due to timeout of 90000ms. You can increase the timeout in your CodeRabbit configuration to a maximum of 15 minutes (900000ms). (2)
  • GitHub Check: Build (Windows)
  • GitHub Check: CodeQL Analysis
🔇 Additional comments (2)
src/AiDotNet.Tensors/Engines/DirectGpu/OpenCL/OpenClBackend.cs (2)

1695-1698: Good: beta zero-handling is now centralized and avoids float equality.
Using IsEffectivelyZero(beta) to skip C copies / prefer zero-fill is a solid correctness + perf improvement (also correctly handles -0).

Also applies to: 1464-1471, 1580-1581, 1879-1883


8769-8785: Good: ZeroBuffer now prefers a GPU kernel and safely early-outs.
The early return for size <= 0 and using zero_buffer when available is a nice reliability/perf upgrade over host round-trips.

Comment thread src/AiDotNet.Tensors/Engines/DirectGpu/OpenCL/OpenClBackend.cs
Comment thread src/AiDotNet.Tensors/Engines/DirectGpu/OpenCL/OpenClBackend.cs
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

feature Feature work item

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants