Summary
Benchmarks show AiDotNet GPU ops are significantly slower than TorchSharp CPU and TensorFlow.NET CPU for micro-ops (10x-235x slower on elementwise and reductions, ~9x-20x on MatMul/Conv2D). This suggests launch + transfer overheads dominate and kernel implementations are not yet competitive.
Evidence
- Benchmark docs: docs/GPU_OPTIMIZATION_REPORT.md
- TorchSharp run (CPU only):
AiDotNetBenchmarkTests/BenchmarkDotNet.Artifacts/results/AiDotNetBenchmarkTests.TorchSharpComparisonBenchmarks-report-github.md
- TensorFlow.NET run (CPU oneDNN):
BENCHMARKS.md
Hypotheses / Bottlenecks
- Per-op allocations and CPU<->GPU transfers (
Allocate1D, CopyFromCPU, CopyToCPU) dominate.
- Frequent
_accelerator.Synchronize() and _gpuLock serialize work and prevent overlap.
- Kernel dispatch overhead and lack of fusion for common op chains.
- Naive GEMM/Conv2D kernels lacking tiling/shared memory.
- High managed allocations in hot paths (Gen0/Gen1 pressure in BDN).
Proposed work (prioritized)
P0
P1
P2
Notes
TorchSharp CUDA is installed, but torch.cuda.is_available() is false on the current AMD GPU (gfx902). GPU-vs-GPU comparisons should be re-run on NVIDIA once available.
Summary
Benchmarks show AiDotNet GPU ops are significantly slower than TorchSharp CPU and TensorFlow.NET CPU for micro-ops (10x-235x slower on elementwise and reductions, ~9x-20x on MatMul/Conv2D). This suggests launch + transfer overheads dominate and kernel implementations are not yet competitive.
Evidence
AiDotNetBenchmarkTests/BenchmarkDotNet.Artifacts/results/AiDotNetBenchmarkTests.TorchSharpComparisonBenchmarks-report-github.mdBENCHMARKS.mdHypotheses / Bottlenecks
Allocate1D,CopyFromCPU,CopyToCPU) dominate._accelerator.Synchronize()and_gpuLockserialize work and prevent overlap.Proposed work (prioritized)
P0
P1
LoadAutoGroupedStreamKernel).P2
Notes
TorchSharp CUDA is installed, but
torch.cuda.is_available()is false on the current AMD GPU (gfx902). GPU-vs-GPU comparisons should be re-run on NVIDIA once available.