Skip to content

GPU performance gap vs TorchSharp/TensorFlow (Phase B) #682

Description

@ooples

Summary

Benchmarks show AiDotNet GPU ops are significantly slower than TorchSharp CPU and TensorFlow.NET CPU for micro-ops (10x-235x slower on elementwise and reductions, ~9x-20x on MatMul/Conv2D). This suggests launch + transfer overheads dominate and kernel implementations are not yet competitive.

Evidence

  • Benchmark docs: docs/GPU_OPTIMIZATION_REPORT.md
  • TorchSharp run (CPU only): AiDotNetBenchmarkTests/BenchmarkDotNet.Artifacts/results/AiDotNetBenchmarkTests.TorchSharpComparisonBenchmarks-report-github.md
  • TensorFlow.NET run (CPU oneDNN): BENCHMARKS.md

Hypotheses / Bottlenecks

  • Per-op allocations and CPU<->GPU transfers (Allocate1D, CopyFromCPU, CopyToCPU) dominate.
  • Frequent _accelerator.Synchronize() and _gpuLock serialize work and prevent overlap.
  • Kernel dispatch overhead and lack of fusion for common op chains.
  • Naive GEMM/Conv2D kernels lacking tiling/shared memory.
  • High managed allocations in hot paths (Gen0/Gen1 pressure in BDN).

Proposed work (prioritized)

P0

  • Add GPU-resident tensor paths and minimize host/device transfers across chained ops.
  • Use pooled GPU buffers in all hot paths and reuse outputs where safe.
  • Reduce or defer synchronization and shrink lock scope to allow async execution.

P1

  • Ensure kernel caching covers all hot paths (remove runtime LoadAutoGroupedStreamKernel).
  • Add kernel fusion for common op sequences (bias+activation, add+activation, etc.).
  • Optimize GEMM/Conv2D with tiling/shared memory and vectorized loads.

P2

  • Optional vendor lib interop (cuBLAS/cuDNN or rocBLAS/MIOpen) for GEMM/Conv2D.
  • Recompute adaptive thresholds from size-sweep benchmarks.
  • Add profiling to split kernel vs transfer time in benchmarks.

Notes

TorchSharp CUDA is installed, but torch.cuda.is_available() is false on the current AMD GPU (gfx902). GPU-vs-GPU comparisons should be re-run on NVIDIA once available.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    featureFeature work iteminferenceRuntime optimizations

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions