Skip to content

cuda : enable i16 and i32 for DUP - #28897

Merged
am17an merged 2 commits into
ggml-org:masterfrom
amankarki151:cuda-dup-i16-i32
Sep 15, 2026
Merged

am17an merged 2 commits into
ggml-org:masterfrom
amankarki151:cuda-dup-i16-i32

Conversation

@amankarki151

Copy link
Copy Markdown
Contributor

Overview

i32 and i16 DUP ops were silently falling back to the CPU. The i32 path already existed under the hood but was blocked by the support gate. This PR adds the missing i16 branch and fixes the gate to allow both.

Additional information

Testing: test-backend-ops -o DUP passes 10/10 on dual T4s. The full suite passes with zero regressions (16097/16097).

test-backend-ops -o DUP
ggml_cuda_init: found 2 CUDA devices (Total VRAM: 29823 MiB):
  Device 0: Tesla T4, compute capability 7.5, VMM: no, VRAM: 14911 MiB
  Device 1: Tesla T4, compute capability 7.5, VMM: no, VRAM: 14911 MiB
Testing 3 devices

Backend 1/3: CUDA0
  Device description: Tesla T4
  Device memory: 14911 MB (14806 MB free)

  DUP(type=f32,ne=[10,10,20,1]): OK
  DUP(type=f16,ne=[10,10,20,1]): OK
  DUP(type=i32,ne=[10,10,20,1]): OK
  DUP(type=i16,ne=[10,10,20,1]): OK
  DUP(type=f32,ne=[10,10,5,1],permute=[0,2,1,3]): OK
  DUP(type=f16,ne=[10,10,5,1],permute=[0,2,1,3]): OK
  DUP(type=f32,ne=[10,10,5,1],permute=[1,0,2,3]): OK
  DUP(type=f16,ne=[10,10,5,1],permute=[1,0,2,3]): OK
  DUP(type=i16,ne=[10,8,3,1],permute=[0,2,1,3]): OK
  DUP(type=i16,ne=[10,8,3,1],permute=[1,2,0,3]): OK
  10/10 tests passed
  Backend CUDA0: OK
Backend 2/3: CUDA1
  Device description: Tesla T4
  Device memory: 14911 MB (14806 MB free)

  DUP(type=f32,ne=[10,10,20,1]): OK
  DUP(type=f16,ne=[10,10,20,1]): OK
  DUP(type=i32,ne=[10,10,20,1]): OK
  DUP(type=i16,ne=[10,10,20,1]): OK
  DUP(type=f32,ne=[10,10,5,1],permute=[0,2,1,3]): OK
  DUP(type=f16,ne=[10,10,5,1],permute=[0,2,1,3]): OK
  DUP(type=f32,ne=[10,10,5,1],permute=[1,0,2,3]): OK
  DUP(type=f16,ne=[10,10,5,1],permute=[1,0,2,3]): OK
  DUP(type=i16,ne=[10,8,3,1],permute=[0,2,1,3]): OK
  DUP(type=i16,ne=[10,8,3,1],permute=[1,2,0,3]): OK
  10/10 tests passed
  Backend CUDA1: OK
Backend 3/3: CPU
  Skipping CPU backend
3/3 backends passed
OK

Docs: Manually updated the 4 DUP rows in docs/ops/CUDA.csv. I skipped a full regen since running it on a T4 (compute 7.5) would incorrectly downgrade rows for ops that require newer hardware.

Ref: #14909

Requirements

  • I have read and agree with the contributing guidelines
  • AI usage disclosure: YES - used AI to help navigate the ggml-cuda backend.

@amankarki151
amankarki151 requested a review from a team as a code owner September 14, 2026 12:36
@github-actions github-actions Bot added documentation Improvements or additions to documentation ggml changes relating to the ggml tensor library for machine learning CUDA Related to the CUDA backend labels Sep 14, 2026
@am17an
am17an merged commit 4c9233c into ggml-org:master Sep 15, 2026
23 of 25 checks passed
@amankarki151
amankarki151 deleted the cuda-dup-i16-i32 branch September 15, 2026 04:49
pl752 pushed a commit to pl752/llama.cpp that referenced this pull request Sep 15, 2026
* cuda : enable i16 and i32 for DUP

* docs : update ops table for DUP on CUDA
dzannotti added a commit to halo-box/llama.cpp that referenced this pull request Sep 15, 2026
* upstream/master: (72 commits)
  HIP: Enable AllReduce for ROCm (ggml-org#27825)
  opencl: choose the MoE expert matmul by batch size for speculative decoding/MTP (ggml-org#27637)
  ci: build MUSA for only 1 arch (ggml-org#28944)
  docs: Rule of thumb for AI review time [no ci] (ggml-org#28945)
  rpc : hash-cache only weights (ggml-org#28789)
  cuda: support row-contiguous SUM_ROWS (ggml-org#26308)
  models : move build_arch_graph() after graph() template specialization (ggml-org#28934)
  vulkan: support sparse Flash Attention (ggml-org#28105)
  OpenVINO: optimize stateful decode and GPU MoE inference (ggml-org#28638)
  opencl: add generic ssm_scan (ggml-org#28881)
  ci: bump kleidiai runners from 22.04 to 24.04 (ggml-org#28885)
  metal : add FA kernels for HSK=96, HSV=64 (MiniCPM3) (ggml-org#28599)
  ci: Bump CUDA Windows x64 builds to 13.4.1 (ggml-org#28930)
  ci : fix android release (ggml-org#28936)
  cuda : enable i16 and i32 for DUP (ggml-org#28897)
  cmake : use PROJECT_SOURCE_DIR instead of CMAKE_SOURCE_DIR (ggml-org#28771)
  webui: stop re-probing disabled /tools endpoint on every message (ggml-org#28646)
  ci : reuse build tag name when used instead of safe one (ggml-org#28911)
  CI: hip-quality-check: ignore spill added in bfdc321 (ggml-org#28909)
  HIP: fattn-mma: use fp32 accumulation on MFMA devices (ggml-org#28576)
  ...
quimmedes pushed a commit to quimmedes/cafe-llama.cpp that referenced this pull request Sep 16, 2026
* cuda : enable i16 and i32 for DUP

* docs : update ops table for DUP on CUDA
zsogitbe pushed a commit to zsogitbe/llama.cpp that referenced this pull request Sep 17, 2026
* cuda : enable i16 and i32 for DUP

* docs : update ops table for DUP on CUDA
adromir pushed a commit to adromir/llama-cpp-turboquant that referenced this pull request Sep 17, 2026
* cuda : enable i16 and i32 for DUP

* docs : update ops table for DUP on CUDA
@BrewTestBot BrewTestBot mentioned this pull request Sep 23, 2026
1 task done
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CUDA Related to the CUDA backend documentation Improvements or additions to documentation ggml changes relating to the ggml tensor library for machine learning

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants