Skip to content

Eval bug: generation performance 3x slower on CUDA compared to llama.cpp #67

Description

@sakharovaan

Name and Version

$ ./build/bin/llama-cli --version
version: 10102 (85e22ea)
built with GNU 15.2.0 for Linux x86_64

Operating systems

Linux

GGML backends

CUDA

Hardware

AMD Ryzen AI Max+ 395
NVIDIA RTX PRO 6000 Blackwell Workstation Edition

Models

https://huggingface.co/unsloth/MiniMax-M2.7-GGUF/tree/main/UD-Q3_K_S

Problem description & steps to reproduce

Use CUDA 13.3 and ubuntu 26.04
Clone upstream llama.cpp and beellama.cpp

Compile both using:

cmake -B build -DGGML_CUDA=ON -DGGML_NATIVE=ON -DGGML_CUDA_FA_ALL_QUANTS=ON -DCMAKE_BUILD_TYPE=Release -DLLAMA_BUILD_UI=OFF -DLLAMA_BUILD_WEBUI=OFF -DCMAKE_CUDA_ARCHITECTURES=120 -DCUDA_PATH=/usr/local/cuda-13.3 -DGGML_CUDA_FA=ON
cmake --build build -j

Then benchmark compiled versions (using parameters below)

Prefill speed is mostly the same, but generation speed is about 3 times slower on beellama.cpp on same parameters. Why? What should I tweak?

First Bad Commit

No response

Relevant log output

Logs
user@faex:~/beellama.cpp$ GGML_VK_VISIBLE_DEVICES=-1 ~/llama.cpp/build/bin/llama-bench -m /apps/models/unsloth/MiniMax-M2.7-GGUF/UD-Q3_K_S/MiniMax-M2.7-UD-Q3_K_S-00001-of-00003.gguf -p 1000 -d 60000 -b 1024 -ub
 1024 -ctk q5_0 -ctv q4_1 -t 4 -ngl 99 -fa on -dio on --device CUDA0
ggml_cuda_init: found 1 CUDA devices (Total VRAM: 97288 MiB):
  Device 0: NVIDIA RTX PRO 6000 Blackwell Workstation Edition, compute capability 12.0, VMM: yes, VRAM: 97288 MiB
ggml_vulkan: Invalid device index 18446744073709551615 in GGML_VK_VISIBLE_DEVICES.
| model                          |       size |     params | backend    | ngl | threads | n_batch | n_ubatch | type_k | type_v |  fa | dev          |            test |                  t/s |
| ------------------------------ | ---------: | ---------: | ---------- | --: | ------: | ------: | -------: | -----: | -----: | --: | ------------ | --------------: | -------------------: |
| minimax-m2 230B.A10B Q3_K - Small |  87.20 GiB |   228.69 B | CUDA       |  99 |       4 |    1024 |     1024 |   q5_0 |   q4_1 |   1 | CUDA0        | pp1000 @ d60000 |        984.10 ± 0.94 |
| minimax-m2 230B.A10B Q3_K - Small |  87.20 GiB |   228.69 B | CUDA       |  99 |       4 |    1024 |     1024 |   q5_0 |   q4_1 |   1 | CUDA0        |  tg128 @ d60000 |         47.38 ± 0.05 |

build: ebc10770a (9614)
user@faex:~/beellama.cpp$ GGML_VK_VISIBLE_DEVICES=-1 ~/beellama.cpp/build/bin/llama-bench -m /apps/models/unsloth/MiniMax-M2.7-GGUF/UD-Q3_K_S/MiniMax-M2.7-UD-Q3_K_S-00001-of-00003.gguf -p 1000 -d 60000 -b 1024
-ub 1024 -ctk q5_0 -ctv q4_1 -t 4 -ngl 99 -fa on -dio on --device CUDA0
ggml_cuda_init: found 1 CUDA devices (Total VRAM: 97288 MiB):
  Device 0: NVIDIA RTX PRO 6000 Blackwell Workstation Edition, compute capability 12.0, VMM: yes, VRAM: 97288 MiB
| model                          |       size |     params | backend    | ngl | threads | n_batch | n_ubatch | type_k | type_v |  fa | dev          |            test |                  t/s |
| ------------------------------ | ---------: | ---------: | ---------- | --: | ------: | ------: | -------: | -----: | -----: | --: | ------------ | --------------: | -------------------: |
TCQ decode: context-adaptive V alpha enabled
| minimax-m2 230B.A10B Q3_K - Small |  87.20 GiB |   228.69 B | CUDA       |  99 |       4 |    1024 |     1024 |   q5_0 |   q4_1 |   1 | CUDA0        | pp1000 @ d60000 |        990.71 ± 0.83 |
| minimax-m2 230B.A10B Q3_K - Small |  87.20 GiB |   228.69 B | CUDA       |  99 |       4 |    1024 |     1024 |   q5_0 |   q4_1 |   1 | CUDA0        |  tg128 @ d60000 |         17.65 ± 0.01 |

build: 85e22ea0b (10102)

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions