Skip to content

HIP: Enable AllReduce for ROCm - #27825

Merged
JohannesGaessler merged 1 commit into
ggml-org:masterfrom
Stastez:master
Sep 15, 2026
Merged

JohannesGaessler merged 1 commit into
ggml-org:masterfrom
Stastez:master

Conversation

@Stastez

@Stastez Stastez commented Aug 27, 2026

Copy link
Copy Markdown
Contributor

Overview

The internal AllReduce implementation is disabled on HIP due to missing functionality (cudaHostAllocPortable, cudaHostAllocMapped, cudaHostGetDevicePointer, __nanosleep).
However, with the exception of __nanosleep, these functions DO exist on HIP (at least by now, see here, and here).
I have emulated __nanosleep via __builtin_amdgcn_s_sleep, which lowers to S_SLEEP and sleeps a wavefront for an approximate number of cycles (see here). This instruction appears to be available on all ISAs listed here, i.e., across RDNA, CDNA, Vega, back to GCN 3.
I have assumed a clock rate of about 2500 MHz (my RX 9070 clocks to ~3000 MHz, my RX 6800XT to ~2400 MHz, an MI 300A to 2100 MHz), which would mean the wavefront has to sleep for 250 cycles, or around 4 * 64 cycles to approximately match the 100 ns of __nanosleep. Even if it does not sleep for the exact same time, I think it should not impact correctness and merely result in busier waiting, or slightly longer sleep times.

On my test system (aforementioned RX 9070 and RX 6800XT), this PR provides a very noticeable performance uplift over the meta butterfly implementation:

./llama-bench -p 2048 -n 512 -d 4096 -b 1024 -ub 512 -t 3 -sm tensor -mg 1 -fa on -dev ROCm0/ROCm1 --load-mode none -m /models/gemma-4-31B-it-Q6_K.gguf
ggml_cuda_init: found 2 ROCm devices (Total VRAM: 32672 MiB):
  Device 0: AMD Radeon RX 6800 XT, gfx1030 (0x1030), VMM: no, Wave Size: 32, VRAM: 16368 MiB
  Device 1: AMD Radeon RX 9070, gfx1201 (0x1201), VMM: no, Wave Size: 32, VRAM: 16304 MiB
model size params backend ngl threads n_batch main_gpu sm fa dev lm test t/s t/s uplift
gemma4 31B Q6_K 23.46 GiB 30.70 B ROCm -1 3 1024 1 tensor 1 ROCm0/ROCm1 none pp2048 @ d4096 384.15 ± 0.83 445.25 ± 0.10 ~15.91 %
gemma4 31B Q6_K 23.46 GiB 30.70 B ROCm -1 3 1024 1 tensor 1 ROCm0/ROCm1 none tg512 @ d4096 23.62 ± 0.01 24.15 ± 0.01 ~2.24 %
build: 6fdd0ac89 (10659)

I have not looked at MUSA due to not having MUSA hardware.

Additional information

RX 9070 @ PCIe 4.0x16, RX 6800XT @ PCIe 4.0x4

Requirements

  • I have read and agree with the contributing guidelines
  • AI usage disclosure: YES, for the actual implementation of the concept.

@Stastez
Stastez requested review from a team and IMbackK as code owners August 27, 2026 19:21
@github-actions github-actions Bot added ggml changes relating to the ggml tensor library for machine learning CUDA Related to the CUDA backend labels Aug 27, 2026
@Stoney49th

Copy link
Copy Markdown

Benchmark with 2x9700, 230W Powercapped, with/without, based on master @ 6fdd0ac :

Modelconfig

[qwen3-8-27b]
n                      = -2
cache-ram              = 14336
ctx-checkpoints        = 4
checkpoint-min-step    = 8192
main-gpu               = 0
parallel               = 3
batch-size             = 4096
ubatch-size            = 512
kv-unified             = false
hf                     = unsloth/Qwen3.8-27B-GGUF:UD-Q4_K_XL
ctx-size               = 491520
temp                   = 1.0
top-p                  = 0.95
top-k                  = 20
min-p                  = 0.0
presence-penalty       = 0.0
repeat-penalty         = 1.0
cache-type-k           = q8_0
cache-type-v           = q8_0
flash-attn             = true
split-mode             = tensor
#tensor-split           = 50,50
jinja                  = true
reasoning-preserve     = true
reasoning-effort       = medium
reasoning-budget       = 20000
image-min-tokens       = 1024
no-mmproj-offload      = true
spec-type              = none
spec-draft-n-max       = 1
spec-draft-p-min       = 0.75

ARG ROCM_VERSION=7.2.4
-DGGML_HIP=ON \
-DGGML_HIP_RCCL=ON \
-DGGML_HIP_GRAPHS=ON \
-DGGML_HIP_ROCWMMA_FATTN=ON \
Startup Log

[58819] 0.00.535.721 I load_tensors: offloading output layer to GPU
[58819] 0.00.535.731 I load_tensors: offloading 64 repeating layers to GPU
[58819] 0.00.535.732 I load_tensors: offloaded 66/66 layers to GPU
[58819] 0.00.535.741 I load_tensors: Meta() model buffer size = 7860.53 MiB
[58819] 0.00.535.742 I load_tensors: ROCm_Host model buffer size = 682.03 MiB
[58819] 0.05.528.316 I cmn common_init_: added <|endoftext|> logit bias = -inf
[58819] 0.05.528.329 I cmn common_init_: added <|im_end|> logit bias = -inf
[58819] 0.05.528.330 I cmn common_init_: added <|fim_pad|> logit bias = -inf
[58819] 0.05.528.331 I cmn common_init_: added <|repo_name|> logit bias = -inf
[58819] 0.05.528.332 I cmn common_init_: added <|file_sep|> logit bias = -inf
[58819] 0.05.528.443 I llama_context: constructing llama_context
[58819] 0.05.528.456 I llama_context: n_seq_max = 3
[58819] 0.05.528.456 I llama_context: n_ctx = 491520
[58819] 0.05.528.457 I llama_context: n_ctx_seq = 163840
[58819] 0.05.528.457 I llama_context: n_batch = 4096
[58819] 0.05.528.457 I llama_context: n_ubatch = 512
[58819] 0.05.528.457 I llama_context: causal_attn = 1
[58819] 0.05.528.459 I llama_context: flash_attn = enabled
[58819] 0.05.528.459 I llama_context: kv_unified = false
[58819] 0.05.528.463 I llama_context: freq_base = 10000000.0
[58819] 0.05.528.464 I llama_context: freq_scale = 1
[58819] 0.05.528.465 I llama_context: n_rs_seq = 0
[58819] 0.05.528.465 I llama_context: n_outputs_max = 3
[58819] 0.05.528.465 I llama_context: n_outputs_max_per_seq = 1
[58819] 0.05.528.466 I llama_context: n_ctx_seq (163840) < n_ctx_train (262144) -- the full capacity of the model will not be utilized
[58819] 0.06.737.744 I llama_context: ROCm_Host output buffer size = 2.84 MiB
[58819] 0.06.775.087 I llama_kv_cache: Meta() KV buffer size = 8160.00 MiB
[58819] 0.06.807.616 I llama_kv_cache: size = 16320.00 MiB (163840 cells, 16 layers, 3/3 seqs), K (q8_0): 8160.00 MiB, V (q8_0): 8160.00 MiB
[58819] 0.06.807.626 I llama_kv_cache: attn_rot_k = 1, n_embd_head_k_all = 256
[58819] 0.06.807.627 I llama_kv_cache: attn_rot_v = 1, n_embd_head_k_all = 256
[58819] 0.06.810.402 I llama_memory_recurrent: Meta() RS buffer size = 224.44 MiB
[58819] 0.06.810.417 I llama_memory_recurrent: size = 448.88 MiB ( 3 cells, 64 layers, 3 seqs 0 rs_seq), R (f32): 16.88 MiB, S (f32): 432.00 MiB
[58819] 0.06.810.428 I sched_reserve: reserving ...
[58819] 0.06.814.128 I resolve_fused_ops: resolving fused Gated Delta Net support:
[58819] 0.06.815.405 I resolve_fused_ops: fused Gated Delta Net (autoregressive) enabled
[58819] 0.06.816.245 I resolve_fused_ops: fused Gated Delta Net (chunked) enabled
[58819] 0.06.816.253 I resolve_fused_ops: resolving fused Lightning Indexer support:
[58819] 0.06.817.052 I resolve_fused_ops: Lightning Indexer enabled
[58819] 0.06.817.061 I resolve_fused_ops: resolving fused DeepSeek V4 HC support:
[58819] 0.06.817.838 I resolve_fused_ops: fused DeepSeek V4 HC pre enabled
[58819] 0.06.818.641 I resolve_fused_ops: fused DeepSeek V4 HC comb enabled
[58819] 0.06.819.448 I resolve_fused_ops: fused DeepSeek V4 HC post enabled
[58819] 0.06.853.565 I sched_reserve: Meta() compute buffer size = 2160.75 MiB
[58819] 0.06.853.576 I sched_reserve: ROCm_Host compute buffer size = 180.63 MiB
[58819] 0.06.853.576 I sched_reserve: graph nodes = 3879
[58819] 0.06.853.577 I sched_reserve: graph splits = 2
[58819] 0.06.853.579 I sched_reserve: reserve took 43.15 ms, sched copies = 1
[58819] 0.06.854.022 I cmn init: llama threadpool init, n_threads = 24
[58819] 0.06.854.061 I cmn common_init_: warming up the model with an empty run - please wait ... (--no-warmup to disable)
[58819] 0.07.185.687 I cmn common_conte: the context does not support partial sequence removal
[58819] 0.07.217.306 I srv load_model: speculative decoding will use checkpoints
[58819] 0.07.217.316 I srv load_model: initializing, n_slots = 3, n_ctx_slot = 163840, kv_unified = 'false'
[58819] 0.07.217.349 I spec common_specu: no implementations specified for speculative decoding
[58819] 0.07.217.352 I slot load_model: id 0 | task -1 | new slot, n_ctx = 163840
[58819] 0.07.217.358 I slot load_model: id 1 | task -1 | new slot, n_ctx = 163840
[58819] 0.07.217.366 I slot load_model: id 2 | task -1 | new slot, n_ctx = 163840
[58819] 0.07.217.520 I srv load_model: prompt cache is enabled, size limit: 14336 MiB
[58819] 0.07.217.528 I srv load_model: use --cache-ram 0 to disable the prompt cache
[58819] 0.07.217.528 I srv load_model: for more info see #16391
[58819] 0.07.217.529 I srv load_model: context checkpoints enabled, max = 4, min spacing = 8192
[58819] 0.07.217.543 I srv init: idle slots will be saved to prompt cache upon starting a new task
[58819] 0.07.222.791 I srv init: init: chat template, example_format: '<|im_start|>system

Bench Tables

Prompt processing (PP) -- tokens/s, independent of parallel level

mtp=none

context length A: PR26419-testing, gfx1201, GGML_HIP_ROCWMMA_FATTN=ON, Code Master B: PR27825-allreduce, gfx1201, FATTN ON % diff
60000 847.7 843.7 -0.5%
75000 750.5 748.2 -0.3%
90000 680.0 678.2 -0.3%
105000 621.8 622.2 +0.1%
120000 574.1 572.7 -0.2%

mtp=1

context length A: PR26419-testing, gfx1201, GGML_HIP_ROCWMMA_FATTN=ON, Code Master B: PR27825-allreduce, gfx1201, FATTN ON % diff
60000 772.5 777.9 +0.7%
75000 700.4 698.0 -0.3%
90000 638.2 638.7 +0.1%
105000 584.3 584.5 +0.0%
120000 540.0 539.9 -0.0%

mtp=2

context length A: PR26419-testing, gfx1201, GGML_HIP_ROCWMMA_FATTN=ON, Code Master B: PR27825-allreduce, gfx1201, FATTN ON % diff
60000 772.5 777.9 +0.7%
75000 700.4 698.0 -0.3%
90000 638.2 638.7 +0.1%
105000 584.3 584.5 +0.0%
120000 540.0 539.9 -0.0%

mtp=3

context length A: PR26419-testing, gfx1201, GGML_HIP_ROCWMMA_FATTN=ON, Code Master B: PR27825-allreduce, gfx1201, FATTN ON % diff
60000 772.5 777.9 +0.7%
75000 700.4 698.0 -0.3%
90000 638.2 638.7 +0.1%
105000 584.3 584.5 +0.0%
120000 540.0 539.9 -0.0%

mtp=4

context length A: PR26419-testing, gfx1201, GGML_HIP_ROCWMMA_FATTN=ON, Code Master B: PR27825-allreduce, gfx1201, FATTN ON % diff
60000 772.5 777.9 +0.7%
75000 700.4 698.0 -0.3%
90000 638.2 638.7 +0.1%
105000 584.3 584.5 +0.0%
120000 540.0 539.9 -0.0%

Decode (TG) -- combined tokens/s across concurrently-decoding slots

mtp=none parallel=1

context length A: PR26419-testing, gfx1201, GGML_HIP_ROCWMMA_FATTN=ON, Code Master B: PR27825-allreduce, gfx1201, FATTN ON % diff
60010 29.31 29.31 +0.0%
75010 27.79 27.79 +0.0%
90010 26.30 26.30 +0.0%
105010 25.06 25.06 +0.0%
120010 23.85 23.85 +0.0%

mtp=none parallel=2

context length A: PR26419-testing, gfx1201, GGML_HIP_ROCWMMA_FATTN=ON, Code Master B: PR27825-allreduce, gfx1201, FATTN ON % diff
60010 43.14 43.14 +0.0%
75010 39.94 39.94 +0.0%
90010 37.09 37.09 +0.0%
105010 34.62 34.62 +0.0%
120010 32.48 32.48 +0.0%

mtp=none parallel=3

context length A: PR26419-testing, gfx1201, GGML_HIP_ROCWMMA_FATTN=ON, Code Master B: PR27825-allreduce, gfx1201, FATTN ON % diff
60010 51.24 51.24 +0.0%
75010 46.81 46.81 +0.0%
90010 43.11 43.11 +0.0%
105010 40.08 40.08 +0.0%
120010 37.29 37.29 +0.0%

mtp=1 parallel=1

context length A: PR26419-testing, gfx1201, GGML_HIP_ROCWMMA_FATTN=ON, Code Master B: PR27825-allreduce, gfx1201, FATTN ON % diff
60010 37.30 37.30 +0.0%
75010 36.53 36.53 +0.0%
90010 43.41 43.41 +0.0%
105010 41.57 41.57 +0.0%
120010 39.88 39.88 +0.0%

mtp=1 parallel=2

context length A: PR26419-testing, gfx1201, GGML_HIP_ROCWMMA_FATTN=ON, Code Master B: PR27825-allreduce, gfx1201, FATTN ON % diff
60010 27.72 27.72 +0.0%
75010 45.77 45.77 +0.0%
90010 39.49 39.49 +0.0%
105010 53.84 53.84 +0.0%
120010 38.84 38.84 +0.0%

mtp=1 parallel=3

context length A: PR26419-testing, gfx1201, GGML_HIP_ROCWMMA_FATTN=ON, Code Master B: PR27825-allreduce, gfx1201, FATTN ON % diff
60010 70.54 70.54 +0.0%
75010 65.12 65.12 +0.0%
90010 56.98 56.98 +0.0%
105010 57.26 57.26 +0.0%
120010 47.48 47.48 +0.0%

mtp=2 parallel=1

context length A: PR26419-testing, gfx1201, GGML_HIP_ROCWMMA_FATTN=ON, Code Master B: PR27825-allreduce, gfx1201, FATTN ON % diff
60010 32.85 32.85 +0.0%
75010 47.25 47.25 +0.0%
90010 29.37 29.37 +0.0%
105010 47.91 47.91 +0.0%
120010 35.76 35.76 +0.0%

mtp=2 parallel=2

context length A: PR26419-testing, gfx1201, GGML_HIP_ROCWMMA_FATTN=ON, Code Master B: PR27825-allreduce, gfx1201, FATTN ON % diff
60010 34.35 34.35 +0.0%
75010 39.35 39.35 +0.0%
90010 34.55 34.55 +0.0%
105010 56.93 56.93 +0.0%
120010 53.41 53.41 +0.0%

mtp=2 parallel=3

context length A: PR26419-testing, gfx1201, GGML_HIP_ROCWMMA_FATTN=ON, Code Master B: PR27825-allreduce, gfx1201, FATTN ON % diff
60010 98.73 98.73 +0.0%
75010 24.96 26.48 +6.1%
90010 62.74 62.74 +0.0%
105010 41.31 41.31 +0.0%
120010 37.27 37.27 +0.0%

mtp=3 parallel=1

context length A: PR26419-testing, gfx1201, GGML_HIP_ROCWMMA_FATTN=ON, Code Master B: PR27825-allreduce, gfx1201, FATTN ON % diff
60010 44.49 44.49 +0.0%
75010 42.29 42.29 +0.0%
90010 29.93 29.93 +0.0%
105010 46.91 46.91 +0.0%
120010 53.99 53.99 +0.0%

mtp=3 parallel=2

context length A: PR26419-testing, gfx1201, GGML_HIP_ROCWMMA_FATTN=ON, Code Master B: PR27825-allreduce, gfx1201, FATTN ON % diff
60010 37.81 37.81 +0.0%
75010 39.98 39.98 +0.0%
90010 32.88 32.88 +0.0%
105010 67.45 67.45 +0.0%
120010 35.61 35.61 +0.0%

mtp=3 parallel=3

context length A: PR26419-testing, gfx1201, GGML_HIP_ROCWMMA_FATTN=ON, Code Master B: PR27825-allreduce, gfx1201, FATTN ON % diff
60010 123.93 123.93 +0.0%
75010 106.30 106.30 +0.0%
90010 102.46 102.46 +0.0%
105010 39.43 39.43 +0.0%
120010 30.52 30.52 +0.0%

mtp=4 parallel=1

context length A: PR26419-testing, gfx1201, GGML_HIP_ROCWMMA_FATTN=ON, Code Master B: PR27825-allreduce, gfx1201, FATTN ON % diff
60010 32.43 32.43 +0.0%
75010 24.72 24.72 +0.0%
90010 32.91 32.91 +0.0%
105010 50.72 50.72 +0.0%
120010 47.40 47.40 +0.0%

mtp=4 parallel=2

context length A: PR26419-testing, gfx1201, GGML_HIP_ROCWMMA_FATTN=ON, Code Master B: PR27825-allreduce, gfx1201, FATTN ON % diff
60010 39.67 39.67 +0.0%
75010 54.17 54.17 +0.0%
90010 42.74 42.74 +0.0%
105010 84.05 84.05 +0.0%
120010 60.33 60.33 +0.0%

mtp=4 parallel=3

context length A: PR26419-testing, gfx1201, GGML_HIP_ROCWMMA_FATTN=ON, Code Master B: PR27825-allreduce, gfx1201, FATTN ON % diff
60010 124.29 124.29 +0.0%
75010 55.03 55.03 +0.0%
90010 94.96 94.96 +0.0%
105010 87.06 87.06 +0.0%
120010 31.36 31.36 +0.0%

TL;DR: Zero / No Change with 2xR9700 and P2P DMA working. But at least I confirmed again the flakyness of MTP with parallel sessions in long contexts :D

@IMbackK

IMbackK commented Aug 28, 2026

Copy link
Copy Markdown
Contributor
ARG ROCM_VERSION=7.2.4
-DGGML_HIP=ON \
-DGGML_HIP_RCCL=ON \
-DGGML_HIP_GRAPHS=ON \
-DGGML_HIP_ROCWMMA_FATTN=ON \

TL;DR: Zero / No Change with 2xR9700 and P2P DMA working. But at least I confirmed again the flakyness of MTP with parallel sessions in long contexts :D

This changes the fallback path used when rccl is not available/enabled so you should not see anything with it enabled. rccl should also be faster anyhow

@Stastez

Stastez commented Aug 28, 2026

Copy link
Copy Markdown
Contributor Author

Yes, I should have clarified for anyone wanting to test the changes:

This only changes the case when using exactly two AMD GPUs via ROCm without RCCL.

@Stastez

Stastez commented Aug 28, 2026

Copy link
Copy Markdown
Contributor Author

@Stoney49th would you be willing to re-run your benchmarks with -DGGML_HIP_RCCL=OFF?

@Stoney49th

Copy link
Copy Markdown

@Stoney49th would you be willing to re-run your benchmarks with -DGGML_HIP_RCCL=OFF?

Sure, just the no-MTP case? That would shorten the run considerably... MTP is so spikey anyway for me...

@Stastez

Stastez commented Aug 28, 2026

Copy link
Copy Markdown
Contributor Author

My understanding is that MTP performance will always be flaky, i.e., depend on the context.
I personally think a non-MTP should be sufficient, which is why I've reported the llama-bench results.

I can however say that in my experience running this patch, the MTP performance actually seems to also have experienced a small uplift, though I have not performed controlled tests in that regard.

@Stoney49th

Copy link
Copy Markdown

Heres the testpoint with MTP off - cannot run the full suite now, but LGTM:

All settings equal, just RCCL=OFF, cards are PCIe4x8/x8, with a 5900X and DDR4-3600:

Details

mtp=none

context length A: PR27825-allreduce, gfx1201, FATTN ON B: PR27825-allreduce, gfx1201, FATTN ON, RCCL OFF % diff
60000 843.7 813.6 -3.6%
75000 748.2 726.3 -2.9%
90000 678.2 610.1 -10.1%
105000 622.2 563.9 -9.4%
120000 572.7 523.9 -8.5%

mtp=none parallel=1

context length A: PR27825-allreduce, gfx1201, FATTN ON B: PR27825-allreduce, gfx1201, FATTN ON, RCCL OFF % diff
60010 29.03 29.03 +0.0%
75010 27.50 27.50 +0.0%
90010 26.10 26.10 +0.0%
105010 24.86 24.86 +0.0%
120010 23.74 23.74 +0.0%

mtp=none parallel=2

context length A: PR27825-allreduce, gfx1201, FATTN ON B: PR27825-allreduce, gfx1201, FATTN ON, RCCL OFF % diff
60010 43.08 43.08 +0.0%
75010 39.90 39.90 +0.0%
90010 37.06 37.06 +0.0%
105010 34.72 34.72 +0.0%
120010 32.56 32.56 +0.0%

mtp=none parallel=3

context length A: PR27825-allreduce, gfx1201, FATTN ON B: PR27825-allreduce, gfx1201, FATTN ON, RCCL OFF % diff
60010 51.60 51.60 +0.0%
75010 39.58 39.58 +0.0%
90010 43.23 43.23 +0.0%
105010 40.29 40.29 +0.0%
120010 37.43 37.43 +0.0%

TL;DR: with longer prompts, RCCL=OFF Is ~10% slower in PP, I could not measure a difference in TG, but I might be limited there elsewhere

@Stastez

Stastez commented Aug 28, 2026

Copy link
Copy Markdown
Contributor Author

@Stoney49th If I understand your tables, this unfortunately just tests the difference between the RCCL AllReduce and the non-RCCL AllReduce, where RCCL is expected to be faster.
The interesting difference would be "main branch, no RCCL" vs "this PR, no RCCL", i.e., the difference between meta butterfly and the non-RCCL AllReduce.

@Stastez

Stastez commented Aug 31, 2026

Copy link
Copy Markdown
Contributor Author

@Stoney49th Would you mind running the benchmarks as I described above?

@Stoney49th

Copy link
Copy Markdown

@Stoney49th Would you mind running the benchmarks as I described above?

Sadly not before the weekend - If you need the test earlier maybe someone else can jump in with a multi card setup.

@seamusmcshane

seamusmcshane commented Sep 3, 2026 •

Copy link
Copy Markdown

I have a similar setup, AM4 x570s with cards on asymmetric PCIe and unable to use RCCL due to only one x16 slot on the CPU. You can in theory bi-furcate the x16 slot on the CPU to x8-x8 then use RCCL but then you need to work out the splitter hardware and custom mounting the cards.

This is this test setup-
AM4 x570s
CPU x5800x3d
Dual RX 7900 XTX

  • card0 PCIe4 x16 (CPU)
  • card1 PCIe4 x4 (Chipset)

With this change there is a noticeable speedup in both pp/tg for my system, but it very much depends on the model, ubatch size and the context length.

Layer split wins pp2048 being fastest at ub512
Tensor split seeming to like ub1024 for most models
Single card has the same guidance, if the model is small enough it can still win.

I dont have the same numbers for longer context (131072), but in use with PI tensor-split+qwen3.8+ub1024 feels much faster as the context grows.

Linux multi 7.1.8+deb14.1-amd64 #1 SMP PREEMPT_DYNAMIC Debian 7.1.8-2 (2026-08-15) x86_64 GNU/Linux
amdgpu-dkms-firmware/now 1:6.12.12.60403-2194681.22.04

I tested a few models with variations of these args
-p 2048 -n 512 -d 0 -b 4096 -ub <varied> -sm <varied> -mg 1 --load-mode none -dev <ROCm0 or ROCm0/ROCm1>

gemma-4-31B-it-Q6_K

unpatched/main (c7bda03) + nightly rocm (therock-dist-linux-multiarch-10.1.0a20260822)

Scenario pp2048 tg512 Result
Tensor split - ub512 533.5 ± 0.1 25.2 ± 1.3 Success
Layer split - ub512 1270.3 ± 2.0 22.3 ± 0.0 Success
Tensor split - ub1024 542.7 ± 75.4 25.3 ± 1.3 Success
Layer split - ub1024 1172.4 ± 1.1 22.3 ± 0.0 Success
Tensor split - ub2048 545.9 ± 57.1 25.2 ± 1.3 Success
Layer split - ub2048 976.7 ± 19.9 22.3 ± 0.0 Success

patched/main (c7bda03) + nightly rocm (therock-dist-linux-multiarch-10.1.0a20260822)

Scenario pp2048 tg512 Result
Tensor split - ub512 982.7 ± 1.0 35.0 ± 1.5 Success
Layer split - ub512 1277.6 ± 2.7 22.3 ± 0.0 Success
Tensor split - ub1024 1009.6 ± 57.9 35.1 ± 1.5 Success
Layer split - ub1024 1172.7 ± 0.9 22.3 ± 0.0 Success
Tensor split - ub2048 1007.6 ± 3.9 35.0 ± 1.5 Success
Layer split - ub2048 977.1 ± 19.6 22.3 ± 0.0 Success

Qwen3.8-27B-UD-Q4_K_XL

unpatched/main (c7bda03) + nightly rocm (therock-dist-linux-multiarch-10.1.0a202608223)

Scenario pp2048 tg512 Result
Single GPU - ub512 1007.7 ± 1.2 35.2 ± 0.1 Success
Tensor split - ub512 558.1 ± 0.2 27.0 ± 1.7 Success
Layer split - ub512 1414.3 ± 71.0 28.0 ± 0.0 Success
Single GPU - ub1024 1027.8 ± 1.6 35.2 ± 0.1 Success
Tensor split - ub1024 548.8 ± 83.5 27.1 ± 1.7 Success
Layer split - ub1024 1273.2 ± 0.7 28.0 ± 0.0 Success
Single GPU - ub2048 1035.6 ± 25.4 35.0 ± 0.1 Success
Tensor split - ub2048 555.6 ± 67.4 27.1 ± 1.6 Success
Layer split - ub2048 1007.1 ± 15.7 28.0 ± 0.0 Success

patched/main (c7bda03) + nightly rocm (therock-dist-linux-multiarch-10.1.0a20260822)

Scenario pp2048 tg512 Result
Single GPU - ub512 1011.8 ± 1.2 35.0 ± 0.0 Success
Tensor split - ub512 1026.4 ± 3.2 39.9 ± 2.4 Success
Layer split - ub512 1452.2 ± 2.2 28.0 ± 0.0 Success
Single GPU - ub1024 1032.1 ± 1.6 34.9 ± 0.1 Success
Tensor split - ub1024 1088.2 ± 6.1 40.1 ± 2.4 Success
Layer split - ub1024 1276.8 ± 0.8 27.9 ± 0.0 Success
Single GPU - ub2048 1038.5 ± 24.6 35.1 ± 0.1 Success
Tensor split - ub2048 1059.5 ± 1.1 40.2 ± 2.4 Success
Layer split - ub2048 1011.4 ± 16.5 27.9 ± 0.0 Success

Muse-Glimmer-30B-UD-Q4_K_XL

unpatched/main (c7bda03) + nightly rocm (therock-dist-linux-multiarch-10.1.0a20260822)

Scenario pp2048 tg512 Result
Single GPU - ub512 1101.2 ± 1.4 38.7 ± 0.0 Success
Tensor split - ub512 576.6 ± 1.3 32.9 ± 1.8 Success
Layer split - ub512 1571.5 ± 2.4 29.6 ± 0.0 Success
Single GPU - ub1024 1131.6 ± 0.8 38.6 ± 0.0 Success
Tensor split - ub1024 558.0 ± 66.7 32.9 ± 1.7 Success
Layer split - ub1024 1387.6 ± 0.9 29.6 ± 0.0 Success
Single GPU - ub2048 1127.0 ± 18.2 38.6 ± 0.0 Success
Tensor split - ub2048 564.0 ± 57.1 33.0 ± 1.8 Success
Layer split - ub2048 1091.5 ± 13.5 29.6 ± 0.0 Success

patched/main (c7bda03) + nightly rocm (therock-dist-linux-multiarch-10.1.0a20260822)

Scenario pp2048 tg512 Result
Single GPU - ub512 1096.9 ± 1.3 38.6 ± 0.0 Success
Tensor split - ub512 1070.8 ± 2.9 48.9 ± 2.5 Success
Layer split - ub512 1576.6 ± 1.4 29.6 ± 0.0 Success
Single GPU - ub1024 1132.4 ± 1.1 38.6 ± 0.0 Success
Tensor split - ub1024 1094.4 ± 4.3 48.9 ± 2.5 Success
Layer split - ub1024 1390.1 ± 1.0 29.6 ± 0.0 Success
Single GPU - ub2048 1129.9 ± 18.4 38.5 ± 0.0 Success
Tensor split - ub2048 1104.2 ± 2.1 48.9 ± 2.5 Success
Layer split - ub2048 1094.6 ± 15.0 29.7 ± 0.0 Success

gpt-oss-20b-UD-Q8_K_XL

unpatched/main (c7bda03) + nightly rocm (therock-dist-linux-multiarch-10.1.0a20260822)

Scenario pp2048 tg512 Result
Single GPU - ub512 4203.0 ± 37.3 152.3 ± 0.1 Success
Tensor split - ub512 2496.0 ± 9.2 71.7 ± 4.0 Success
Layer split - ub512 5396.4 ± 19.9 91.8 ± 0.1 Success
Single GPU - ub1024 5243.4 ± 18.5 151.5 ± 0.1 Success
Tensor split - ub1024 2431.1 ± 554.9 71.6 ± 4.0 Success
Layer split - ub1024 5874.8 ± 17.5 92.1 ± 0.1 Success
Single GPU - ub2048 5597.7 ± 216.3 151.2 ± 0.1 Success
Tensor split - ub2048 2564.7 ± 602.3 71.4 ± 3.9 Success
Layer split - ub2048 5141.9 ± 179.4 92.1 ± 0.0 Success

patched/main (c7bda03) + nightly rocm (therock-dist-linux-multiarch-10.1.0a20260822)

Scenario pp2048 tg512 Result
Single GPU - ub512 4181.2 ± 11.5 152.5 ± 0.1 Success
Tensor split - ub512 4141.5 ± 36.0 150.2 ± 11.0 Success
Layer split - ub512 5383.3 ± 32.7 90.7 ± 0.0 Success
Single GPU - ub1024 5253.5 ± 28.4 151.7 ± 0.1 Success
Tensor split - ub1024 4510.1 ± 894.4 149.9 ± 11.0 Success
Layer split - ub1024 5897.0 ± 9.5 92.0 ± 0.0 Success
Single GPU - ub2048 5613.5 ± 214.8 151.4 ± 0.1 Success
Tensor split - ub2048 5215.6 ± 54.5 149.7 ± 11.1 Success
Layer split - ub2048 5146.1 ± 162.9 92.2 ± 0.0 Success

@SinclairKS

SinclairKS commented Sep 4, 2026 •

Copy link
Copy Markdown

I have no end of GPU hangs and random issues with RCCL enabled (perfect without), so this was right in my area of interest currently.

I run the RDNA Boost Patches (except patch 12 which touches AllReduce), so this is with those stock and with those + this PR, so it should be relevant.

Device 0: AMD Radeon AI PRO R9700, gfx1201 (0x1201), VMM: no, Wave Size: 32, VRAM: 32624 MiB
Device 1: AMD Radeon AI PRO R9700, gfx1201 (0x1201), VMM: no, Wave Size: 32, VRAM: 32624 MiB

ROCm 7.14.0

5900X / x570 system.

The two tested GPUs are on PCIe 4.0/8x.

RDNA Boost Patches Only:

model size params backend ngl threads n_batch main_gpu sm fa dev lm test t/s
gemma4 31B Q6_K 23.46 GiB 30.70 B ROCm -1 3 1024 1 tensor 1 ROCm0/ROCm1 none pp2048 @ d4096 959.91 ± 0.93
gemma4 31B Q6_K 23.46 GiB 30.70 B ROCm -1 3 1024 1 tensor 1 ROCm0/ROCm1 none tg512 @ d4096 30.02 ± 0.03

RDNA Boost Patches + This PR:

model size params backend ngl threads n_batch main_gpu sm fa dev lm test t/s
gemma4 31B Q6_K 23.46 GiB 30.70 B ROCm -1 3 1024 1 tensor 1 ROCm0/ROCm1 none pp2048 @ d4096 1127.49 ± 0.73
gemma4 31B Q6_K 23.46 GiB 30.70 B ROCm -1 3 1024 1 tensor 1 ROCm0/ROCm1 none tg512 @ d4096 32.51 ± 0.03

So about a 17.45% increase in PP and 8.29% increase in TG. Not too shabby.

@asa-degroff

Copy link
Copy Markdown

I hit the same problem and landed on essentially your approach, so this is +1 from another data point.

RDNA3 setup:
X670E system, 2x RX 7700, PCIe 4.0 x8/x8. On Qwen3.8-27B (-sm tensor, q8_0 KV cache,
ub 1024), butterfly vs. this PR:

test butterfly internal AllReduce uplift
pp16384 391.7 457.9 +16.9 %
pp65536 292.9 331.3 +13.1 %

@Stastez

Stastez commented Sep 4, 2026

Copy link
Copy Markdown
Contributor Author

@IMbackK
Can I do anything else to help this PR?

@IMbackK

IMbackK commented Sep 4, 2026

Copy link
Copy Markdown
Contributor

No i think its great, i just have several other PRS to attend to ahead of reviewing this one, its relatively low impact (i expect most people usw rccl) so it got pushed down the que

@Stastez

Stastez commented Sep 4, 2026

Copy link
Copy Markdown
Contributor Author

No worries, just wanted to make sure you've got all you need from me for when you get to it.
Thanks!

@Stastez
Stastez force-pushed the master branch 2 times, most recently from 517472f to ee5b990 Compare September 6, 2026 14:58
@IMbackK

IMbackK commented Sep 15, 2026 •

Copy link
Copy Markdown
Contributor

@Stastez could you take a look at https://github.com/ggml-org/llama.cpp/actions/runs/34040825804/job/104104966391?pr=27825 and wrap the nodiscard functions in CUDA_CHECK macros?

otherwise i can confirm the performance improvement on 2xgfx908 which are pretty much as far out of band clock speed wise as you are going to have on a supported hip device.

@Stastez

Stastez commented Sep 15, 2026

Copy link
Copy Markdown
Contributor Author

@IMbackK Done.

@JohannesGaessler
JohannesGaessler merged commit 38a5b42 into ggml-org:master Sep 15, 2026
28 of 31 checks passed
dzannotti added a commit to halo-box/llama.cpp that referenced this pull request Sep 15, 2026
* upstream/master: (72 commits)
  HIP: Enable AllReduce for ROCm (ggml-org#27825)
  opencl: choose the MoE expert matmul by batch size for speculative decoding/MTP (ggml-org#27637)
  ci: build MUSA for only 1 arch (ggml-org#28944)
  docs: Rule of thumb for AI review time [no ci] (ggml-org#28945)
  rpc : hash-cache only weights (ggml-org#28789)
  cuda: support row-contiguous SUM_ROWS (ggml-org#26308)
  models : move build_arch_graph() after graph() template specialization (ggml-org#28934)
  vulkan: support sparse Flash Attention (ggml-org#28105)
  OpenVINO: optimize stateful decode and GPU MoE inference (ggml-org#28638)
  opencl: add generic ssm_scan (ggml-org#28881)
  ci: bump kleidiai runners from 22.04 to 24.04 (ggml-org#28885)
  metal : add FA kernels for HSK=96, HSV=64 (MiniCPM3) (ggml-org#28599)
  ci: Bump CUDA Windows x64 builds to 13.4.1 (ggml-org#28930)
  ci : fix android release (ggml-org#28936)
  cuda : enable i16 and i32 for DUP (ggml-org#28897)
  cmake : use PROJECT_SOURCE_DIR instead of CMAKE_SOURCE_DIR (ggml-org#28771)
  webui: stop re-probing disabled /tools endpoint on every message (ggml-org#28646)
  ci : reuse build tag name when used instead of safe one (ggml-org#28911)
  CI: hip-quality-check: ignore spill added in bfdc321 (ggml-org#28909)
  HIP: fattn-mma: use fp32 accumulation on MFMA devices (ggml-org#28576)
  ...
fencerJP pushed a commit to fencerJP/llama-apu that referenced this pull request Sep 16, 2026
angt pushed a commit to angt/llama.cpp that referenced this pull request Sep 16, 2026
@a-n-t-0

a-n-t-0 commented Sep 16, 2026 •

Copy link
Copy Markdown

I tested the new internal ROCm 2-GPU AllReduce from this PR against RCCL on a dual RDNA3 setup.

Hardware:

  • 2x Radeon RX 7900 XTX, gfx1100, 24 GB each
  • PCIe 4.0 x8/x8
  • Both GPUs connected through CPU PCIe lanes
  • Ryzen 5 3600 / X570

Model:

  • Qwen3.8-27B Q8_0
  • llama.cpp b10998 / commit 37b53fd
  • -ngl 999
  • -sm tensor

I built exactly the same commit twice:

  • GGML_HIP_RCCL=ON
  • GGML_HIP_RCCL=OFF

Results:

                     Internal AR     RCCL
                     RCCL OFF        RCCL ON

pp512                1445.95         1579.61
tg256                   36.54           35.67

pp8192               1356.76         1479.34
tg512                   36.60           35.62

pp512 @ d100000        791.56          823.04
tg256 @ d100000         31.59           30.81

So on this system:

  • RCCL is about 9% faster for large prompt processing.
  • The internal AllReduce is consistently about 2.5-2.8% faster for token generation.
  • The decode advantage remains with a 100k-token KV depth.

The TG difference is small, but it seems very reproducible.

For example:

short context:
internal AR: 36.60 t/s
RCCL:       35.62 t/s

100k KV depth:
internal AR: 31.59 t/s
RCCL:       30.81 t/s

This seems especially interesting for long-reasoning workloads, where thousands or tens of thousands of generated tokens can make a small TG advantage more important than the one-time PP advantage.

It may also be an interesting data point for future AMD/HRX multi-GPU work: on dual gfx1100, a lightweight specialized 2-GPU reduction path appears able to outperform the more general RCCL path for decode, while RCCL remains clearly better for prefill.

quimmedes pushed a commit to quimmedes/cafe-llama.cpp that referenced this pull request Sep 16, 2026
zsogitbe pushed a commit to zsogitbe/llama.cpp that referenced this pull request Sep 17, 2026
Te-eMster pushed a commit to Te-eMster/mx-llama.cpp that referenced this pull request Sep 18, 2026
@BrewTestBot BrewTestBot mentioned this pull request Sep 23, 2026
1 task done
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CUDA Related to the CUDA backend ggml changes relating to the ggml tensor library for machine learning

Projects

None yet

Development

Successfully merging this pull request may close these issues.

8 participants