HIP: Enable AllReduce for ROCm - #27825
Conversation
|
Benchmark with 2x9700, 230W Powercapped, with/without, based on master @ 6fdd0ac : Modelconfig
Startup Log
[58819] 0.00.535.721 I load_tensors: offloading output layer to GPU Bench Tables
Prompt processing (PP) -- tokens/s, independent of parallel levelmtp=none
mtp=1
mtp=2
mtp=3
mtp=4
Decode (TG) -- combined tokens/s across concurrently-decoding slotsmtp=none parallel=1
mtp=none parallel=2
mtp=none parallel=3
mtp=1 parallel=1
mtp=1 parallel=2
mtp=1 parallel=3
mtp=2 parallel=1
mtp=2 parallel=2
mtp=2 parallel=3
mtp=3 parallel=1
mtp=3 parallel=2
mtp=3 parallel=3
mtp=4 parallel=1
mtp=4 parallel=2
mtp=4 parallel=3
TL;DR: Zero / No Change with 2xR9700 and P2P DMA working. But at least I confirmed again the flakyness of MTP with parallel sessions in long contexts :D |
This changes the fallback path used when rccl is not available/enabled so you should not see anything with it enabled. rccl should also be faster anyhow |
|
Yes, I should have clarified for anyone wanting to test the changes: This only changes the case when using exactly two AMD GPUs via ROCm without RCCL. |
|
@Stoney49th would you be willing to re-run your benchmarks with |
Sure, just the no-MTP case? That would shorten the run considerably... MTP is so spikey anyway for me... |
|
My understanding is that MTP performance will always be flaky, i.e., depend on the context. I can however say that in my experience running this patch, the MTP performance actually seems to also have experienced a small uplift, though I have not performed controlled tests in that regard. |
|
Heres the testpoint with MTP off - cannot run the full suite now, but LGTM: All settings equal, just RCCL=OFF, cards are PCIe4x8/x8, with a 5900X and DDR4-3600: Details
mtp=none
mtp=none parallel=1
mtp=none parallel=2
mtp=none parallel=3
TL;DR: with longer prompts, RCCL=OFF Is ~10% slower in PP, I could not measure a difference in TG, but I might be limited there elsewhere |
|
@Stoney49th If I understand your tables, this unfortunately just tests the difference between the RCCL AllReduce and the non-RCCL AllReduce, where RCCL is expected to be faster. |
|
@Stoney49th Would you mind running the benchmarks as I described above? |
Sadly not before the weekend - If you need the test earlier maybe someone else can jump in with a multi card setup. |
|
I have a similar setup, AM4 x570s with cards on asymmetric PCIe and unable to use RCCL due to only one x16 slot on the CPU. You can in theory bi-furcate the x16 slot on the CPU to x8-x8 then use RCCL but then you need to work out the splitter hardware and custom mounting the cards. This is this test setup-
With this change there is a noticeable speedup in both pp/tg for my system, but it very much depends on the model, ubatch size and the context length. Layer split wins pp2048 being fastest at ub512 I dont have the same numbers for longer context (131072), but in use with PI tensor-split+qwen3.8+ub1024 feels much faster as the context grows. Linux multi 7.1.8+deb14.1-amd64 #1 SMP PREEMPT_DYNAMIC Debian 7.1.8-2 (2026-08-15) x86_64 GNU/Linux I tested a few models with variations of these args gemma-4-31B-it-Q6_Kunpatched/main (c7bda03) + nightly rocm (therock-dist-linux-multiarch-10.1.0a20260822)
patched/main (c7bda03) + nightly rocm (therock-dist-linux-multiarch-10.1.0a20260822)
Qwen3.8-27B-UD-Q4_K_XLunpatched/main (c7bda03) + nightly rocm (therock-dist-linux-multiarch-10.1.0a202608223)
patched/main (c7bda03) + nightly rocm (therock-dist-linux-multiarch-10.1.0a20260822)
Muse-Glimmer-30B-UD-Q4_K_XLunpatched/main (c7bda03) + nightly rocm (therock-dist-linux-multiarch-10.1.0a20260822)
patched/main (c7bda03) + nightly rocm (therock-dist-linux-multiarch-10.1.0a20260822)
gpt-oss-20b-UD-Q8_K_XLunpatched/main (c7bda03) + nightly rocm (therock-dist-linux-multiarch-10.1.0a20260822)
patched/main (c7bda03) + nightly rocm (therock-dist-linux-multiarch-10.1.0a20260822)
|
|
I have no end of GPU hangs and random issues with RCCL enabled (perfect without), so this was right in my area of interest currently. I run the RDNA Boost Patches (except patch 12 which touches AllReduce), so this is with those stock and with those + this PR, so it should be relevant. Device 0: AMD Radeon AI PRO R9700, gfx1201 (0x1201), VMM: no, Wave Size: 32, VRAM: 32624 MiB ROCm 7.14.0 5900X / x570 system. The two tested GPUs are on PCIe 4.0/8x. RDNA Boost Patches Only:
RDNA Boost Patches + This PR:
So about a 17.45% increase in PP and 8.29% increase in TG. Not too shabby. |
|
I hit the same problem and landed on essentially your approach, so this is +1 from another data point. RDNA3 setup:
|
|
@IMbackK |
|
No i think its great, i just have several other PRS to attend to ahead of reviewing this one, its relatively low impact (i expect most people usw rccl) so it got pushed down the que |
|
No worries, just wanted to make sure you've got all you need from me for when you get to it. |
517472f to
ee5b990
Compare
|
@Stastez could you take a look at https://github.com/ggml-org/llama.cpp/actions/runs/34040825804/job/104104966391?pr=27825 and wrap the nodiscard functions in CUDA_CHECK macros? otherwise i can confirm the performance improvement on 2xgfx908 which are pretty much as far out of band clock speed wise as you are going to have on a supported hip device. |
|
@IMbackK Done. |
* upstream/master: (72 commits) HIP: Enable AllReduce for ROCm (ggml-org#27825) opencl: choose the MoE expert matmul by batch size for speculative decoding/MTP (ggml-org#27637) ci: build MUSA for only 1 arch (ggml-org#28944) docs: Rule of thumb for AI review time [no ci] (ggml-org#28945) rpc : hash-cache only weights (ggml-org#28789) cuda: support row-contiguous SUM_ROWS (ggml-org#26308) models : move build_arch_graph() after graph() template specialization (ggml-org#28934) vulkan: support sparse Flash Attention (ggml-org#28105) OpenVINO: optimize stateful decode and GPU MoE inference (ggml-org#28638) opencl: add generic ssm_scan (ggml-org#28881) ci: bump kleidiai runners from 22.04 to 24.04 (ggml-org#28885) metal : add FA kernels for HSK=96, HSV=64 (MiniCPM3) (ggml-org#28599) ci: Bump CUDA Windows x64 builds to 13.4.1 (ggml-org#28930) ci : fix android release (ggml-org#28936) cuda : enable i16 and i32 for DUP (ggml-org#28897) cmake : use PROJECT_SOURCE_DIR instead of CMAKE_SOURCE_DIR (ggml-org#28771) webui: stop re-probing disabled /tools endpoint on every message (ggml-org#28646) ci : reuse build tag name when used instead of safe one (ggml-org#28911) CI: hip-quality-check: ignore spill added in bfdc321 (ggml-org#28909) HIP: fattn-mma: use fp32 accumulation on MFMA devices (ggml-org#28576) ...
|
I tested the new internal ROCm 2-GPU AllReduce from this PR against RCCL on a dual RDNA3 setup. Hardware:
Model:
I built exactly the same commit twice:
Results: So on this system:
The TG difference is small, but it seems very reproducible. For example: This seems especially interesting for long-reasoning workloads, where thousands or tens of thousands of generated tokens can make a small TG advantage more important than the one-time PP advantage. It may also be an interesting data point for future AMD/HRX multi-GPU work: on dual gfx1100, a lightweight specialized 2-GPU reduction path appears able to outperform the more general RCCL path for decode, while RCCL remains clearly better for prefill. |
Overview
The internal AllReduce implementation is disabled on HIP due to missing functionality (
cudaHostAllocPortable,cudaHostAllocMapped,cudaHostGetDevicePointer,__nanosleep).However, with the exception of
__nanosleep, these functions DO exist on HIP (at least by now, see here, and here).I have emulated
__nanosleepvia__builtin_amdgcn_s_sleep, which lowers toS_SLEEPand sleeps a wavefront for an approximate number of cycles (see here). This instruction appears to be available on all ISAs listed here, i.e., across RDNA, CDNA, Vega, back to GCN 3.I have assumed a clock rate of about 2500 MHz (my RX 9070 clocks to ~3000 MHz, my RX 6800XT to ~2400 MHz, an MI 300A to 2100 MHz), which would mean the wavefront has to sleep for 250 cycles, or around 4 * 64 cycles to approximately match the 100 ns of
__nanosleep. Even if it does not sleep for the exact same time, I think it should not impact correctness and merely result in busier waiting, or slightly longer sleep times.On my test system (aforementioned RX 9070 and RX 6800XT), this PR provides a very noticeable performance uplift over the meta butterfly implementation:
I have not looked at MUSA due to not having MUSA hardware.
Additional information
RX 9070 @ PCIe 4.0x16, RX 6800XT @ PCIe 4.0x4
Requirements