Skip to content
 
 

Latest commit

 

History

8,738 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Turbo4 on Gemma 4 26B A4B: 120 t/s with 3.8x KV compression (RTX 3090)

Got turbo4 working on Gemma 4 26B A4B with zero speed penalty vs f16 KV. The key challenge was Gemma 4's variable head dimensions (256-dim on SWA layers, 512-dim on global layers) — the existing turbo4 FA kernels only supported D=128.

Hardware: RTX 3090 (24GB), Ryzen 9 5950X, 96GB RAM, Windows, MSVC + CUDA 13.0

Results

Config Speed (32K ctx) KV VRAM (32K) KV VRAM (256K) Fits 24GB at 256K?
f16 KV + FA 120 t/s ~1.2 GB ~9.4 GB No (25.4GB total)
q8_0 KV + FA 122 t/s ~0.6 GB ~4.7 GB Yes
turbo4 KV + FA 120 t/s ~0.4 GB ~2.5 GB Yes (18.5GB total)

Turbo4 is the only config that runs Gemma 4 at full 256K context on a 24GB card at full speed.

The optimization journey

Starting from the existing turbo4 FA kernels (D=128 only), extending to D=256/512 gave 63 t/s — about half of f16. Six optimizations brought it to parity:

Optimization Speed vs f16
Baseline turbo4 K+V (full WHT per position) 63 t/s 53%
+ Lazy V: defer inverse WHT to single post-loop pass 72 t/s 60%
+ Lazy K: forward WHT on Q once, dot raw centroids 80 t/s 67%
+ Batch centroid decode: uint32/uint64 loads 96 t/s 80%
+ Optimized write path: unrolled WHT + batch 3-bit packing 104 t/s 87%
+ Warp-cooperative write kernel: 16-thread shuffle WHT 120 t/s 100%

Key ideas

Lazy V (deferred WHT): Instead of inverse WHT on every V vector during attention, accumulate attention-weighted centroids in WHT space. Apply one inverse WHT to the accumulated VKQ after the loop. Exploits linearity: WHT_inv(Σ aᵢvᵢ) = Σ aᵢ·WHT_inv(vᵢ). Reduces V WHT cost from O(context_length) to O(1).

Lazy K (pre-transformed Q): Instead of inverse WHT on every K vector, apply forward WHT to Q once before the loop. Math: Q · WHT_inv(K_raw) = WHT_fwd(Q) · K_raw. Per-position K cost becomes just centroid unpack + dot + norm multiply. Also O(1).

Batch centroid decode: For K (8 elements, byte-aligned): load 3 bytes into uint32_t → 8 centroids via shifts. For V: uint16_t (ne=4), uint32_t (ne=8), uint64_t (ne=16). Eliminates per-element bit_offset/byte_idx computation.

Warp-cooperative write kernel: The original KV write path ran one CUDA thread per 128-element turbo4 block with a serial WHT butterfly. Replaced with 16 threads per block using warp-shuffle WHT — same approach as the FA read path. ~10-16x faster per block.

Files modified (all in ggml/src/ggml-cuda/)

  • fattn-turbo4.cuh — Lazy K (Q pretransform) + lazy V (deferred WHT) + batch centroid decode
  • fattn-vec.cuh — Q pretransform call before KV loop, VKQ post-process after loop
  • fattn-vec-turbo4.cu — D=256/512 kernel instantiations
  • fattn.cu — Dispatch table + kernel selection for D=256/512 turbo4
  • cpy-utils.cuh — Fused butterfly + unrolled WHT + batch 3-bit packing
  • set-rows.cu — Warp-cooperative turbo4 quantize kernel (16 threads/block)

Correctness

Model produces identical-quality output between f16 and turbo4 — verified on math (17×23=391, three methods shown), reasoning (general relativity explanations), and general chat. No quality degradation observed.

Transparency note

The CUDA kernel development was done mostly using Claude 4.6 Opus. I directed the project, chose the model and optimization targets, ran all builds and benchmarks, and validated correctness, but the actual kernel code was written by the AI. I found it genuinely interesting that every time the current iteration of turbo4 wasn't as fast as the f16 version, all I had to do was ask Claude if there was a clever way to make it faster, and it always found a way. So all credit for this work (and for most of this post except for this Transparency note) goes to Claude.

Fork

Built on top of the existing TurboQuant llama.cpp fork. To test:

cmake -B build -S . -DGGML_CUDA=ON -DCMAKE_CUDA_ARCHITECTURES=86
cmake --build build --config Release

build/bin/llama-server -m gemma-4-26B-A4B-it-Q4_K_M.gguf \
  --mmproj mmproj-gemma-4-26B-A4B-it-f16.gguf \
  --cache-type-k turbo4 --cache-type-v turbo4 \
  --flash-attn on --ctx-size 262144 \
  --n-gpu-layers 99 --host 0.0.0.0 --port 8080

The lazy K/V approach should generalize to any model with head_dim > 128 — it's not Gemma 4 specific.


RTX 5080 / Blackwell (SM 120a) — Verified Build & Benchmarks

Independently verified on RTX 5080 16GB (Blackwell SM 120a, CUDA 13.2), Ryzen 9 9950X, WSL2, Ubuntu. Results below are measured, not extrapolated.

Why turbo4 matters on 16GB

At 32K context, f16 KV cache exceeds the 16GB budget:

Component f16 / f16 turbo4 / turbo4
Model weights (Q4_0 imatrix) 13,972 MiB 13,972 MiB
KV non-SWA (5 global layers) 640 MiB 170 MiB
KV SWA (25 sliding layers) 600 MiB 159 MiB
Compute buffer 527 MiB 527 MiB
Total ~15,739 MiB ⚠️ over budget ~15,077 MiB ✅ 901 MiB headroom

f16 at 32K pushes past the 16GB limit and relies on VRAM overcommit (WSL2 allows this, bare metal will OOM). turbo4 fits cleanly with 900MB headroom.

KV compression: 1,240 MiB → 329 MiB = 3.77x (measured; the SWA layers only cache their 1024-token sliding window, not the full 32K, which is why the ratio is lower than the raw 4.25-bit figure would suggest).

Benchmark (llama-bench, pp512/tg128, flash-attn on, Q4_0 imatrix)

type_k type_v pp512 (t/s) tg128 (t/s) tg vs f16
f16 f16 7,154 ± 24 181.8 ± 3.6 baseline
f16 turbo4 4,503 ± 686 174.2 ± 18 95.8%
turbo4 f16 2,632 ± 14 176.8 ± 0.3 97.2%
turbo4 turbo4 3,802 ± 49 163.9 ± 4.8 90.2%

Generation speed (tg) is what users feel in practice. 163.9 vs 181.8 t/s = 10% slower, 3.77x less VRAM. The prefill (pp) hit is larger but prefill is a one-time cost per conversation turn.

Note: the high pp variance on f16/turbo4 (±686) suggests asymmetric configs are less stable. Symmetric turbo4/turbo4 (±49) is more consistent.

Build for Blackwell (RTX 5080 / 4080 / 5090)

Critical: cmake auto-detects /usr/bin/nvcc from PATH on most Linux systems. On Ubuntu with multiple CUDA versions installed, this is typically CUDA 11.x — too old for SM 120. Must specify explicitly:

git clone https://github.com/test1111111111111112/llama-cpp-turboquant-gemma4.git
cd llama-cpp-turboquant-gemma4

cmake -B build -S . \
  -DCMAKE_CUDA_COMPILER=/usr/local/cuda-13.2/bin/nvcc \
  -DGGML_CUDA=ON \
  -DCMAKE_CUDA_ARCHITECTURES=120 \
  -DCMAKE_BUILD_TYPE=Release

cmake --build build --config Release -j$(nproc)
# ~15 min. Binary: build/bin/llama-server (~90 MB with SM120a PTX)

cmake 3.28+ auto-promotes 120 → 120a (Blackwell-a native FP4 tensor core path). Confirm in cmake output: --generate-code=arch=compute_120a,code=[compute_120a,sm_120a].

Serve script (RTX 5080, 16GB, Q4_0 imatrix, 32K ctx)

#!/bin/bash
exec build/bin/llama-server \
  -m /path/to/gemma4-Q4_0_imatrix.gguf \
  --cache-type-k turbo4 \
  --cache-type-v turbo4 \
  --flash-attn on \
  --ctx-size 32768 \
  --n-gpu-layers 99 \
  --host 0.0.0.0 \
  --port 8090 \
  --parallel 2 \
  --no-mmap

Known limitation: asymmetric KV on CUDA

--cache-type-k q8_0 --cache-type-v turbo4 crashes on CUDA at startup. The CUDA VEC flash-attn kernel dispatches via FATTN_VEC_CASE(head_dim, K_type, V_type). Entries exist for (256, TURBO4_0, TURBO4_0), (256, F16, TURBO4_0), (256, TURBO4_0, F16), but not (256, Q8_0, TURBO4_0). Falls through to GGML_ABORT("fatal error") at fattn.cu:331.

The asymmetric q8_0-K + turbo4-V path works on CPU (AVX-512) and Metal but not CUDA. Use symmetric turbo4/turbo4 on GPU. To add CUDA support, add FATTN_VEC_CASE(256, GGML_TYPE_Q8_0, GGML_TYPE_TURBO4_0) and FATTN_VEC_CASE(512, ...) to fattn.cu, plus a q8_0-K decode path in fattn-turbo4.cuh.

In development: PlanarQuant (pq4_0)

A parallel implementation using 64 independent 2D Givens rotations instead of WHT is in development on a separate branch. Dequant outputs original-space f32, which would allow asymmetric q8_0-K + pq4-V without new CUDA flash-attn kernels. Currently blocked on a missing k_set_rows_pq4_0 CUDA write kernel (same role as k_set_rows_turbo4 in set-rows.cu). Tracked in [Claude A's turboquant-kv branch].

Verified on

Hardware CUDA cmake ARCH Status
RTX 3090 24GB, SM 86 13.0 86 ✅ original author
RTX 5080 16GB, SM 120a 13.2 120 (→ 120a) ✅ verified 2026-04-11

About

TurboQuant llama.cpp fork with optimized turbo4 kernels for Gemma 4 D=256/512 heads — lazy K/V, batch decode, warp-cooperative write. 120 t/s with 3.8x KV compression on RTX 3090.

Resources

Contributing

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages