Got turbo4 working on Gemma 4 26B A4B with zero speed penalty vs f16 KV. The key challenge was Gemma 4's variable head dimensions (256-dim on SWA layers, 512-dim on global layers) — the existing turbo4 FA kernels only supported D=128.
Hardware: RTX 3090 (24GB), Ryzen 9 5950X, 96GB RAM, Windows, MSVC + CUDA 13.0
| Config | Speed (32K ctx) | KV VRAM (32K) | KV VRAM (256K) | Fits 24GB at 256K? |
|---|---|---|---|---|
| f16 KV + FA | 120 t/s | ~1.2 GB | ~9.4 GB | No (25.4GB total) |
| q8_0 KV + FA | 122 t/s | ~0.6 GB | ~4.7 GB | Yes |
| turbo4 KV + FA | 120 t/s | ~0.4 GB | ~2.5 GB | Yes (18.5GB total) |
Turbo4 is the only config that runs Gemma 4 at full 256K context on a 24GB card at full speed.
Starting from the existing turbo4 FA kernels (D=128 only), extending to D=256/512 gave 63 t/s — about half of f16. Six optimizations brought it to parity:
| Optimization | Speed | vs f16 |
|---|---|---|
| Baseline turbo4 K+V (full WHT per position) | 63 t/s | 53% |
| + Lazy V: defer inverse WHT to single post-loop pass | 72 t/s | 60% |
| + Lazy K: forward WHT on Q once, dot raw centroids | 80 t/s | 67% |
| + Batch centroid decode: uint32/uint64 loads | 96 t/s | 80% |
| + Optimized write path: unrolled WHT + batch 3-bit packing | 104 t/s | 87% |
| + Warp-cooperative write kernel: 16-thread shuffle WHT | 120 t/s | 100% |
Lazy V (deferred WHT): Instead of inverse WHT on every V vector during attention, accumulate attention-weighted centroids in WHT space. Apply one inverse WHT to the accumulated VKQ after the loop. Exploits linearity: WHT_inv(Σ aᵢvᵢ) = Σ aᵢ·WHT_inv(vᵢ). Reduces V WHT cost from O(context_length) to O(1).
Lazy K (pre-transformed Q): Instead of inverse WHT on every K vector, apply forward WHT to Q once before the loop. Math: Q · WHT_inv(K_raw) = WHT_fwd(Q) · K_raw. Per-position K cost becomes just centroid unpack + dot + norm multiply. Also O(1).
Batch centroid decode: For K (8 elements, byte-aligned): load 3 bytes into uint32_t → 8 centroids via shifts. For V: uint16_t (ne=4), uint32_t (ne=8), uint64_t (ne=16). Eliminates per-element bit_offset/byte_idx computation.
Warp-cooperative write kernel: The original KV write path ran one CUDA thread per 128-element turbo4 block with a serial WHT butterfly. Replaced with 16 threads per block using warp-shuffle WHT — same approach as the FA read path. ~10-16x faster per block.
fattn-turbo4.cuh— Lazy K (Q pretransform) + lazy V (deferred WHT) + batch centroid decodefattn-vec.cuh— Q pretransform call before KV loop, VKQ post-process after loopfattn-vec-turbo4.cu— D=256/512 kernel instantiationsfattn.cu— Dispatch table + kernel selection for D=256/512 turbo4cpy-utils.cuh— Fused butterfly + unrolled WHT + batch 3-bit packingset-rows.cu— Warp-cooperative turbo4 quantize kernel (16 threads/block)
Model produces identical-quality output between f16 and turbo4 — verified on math (17×23=391, three methods shown), reasoning (general relativity explanations), and general chat. No quality degradation observed.
The CUDA kernel development was done mostly using Claude 4.6 Opus. I directed the project, chose the model and optimization targets, ran all builds and benchmarks, and validated correctness, but the actual kernel code was written by the AI. I found it genuinely interesting that every time the current iteration of turbo4 wasn't as fast as the f16 version, all I had to do was ask Claude if there was a clever way to make it faster, and it always found a way. So all credit for this work (and for most of this post except for this Transparency note) goes to Claude.
Built on top of the existing TurboQuant llama.cpp fork. To test:
cmake -B build -S . -DGGML_CUDA=ON -DCMAKE_CUDA_ARCHITECTURES=86
cmake --build build --config Release
build/bin/llama-server -m gemma-4-26B-A4B-it-Q4_K_M.gguf \
--mmproj mmproj-gemma-4-26B-A4B-it-f16.gguf \
--cache-type-k turbo4 --cache-type-v turbo4 \
--flash-attn on --ctx-size 262144 \
--n-gpu-layers 99 --host 0.0.0.0 --port 8080The lazy K/V approach should generalize to any model with head_dim > 128 — it's not Gemma 4 specific.
Independently verified on RTX 5080 16GB (Blackwell SM 120a, CUDA 13.2), Ryzen 9 9950X, WSL2, Ubuntu. Results below are measured, not extrapolated.
At 32K context, f16 KV cache exceeds the 16GB budget:
| Component | f16 / f16 | turbo4 / turbo4 |
|---|---|---|
| Model weights (Q4_0 imatrix) | 13,972 MiB | 13,972 MiB |
| KV non-SWA (5 global layers) | 640 MiB | 170 MiB |
| KV SWA (25 sliding layers) | 600 MiB | 159 MiB |
| Compute buffer | 527 MiB | 527 MiB |
| Total | ~15,739 MiB |
~15,077 MiB ✅ 901 MiB headroom |
f16 at 32K pushes past the 16GB limit and relies on VRAM overcommit (WSL2 allows this, bare metal will OOM). turbo4 fits cleanly with 900MB headroom.
KV compression: 1,240 MiB → 329 MiB = 3.77x (measured; the SWA layers only cache their 1024-token sliding window, not the full 32K, which is why the ratio is lower than the raw 4.25-bit figure would suggest).
| type_k | type_v | pp512 (t/s) | tg128 (t/s) | tg vs f16 |
|---|---|---|---|---|
| f16 | f16 | 7,154 ± 24 | 181.8 ± 3.6 | baseline |
| f16 | turbo4 | 4,503 ± 686 | 174.2 ± 18 | 95.8% |
| turbo4 | f16 | 2,632 ± 14 | 176.8 ± 0.3 | 97.2% |
| turbo4 | turbo4 | 3,802 ± 49 | 163.9 ± 4.8 | 90.2% |
Generation speed (tg) is what users feel in practice. 163.9 vs 181.8 t/s = 10% slower, 3.77x less VRAM. The prefill (pp) hit is larger but prefill is a one-time cost per conversation turn.
Note: the high pp variance on f16/turbo4 (±686) suggests asymmetric configs are less stable. Symmetric turbo4/turbo4 (±49) is more consistent.
Critical: cmake auto-detects /usr/bin/nvcc from PATH on most Linux systems. On Ubuntu with multiple CUDA versions installed, this is typically CUDA 11.x — too old for SM 120. Must specify explicitly:
git clone https://github.com/test1111111111111112/llama-cpp-turboquant-gemma4.git
cd llama-cpp-turboquant-gemma4
cmake -B build -S . \
-DCMAKE_CUDA_COMPILER=/usr/local/cuda-13.2/bin/nvcc \
-DGGML_CUDA=ON \
-DCMAKE_CUDA_ARCHITECTURES=120 \
-DCMAKE_BUILD_TYPE=Release
cmake --build build --config Release -j$(nproc)
# ~15 min. Binary: build/bin/llama-server (~90 MB with SM120a PTX)cmake 3.28+ auto-promotes 120 → 120a (Blackwell-a native FP4 tensor core path). Confirm in cmake output: --generate-code=arch=compute_120a,code=[compute_120a,sm_120a].
#!/bin/bash
exec build/bin/llama-server \
-m /path/to/gemma4-Q4_0_imatrix.gguf \
--cache-type-k turbo4 \
--cache-type-v turbo4 \
--flash-attn on \
--ctx-size 32768 \
--n-gpu-layers 99 \
--host 0.0.0.0 \
--port 8090 \
--parallel 2 \
--no-mmap--cache-type-k q8_0 --cache-type-v turbo4 crashes on CUDA at startup. The CUDA VEC flash-attn kernel dispatches via FATTN_VEC_CASE(head_dim, K_type, V_type). Entries exist for (256, TURBO4_0, TURBO4_0), (256, F16, TURBO4_0), (256, TURBO4_0, F16), but not (256, Q8_0, TURBO4_0). Falls through to GGML_ABORT("fatal error") at fattn.cu:331.
The asymmetric q8_0-K + turbo4-V path works on CPU (AVX-512) and Metal but not CUDA. Use symmetric turbo4/turbo4 on GPU. To add CUDA support, add FATTN_VEC_CASE(256, GGML_TYPE_Q8_0, GGML_TYPE_TURBO4_0) and FATTN_VEC_CASE(512, ...) to fattn.cu, plus a q8_0-K decode path in fattn-turbo4.cuh.
A parallel implementation using 64 independent 2D Givens rotations instead of WHT is in development on a separate branch. Dequant outputs original-space f32, which would allow asymmetric q8_0-K + pq4-V without new CUDA flash-attn kernels. Currently blocked on a missing k_set_rows_pq4_0 CUDA write kernel (same role as k_set_rows_turbo4 in set-rows.cu). Tracked in [Claude A's turboquant-kv branch].
| Hardware | CUDA | cmake ARCH | Status |
|---|---|---|---|
| RTX 3090 24GB, SM 86 | 13.0 | 86 |
✅ original author |
| RTX 5080 16GB, SM 120a | 13.2 | 120 (→ 120a) |
✅ verified 2026-04-11 |