Follow-up to #458: the qwen35 Vulkan kernels are correct (bit-identical to cpu_scalar on an FP32 KV cache) but not tuned.
Not measured / not tuned
- Q4_0 / Q4_1 / Q8_0 / Q5_K / TQ2_0 matvec + GEMM: first-cut kernels (size-agnostic matvec, 8 rows × 32 batch rows GEMM). No comparison with llama.cpp Vulkan on the same GGUFs, no occupancy/vectorization work; Q4_0 SoA loads not benchmarked against the native block layout.
- Gated-DeltaNet mixer (
deltanet_conv_f32, deltanet_delta_f32): one workgroup per v-head, thread j owns column j of S in global memory, shared-memory tree reductions with ~19 barriers per token per head; prefill runs the delta rule serially over the sequence. Candidates: keep S in registers/shared memory, subgroup reductions with a size specialization constant, chunked delta rule for prefill.
caps.dn_subchunk is not declared for Vulkan; the effect of chunking at GEIST_M_MAX 64/128 on prefill is unmeasured.
- Only the 0.8B was benchmarked end to end (pp64 ≈ 1100 t/s, tg ≈ 256 t/s on the 2080 Ti). No numbers for 4B (2080 Ti) or 27B (RADV) beyond "works".
Proposal
Add the qwen3.5-0.8B/4B/27B rows to benchmark/results/VULKAN.md (same protocol as #364), profile with GEIST_VK_PROFILE=1, and attack the top pipes.
Acceptance
- Documented pp/tg per model and device; a ratio against llama.cpp Vulkan for each; PRs per optimisation with before/after.
Follow-up to #458: the qwen35 Vulkan kernels are correct (bit-identical to
cpu_scalaron an FP32 KV cache) but not tuned.Not measured / not tuned
deltanet_conv_f32,deltanet_delta_f32): one workgroup per v-head, thread j owns column j of S in global memory, shared-memory tree reductions with ~19 barriers per token per head; prefill runs the delta rule serially over the sequence. Candidates: keep S in registers/shared memory, subgroup reductions with a size specialization constant, chunked delta rule for prefill.caps.dn_subchunkis not declared for Vulkan; the effect of chunking atGEIST_M_MAX64/128 on prefill is unmeasured.Proposal
Add the qwen3.5-0.8B/4B/27B rows to
benchmark/results/VULKAN.md(same protocol as #364), profile withGEIST_VK_PROFILE=1, and attack the top pipes.Acceptance