Skip to content

vulkan: tune Q4_0/Q5_K/TQ2_0 and DeltaNet kernels; benchmark qwen3.5/3.8 (0.8B, 4B, 27B) #467

Description

@geisten

Follow-up to #458: the qwen35 Vulkan kernels are correct (bit-identical to cpu_scalar on an FP32 KV cache) but not tuned.

Not measured / not tuned

  • Q4_0 / Q4_1 / Q8_0 / Q5_K / TQ2_0 matvec + GEMM: first-cut kernels (size-agnostic matvec, 8 rows × 32 batch rows GEMM). No comparison with llama.cpp Vulkan on the same GGUFs, no occupancy/vectorization work; Q4_0 SoA loads not benchmarked against the native block layout.
  • Gated-DeltaNet mixer (deltanet_conv_f32, deltanet_delta_f32): one workgroup per v-head, thread j owns column j of S in global memory, shared-memory tree reductions with ~19 barriers per token per head; prefill runs the delta rule serially over the sequence. Candidates: keep S in registers/shared memory, subgroup reductions with a size specialization constant, chunked delta rule for prefill.
  • caps.dn_subchunk is not declared for Vulkan; the effect of chunking at GEIST_M_MAX 64/128 on prefill is unmeasured.
  • Only the 0.8B was benchmarked end to end (pp64 ≈ 1100 t/s, tg ≈ 256 t/s on the 2080 Ti). No numbers for 4B (2080 Ti) or 27B (RADV) beyond "works".

Proposal

Add the qwen3.5-0.8B/4B/27B rows to benchmark/results/VULKAN.md (same protocol as #364), profile with GEIST_VK_PROFILE=1, and attack the top pipes.

Acceptance

  • Documented pp/tg per model and device; a ratio against llama.cpp Vulkan for each; PRs per optimisation with before/after.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions