Skip to content

cpu_x86: Q4_K reads the GGUF bytes below AVX-512, one copy of the weights (#577) - #584

Merged
geisten merged 1 commit into
mainfrom
claude/project-thread-ddgas0
Oct 3, 2026
Merged

geisten merged 1 commit into
mainfrom
claude/project-thread-ddgas0

Conversation

@geisten

@geisten geisten commented Oct 3, 2026

Copy link
Copy Markdown
Owner

Summary

Before: below AVX-512, cpu_x86 repacked every Q4_K tensor into the x8 layout. That kept a second full copy of the weights next to the mapped GGUF and gave no speed advantage on that tier.

After: below AVX-512, Q4_K linears dot the GGUF blocks directly, so the weights exist once, in the mapped file. On AVX-512 the repack stays the default, because the raw kernel's prefill is about 2× slower there.

Part of #577.

Changes

  • linear_q4k_raw.{c,h} (new) adds the raw Q4_K kernel:
    • Activations are quantized to int8 once per call, with one scale per 256 elements and an int32 sum per 32 elements.
    • The nibbles go straight into maddubs/madd, and the dmin offset is taken from the activation sums.
    • M>1 processes 4 tokens per weight-row pass, parallelised with OpenMP.
  • backend.c picks the path automatically: raw unless q4kx8_avx512_usable(). GEIST_Q4K_RAW=1/0 overrides the choice.
  • q4kx8_avx512_usable() is factored out of the AVX-512 kernel so the gate and the kernel agree.
  • tests/test_x86_q4k_raw_unit.c compares against cpu_scalar with a derived activation-rounding bound. The worst ratio is 0.20, and a mutation check makes 47 cases fail.
  • CHANGELOG entry.

Testing

  • make format-check passes.

  • make test-unit passes (84 passed, 31 skipped, 0 failed), both natively (AVX-512 VNNI) and with GEIST_FORCE_ISA=avx2.

  • Peak RSS on a synthetic 1B Q4_K model (spm tokenizer, set_prompt plus a forward pass), AVX2 tier:

    Path Peak RSS
    Repack 1308.7 MiB
    Raw 682.7 MiB (−48 %)
  • bench_revision_ab.py comparing main against this branch with GEIST_FORCE_ISA=avx2 and 6 cycles: prefill −5.1 % / −3.6 % and decode −4.7 % / −2.3 %, all within noise.

API impact

None. There is one new optional env var, GEIST_Q4K_RAW.

🤖 Generated with Claude Code

https://claude.ai/code/session_01AD86DiU1PpDqYprALdARAw


Generated by Claude Code

…ghts (#577)

Below the AVX-512 tier the Q4_K x8 repack bought no speed but kept a
second full copy of every Q4_K tensor. A raw kernel now dots the GGUF
blocks directly against int8 activations (maddubs on the nibbles, the
dmin offset from per-32 activation sums), so the model's weights stay
in the mapped file. GEIST_Q4K_RAW=1/0 overrides the automatic choice.

Synthetic 1B Q4_K model, AVX2 tier: peak RSS 1308.7 -> 682.7 MiB;
prefill and decode within noise of main.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01AD86DiU1PpDqYprALdARAw
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants