Repository navigation
cpu_x86: Q4_K reads the GGUF bytes below AVX-512, one copy of the weights (#577) - #584
Merged
Merged
Conversation
…ghts (#577) Below the AVX-512 tier the Q4_K x8 repack bought no speed but kept a second full copy of every Q4_K tensor. A raw kernel now dots the GGUF blocks directly against int8 activations (maddubs on the nibbles, the dmin offset from per-32 activation sums), so the model's weights stay in the mapped file. GEIST_Q4K_RAW=1/0 overrides the automatic choice. Synthetic 1B Q4_K model, AVX2 tier: peak RSS 1308.7 -> 682.7 MiB; prefill and decode within noise of main. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AD86DiU1PpDqYprALdARAw
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Before: below AVX-512, cpu_x86 repacked every Q4_K tensor into the x8 layout. That kept a second full copy of the weights next to the mapped GGUF and gave no speed advantage on that tier.
After: below AVX-512, Q4_K linears dot the GGUF blocks directly, so the weights exist once, in the mapped file. On AVX-512 the repack stays the default, because the raw kernel's prefill is about 2× slower there.
Part of #577.
Changes
linear_q4k_raw.{c,h}(new) adds the raw Q4_K kernel:maddubs/madd, and the dmin offset is taken from the activation sums.backend.cpicks the path automatically: raw unlessq4kx8_avx512_usable().GEIST_Q4K_RAW=1/0overrides the choice.q4kx8_avx512_usable()is factored out of the AVX-512 kernel so the gate and the kernel agree.tests/test_x86_q4k_raw_unit.ccompares against cpu_scalar with a derived activation-rounding bound. The worst ratio is 0.20, and a mutation check makes 47 cases fail.Testing
make format-checkpasses.make test-unitpasses (84 passed, 31 skipped, 0 failed), both natively (AVX-512 VNNI) and withGEIST_FORCE_ISA=avx2.Peak RSS on a synthetic 1B Q4_K model (spm tokenizer, set_prompt plus a forward pass), AVX2 tier:
bench_revision_ab.pycomparing main against this branch withGEIST_FORCE_ISA=avx2and 6 cycles: prefill −5.1 % / −3.6 % and decode −4.7 % / −2.3 %, all within noise.API impact
None. There is one new optional env var,
GEIST_Q4K_RAW.🤖 Generated with Claude Code
https://claude.ai/code/session_01AD86DiU1PpDqYprALdARAw
Generated by Claude Code