Skip to content

transformer: size session KV from the session's max_seq_len; document resident memory per backend (#577) - #611

Merged
geisten merged 2 commits into
mainfrom
claude/project-thread-xm9brj-577
Oct 4, 2026
Merged

geisten merged 2 commits into
mainfrom
claude/project-thread-xm9brj-577

Conversation

@geisten

@geisten geisten commented Oct 4, 2026

Copy link
Copy Markdown
Owner

Summary

Before: every session's KV cache was sized from the model's cap. On a model loaded with max_seq_len = 32768, a session created with max_seq_len = 64 still allocated and zeroed 32768 rows of KV. That is 64 MiB in FP32 for a two-layer test model; for a real model it is gigabytes.

After: each session's KV is sized from its own max_seq_len, which is what docs/API_CONTRACT.md already promised. The docs now say what stays resident per backend, which was the open documentation item on #577.

Every forward path already rejects writes past sess->max_seq_len (forward.h, layer.c:308, mtp.c), so only the allocation changes. Snapshot/restore (#603) checks kv_len against the session's cap too. I built #603 merged with this change, and its snapshot test passes.

Changes

  • arch_state.c: the KV, INT8/INT4 scale and KIVI buffers are sized from sess->max_seq_len. The field comment in arch_state.h now says so.

  • docs/BACKENDS.md: new section Resident memory per backend, covering:

    • how load(path) and load_from_memory hold the weights;
    • a table per backend of the repacked copies kept at load (cpu_x86 Q4_K→Q4_Kx8 only with AVX-512 panels, Q6_K→W8A8 with VNNI, I2_S, the F16 lm_head; the cpu_neon x8/predecode panels; Metal NoCopy vs. copy for writable caller memory; the Vulkan host arena and Q/K copy), with the env switch that turns each one off;
    • the per-session KV formula per KV mode, and the scratch pool.
  • README.md: the "zero-copy weights" line links to that section.

  • tensor_views.c: corrects a stale comment that called β mode the default. It is the default only on Vulkan; CPU and Metal default to mmap-alias.

  • New tests/test_session_kv_sizing_unit.c (cpu_scalar). For FP32, INT8, INT4 and KIVI it checks:

    • a 64-row session adds less than ⅛ of the KV that a 32768-row session adds (VmRSS; glibc mmap thresholds are pinned so freed pages are not reused unseen);
    • both sessions decode the same tokens;
    • the small session refuses a 65-token prompt.

    Without the fix the test fails: 68 MiB at 64 rows vs. 68 MiB at 32768 rows in FP32.

Not part of this PR (still open on #577): the CPU repack items (tracked in #584/#590) and the peak-RSS row for #364.

Testing

  • make test-unit (debug, BACKENDS="vulkan cpu_x86 cpu_scalar"): 101 passed, 26 skipped. test_backend_vulkan_ops_unit also fails on main on lavapipe and is not in CI.
  • make MODE=asan test-unit FILTER=session passes.
  • make format-check clean.

API impact

  • No change to include/geist.h

🤖 Generated with Claude Code

https://claude.ai/code/session_01Nb4xVBy5t3fTdSStHPPKPg


Generated by Claude Code

…ment resident memory per backend

A session's KV caches were sized from the model's cap, not the session's
max_seq_len, so a 64-token session on a model loaded with a 32768 cap held
32768 rows of KV. Size them from sess->max_seq_len, as API_CONTRACT.md
promises; every reader already bounds by it.

docs/BACKENDS.md gains a resident-memory section: which backends keep
repacked weight copies, the switch for each, how load_from_memory and the
mapping count, and the per-session KV formula (#577). README's zero-copy
line points there; a stale storage-mode comment is corrected.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Nb4xVBy5t3fTdSStHPPKPg
@geisten
geisten enabled auto-merge October 4, 2026 14:30
geisten pushed a commit that referenced this pull request Oct 4, 2026
The KV cache follows the session's max_seq_len once #611 lands, so a
64-row session no longer holds enough KV to separate the budgets.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Nb4xVBy5t3fTdSStHPPKPg
@geisten
geisten merged commit 56f0bd3 into main Oct 4, 2026
24 checks passed
@geisten
geisten deleted the claude/project-thread-xm9brj-577 branch October 4, 2026 14:52
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants