Repository navigation
transformer: size session KV from the session's max_seq_len; document resident memory per backend (#577) - #611
Merged
Merged
Conversation
…ment resident memory per backend A session's KV caches were sized from the model's cap, not the session's max_seq_len, so a 64-token session on a model loaded with a 32768 cap held 32768 rows of KV. Size them from sess->max_seq_len, as API_CONTRACT.md promises; every reader already bounds by it. docs/BACKENDS.md gains a resident-memory section: which backends keep repacked weight copies, the switch for each, how load_from_memory and the mapping count, and the per-session KV formula (#577). README's zero-copy line points there; a stale storage-mode comment is corrected. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Nb4xVBy5t3fTdSStHPPKPg
4 tasks
geisten
enabled auto-merge
October 4, 2026 14:30
geisten
pushed a commit
that referenced
this pull request
Oct 4, 2026
The KV cache follows the session's max_seq_len once #611 lands, so a 64-row session no longer holds enough KV to separate the budgets. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Nb4xVBy5t3fTdSStHPPKPg
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Before: every session's KV cache was sized from the model's cap. On a model loaded with
max_seq_len = 32768, a session created withmax_seq_len = 64still allocated and zeroed 32768 rows of KV. That is 64 MiB in FP32 for a two-layer test model; for a real model it is gigabytes.After: each session's KV is sized from its own
max_seq_len, which is whatdocs/API_CONTRACT.mdalready promised. The docs now say what stays resident per backend, which was the open documentation item on #577.Every forward path already rejects writes past
sess->max_seq_len(forward.h,layer.c:308,mtp.c), so only the allocation changes. Snapshot/restore (#603) checkskv_lenagainst the session's cap too. I built #603 merged with this change, and its snapshot test passes.Changes
arch_state.c: the KV, INT8/INT4 scale and KIVI buffers are sized fromsess->max_seq_len. The field comment inarch_state.hnow says so.docs/BACKENDS.md: new section Resident memory per backend, covering:load(path)andload_from_memoryhold the weights;README.md: the "zero-copy weights" line links to that section.tensor_views.c: corrects a stale comment that called β mode the default. It is the default only on Vulkan; CPU and Metal default to mmap-alias.New
tests/test_session_kv_sizing_unit.c(cpu_scalar). For FP32, INT8, INT4 and KIVI it checks:Without the fix the test fails: 68 MiB at 64 rows vs. 68 MiB at 32768 rows in FP32.
Not part of this PR (still open on #577): the CPU repack items (tracked in #584/#590) and the peak-RSS row for #364.
Testing
make test-unit(debug,BACKENDS="vulkan cpu_x86 cpu_scalar"): 101 passed, 26 skipped.test_backend_vulkan_ops_unitalso fails on main on lavapipe and is not in CI.make MODE=asan test-unit FILTER=sessionpasses.make format-checkclean.API impact
include/geist.h🤖 Generated with Claude Code
https://claude.ai/code/session_01Nb4xVBy5t3fTdSStHPPKPg
Generated by Claude Code