Summary
geist_session_pin_prefix on Gemma 4 E4B Q4_K_M with the Metal backend makes requests slower, not faster.
HELIO kept one session per figure. Each session pinned its constant system turn once (BOS plus about 40 % of the prompt). After that, every request called geist_session_reset (which truncates to the pin) and prefilled only the rest, in 64-token chunks via geist_session_prefill_tokens.
M1 Max, geistlib 7a0b421, --ai-bench 20 across three figures:
|
P50 |
P95 |
prefill P50 |
prefill P95 |
peak RSS |
| no pinning (full prompt every time) |
1940 ms |
2417 ms |
1566 ms |
2016 ms |
9.93 GiB |
| pinning |
2000 ms |
5174 ms |
1737 ms |
4795 ms |
11.23 GiB |
- Pinning itself costs about 4.7 s per session on Metal. A cold full prefill of the whole prompt takes about 2.1 s.
- After pinning, prefilling the shorter remainder is not faster than prefilling the whole prompt.
- On NVIDIA Vulkan the same code gives −11 % P50 (1536 → 1360 ms).
Requests
Related: #548 (snapshot/restore), #577 (memory).
Summary
geist_session_pin_prefixon Gemma 4 E4B Q4_K_M with the Metal backend makes requests slower, not faster.HELIO kept one session per figure. Each session pinned its constant system turn once (BOS plus about 40 % of the prompt). After that, every request called
geist_session_reset(which truncates to the pin) and prefilled only the rest, in 64-token chunks viageist_session_prefill_tokens.M1 Max, geistlib
7a0b421,--ai-bench 20across three figures:Requests
pin_prefixtakes a slow path on Metal (token by token instead of batched).Related: #548 (snapshot/restore), #577 (memory).