Skip to content

perf(metal): pin_prefix on Gemma 4 E4B is slower than re-prefilling (4.7 s pin, no gain after) #578

Description

@geisten

Summary

geist_session_pin_prefix on Gemma 4 E4B Q4_K_M with the Metal backend makes requests slower, not faster.

HELIO kept one session per figure. Each session pinned its constant system turn once (BOS plus about 40 % of the prompt). After that, every request called geist_session_reset (which truncates to the pin) and prefilled only the rest, in 64-token chunks via geist_session_prefill_tokens.

M1 Max, geistlib 7a0b421, --ai-bench 20 across three figures:

P50 P95 prefill P50 prefill P95 peak RSS
no pinning (full prompt every time) 1940 ms 2417 ms 1566 ms 2016 ms 9.93 GiB
pinning 2000 ms 5174 ms 1737 ms 4795 ms 11.23 GiB
  • Pinning itself costs about 4.7 s per session on Metal. A cold full prefill of the whole prompt takes about 2.1 s.
  • After pinning, prefilling the shorter remainder is not faster than prefilling the whole prompt.
  • On NVIDIA Vulkan the same code gives −11 % P50 (1536 → 1360 ms).

Requests

  • Check whether pin_prefix takes a slow path on Metal (token by token instead of batched).
  • Check why prefilling after a pinned prefix costs as much as the full prompt (does reset re-run the prefix?).

Related: #548 (snapshot/restore), #577 (memory).

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't working

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions