Request: a public, architecture-independent way to save and restore a session's decoding state, e.g.
size_t geist_session_snapshot_size(const struct geist_session *s);
enum geist_status geist_session_snapshot(struct geist_session *s, void *buf, size_t n); // KV + recurrent state + position
enum geist_status geist_session_restore(struct geist_session *s, const void *buf, size_t n);
(or an in-memory geist_session_fork that copies the state into a second session).
Why: game and agent runtimes re-run many short requests against one long, mostly constant context (character sheet, situation, recent turns). On a 0.8B qwen35 model, 90–95 % of each decision is prefill of ~700 tokens (≈2.5 s on M1 Max, ≈16 s on Zen 5 with Q8_0, see #504). With a snapshot after the constant part, each request would only prefill its ~100 new tokens. It would also allow scoring several candidate continuations from the same state (log-prob of a whole label instead of only its first token), which today needs pin_prefix + reset and therefore only works on attention-only models.
pin_prefix covers part of this for attention-only models but is wrong on the qwen35 hybrid (separate bug report). The internals seem to be there already: the speculative-decode path checkpoints and rolls back the DeltaNet state (deltanet_txn_*, kv_truncate, src/archs/transformer/arch_ops.c ~141–221). The state of a 0.8B qwen35 session is a few MB, so a copy is cheap compared to a prefill.
A minimal version that only supports restore into the same session/model (no on-disk format, no cross-version stability) would already cover the use case.
Request: a public, architecture-independent way to save and restore a session's decoding state, e.g.
(or an in-memory
geist_session_forkthat copies the state into a second session).Why: game and agent runtimes re-run many short requests against one long, mostly constant context (character sheet, situation, recent turns). On a 0.8B qwen35 model, 90–95 % of each decision is prefill of ~700 tokens (≈2.5 s on M1 Max, ≈16 s on Zen 5 with Q8_0, see #504). With a snapshot after the constant part, each request would only prefill its ~100 new tokens. It would also allow scoring several candidate continuations from the same state (log-prob of a whole label instead of only its first token), which today needs
pin_prefix+resetand therefore only works on attention-only models.pin_prefixcovers part of this for attention-only models but is wrong on the qwen35 hybrid (separate bug report). The internals seem to be there already: the speculative-decode path checkpoints and rolls back the DeltaNet state (deltanet_txn_*,kv_truncate,src/archs/transformer/arch_ops.c~141–221). The state of a 0.8B qwen35 session is a few MB, so a copy is cheap compared to a prefill.A minimal version that only supports restore into the same session/model (no on-disk format, no cross-version stability) would already cover the use case.