Skip to content

server : consolidate slot selection into get_available_slot - #24755

Merged
ggerganov merged 1 commit into
masterfrom
gg/server-consolidate-slot-selection
Jun 19, 2026
Merged

ggerganov merged 1 commit into
masterfrom
gg/server-consolidate-slot-selection

Conversation

@ggerganov

@ggerganov ggerganov commented Jun 18, 2026 •

Copy link
Copy Markdown
Member

Overview

fix #24746

Absorb get_slot_by_id logic into get_available_slot so slot selection is handled by a single function call. When a specific slot id is requested, the LCP similarity check still runs to enable proper prompt cache updates.

Requirements

Absorb get_slot_by_id logic into get_available_slot so slot selection
is handled by a single function call. When a specific slot id is
requested, the LCP similarity check still runs to enable proper
prompt cache updates.

Assisted-by: pi:llama.cpp/Qwen3.6-27B
@ggerganov
ggerganov marked this pull request as ready for review June 19, 2026 06:21
@ggerganov
ggerganov requested a review from a team as a code owner June 19, 2026 06:21
@ggerganov
ggerganov merged commit 80452d6 into master Jun 19, 2026
23 of 25 checks passed
@ggerganov
ggerganov deleted the gg/server-consolidate-slot-selection branch June 19, 2026 06:22
adrianhoehne pushed a commit to adrianhoehne/llama.cpp that referenced this pull request Jul 5, 2026
…#24755)

Absorb get_slot_by_id logic into get_available_slot so slot selection
is handled by a single function call. When a specific slot id is
requested, the LCP similarity check still runs to enable proper
prompt cache updates.

Assisted-by: pi:llama.cpp/Qwen3.6-27B
zommiommy pushed a commit to zommiommy/llama.cpp that referenced this pull request Aug 18, 2026
…#24755)

Absorb get_slot_by_id logic into get_available_slot so slot selection
is handled by a single function call. When a specific slot id is
requested, the LCP similarity check still runs to enable proper
prompt cache updates.

Assisted-by: pi:llama.cpp/Qwen3.6-27B
zbrad pushed a commit to zbrad/llama.cpp that referenced this pull request Sep 10, 2026
…#24755)

Absorb get_slot_by_id logic into get_available_slot so slot selection
is handled by a single function call. When a specific slot id is
requested, the LCP similarity check still runs to enable proper
prompt cache updates.

Assisted-by: pi:llama.cpp/Qwen3.6-27B
marcospaulo added a commit to torad-labs/llama.cpp that referenced this pull request Sep 26, 2026
…hen empty, and is left alone while busy

Since ggml-org#24755 a request's id_slot goes through get_available_slot, and the cache update ran on the named slot as if the similarity loop had picked it:

- An idle slot saved and cleared (--cache-idle-slots, the default, with a unified KV cache) holds no tokens, so f_keep was 0 / 0. NaN compares false against 0.5, the prompt cache was never consulted, and a conversation returning to its slot was processed again from its first token: 74.1 s for a 178K-token conversation on an RTX 5090 (8 x 786,432 unified). An empty named slot now always goes through the prompt cache, also with --slot-prompt-similarity 0. A request that names no slot takes the similarity or LRU path, which already did.
- The similarity loop skips a named slot that is processing, so f_keep was 0 and the cache update ran on it: the running request's state was saved, and the best cached match for the new prompt was restored into its sequence mid-generation. A busy named slot is now returned as is, and the caller defers the task as it did before ggml-org#24755.

test_kv_keep_only_active: a named slot that was saved and cleared restores from cache-ram, and a stream on slot 0 while another request asks for slot 0 generates the same text as alone.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Misc. bug: Explicit Slot Requests Bypass Prompt Cache Restore

1 participant