Skip to content

Misc. bug: Substantial prefill speed regression when resuming long session from a coding agent #25213

Description

@pinkfluid

Name and Version

❯ ./bin/llama-cli --version
version: 9859 (4fc4ec5)
built with GNU 16.1.1 for Linux x86_64

Operating systems

Linux

Which llama.cpp modules do you know to be affected?

llama-server

Command line

build/bin/llama-server --chat-template-kwargs {"preserve_thinking": true} --host 127.0.0.1 --image-max-tokens 1024 --image-min-tokens 256 --jinja --min-p 0.0 --port 56091 --presence-penalty 1.5 --repeat-penalty 1.0 --sleep-idle-seconds 610 --temperature 0.6 --threads-http 128 --top-k 20 --top-p 0.95 --webui-mcp-proxy --no-warmup --alias coding --ctx-size 131072 --cache-ram 0 --cache-type-k q8_0 --cache-type-v q5_1 --flash-attn on --fit on --fit-ctx 131072 --fit-target 1024 --model models/autoround/Qwen3.6-35B-A3B-Q8_0.gguf --mmproj models/autoround/mmproj-model-bf16.gguf --parallel 1 --ubatch-size 2048


The llama-server above was started from router mode, but I don't believe it matters.

Problem description & steps to reproduce

Resume a long context (70-80% out of 128k) using a coding agent. I have reproduced this using the latest pi agent (0.80.3).

Basically do a coding task until you fill up 70-80% of the context. Exit pi, and restart llama-server. Resume the previous session using pi --resume. The prefill rate tanks.

Currently seems to affect all architectures (or at least AMD/Vulkan + Nvidia/Cuda), as confirmed by #25062 (comment)

First Bad Commit

73618f2 (#24176)

Relevant log output

Resuming the same context on master vs master + fix (see follow up comment).

master

[51001] 13.30.533.725 I slot print_timing: id 0 | task 18 | prompt eval time = 698035.65 ms / 91463 tokens ( 7.63 ms per token, 131.03 tokens per second)

mater + fix

[60149] 5.59.819.066 I slot print_timing: id 0 | task 12 | prompt eval time = 249618.89 ms / 79518 tokens ( 3.14 ms per token, 318.56 tokens per second)

The log below was captured from another run without the fix:

❯ ~/llama/bin/llama-router
WARNING: radv is not a conformant Vulkan implementation, testing use only.
0.00.013.348 I cmn  common_param: common_params_print_info: verbosity = 3 (adjust with the `-lv N` CLI arg)
0.00.014.348 I srv   load_models: Loaded 0 cached model presets
0.00.014.957 I srv   load_models: Loaded 2 custom model presets from /home/mitja/llama/models.ini
0.00.016.080 I srv    operator(): Available models (2) (*: custom preset)
0.00.016.084 I srv    operator():   * coding (aliases: qwen36-q8)
0.00.016.085 I srv    operator():   * default
0.00.016.178 W srv  llama_server: -----------------
0.00.016.180 W srv  llama_server: CORS proxy is enabled, do not expose server to untrusted environments
0.00.016.180 W srv  llama_server: This feature is EXPERIMENTAL and may be removed or changed in future versions
0.00.016.180 W srv  llama_server: -----------------
0.00.016.185 I srv  llama_server: starting server in router mode. models will be automatically loaded on-demand
0.00.017.261 I srv  llama_server: listening on http://0.0.0.0:8080
0.00.017.262 W srv  llama_server: NOTE: router mode is experimental
0.00.017.263 W srv  llama_server:       it is not recommended to use this mode in untrusted environments
0.12.490.702 I srv  ensure_model: model name=coding is not loaded, loading...
0.12.490.741 I srv          load: spawning server instance with name=coding on port 52513
0.12.490.757 I srv          load: spawning server instance with args:
0.12.490.758 I srv          load:   /home/mitja/src/llama.cpp/build/bin/llama-server
0.12.490.758 I srv          load:   --chat-template-kwargs
0.12.490.758 I srv          load:   {"preserve_thinking": true}
0.12.490.759 I srv          load:   --host
0.12.490.759 I srv          load:   127.0.0.1
0.12.490.759 I srv          load:   --image-max-tokens
0.12.490.759 I srv          load:   1024
0.12.490.759 I srv          load:   --image-min-tokens
0.12.490.759 I srv          load:   256
0.12.490.759 I srv          load:   --jinja
0.12.490.759 I srv          load:   --min-p
0.12.490.760 I srv          load:   0.0
0.12.490.760 I srv          load:   --port
0.12.490.760 I srv          load:   52513
0.12.490.760 I srv          load:   --presence-penalty
0.12.490.760 I srv          load:   1.5
0.12.490.760 I srv          load:   --repeat-penalty
0.12.490.760 I srv          load:   1.0
0.12.490.760 I srv          load:   --sleep-idle-seconds
0.12.490.761 I srv          load:   610
0.12.490.761 I srv          load:   --temperature
0.12.490.761 I srv          load:   0.6
0.12.490.761 I srv          load:   --threads-http
0.12.490.761 I srv          load:   128
0.12.490.761 I srv          load:   --top-k
0.12.490.761 I srv          load:   20
0.12.490.761 I srv          load:   --top-p
0.12.490.762 I srv          load:   0.95
0.12.490.762 I srv          load:   --webui-mcp-proxy
0.12.490.762 I srv          load:   --no-warmup
0.12.490.762 I srv          load:   --alias
0.12.490.762 I srv          load:   coding
0.12.490.763 I srv          load:   --ctx-size
0.12.490.763 I srv          load:   131072
0.12.490.763 I srv          load:   --cache-ram
0.12.490.763 I srv          load:   0
0.12.490.763 I srv          load:   --cache-type-k
0.12.490.763 I srv          load:   q8_0
0.12.490.764 I srv          load:   --cache-type-v
0.12.490.764 I srv          load:   q5_1
0.12.490.764 I srv          load:   --flash-attn
0.12.490.764 I srv          load:   on
0.12.490.764 I srv          load:   --fit
0.12.490.764 I srv          load:   on
0.12.490.765 I srv          load:   --fit-ctx
0.12.490.765 I srv          load:   131072
0.12.490.765 I srv          load:   --fit-target
0.12.490.765 I srv          load:   1024
0.12.490.765 I srv          load:   --model
0.12.490.765 I srv          load:   models/autoround/Qwen3.6-35B-A3B-Q8_0.gguf
0.12.490.766 I srv          load:   --mmproj
0.12.490.766 I srv          load:   models/autoround/mmproj-model-bf16.gguf
0.12.490.766 I srv          load:   --parallel
0.12.490.766 I srv          load:   1
0.12.490.767 I srv          load:   --ubatch-size
0.12.490.767 I srv          load:   2048
0.12.490.953 I srv  ensure_model: waiting until model name=coding is fully loaded...
[52513] WARNING: radv is not a conformant Vulkan implementation, testing use only.
[52513] 0.00.013.632 I cmn  common_param: common_params_print_info: verbosity = 3 (adjust with the `-lv N` CLI arg)
[52513] 0.00.013.898 W srv  llama_server: -----------------
[52513] 0.00.013.899 W srv  llama_server: CORS proxy is enabled, do not expose server to untrusted environments
[52513] 0.00.013.900 W srv  llama_server: This feature is EXPERIMENTAL and may be removed or changed in future versions
[52513] 0.00.013.900 W srv  llama_server: -----------------
[52513] cmd_child_to_router:state:{"state":"loading","payload":{"stages":["text_model","mmproj_model"],"current":"text_model","value":0.0}}
[52513] 0.00.015.069 I srv    load_model: loading model 'models/autoround/Qwen3.6-35B-A3B-Q8_0.gguf'
[52513] 0.02.976.017 W llama_model_loader: tensor overrides to CPU are used with mmap enabled - consider using --no-mmap for better performance
[52513] cmd_child_to_router:state:{"state":"loading","payload":{"stages":["text_model","mmproj_model"],"current":"text_model","value":0.0}}
[52513] cmd_child_to_router:state:{"state":"loading","payload":{"stages":["text_model","mmproj_model"],"current":"text_model","value":0.7107987403869629}}
[52513] cmd_child_to_router:state:{"state":"loading","payload":{"stages":["text_model","mmproj_model"],"current":"text_model","value":0.7262607216835022}}
[52513] cmd_child_to_router:state:{"state":"loading","payload":{"stages":["text_model","mmproj_model"],"current":"text_model","value":0.7428445219993591}}
[52513] cmd_child_to_router:state:{"state":"loading","payload":{"stages":["text_model","mmproj_model"],"current":"text_model","value":0.7594283223152161}}
[52513] cmd_child_to_router:state:{"state":"loading","payload":{"stages":["text_model","mmproj_model"],"current":"text_model","value":0.7748903036117554}}
[52513] cmd_child_to_router:state:{"state":"loading","payload":{"stages":["text_model","mmproj_model"],"current":"text_model","value":0.7912856340408325}}
[52513] cmd_child_to_router:state:{"state":"loading","payload":{"stages":["text_model","mmproj_model"],"current":"text_model","value":0.8000815510749817}}
[52513] cmd_child_to_router:state:{"state":"loading","payload":{"stages":["text_model","mmproj_model"],"current":"text_model","value":0.8156003952026367}}
[52513] cmd_child_to_router:state:{"state":"loading","payload":{"stages":["text_model","mmproj_model"],"current":"text_model","value":0.8321841955184937}}
[52513] cmd_child_to_router:state:{"state":"loading","payload":{"stages":["text_model","mmproj_model"],"current":"text_model","value":0.847646176815033}}
[52513] cmd_child_to_router:state:{"state":"loading","payload":{"stages":["text_model","mmproj_model"],"current":"text_model","value":0.8564989566802979}}
[52513] cmd_child_to_router:state:{"state":"loading","payload":{"stages":["text_model","mmproj_model"],"current":"text_model","value":0.8719609379768372}}
[52513] cmd_child_to_router:state:{"state":"loading","payload":{"stages":["text_model","mmproj_model"],"current":"text_model","value":0.8883562684059143}}
[52513] cmd_child_to_router:state:{"state":"loading","payload":{"stages":["text_model","mmproj_model"],"current":"text_model","value":0.8960872292518616}}
[52513] cmd_child_to_router:state:{"state":"loading","payload":{"stages":["text_model","mmproj_model"],"current":"text_model","value":0.9049400687217712}}
[52513] cmd_child_to_router:state:{"state":"loading","payload":{"stages":["text_model","mmproj_model"],"current":"text_model","value":0.9204020500183105}}
[52513] cmd_child_to_router:state:{"state":"loading","payload":{"stages":["text_model","mmproj_model"],"current":"text_model","value":0.9369857907295227}}
[52513] cmd_child_to_router:state:{"state":"loading","payload":{"stages":["text_model","mmproj_model"],"current":"text_model","value":0.9535696506500244}}
[52513] cmd_child_to_router:state:{"state":"loading","payload":{"stages":["text_model","mmproj_model"],"current":"text_model","value":0.9690316319465637}}
[52513] cmd_child_to_router:state:{"state":"loading","payload":{"stages":["text_model","mmproj_model"],"current":"text_model","value":0.9780560731887817}}
[52513] cmd_child_to_router:state:{"state":"loading","payload":{"stages":["text_model","mmproj_model"],"current":"text_model","value":0.986653745174408}}
[52513] cmd_child_to_router:state:{"state":"loading","payload":{"stages":["text_model","mmproj_model"],"current":"text_model","value":0.9941295981407166}}
[52513] cmd_child_to_router:state:{"state":"loading","payload":{"stages":["text_model","mmproj_model"],"current":"text_model","value":1.0}}
[52513] cmd_child_to_router:state:{"state":"loading","payload":{"stage":"mmproj_model"}}
[52513] 0.37.322.129 W load_hparams: Qwen-VL models require at minimum 1024 image tokens to function correctly on grounding tasks
[52513] 0.37.322.131 W load_hparams: if you encounter problems with accuracy, try adding --image-min-tokens 1024
[52513] 0.37.322.132 W load_hparams: more info: https://github.com/ggml-org/llama.cpp/issues/16842
[52513]
[52513] cmd_child_to_router:state:{"state":"loading","payload":{"stages":["text_model","mmproj_model"],"current":"mmproj_model","value":0.0}}
[52513] cmd_child_to_router:state:{"state":"loading","payload":{"stages":["text_model","mmproj_model"],"current":"mmproj_model","value":0.3465363383293152}}
[52513] cmd_child_to_router:state:{"state":"loading","payload":{"stages":["text_model","mmproj_model"],"current":"mmproj_model","value":0.6954689025878906}}
[52513] cmd_child_to_router:state:{"state":"loading","payload":{"stages":["text_model","mmproj_model"],"current":"mmproj_model","value":0.9999909400939941}}
[52513] cmd_child_to_router:state:{"state":"loading","payload":{"stages":["text_model","mmproj_model"],"current":"mmproj_model","value":1.0}}
[52513] 0.37.962.415 I srv    load_model: loaded multimodal model, 'models/autoround/mmproj-model-bf16.gguf'
[52513] 0.38.851.748 I srv    load_model: initializing, n_slots = 1, n_ctx_slot = 131072, kv_unified = 'false'
[52513] 0.38.852.876 W srv          init: --cache-idle-slots requires --cache-ram, disabling
[52513] 0.38.877.426 I srv          init: chat template supports preserving reasoning, consider enabling it via --reasoning-preserve
[52513] 0.38.877.467 I srv  llama_server: model loaded
[52513] 0.38.877.492 I srv  llama_server: listening on http://127.0.0.1:52513
[52513] 0.38.877.539 I srv    operator(): child server monitoring thread started, waiting for EOF on stdin...
[52513] cmd_child_to_router:state:{"state":"ready","payload":{"id":"coding","aliases":["coding"],"tags":[],"object":"model","created":1782937613,"owned_by":"llamacpp","meta":{"vocab_type":2,"n_vocab":248320,"n_ctx":131072,"n_ctx_train":262144,"n_embd":2048,"n_params":34660610688,"size":36892150272}}}
0.51.374.973 I srv  proxy_reques: proxying request to model coding on port 52513
[52513] 0.39.475.264 I slot get_availabl: id  0 | task -1 | selected slot by LRU, t_last = -1
[52513] 0.39.476.307 I slot launch_slot_: id  0 | task 0 | processing task, is_child = 0
[52513] 0.57.285.932 I slot print_timing: id  0 | task 0 | prompt processing, n_tokens =   2048, progress = 0.02, t =  17.81 s / 115.00 tokens per second
[52513] 1.07.582.518 I slot print_timing: id  0 | task 0 | prompt processing, n_tokens =   3828, progress = 0.04, t =  28.11 s / 136.20 tokens per second
[52513] 1.11.946.000 I slot print_timing: id  0 | task 0 | prompt processing, n_tokens =   5876, progress = 0.05, t =  32.47 s / 180.97 tokens per second
[52513] 1.14.406.946 I slot print_timing: id  0 | task 0 | prompt processing, n_tokens =   6104, progress = 0.06, t =  34.93 s / 174.75 tokens per second
[52513] 1.18.176.603 I slot print_timing: id  0 | task 0 | prompt processing, n_tokens =   6431, progress = 0.06, t =  38.70 s / 166.17 tokens per second
[52513] 1.20.770.892 I slot print_timing: id  0 | task 0 | prompt processing, n_tokens =   6594, progress = 0.06, t =  41.29 s / 159.68 tokens per second
[52513] 1.23.464.505 I slot print_timing: id  0 | task 0 | prompt processing, n_tokens =   6746, progress = 0.06, t =  43.99 s / 153.36 tokens per second
[52513] 1.25.775.663 I slot print_timing: id  0 | task 0 | prompt processing, n_tokens =   6845, progress = 0.06, t =  46.30 s / 147.84 tokens per second
[52513] 1.28.506.914 I slot print_timing: id  0 | task 0 | prompt processing, n_tokens =   7052, progress = 0.06, t 

...

t = 803.51 s / 128.38 tokens per second
[52513] 14.05.899.805 I slot print_timing: id  0 | task 0 | prompt processing, n_tokens = 103470, progress = 0.95, t = 806.42 s / 128.31 tokens per second
[52513] 14.09.737.584 I slot print_timing: id  0 | task 0 | prompt processing, n_tokens = 104038, progress = 0.96, t = 810.26 s / 128.40 tokens per second
[52513] 14.11.828.466 I slot print_timing: id  0 | task 0 | prompt processing, n_tokens = 104127, progress = 0.96, t = 812.35 s / 128.18 tokens per second
[52513] 14.15.206.669 I slot print_timing: id  0 | task 0 | prompt processing, n_tokens = 104516, progress = 0.96, t = 815.72 s / 128.13 tokens per second
[52513] 14.18.624.747 I slot print_timing: id  0 | task 0 | prompt processing, n_tokens = 104850, progress = 0.96, t = 819.15 s / 128.00 tokens per second
[52513] 14.22.063.804 I slot print_timing: id  0 | task 0 | prompt processing, n_tokens = 105231, progress = 0.97, t = 822.59 s / 127.93 tokens per second
[52513] 14.25.296.841 I slot print_timing: id  0 | task 0 | prompt processing, n_tokens = 105588, progress = 0.97, t = 825.82 s / 127.86 tokens per second
[52513] 14.28.918.699 I slot print_timing: id  0 | task 0 | prompt processing, n_tokens = 106171, progress = 0.98, t = 829.44 s / 128.00 tokens per second
[52513] 14.32.540.353 I slot print_timing: id  0 | task 0 | prompt processing, n_tokens = 106696, progress = 0.98, t = 833.06 s / 128.08 tokens per second
[52513] 14.34.377.307 I slot print_timing: id  0 | task 0 | prompt processing, n_tokens = 106779, progress = 0.98, t = 834.90 s / 127.89 tokens per second
[52513] 14.38.274.233 I slot print_timing: id  0 | task 0 | prompt processing, n_tokens = 107353, progress = 0.99, t = 838.80 s / 127.98 tokens per second
[52513] 14.42.161.357 I slot print_timing: id  0 | task 0 | prompt processing, n_tokens = 107823, progress = 0.99, t = 842.69 s / 127.95 tokens per second
[52513] 14.46.151.042 I slot print_timing: id  0 | task 0 | prompt processing, n_tokens = 108431, progress = 1.00, t = 846.67 s / 128.07 tokens per second
[52513] 14.49.644.688 I slot print_timing: id  0 | task 0 | prompt processing, n_tokens = 108816, progress = 1.00, t = 850.17 s / 127.99 tokens per second
[52513] 14.49.821.158 I slot print_timing: id  0 | task 0 | prompt processing, n_tokens = 108823, progress = 1.00, t = 850.34 s / 127.98 tokens per second
[52513] 14.51.769.595 I slot print_timing: id  0 | task 0 | prompt eval time =  850468.42 ms / 108827 tokens (    7.81 ms per token,   127.96 tokens per second)
[52513] 14.51.769.597 I slot print_timing: id  0 | task 0 |        eval time =    1824.40 ms /    37 tokens (   49.31 ms per token,    20.28 tokens per second)
[52513] 14.51.769.598 I slot print_timing: id  0 | task 0 |       total time =  852292.82 ms / 108864 tokens
[52513] 14.51.769.697 I slot print_timing: id  0 | task 0 |    graphs reused =         35
[52513] 14.51.864.479 I slot      release: id  0 | task 0 | stop processing: n_tokens = 108863, truncated = 0

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions