Skip to content

Misc. bug: Context checkpoints always invalidated on hybrid/recurrent models  #24055

Description

@ziggy416

Name and Version

started with docker image: server-cuda12-b9354

Operating systems

Linux

Which llama.cpp modules do you know to be affected?

llama-server

Command line

services:
  llama:
    image: ghcr.io/ggml-org/llama.cpp:server-cuda12-b9354
    runtime: nvidia
    environment:
      - NVIDIA_VISIBLE_DEVICES=all
    volumes:
      - ./models:/models
    ports:
      - "8000:8080"
    command: >
      --model /models/Qwopus3.6-27B-v2-MTP-Q4_K_M.gguf
      --n-gpu-layers 999
      --tensor-split 0.5,0.5
      --jinja
      --split-mode tensor
      --swa-full
      --ctx-size 32768
      --ctx-checkpoints 64
      --checkpoint-min-step 512
      --cache-prompt
      --cache-reuse 256
      --cache-ram 2048
      --threads 4
      --threads-batch 4
      --flash-attn on
      --batch-size 2048
      --ubatch-size 512
      --cache-type-k f16
      --cache-type-v f16
      --fit off
      --kv-unified
      --kv-offload
      --cont-batching
      --op-offload
      --repack
      --no-mmap
      --temp 0.6
      --top-k 20
      --top-p 0.95
      --min-p 0
      --reasoning-budget 1024
      --host 0.0.0.0
      --port 8080
    ipc: host
    restart: unless-stopped

Problem description & steps to reproduce

the --checkpoint-min-step flag doesn't work and forces full prompt reprocessing every time. When i switch to the docker image before that update and use --checkpoint-every-n-tokens instead, the checkpoints work just fine. Both logs use the same compose except I'm forced to use the new checkpoint flag with b9354. It also doesn't work with any of the newer version I've tried.

First Bad Commit

e98cb51

Relevant log output

Logs (b9354)
llama-1  | warn: LLAMA_ARG_HOST environment variable is set, but will be overwritten by command line argument --host
llama-1  | 0.00.454.846 I log_info: verbosity = 3 (adjust with the `-lv N` CLI arg)
llama-1  | 0.00.454.854 I device_info:
llama-1  | 0.00.590.769 I   - CUDA0   : Tesla P100-PCIE-16GB (16269 MiB, 16011 MiB free)
llama-1  | 0.00.703.513 I   - CUDA1   : Tesla P100-PCIE-16GB (16269 MiB, 16011 MiB free)
llama-1  | 0.00.703.528 I   - CPU     : Intel(R) Xeon(R) CPU E5-2690 v3 @ 2.60GHz (15992 MiB, 15992 MiB free)
llama-1  | 0.00.703.673 I system_info: n_threads = 4 (n_threads_batch = 4) / 6 | CUDA : ARCHS = 500,610,700,750,800,860,890,900,1200 | USE_GRAPHS = 1 | PEER_MAX_BATCH_SIZE = 128 | CPU : SSE3 = 1 | SSSE3 = 1 | AVX = 1 | AVX2 = 1 | F16C = 1 | FMA = 1 | BMI2 = 1 | LLAMAFILE = 1 | OPENMP = 1 | REPACK = 1 | 
llama-1  | 0.00.703.683 I srv  llama_server: n_parallel is set to auto, using n_parallel = 4 and kv_unified = true
llama-1  | 0.00.703.800 I srv          init: running without SSL
llama-1  | 0.00.703.853 I srv          init: using 8 threads for HTTP server
llama-1  | 0.00.704.063 I srv         start: binding port with default address family
llama-1  | 0.00.705.579 I srv  llama_server: loading model
llama-1  | 0.00.705.599 I srv    load_model: loading model '/models/Qwopus3.6-27B-v2-MTP-Q4_K_M.gguf'
llama-1  | 0.33.193.984 W llama_context: n_ctx_seq (32768) < n_ctx_train (262144) -- the full capacity of the model will not be utilized
llama-1  | 0.33.613.540 I common_init_from_params: warming up the model with an empty run - please wait ... (--no-warmup to disable)
llama-1  | 1.05.224.075 W srv    load_model: cache_reuse is not supported by this context, it will be disabled
llama-1  | 1.05.224.088 W srv    load_model: swa_full is not supported by this model, it will be disabled
llama-1  | 1.05.224.089 I srv    load_model: initializing slots, n_slots = 4
llama-1  | 1.05.297.363 W srv    load_model: speculative decoding will use checkpoints
llama-1  | 1.05.297.427 W common_speculative_init: no implementations specified for speculative decoding
llama-1  | 1.05.297.429 I slot   load_model: id  0 | task -1 | new slot, n_ctx = 32768
llama-1  | 1.05.297.435 I slot   load_model: id  1 | task -1 | new slot, n_ctx = 32768
llama-1  | 1.05.297.435 I slot   load_model: id  2 | task -1 | new slot, n_ctx = 32768
llama-1  | 1.05.297.436 I slot   load_model: id  3 | task -1 | new slot, n_ctx = 32768
llama-1  | 1.05.297.520 I srv    load_model: prompt cache is enabled, size limit: 2048 MiB
llama-1  | 1.05.297.526 I srv    load_model: use `--cache-ram 0` to disable the prompt cache
llama-1  | 1.05.297.527 I srv    load_model: for more info see https://github.com/ggml-org/llama.cpp/pull/16391
llama-1  | 1.05.297.527 I srv    load_model: context checkpoints enabled, max = 64, min spacing = 512
llama-1  | 1.05.297.579 I srv          init: idle slots will be saved to prompt cache and cleared upon starting a new task
llama-1  | 1.05.331.162 I init: chat template, example_format: '<|im_start|>system
llama-1  | You are a helpful assistant<|im_end|>
llama-1  | <|im_start|>user
llama-1  | Hello<|im_end|>
llama-1  | <|im_start|>assistant
llama-1  | Hi there<|im_end|>
llama-1  | <|im_start|>user
llama-1  | How are you?<|im_end|>
llama-1  | <|im_start|>assistant
llama-1  | <think>
llama-1  | '
llama-1  | 1.05.353.443 I srv          init: init: chat template, thinking = 1
llama-1  | 1.05.353.862 I srv  llama_server: model loaded
llama-1  | 1.05.353.872 I srv  llama_server: server is listening on http://0.0.0.0:8080
llama-1  | 1.05.353.890 I srv  update_slots: all slots are idle
llama-1  | 1.13.064.120 I srv  params_from_: Chat format: peg-native
llama-1  | 1.13.072.419 I slot get_availabl: id  3 | task -1 | selected slot by LRU, t_last = -1
llama-1  | 1.13.072.425 I srv  get_availabl: updating prompt cache
llama-1  | 1.13.072.453 I srv          load:  - looking for better prompt, base f_keep = -1.000, sim = 0.000
llama-1  | 1.13.072.476 I srv        update:  - cache state: 0 prompts, 0.000 MiB (limits: 2048.000 MiB, 32768 tokens, 2147483648 est)
llama-1  | 1.13.072.477 I srv  get_availabl: prompt cache update took 0.05 ms
llama-1  | 1.13.073.320 I reasoning-budget: activated, budget=1024 tokens
llama-1  | 1.13.073.365 I slot launch_slot_: id  3 | task 0 | processing task, is_child = 0
llama-1  | 1.21.658.892 I slot print_timing: id  3 | task 0 | prompt processing, n_tokens =   2048, progress = 0.73, t =   8.59 s / 238.54 tokens per second
llama-1  | 1.22.987.039 I slot print_timing: id  3 | task 0 | prompt processing, n_tokens =   2295, progress = 0.82, t =   9.91 s / 231.50 tokens per second
llama-1  | 1.24.832.409 I slot print_timing: id  3 | task 0 | prompt processing, n_tokens =   2800, progress = 1.00, t =  11.76 s / 238.12 tokens per second
llama-1  | 1.25.561.648 I slot create_check: id  3 | task 0 | created context checkpoint 1 of 64 (pos_min = 2799, pos_max = 2799, n_tokens = 2800, size = 149.626 MiB)
llama-1  | 1.25.672.658 I slot print_timing: id  3 | task 0 | prompt processing, n_tokens =   2807, progress = 1.00, t =  12.60 s / 222.79 tokens per second
llama-1  | 1.27.911.875 I reasoning-budget: deactivated (natural end)
llama-1  | 1.28.878.535 I slot print_timing: id  3 | task 0 | prompt eval time =   12758.15 ms /  2811 tokens (    4.54 ms per token,   220.33 tokens per second)
llama-1  | 1.28.878.541 I slot print_timing: id  3 | task 0 |        eval time =    3046.97 ms /    54 tokens (   56.43 ms per token,    17.72 tokens per second)
llama-1  | 1.28.878.542 I slot print_timing: id  3 | task 0 |       total time =   15805.11 ms /  2865 tokens
llama-1  | 1.28.878.565 I slot print_timing: id  3 | task 0 |    graphs reused =         52
llama-1  | 1.28.878.951 I slot      release: id  3 | task 0 | stop processing: n_tokens = 2864, truncated = 0
llama-1  | 1.28.878.963 I srv  update_slots: all slots are idle
llama-1  | 1.34.286.919 I srv  params_from_: Chat format: peg-native
llama-1  | 1.34.295.157 I slot get_availabl: id  3 | task -1 | selected slot by LCP similarity, sim_best = 0.928 (> 0.100 thold), f_keep = 0.911
llama-1  | 1.34.295.809 I reasoning-budget: activated, budget=1024 tokens
llama-1  | 1.34.295.886 I slot launch_slot_: id  3 | task 59 | processing task, is_child = 0
llama-1  | 1.34.295.902 I slot update_slots: id  3 | task 59 | Checking checkpoint with [2799, 2799] against 2608...
llama-1  | 1.34.295.903 W slot update_slots: id  3 | task 59 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
llama-1  | 1.34.295.906 W slot update_slots: id  3 | task 59 | erased invalidated context checkpoint (pos_min = 2799, pos_max = 2799, n_tokens = 2800, n_swa = 0, pos_next = 0, size = 149.626 MiB)
llama-1  | 1.42.539.793 I slot print_timing: id  3 | task 59 | prompt processing, n_tokens =   2048, progress = 0.73, t =   8.24 s / 248.43 tokens per second
llama-1  | 1.43.871.974 I slot print_timing: id  3 | task 59 | prompt processing, n_tokens =   2295, progress = 0.82, t =   9.58 s / 239.66 tokens per second
llama-1  | 1.45.713.099 I slot print_timing: id  3 | task 59 | prompt processing, n_tokens =   2798, progress = 1.00, t =  11.42 s / 245.07 tokens per second
llama-1  | 1.46.442.846 I slot create_check: id  3 | task 59 | created context checkpoint 1 of 64 (pos_min = 2797, pos_max = 2797, n_tokens = 2798, size = 149.626 MiB)
llama-1  | 1.46.757.723 I slot print_timing: id  3 | task 59 | prompt processing, n_tokens =   2807, progress = 1.00, t =  12.46 s / 225.25 tokens per second
llama-1  | 1.49.042.920 I reasoning-budget: deactivated (natural end)
llama-1  | 1.49.655.135 I slot print_timing: id  3 | task 59 | prompt eval time =   12680.06 ms /  2811 tokens (    4.51 ms per token,   221.69 tokens per second)
llama-1  | 1.49.655.140 I slot print_timing: id  3 | task 59 |        eval time =    2679.16 ms /    52 tokens (   51.52 ms per token,    19.41 tokens per second)
llama-1  | 1.49.655.141 I slot print_timing: id  3 | task 59 |       total time =   15359.22 ms /  2863 tokens
llama-1  | 1.49.655.142 I slot print_timing: id  3 | task 59 |    graphs reused =        101
llama-1  | 1.49.655.437 I slot      release: id  3 | task 59 | stop processing: n_tokens = 2862, truncated = 0
llama-1  | 1.49.655.515 I srv  update_slots: all slots are idle
Logs (b9309)
llama-1  | warn: LLAMA_ARG_HOST environment variable is set, but will be overwritten by command line argument --host
llama-1  | 0.00.480.215 I log_info: verbosity = 3 (adjust with the `-lv N` CLI arg)
llama-1  | 0.00.480.222 I device_info:
llama-1  | 0.00.609.625 I   - CUDA0   : Tesla P100-PCIE-16GB (16269 MiB, 16011 MiB free)
llama-1  | 0.00.722.524 I   - CUDA1   : Tesla P100-PCIE-16GB (16269 MiB, 16011 MiB free)
llama-1  | 0.00.722.555 I   - CPU     : Intel(R) Xeon(R) CPU E5-2690 v3 @ 2.60GHz (15992 MiB, 15992 MiB free)
llama-1  | 0.00.723.757 I system_info: n_threads = 4 (n_threads_batch = 4) / 6 | CUDA : ARCHS = 500,610,700,750,800,860,890,900,1200 | USE_GRAPHS = 1 | PEER_MAX_BATCH_SIZE = 128 | CPU : SSE3 = 1 | SSSE3 = 1 | AVX = 1 | AVX2 = 1 | F16C = 1 | FMA = 1 | BMI2 = 1 | LLAMAFILE = 1 | OPENMP = 1 | REPACK = 1 | 
llama-1  | 0.00.723.771 I srv  llama_server: n_parallel is set to auto, using n_parallel = 4 and kv_unified = true
llama-1  | 0.00.724.330 I srv          init: running without SSL
llama-1  | 0.00.725.526 I srv          init: using 8 threads for HTTP server
llama-1  | 0.00.726.136 I srv         start: binding port with default address family
llama-1  | 0.00.727.361 I srv  llama_server: loading model
llama-1  | 0.00.727.710 I srv    load_model: loading model '/models/Qwopus3.6-27B-v2-MTP-Q4_K_M.gguf'
llama-1  | 0.33.784.626 W llama_context: n_ctx_seq (32768) < n_ctx_train (262144) -- the full capacity of the model will not be utilized
llama-1  | 0.34.233.650 I common_init_from_params: warming up the model with an empty run - please wait ... (--no-warmup to disable)
llama-1  | 1.05.948.146 W srv    load_model: cache_reuse is not supported by this context, it will be disabled
llama-1  | 1.05.948.157 W srv    load_model: swa_full is not supported by this model, it will be disabled
llama-1  | 1.05.948.158 I srv    load_model: initializing slots, n_slots = 4
llama-1  | 1.06.021.651 W srv    load_model: speculative decoding will use checkpoints
llama-1  | 1.06.021.702 W common_speculative_init: no implementations specified for speculative decoding
llama-1  | 1.06.021.705 I slot   load_model: id  0 | task -1 | new slot, n_ctx = 32768
llama-1  | 1.06.021.712 I slot   load_model: id  1 | task -1 | new slot, n_ctx = 32768
llama-1  | 1.06.021.713 I slot   load_model: id  2 | task -1 | new slot, n_ctx = 32768
llama-1  | 1.06.021.713 I slot   load_model: id  3 | task -1 | new slot, n_ctx = 32768
llama-1  | 1.06.021.789 I srv    load_model: prompt cache is enabled, size limit: 2048 MiB
llama-1  | 1.06.021.795 I srv    load_model: use `--cache-ram 0` to disable the prompt cache
llama-1  | 1.06.021.796 I srv    load_model: for more info see https://github.com/ggml-org/llama.cpp/pull/16391
llama-1  | 1.06.022.697 I srv          init: idle slots will be saved to prompt cache and cleared upon starting a new task
llama-1  | 1.06.056.550 I init: chat template, example_format: '<|im_start|>system
llama-1  | You are a helpful assistant<|im_end|>
llama-1  | <|im_start|>user
llama-1  | Hello<|im_end|>
llama-1  | <|im_start|>assistant
llama-1  | Hi there<|im_end|>
llama-1  | <|im_start|>user
llama-1  | How are you?<|im_end|>
llama-1  | <|im_start|>assistant
llama-1  | <think>
llama-1  | '
llama-1  | 1.06.075.840 I srv          init: init: chat template, thinking = 1
llama-1  | 1.06.076.391 I srv  llama_server: model loaded
llama-1  | 1.06.076.404 I srv  llama_server: server is listening on http://0.0.0.0:8080
llama-1  | 1.06.076.439 I srv  update_slots: all slots are idle
llama-1  | 1.26.044.184 I srv  params_from_: Chat format: peg-native
llama-1  | 1.26.048.001 I slot get_availabl: id  3 | task -1 | selected slot by LRU, t_last = -1
llama-1  | 1.26.048.007 I srv  get_availabl: updating prompt cache
llama-1  | 1.26.048.050 I srv          load:  - looking for better prompt, base f_keep = -1.000, sim = 0.000
llama-1  | 1.26.048.064 I srv        update:  - cache state: 0 prompts, 0.000 MiB (limits: 2048.000 MiB, 32768 tokens, 2147483648 est)
llama-1  | 1.26.048.065 I srv  get_availabl: prompt cache update took 0.06 ms
llama-1  | 1.26.049.458 I reasoning-budget: activated, budget=1024 tokens
llama-1  | 1.26.049.520 I slot launch_slot_: id  3 | task 0 | processing task, is_child = 0
llama-1  | 1.34.617.418 I slot print_timing: id  3 | task 0 | prompt processing, n_tokens =   2048, progress = 0.73, t =   8.57 s / 239.04 tokens per second
llama-1  | 1.34.617.473 I slot update_slots: id  3 | task 0 | 512 tokens since last checkpoint at 0, creating new checkpoint during processing at position 2300
llama-1  | 1.35.522.979 I slot create_check: id  3 | task 0 | created context checkpoint 1 of 64 (pos_min = 2047, pos_max = 2047, n_tokens = 2048, size = 149.626 MiB)
llama-1  | 1.36.252.464 I slot print_timing: id  3 | task 0 | prompt processing, n_tokens =   2300, progress = 0.82, t =  10.20 s / 225.43 tokens per second
llama-1  | 1.36.563.631 I slot create_check: id  3 | task 0 | created context checkpoint 2 of 64 (pos_min = 2299, pos_max = 2299, n_tokens = 2300, size = 149.626 MiB)
llama-1  | 1.38.187.796 I slot print_timing: id  3 | task 0 | prompt processing, n_tokens =   2812, progress = 1.00, t =  12.14 s / 231.67 tokens per second
llama-1  | 1.38.868.129 I slot create_check: id  3 | task 0 | created context checkpoint 3 of 64 (pos_min = 2811, pos_max = 2811, n_tokens = 2812, size = 149.626 MiB)
llama-1  | 1.40.820.744 I reasoning-budget: deactivated (natural end)
llama-1  | 1.41.635.222 I slot print_timing: id  3 | task 0 | prompt eval time =   12937.92 ms /  2816 tokens (    4.59 ms per token,   217.65 tokens per second)
llama-1  | 1.41.635.228 I slot print_timing: id  3 | task 0 |        eval time =    2647.45 ms /    52 tokens (   50.91 ms per token,    19.64 tokens per second)
llama-1  | 1.41.635.229 I slot print_timing: id  3 | task 0 |       total time =   15585.36 ms /  2868 tokens
llama-1  | 1.41.635.247 I slot print_timing: id  3 | task 0 |    graphs reused =         51
llama-1  | 1.41.635.648 I slot      release: id  3 | task 0 | stop processing: n_tokens = 2867, truncated = 0
llama-1  | 1.41.635.663 I srv  update_slots: all slots are idle
llama-1  | 1.46.855.082 I srv  params_from_: Chat format: peg-native
llama-1  | 1.46.858.399 I slot get_availabl: id  3 | task -1 | selected slot by LCP similarity, sim_best = 0.926 (> 0.100 thold), f_keep = 0.910
llama-1  | 1.46.858.991 I reasoning-budget: activated, budget=1024 tokens
llama-1  | 1.46.859.070 I slot launch_slot_: id  3 | task 56 | processing task, is_child = 0
llama-1  | 1.46.859.086 W slot update_slots: id  3 | task 56 | n_past = 2609, slot.prompt.tokens.size() = 2867, seq_id = 3, pos_min = 2866, n_swa = 0
llama-1  | 1.46.859.086 I slot update_slots: id  3 | task 56 | Checking checkpoint with [2811, 2811] against 2608...
llama-1  | 1.46.859.087 I slot update_slots: id  3 | task 56 | Checking checkpoint with [2299, 2299] against 2608...
llama-1  | 1.46.897.320 W slot update_slots: id  3 | task 56 | restored context checkpoint (pos_min = 2299, pos_max = 2299, n_tokens = 2300, n_past = 2300, size = 149.626 MiB)
llama-1  | 1.46.897.334 W slot update_slots: id  3 | task 56 | erased invalidated context checkpoint (pos_min = 2811, pos_max = 2811, n_tokens = 2812, n_swa = 0, pos_next = 2300, size = 149.626 MiB)
llama-1  | 1.49.296.281 I slot create_check: id  3 | task 56 | created context checkpoint 3 of 64 (pos_min = 2812, pos_max = 2812, n_tokens = 2813, size = 149.626 MiB)
llama-1  | 1.51.204.185 I reasoning-budget: deactivated (natural end)
llama-1  | 1.52.071.409 I slot print_timing: id  3 | task 56 | prompt eval time =    2581.81 ms /   517 tokens (    4.99 ms per token,   200.25 tokens per second)
llama-1  | 1.52.071.416 I slot print_timing: id  3 | task 56 |        eval time =    2630.47 ms /    52 tokens (   50.59 ms per token,    19.77 tokens per second)
llama-1  | 1.52.071.417 I slot print_timing: id  3 | task 56 |       total time =    5212.28 ms /   569 tokens
llama-1  | 1.52.071.418 I slot print_timing: id  3 | task 56 |    graphs reused =        101
llama-1  | 1.52.071.681 I slot      release: id  3 | task 56 | stop processing: n_tokens = 2868, truncated = 0
llama-1  | 1.52.071.757 I srv  update_slots: all slots are idle

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions