llama-1 | warn: LLAMA_ARG_HOST environment variable is set, but will be overwritten by command line argument --host
llama-1 | 0.00.454.846 I log_info: verbosity = 3 (adjust with the `-lv N` CLI arg)
llama-1 | 0.00.454.854 I device_info:
llama-1 | 0.00.590.769 I - CUDA0 : Tesla P100-PCIE-16GB (16269 MiB, 16011 MiB free)
llama-1 | 0.00.703.513 I - CUDA1 : Tesla P100-PCIE-16GB (16269 MiB, 16011 MiB free)
llama-1 | 0.00.703.528 I - CPU : Intel(R) Xeon(R) CPU E5-2690 v3 @ 2.60GHz (15992 MiB, 15992 MiB free)
llama-1 | 0.00.703.673 I system_info: n_threads = 4 (n_threads_batch = 4) / 6 | CUDA : ARCHS = 500,610,700,750,800,860,890,900,1200 | USE_GRAPHS = 1 | PEER_MAX_BATCH_SIZE = 128 | CPU : SSE3 = 1 | SSSE3 = 1 | AVX = 1 | AVX2 = 1 | F16C = 1 | FMA = 1 | BMI2 = 1 | LLAMAFILE = 1 | OPENMP = 1 | REPACK = 1 |
llama-1 | 0.00.703.683 I srv llama_server: n_parallel is set to auto, using n_parallel = 4 and kv_unified = true
llama-1 | 0.00.703.800 I srv init: running without SSL
llama-1 | 0.00.703.853 I srv init: using 8 threads for HTTP server
llama-1 | 0.00.704.063 I srv start: binding port with default address family
llama-1 | 0.00.705.579 I srv llama_server: loading model
llama-1 | 0.00.705.599 I srv load_model: loading model '/models/Qwopus3.6-27B-v2-MTP-Q4_K_M.gguf'
llama-1 | 0.33.193.984 W llama_context: n_ctx_seq (32768) < n_ctx_train (262144) -- the full capacity of the model will not be utilized
llama-1 | 0.33.613.540 I common_init_from_params: warming up the model with an empty run - please wait ... (--no-warmup to disable)
llama-1 | 1.05.224.075 W srv load_model: cache_reuse is not supported by this context, it will be disabled
llama-1 | 1.05.224.088 W srv load_model: swa_full is not supported by this model, it will be disabled
llama-1 | 1.05.224.089 I srv load_model: initializing slots, n_slots = 4
llama-1 | 1.05.297.363 W srv load_model: speculative decoding will use checkpoints
llama-1 | 1.05.297.427 W common_speculative_init: no implementations specified for speculative decoding
llama-1 | 1.05.297.429 I slot load_model: id 0 | task -1 | new slot, n_ctx = 32768
llama-1 | 1.05.297.435 I slot load_model: id 1 | task -1 | new slot, n_ctx = 32768
llama-1 | 1.05.297.435 I slot load_model: id 2 | task -1 | new slot, n_ctx = 32768
llama-1 | 1.05.297.436 I slot load_model: id 3 | task -1 | new slot, n_ctx = 32768
llama-1 | 1.05.297.520 I srv load_model: prompt cache is enabled, size limit: 2048 MiB
llama-1 | 1.05.297.526 I srv load_model: use `--cache-ram 0` to disable the prompt cache
llama-1 | 1.05.297.527 I srv load_model: for more info see https://github.com/ggml-org/llama.cpp/pull/16391
llama-1 | 1.05.297.527 I srv load_model: context checkpoints enabled, max = 64, min spacing = 512
llama-1 | 1.05.297.579 I srv init: idle slots will be saved to prompt cache and cleared upon starting a new task
llama-1 | 1.05.331.162 I init: chat template, example_format: '<|im_start|>system
llama-1 | You are a helpful assistant<|im_end|>
llama-1 | <|im_start|>user
llama-1 | Hello<|im_end|>
llama-1 | <|im_start|>assistant
llama-1 | Hi there<|im_end|>
llama-1 | <|im_start|>user
llama-1 | How are you?<|im_end|>
llama-1 | <|im_start|>assistant
llama-1 | <think>
llama-1 | '
llama-1 | 1.05.353.443 I srv init: init: chat template, thinking = 1
llama-1 | 1.05.353.862 I srv llama_server: model loaded
llama-1 | 1.05.353.872 I srv llama_server: server is listening on http://0.0.0.0:8080
llama-1 | 1.05.353.890 I srv update_slots: all slots are idle
llama-1 | 1.13.064.120 I srv params_from_: Chat format: peg-native
llama-1 | 1.13.072.419 I slot get_availabl: id 3 | task -1 | selected slot by LRU, t_last = -1
llama-1 | 1.13.072.425 I srv get_availabl: updating prompt cache
llama-1 | 1.13.072.453 I srv load: - looking for better prompt, base f_keep = -1.000, sim = 0.000
llama-1 | 1.13.072.476 I srv update: - cache state: 0 prompts, 0.000 MiB (limits: 2048.000 MiB, 32768 tokens, 2147483648 est)
llama-1 | 1.13.072.477 I srv get_availabl: prompt cache update took 0.05 ms
llama-1 | 1.13.073.320 I reasoning-budget: activated, budget=1024 tokens
llama-1 | 1.13.073.365 I slot launch_slot_: id 3 | task 0 | processing task, is_child = 0
llama-1 | 1.21.658.892 I slot print_timing: id 3 | task 0 | prompt processing, n_tokens = 2048, progress = 0.73, t = 8.59 s / 238.54 tokens per second
llama-1 | 1.22.987.039 I slot print_timing: id 3 | task 0 | prompt processing, n_tokens = 2295, progress = 0.82, t = 9.91 s / 231.50 tokens per second
llama-1 | 1.24.832.409 I slot print_timing: id 3 | task 0 | prompt processing, n_tokens = 2800, progress = 1.00, t = 11.76 s / 238.12 tokens per second
llama-1 | 1.25.561.648 I slot create_check: id 3 | task 0 | created context checkpoint 1 of 64 (pos_min = 2799, pos_max = 2799, n_tokens = 2800, size = 149.626 MiB)
llama-1 | 1.25.672.658 I slot print_timing: id 3 | task 0 | prompt processing, n_tokens = 2807, progress = 1.00, t = 12.60 s / 222.79 tokens per second
llama-1 | 1.27.911.875 I reasoning-budget: deactivated (natural end)
llama-1 | 1.28.878.535 I slot print_timing: id 3 | task 0 | prompt eval time = 12758.15 ms / 2811 tokens ( 4.54 ms per token, 220.33 tokens per second)
llama-1 | 1.28.878.541 I slot print_timing: id 3 | task 0 | eval time = 3046.97 ms / 54 tokens ( 56.43 ms per token, 17.72 tokens per second)
llama-1 | 1.28.878.542 I slot print_timing: id 3 | task 0 | total time = 15805.11 ms / 2865 tokens
llama-1 | 1.28.878.565 I slot print_timing: id 3 | task 0 | graphs reused = 52
llama-1 | 1.28.878.951 I slot release: id 3 | task 0 | stop processing: n_tokens = 2864, truncated = 0
llama-1 | 1.28.878.963 I srv update_slots: all slots are idle
llama-1 | 1.34.286.919 I srv params_from_: Chat format: peg-native
llama-1 | 1.34.295.157 I slot get_availabl: id 3 | task -1 | selected slot by LCP similarity, sim_best = 0.928 (> 0.100 thold), f_keep = 0.911
llama-1 | 1.34.295.809 I reasoning-budget: activated, budget=1024 tokens
llama-1 | 1.34.295.886 I slot launch_slot_: id 3 | task 59 | processing task, is_child = 0
llama-1 | 1.34.295.902 I slot update_slots: id 3 | task 59 | Checking checkpoint with [2799, 2799] against 2608...
llama-1 | 1.34.295.903 W slot update_slots: id 3 | task 59 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
llama-1 | 1.34.295.906 W slot update_slots: id 3 | task 59 | erased invalidated context checkpoint (pos_min = 2799, pos_max = 2799, n_tokens = 2800, n_swa = 0, pos_next = 0, size = 149.626 MiB)
llama-1 | 1.42.539.793 I slot print_timing: id 3 | task 59 | prompt processing, n_tokens = 2048, progress = 0.73, t = 8.24 s / 248.43 tokens per second
llama-1 | 1.43.871.974 I slot print_timing: id 3 | task 59 | prompt processing, n_tokens = 2295, progress = 0.82, t = 9.58 s / 239.66 tokens per second
llama-1 | 1.45.713.099 I slot print_timing: id 3 | task 59 | prompt processing, n_tokens = 2798, progress = 1.00, t = 11.42 s / 245.07 tokens per second
llama-1 | 1.46.442.846 I slot create_check: id 3 | task 59 | created context checkpoint 1 of 64 (pos_min = 2797, pos_max = 2797, n_tokens = 2798, size = 149.626 MiB)
llama-1 | 1.46.757.723 I slot print_timing: id 3 | task 59 | prompt processing, n_tokens = 2807, progress = 1.00, t = 12.46 s / 225.25 tokens per second
llama-1 | 1.49.042.920 I reasoning-budget: deactivated (natural end)
llama-1 | 1.49.655.135 I slot print_timing: id 3 | task 59 | prompt eval time = 12680.06 ms / 2811 tokens ( 4.51 ms per token, 221.69 tokens per second)
llama-1 | 1.49.655.140 I slot print_timing: id 3 | task 59 | eval time = 2679.16 ms / 52 tokens ( 51.52 ms per token, 19.41 tokens per second)
llama-1 | 1.49.655.141 I slot print_timing: id 3 | task 59 | total time = 15359.22 ms / 2863 tokens
llama-1 | 1.49.655.142 I slot print_timing: id 3 | task 59 | graphs reused = 101
llama-1 | 1.49.655.437 I slot release: id 3 | task 59 | stop processing: n_tokens = 2862, truncated = 0
llama-1 | 1.49.655.515 I srv update_slots: all slots are idle
Name and Version
started with docker image: server-cuda12-b9354
Operating systems
Linux
Which llama.cpp modules do you know to be affected?
llama-server
Command line
services: llama: image: ghcr.io/ggml-org/llama.cpp:server-cuda12-b9354 runtime: nvidia environment: - NVIDIA_VISIBLE_DEVICES=all volumes: - ./models:/models ports: - "8000:8080" command: > --model /models/Qwopus3.6-27B-v2-MTP-Q4_K_M.gguf --n-gpu-layers 999 --tensor-split 0.5,0.5 --jinja --split-mode tensor --swa-full --ctx-size 32768 --ctx-checkpoints 64 --checkpoint-min-step 512 --cache-prompt --cache-reuse 256 --cache-ram 2048 --threads 4 --threads-batch 4 --flash-attn on --batch-size 2048 --ubatch-size 512 --cache-type-k f16 --cache-type-v f16 --fit off --kv-unified --kv-offload --cont-batching --op-offload --repack --no-mmap --temp 0.6 --top-k 20 --top-p 0.95 --min-p 0 --reasoning-budget 1024 --host 0.0.0.0 --port 8080 ipc: host restart: unless-stoppedProblem description & steps to reproduce
the
--checkpoint-min-stepflag doesn't work and forces full prompt reprocessing every time. When i switch to the docker image before that update and use--checkpoint-every-n-tokensinstead, the checkpoints work just fine. Both logs use the same compose except I'm forced to use the new checkpoint flag with b9354. It also doesn't work with any of the newer version I've tried.First Bad Commit
e98cb51
Relevant log output
Logs (b9354)
Logs (b9309)