Skip to content

llama : fix K/V and recurrent state cleanup after failed restores - #27530

Merged
ggerganov merged 11 commits into
ggml-org:masterfrom
CHIPMUNK-T0T:fix/state-restore-failure-cleanup
Sep 26, 2026
Merged

ggerganov merged 11 commits into
ggml-org:masterfrom
CHIPMUNK-T0T:fix/state-restore-failure-cleanup

Conversation

@CHIPMUNK-T0T

@CHIPMUNK-T0T CHIPMUNK-T0T commented Aug 22, 2026 •

Copy link
Copy Markdown
Contributor

Overview

A failed per-sequence state restore can remove sequence metadata while leaving K/V or recurrent-state tensor data written by the failed restore.

Buffer-backed restores can also leave deferred writes pending, allowing them to be applied after cleanup.

This PR contains failed restores by:

  • clearing affected K/V or recurrent-state tensor data;
  • discarding deferred writes after cleanup and before the reader is destroyed;
  • applying the same deferred-write cleanup to whole-context buffer restores;
  • validating restored cell_count values before they are used for allocation or target-range setup;
  • cleaning up an already-restored attention state when a later recurrent-memory restore fails in llama_memory_hybrid.

Failures before modifying the target sequence leave it unchanged. After modification, the affected state is removed rather than attempting transactional rollback. For llama_memory_hybrid, if the attention component has already been restored and the recurrent component then fails, the restored attention state is also cleared.

Fixes #27068.

Additional information

The original issue was observed in the K/V cache, but the same failure-cleanup behavior is needed for recurrent memory.

Tensor cleanup is only performed after a valid target range has been established, so failures during metadata parsing do not clear unrelated cells.

Per-sequence restore now also validates cell_count against the available cache or recurrent-memory capacity before using it.

For llama_memory_hybrid, restore is sequential across the attention and recurrent components. If the recurrent restore throws after the attention restore has succeeded, the attention state restored for that sequence is cleared before the exception is rethrown.

The cleanup runs only on restore failure and does not affect successful restore or normal decoding.

Tests

Added failed-restore regression coverage to test-save-load-state. Test 9 exercises corrupted buffer and file restores and runs through the existing architecture coverage, including hybrid models.

The test verifies that a failed restore leaves the affected sequence empty and does not leave tensor data that changes the logits of another sequence.

Local validation:

  • ctest: 62/62 passed.
  • Sequential run across 127 generated models: Test 9 PASS 123 / SKIP 4 / FAIL 0. The four skipped models (dream, llada, llada-moe, and rnd1) do not have memory.

CUDA validation and server-level fault injection with K/V and hybrid recurrent memory were also performed on an earlier base.

Non-goal

This does not provide transactional restore. If a failed restore has already modified the target sequence, its previous contents are not reconstructed.

Cross-component cleanup is handled for llama_memory_hybrid, where attention restore can precede a failing recurrent restore. Equivalent cleanup for the other composite memory implementations (llama_kv_cache_iswa, llama_memory_hybrid_iswa, llama_kv_cache_dsa_iswa, and llama_kv_cache_msa) is not addressed by this PR.

The separate ON_DEVICE pre-validation issue is tracked in #27439.

Requirements

  • I have read and agree with the contributing guidelines

  • AI usage disclosure: YES

    • Motivation & Design: I found the failure through slot save / restore fault injection, reported it in Eval bug: failed slot restore leaves corrupted K/V data that breaks subsequent inference #27068, and defined the scope as failure containment rather than transactional restore.
    • Investigation & Code Review (Fable 5 / Opus 5 / GPT-5.6 Sol): AI agents assisted investigation and code review under my instructions.
    • Implementation (Opus 5 / GPT-5.6 Sol): AI assistance was used to draft parts of the implementation and tests.
    • Draft Translation (Opus 5): AI assisted with English translation of my Japanese draft.
    • Validation & Responsibility: I manually reviewed the changes, ran local tests, mutation tests, CUDA validation, and server-level fault injection, and take responsibility for the submitted changes.

@github-actions github-actions Bot added the testing Everything test related label Aug 22, 2026
@ggerganov ggerganov self-assigned this Aug 22, 2026
@CHIPMUNK-T0T
CHIPMUNK-T0T force-pushed the fix/state-restore-failure-cleanup branch 2 times, most recently from 375cbde to f931b4d Compare August 28, 2026 14:51
@CHIPMUNK-T0T

CHIPMUNK-T0T commented Aug 28, 2026 •

Copy link
Copy Markdown
Contributor Author

@ggerganov
I have rebased onto the latest upstream master and resolved the conflict.
Since this PR is from a fork, the CI workflows are stopped at action_required.
Could you approve the workflow runs when you have time?

@CHIPMUNK-T0T
CHIPMUNK-T0T force-pushed the fix/state-restore-failure-cleanup branch from f931b4d to 5313ec0 Compare September 13, 2026 00:03
@CHIPMUNK-T0T

Copy link
Copy Markdown
Contributor Author

@ggerganov @ngxson

After rebasing this PR onto the latest upstream master, I have one question about a performance asymmetry that became apparent after #27991 was merged.

clear_cells_data() clears the target cells of a failed restore. #27991 now batches state_read_data() by contiguous runs, while the cleanup still clears non-contiguous cells one by one.

This does not affect correctness and only affects the restore-failure path, but fragmented state can cause many more backend writes during cleanup.

Would you prefer that I:

  1. update clear_cells_data() in this PR to use the same contiguous-run grouping, duplicating the small runs calculation;
  2. extract the run calculation into a shared helper used by both state_read_data() and clear_cells_data(), which would touch the newly-upstreamed restore code as well; or
  3. keep this PR focused on correctness/failure containment and handle the fragmented cleanup optimization separately?

I'm leaning toward (3) to keep the scope focused, although (1) would be a small local change.

@CHIPMUNK-T0T

Copy link
Copy Markdown
Contributor Author

I will review the test cases myself, and update later.

@CHIPMUNK-T0T
CHIPMUNK-T0T force-pushed the fix/state-restore-failure-cleanup branch 2 times, most recently from 34ccdb4 to 1348cb0 Compare September 13, 2026 06:47
@CHIPMUNK-T0T

CHIPMUNK-T0T commented Sep 13, 2026 •

Copy link
Copy Markdown
Contributor Author

@ggerganov @ngxson

I'd like to confirm one thing about the tests in this PR.

The existing state-related tests mainly cover the normal save/load paths, so the failed-restore cleanup paths added in this PR are not exercised by the current CI.

On the other hand, if I add this test to test-save-load-state, which runs across all architectures since #27755, triggering restore failures with a corrupted state exposes existing issues in some hybrid architectures that are separate from this PR:

  • a corrupted cell_count is used for allocation before it is sufficiently validated, which can lead to excessive memory allocation or OOM
  • rollback / cleanup on restore failure is not consistent across the memory components of hybrid models

I consider both out of scope for this PR, and I don't think adding per-architecture exclusions to test-save-load-state just to avoid them is the right approach.
For now, this PR adds a small standalone test using llama-dense and mamba-dense to exercise the fix.
However, adding a new test file requires maintainer approval, and I'm not sure a dedicated test just for this case is justified.
Should I add only a minimal case to an existing test, or should this PR not have a dedicated regression test?

I'd appreciate your guidance on how to handle the tests.

@ggerganov

Copy link
Copy Markdown
Member

Ideally, this test should be part of the test-save-load-state suite. How difficult do you think it would be to resolve the issues that were surfaced when adding it there?

@CHIPMUNK-T0T

CHIPMUNK-T0T commented Sep 18, 2026 •

Copy link
Copy Markdown
Contributor Author

@ggerganov

Adding the regression test to test-save-load-state surfaced two issues that are separate from the original PR.

1. Missing cell_count validation

A corrupted cell_count in the state is not sufficiently validated before allocation, which leads to an OOM / bad_alloc.

The cause is that the per-sequence restore does not have the same bounds check as the whole-cache path.
Adding a bounds check to llama_kv_cache and llama_memory_recurrent is enough, and the change is about 10 lines.

With this fix applied, the 111 generated models run without any bad_alloc.

fix code
--- a/src/llama-kv-cache.cpp
+++ b/src/llama-kv-cache.cpp
 bool llama_kv_cache::state_read_meta(...)
         // single sequence
+        if (cell_count > cells.size()) {
+            LLAMA_LOG_ERROR("%s: not enough cells in kv cache\n", __func__);
+            return false;
+        }
+
         seq_rm(dest_seq_id, -1, -1);

--- a/src/llama-memory-recurrent.cpp
+++ b/src/llama-memory-recurrent.cpp
 bool llama_memory_recurrent::state_read_meta(...)
         // single sequence
+        if (cell_count > size) {
+            LLAMA_LOG_ERROR("%s: not enough cells in kv cache\n", __func__);
+            return false;
+        }
+
         seq_rm(dest_seq_id, -1, -1);

2. Failed restore in composite / hybrid memory

When the first memory component is restored and a later component then fails, the components are left inconsistent with each other.

With the fix above applied, the only architectures that fail in the current test are bailingmoe3, kimi-k3 and kimi-linear, and all of them use llama_memory_hybrid. The other hybrid / composite memory types pass in this test.

So the minimal fix needed to make the test pass looks like it can be limited to llama_memory_hybrid. That said, there are other composite memory implementations with the same sequential restore structure.

I have not prototyped this fix yet, so I have not confirmed the exact size of the change.

On the scope of this PR

These two have a different cause and touch different code than this PR, so I think it would be more natural to handle them in separate PRs.

However, if they are split out, adding the regression test to test-save-load-state would fail on 3 architectures, so this PR alone cannot integrate the test.

On the other hand, if this PR also fixes the two issues above, the regression test can be integrated into test-save-load-state. I do not think it makes much sense to add a standalone test permanently just for this case.

So, which would you prefer:

  • include the fixes needed for the test-save-load-state integration in this PR, or
  • handle each issue in its own issue/PR, and add the regression test in a follow-up PR?

Since bug-fix PRs are expected to have a reproducible issue and a regression test, if they are split out, should I open an issue for each of them first and then submit the PRs?

Comment thread src/llama-context.cpp Outdated
size_t buf_size = 0;
size_t size_read = 0;

bool discarded = false;

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Instead of this stateful flag for keeping track of discard(), can we simply clear the rinfos and the other data such that the destructor remains a noop.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@ggerganov
Thank you, updated discard() to clear rinfos and the buffer state directly.

@ggerganov ggerganov left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

include the fixes needed for the test-save-load-state integration in this PR,

Let's try doing it this way - prefer to not introduce a new test binary for that yet.

Comment thread src/llama-memory-recurrent.h Outdated
Comment thread src/llama-impl.cpp
@CHIPMUNK-T0T
CHIPMUNK-T0T force-pushed the fix/state-restore-failure-cleanup branch from 1348cb0 to e8fcf8f Compare September 21, 2026 05:40
@CHIPMUNK-T0T

Copy link
Copy Markdown
Contributor Author

@ggerganov
I have addressed your review comments and fixed the two issues surfaced by the test, and committed the changes.

For issue 2, when the recurrent part fails during a hybrid restore, only the attention part was left in its restored state, so the attention state is now cleared on failure.

I would appreciate it if you could take another look when you have time.

@ggerganov

Copy link
Copy Markdown
Member

Thanks, looks good. Let's wait for #29133 to merge, rebase this PR and run CI to confirm all is good.

@CHIPMUNK-T0T
CHIPMUNK-T0T force-pushed the fix/state-restore-failure-cleanup branch from e8fcf8f to e8531db Compare September 21, 2026 12:45
@CHIPMUNK-T0T

Copy link
Copy Markdown
Contributor Author

removed llama_kv_cache_dsa from non-goal

@CHIPMUNK-T0T

Copy link
Copy Markdown
Contributor Author

@ggerganov
After the rebase, the same problem showed up in llama_kv_cache_dsa, so I fixed it in the same way.

However, this problem can occur in every composite memory class that restores its components sequentially, so fixing it class by class becomes whack-a-mole.

It may be better to handle this at a common level rather than per class.

@ggerganov

Copy link
Copy Markdown
Member

With the test-save-load-state running across all supported architectures, the new test should be able to flag such problems early on, so for now I think we are convered.

If you have ideas how to solve this at a common level, we can explore. But I don't think we want so very complicated mechanism for this - should be something simple if possible at all.

@ggerganov

Copy link
Copy Markdown
Member

@CHIPMUNK-T0T I made some changes to the test-save-load-state.cpp - let's rebase.

@CHIPMUNK-T0T
CHIPMUNK-T0T force-pushed the fix/state-restore-failure-cleanup branch from e8531db to dd10bf3 Compare September 26, 2026 02:17
@CHIPMUNK-T0T

CHIPMUNK-T0T commented Sep 26, 2026 •

Copy link
Copy Markdown
Contributor Author

@ggerganov

Rebased onto the latest upstream, and updated Test 9 to the new results table.

Test 9 does not apply to models without memory, since there is no state to restore.
Since all_passed() requires every cell to be PASS, the test returns true for them (logged as "PASS (model has no memory)"), as before the rebase.

ctest 62/62 and --models over the 127 generated models: 127 passed, 0 failed.

Test output (rebased tree, 127 generated models)

1. Single model without memory (dream-dense), Test 9 log

=== Test 9: state restore failure ===
PASS (model has no memory)

All tests passed.

2. --models table (excerpt of the 127 generated models)

Model                      baseline  seq_rm  state_load  cp_h  cp_d  cp_h_s  cp_d_s  rt    rf
bailingmoe3-moe.gguf       PASS      PASS    PASS        PASS  PASS  PASS    PASS    PASS  PASS
dots3note-moe.gguf         PASS      PASS    PASS        PASS  PASS  PASS    PASS    PASS  PASS
dream-dense.gguf           PASS      PASS    PASS        PASS  PASS  PASS    PASS    PASS  PASS
kimi-k3-moe.gguf           PASS      PASS    PASS        PASS  PASS  PASS    PASS    PASS  PASS
kimi-linear-moe.gguf       PASS      PASS    PASS        PASS  PASS  PASS    PASS    PASS  PASS
llada-dense.gguf           PASS      PASS    PASS        PASS  PASS  PASS    PASS    PASS  PASS
llada-moe-moe.gguf         PASS      PASS    PASS        PASS  PASS  PASS    PASS    PASS  PASS
rnd1-moe.gguf              PASS      PASS    PASS        PASS  PASS  PASS    PASS    PASS  PASS
...
main: summary: 127 passed, 0 failed (of 127)

I would appreciate it if you could take another look when you have time.

Comment thread src/llama-memory-recurrent.cpp Outdated
@ggerganov
ggerganov merged commit 08618ff into ggml-org:master Sep 26, 2026
1 check passed
@CHIPMUNK-T0T
CHIPMUNK-T0T deleted the fix/state-restore-failure-cleanup branch September 26, 2026 07:26
@CISC

CISC commented Sep 26, 2026

Copy link
Copy Markdown
Member

@CHIPMUNK-T0T

Copy link
Copy Markdown
Contributor Author

The new test 9 exposed a latent HRM-Text issue that is normally hidden by reallocation.

hrm.z_l_init is placed on the input layer, so the ggml_add() with the embeddings is assigned to a different backend depending on the batch size. The allocation from the reserve graph therefore does not cover the following small decodes, and GGML_SCHED_NO_REALLOC=ON turns the resulting reallocation into an abort.

I'll document the exact location and cause and report it as a separate issue.

sky-mighty pushed a commit to sky-mighty/llama.cpp that referenced this pull request Sep 26, 2026
…ml-org#27530)

* llama : add discard for deferred state writes

* llama : add tensor zeroing helper for backends without tensor memset

* llama : clear K/V data after failed sequence restore

* llama : clear recurrent state data after failed sequence restore

* llama : simplify discard and restore cleanup

* llama : report error when abnormal cell count is found in state_read_meta

* llama : clear attention state on hybrid restore failure

* tests : cover failed state restore cleanup

* llama : clear MLA state on dsa restore failure

* tests : update test for rebased test suite

* llama : clarify comment in llama_memory_recurrent::state_read
Wizard815 pushed a commit to Wizard815/mx-llama.cpp-Rocm10 that referenced this pull request Sep 29, 2026
…ml-org#27530)

* llama : add discard for deferred state writes

* llama : add tensor zeroing helper for backends without tensor memset

* llama : clear K/V data after failed sequence restore

* llama : clear recurrent state data after failed sequence restore

* llama : simplify discard and restore cleanup

* llama : report error when abnormal cell count is found in state_read_meta

* llama : clear attention state on hybrid restore failure

* tests : cover failed state restore cleanup

* llama : clear MLA state on dsa restore failure

* tests : update test for rebased test suite

* llama : clarify comment in llama_memory_recurrent::state_read

(cherry picked from commit 08618ff)
pierreguillot pushed a commit to Ircam-Partiels/llama.cpp that referenced this pull request Oct 1, 2026
…ml-org#27530)

* llama : add discard for deferred state writes

* llama : add tensor zeroing helper for backends without tensor memset

* llama : clear K/V data after failed sequence restore

* llama : clear recurrent state data after failed sequence restore

* llama : simplify discard and restore cleanup

* llama : report error when abnormal cell count is found in state_read_meta

* llama : clear attention state on hybrid restore failure

* tests : cover failed state restore cleanup

* llama : clear MLA state on dsa restore failure

* tests : update test for rebased test suite

* llama : clarify comment in llama_memory_recurrent::state_read
frostyautumnleaf pushed a commit to frostyautumnleaf/llama.cpp that referenced this pull request Oct 5, 2026
…ml-org#27530)

* llama : add discard for deferred state writes

* llama : add tensor zeroing helper for backends without tensor memset

* llama : clear K/V data after failed sequence restore

* llama : clear recurrent state data after failed sequence restore

* llama : simplify discard and restore cleanup

* llama : report error when abnormal cell count is found in state_read_meta

* llama : clear attention state on hybrid restore failure

* tests : cover failed state restore cleanup

* llama : clear MLA state on dsa restore failure

* tests : update test for rebased test suite

* llama : clarify comment in llama_memory_recurrent::state_read
MarkovInequality pushed a commit to MarkovInequality/koboldcpp that referenced this pull request Oct 6, 2026
…ml-org#27530)

* llama : add discard for deferred state writes

* llama : add tensor zeroing helper for backends without tensor memset

* llama : clear K/V data after failed sequence restore

* llama : clear recurrent state data after failed sequence restore

* llama : simplify discard and restore cleanup

* llama : report error when abnormal cell count is found in state_read_meta

* llama : clear attention state on hybrid restore failure

* tests : cover failed state restore cleanup

* llama : clear MLA state on dsa restore failure

* tests : update test for rebased test suite

* llama : clarify comment in llama_memory_recurrent::state_read

(cherry picked from commit 08618ff)
edwardyoon pushed a commit to edwardyoon/focus-llama that referenced this pull request Oct 7, 2026
…ml-org#27530)

* llama : add discard for deferred state writes

* llama : add tensor zeroing helper for backends without tensor memset

* llama : clear K/V data after failed sequence restore

* llama : clear recurrent state data after failed sequence restore

* llama : simplify discard and restore cleanup

* llama : report error when abnormal cell count is found in state_read_meta

* llama : clear attention state on hybrid restore failure

* tests : cover failed state restore cleanup

* llama : clear MLA state on dsa restore failure

* tests : update test for rebased test suite

* llama : clarify comment in llama_memory_recurrent::state_read

(cherry picked from commit 08618ff)
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

testing Everything test related

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Eval bug: failed slot restore leaves corrupted K/V data that breaks subsequent inference

3 participants