Skip to content

metal : allow forgetting model and context - #28395

Closed
madsmtm wants to merge 1 commit into
ggml-org:masterfrom
nobodywho-ooo:allow-placing-model-in-static
Closed

madsmtm wants to merge 1 commit into
ggml-org:masterfrom
nobodywho-ooo:allow-placing-model-in-static

Conversation

@madsmtm

@madsmtm madsmtm commented Sep 4, 2026 •

Copy link
Copy Markdown
Contributor

Overview

Code like this currently hits a GGML_ABORT in the Metal backend, because it detects on exit that the model wasn't unloaded:

int main() {
    llama_backend_init();
    llama_model * model = llama_model_load_from_file("./model.gguf", llama_model_default_params());
    // llama_model_free(model);
    // llama_backend_free();
    return 0;
}

Such leaks are usually undesirable, but there are situations where it's useful, such as in simple applications or tests where storing the model (or context) in a static is more feasible.

So in this PR, I've removed the assertion, and added a test to ensure that this use-case remains supported.

Additional information

Introduced in #17766.
Fixes #22593.
Fixes #19137.
Replaces #22595.
Replaces #26857.
Replaces #22595.
Replaces #19206.

I decided to write a new PR because the others lacked a test, and were still doing the wrong thing IMO.

Requirements

This is usually undesirable, as it indicates a leak, but can be useful
in simple applications or tests where storing the model / context in a
static is more feasible.
@madsmtm
madsmtm requested review from a team and ggerganov as code owners September 4, 2026 14:49
@github-actions github-actions Bot added testing Everything test related ggml changes relating to the ggml tensor library for machine learning Apple Metal https://en.wikipedia.org/wiki/Metal_(API) labels Sep 4, 2026
// We _could_ end the residency of any extant residency sets here, but
// let's leave that to `ggml_metal_buffer_free`, the user is still in
// control of those buffers (and might free them later).

@madsmtm madsmtm Sep 4, 2026 •

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

and might free them later

Specifically, since the order of global destructors across different translation units is complex to get right, I think it's possible for the user to not even have a leak with something like:

#include "llama-cpp.h"

static llama_model_ptr model;

int main() {
    llama_backend_init();
    model = llama_model_load_from_file("./model.gguf", llama_model_default_params());
    llama_backend_free();
    return 0;
}

And still hit the assertion in case the Metal destructors runs before the user's model destructor.

@nikwen

nikwen commented Sep 4, 2026

Copy link
Copy Markdown
Member

/bot review

@ggml-gh-bot

ggml-gh-bot Bot commented Sep 4, 2026 •

Copy link
Copy Markdown
Automated code review

The diff is small and self-contained: removal of an exit-time GGML_ABORT in the Metal residency-set collection teardown, plus a new regression test. I've reviewed the teardown path, the surrounding residency-set/buffer lifetime code, and the test/CMake wiring.

Summary of findings

No blocking issues. The change is correct and minimal:

  • Ordering in ggml_metal_rsets_free is right: the heartbeat thread is stopped (d_stop + dispatch_group_wait) before rsets->data/rsets->lock are released and the struct freed, so the background thread can no longer touch freed state.
  • In the "forgotten model" scenario, the device (and thus rsets) is only destroyed at process exit via the static devs vector in ggml-metal-device.cpp, and the leaked buffers are never freed afterward, so releasing the array containing extant residency sets does not create a use-after-free.
  • The test follows the existing pattern of test-autorelease.cpp / test-model-load-cancel.cpp (common_get_model_or_exit, LABEL "model", CMake column alignment). The Metal-only nature of the bug is fine: the test is cheap and passes trivially on other backends.

Suggestions:

  • (point 1) ggml/src/ggml-metal/ggml-metal-device.m:1026: consider keeping the diagnostic value of the removed assert without aborting, e.g. log a GGML_LOG_WARN when [rsets->data count] > 0 at teardown. The old assert was the only signal a user got that Metal resources were leaked; with it fully removed, an accidental leak (e.g. a dropped context on an error path) is now silent. A warning preserves debuggability while still allowing intentional leaks at exit, which is the stated goal of this PR.
  • (point 2) tests/test-forget.cpp: the test passes vacuously if the model fails to load - llama_model_load_from_file returns NULL, llama_init_from_model(NULL) returns NULL, and the process exits 0 without exercising the teardown path. Add a sanity check (if (!model || !ctx) { ...; return 1; } or GGML_ASSERT(model && ctx)), otherwise the regression coverage can silently degrade to a no-op.
  • (point 3) Nit, ggml/src/ggml-metal/ggml-metal-device.m:1025: the comment says residency is left to ggml_metal_buffer_free, but in the exact scenario this PR enables (leaked buffers), that function never runs. A shorter phrasing stating the invariant directly (e.g. "residency sets owned by the user's buffers; we do not end them here") would read better and matches the project's comment guidelines.

Everything else (style, ASCII, naming, CMake integration) looks clean and consistent with the surrounding code.

This review was generated automatically by pi coding agent using zai-org/GLM-5.3. It may contain mistakes. Maintainers make the final call.

@nikwen nikwen left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

To be honest, I'm not sure if dropping the free calls is something we should support. It comes at the cost of potentially not catching real errors.

That said, if other people feel differently, I'm okay with my opinion being ignored here.

@madsmtm

madsmtm commented Sep 4, 2026 •

Copy link
Copy Markdown
Contributor Author

It comes at the cost of potentially not catching real errors.

Yeah, and I do understand wanting to be conservative in such cases.

I think my main arguments are:

  • The check has false positives.
  • The check is Metal-only, which makes it harder to debug.
  • There exist much better tools for detecting leaks, namely sanitizers.

@madsmtm

madsmtm commented Sep 10, 2026

Copy link
Copy Markdown
Contributor Author

CC @ggerganov WDYT?

@madsmtm

madsmtm commented Sep 18, 2026

Copy link
Copy Markdown
Contributor Author

@nikwen I'm still quite convinced that doing this is the right solution, what would you propose that I do to move it forwards? Are there other people that I need to ping, or a Discord I can join to pester y'all?

Alternatively, would you accept a PR doing the opposite; adding this assertion on all backends / somewhere global, so that it's no longer platform-specific?

@nikwen

nikwen commented Sep 18, 2026

Copy link
Copy Markdown
Member

Personally, if I were you, I'd focus my energy on more important problems.

I don't think this one is worth the effort/energy.

And, due to the concerns I expressed earlier, I'd be against merging it.

@madsmtm

madsmtm commented Sep 18, 2026

Copy link
Copy Markdown
Contributor Author

Yeah, fair enough, thanks for the advice!

@madsmtm madsmtm closed this Sep 28, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Apple Metal https://en.wikipedia.org/wiki/Metal_(API) ggml changes relating to the ggml tensor library for machine learning testing Everything test related

Projects

None yet

2 participants