Skip to content

Shared, caller-owned embedder (one model for many stores) with memory-based model loading #8

Description

@geisten

Problem

Every gm_open(dir, model_path, …) loads its own embedding model. A host that needs several independent stores (one per user/agent, or one per timeline) loads the same model N times: N × model RSS, N × load time. The model path is also a second, independent model-loading boundary next to the host's own verified loader.

Request

Additive API to open a store with a caller-owned, shared embedder, with explicit ownership and threading rules, for example:

struct gm_embedder;   /* opaque; created once, borrowed by many stores */
[[nodiscard]] enum gm_status gm_embedder_open(const char *model_path, const struct gm_opts *opts,
                                              struct gm_embedder **out);
[[nodiscard]] enum gm_status gm_embedder_open_from_memory(const void *bytes, size_t size,
                                                          const struct gm_opts *opts,
                                                          struct gm_embedder **out);
void gm_embedder_close(struct gm_embedder *e);   /* after all stores using it are closed */
[[nodiscard]] enum gm_status gm_open_with(const char *dir, struct gm_embedder *borrowed,
                                          const struct gm_opts *opts, struct gm **out);
  • *_from_memory lets the host load verified bytes (hash/size checked by the host) instead of letting the library open a second path.
  • The store identity must still record the embedder fingerprint, so reopening a store with a different model is detected (existing behaviour for gm_open).
  • Document concurrency: e.g. one embedder may be used by several stores on one thread; cross-thread use requires external serialisation (or define an internal lock).

Acceptance

  • Two stores opened with one embedder: one model load, identical recall results to two independently opened stores.
  • Closing order enforced/documented (closing the embedder while stores are open is GM_E_BUSY or UB explicitly documented).
  • Model mismatch on reopen still rejected; memory-based load passes the same E2E test as the path-based load.
  • gm_open keeps working unchanged.

Downstream: git.geisten.net/geisten/helio#7 (no additional model per agent; verified model byte path).

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions