Skip to content

memory: Gemma 4 E4B Q4_K_M peaks at 11.7 GiB on cpu_x86 and 10.6 GiB on Metal (4.6 GiB file, load_from_memory) #577

Description

@geisten

Summary

With Gemma 4 E4B Q4_K_M (4.64 GiB GGUF) loaded through geist_model_load_from_memory, peak RSS goes far beyond one copy of the weights on the CPU and Metal backends. The model bytes are a single verified, read-only, page-aligned anonymous mapping owned by the caller. The header says these bytes are aliased zero-copy. The extra memory therefore has to come from backend-side copies.

Backend Host Peak RSS (whole process)
file size 4.64 GiB
vulkan Ryzen 9 9950X + RTX 2080 Ti, Linux 6.74 GiB
auto (cpu_x86, 16 threads) same host 11.66 GiB
metal M1 Max (64 GB) 10.6 GiB

Measured with HELIO (--ai-bench, peak ru_maxrss of the process including its ~1 GiB game state) at geistlib 7a0b421.

Impact

HELIO targets the Steam Deck (16 GB, shared with the GPU) and Macs with 16 GB. Above ~6 GiB the game cannot ship this model there. The CPU path is the Deck fallback.

Hypotheses (unverified)

Requests

  • Document the expected resident footprint per backend for load_from_memory and for load(path).
  • CPU: either drop the source tensors after repacking, or repack lazily or in place, so one layout stays resident.
  • Metal: wrap page-aligned caller memory with newBufferWithBytesNoCopy where alignment allows.
  • Add peak RSS for Gemma 4 E4B to the performance matrix (bench: populate a same-protocol model × system performance matrix #364).

Related: #468 (vulkan resident mmap), #504 (cpu_x86 prefill).

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't working

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions