Skip to content

vulkan: models larger than device memory (27B Q4_0 on an 11 GiB card) — no host spill #466

Description

@geisten

Follow-up to #458 (known limitation).

Problem

The Vulkan path needs the whole model in device memory. The qwen3.8-27B Q4_0 (16 GB) loads and runs on a 21 GiB device (RADV iGPU, GEIST_VK_DEVICE=1; bit-identical to cpu_scalar), but on an 11 GiB RTX 2080 Ti the load fails with an out-of-memory error. There is no spill to host memory.

Proposal (needs a decision)

  • Per-tensor placement: keep what fits in VRAM, put the rest into host-visible device memory (HOST_VISIBLE|HOST_COHERENT, read over PCIe by the shaders) — simplest, decode is bandwidth-bound by the spilled share.
  • Layer-granular offload: first N layers in VRAM, remaining layers streamed or computed from host memory. More work, better bandwidth use.
  • At minimum: a clear, early capacity error (weights + KV + scratch vs. VkPhysicalDeviceMemoryProperties budget) instead of an allocation failure mid-load.
  • GEIST_VK_VRAM_BUDGET knob for tests.

Related: weights_device_copy (#458) already removed the duplicate host arena, so the device budget is the only remaining constraint.

Acceptance

  • 27B Q4_0 loads and decodes on the 2080 Ti (with a reported tok/s), or fails at load with a message naming the needed vs. available bytes.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions