Follow-up to #458 (known limitation).
Problem
The Vulkan path needs the whole model in device memory. The qwen3.8-27B Q4_0 (16 GB) loads and runs on a 21 GiB device (RADV iGPU, GEIST_VK_DEVICE=1; bit-identical to cpu_scalar), but on an 11 GiB RTX 2080 Ti the load fails with an out-of-memory error. There is no spill to host memory.
Proposal (needs a decision)
- Per-tensor placement: keep what fits in VRAM, put the rest into host-visible device memory (
HOST_VISIBLE|HOST_COHERENT, read over PCIe by the shaders) — simplest, decode is bandwidth-bound by the spilled share.
- Layer-granular offload: first N layers in VRAM, remaining layers streamed or computed from host memory. More work, better bandwidth use.
- At minimum: a clear, early capacity error (weights + KV + scratch vs.
VkPhysicalDeviceMemoryProperties budget) instead of an allocation failure mid-load.
GEIST_VK_VRAM_BUDGET knob for tests.
Related: weights_device_copy (#458) already removed the duplicate host arena, so the device budget is the only remaining constraint.
Acceptance
- 27B Q4_0 loads and decodes on the 2080 Ti (with a reported tok/s), or fails at load with a message naming the needed vs. available bytes.
Follow-up to #458 (known limitation).
Problem
The Vulkan path needs the whole model in device memory. The qwen3.8-27B Q4_0 (16 GB) loads and runs on a 21 GiB device (RADV iGPU,
GEIST_VK_DEVICE=1; bit-identical tocpu_scalar), but on an 11 GiB RTX 2080 Ti the load fails with an out-of-memory error. There is no spill to host memory.Proposal (needs a decision)
HOST_VISIBLE|HOST_COHERENT, read over PCIe by the shaders) — simplest, decode is bandwidth-bound by the spilled share.VkPhysicalDeviceMemoryPropertiesbudget) instead of an allocation failure mid-load.GEIST_VK_VRAM_BUDGETknob for tests.Related:
weights_device_copy(#458) already removed the duplicate host arena, so the device budget is the only remaining constraint.Acceptance