Found while tuning the PQ2_0 tensor-core GEMM (#487).
Problem
Prefill throughput collapses with larger chunks: on the RTX 2080 Ti, Ternary-Bonsai-2-27B pp512 is 350 t/s at the default m_max 64 but 139 t/s at GEIST_M_MAX=256 (128 t/s at 128). A GEIST_VK_PROFILE=1 run at 256 shows the elementwise/copy ops exploding while the GEMM stays flat: silu_mul 12 → 522 µs per call, hadamard 11 → 135 µs, copy 4970 calls at 98 µs, attention_f16 0.3 → 3.8 ms. Those buffers are the per-session scratch pool (GEIST_BUFFER_SCRATCH, created host-visible). With a 256 MB BAR heap (no resizable BAR) the pool for m = 256 no longer fits, vk_buffer_create falls back to plain host-visible system memory, and every dispatch that touches the scratch reads/writes it over PCIe.
Proposal
Scratch that only GPU ops touch can live in ordinary device-local VRAM (unmappable): the arch/backend knows which buffers a CPU fallback may map. Options: a GEIST_BUFFER_SCRATCH_DEVICE role for the activation slabs of a fully GPU-resident model (decided from the fused-op probes at plan build), or an explicit "no host access" flag; a CPU fallback then fails loudly (vk_tensor_host already returns nullptr for device-local memory) instead of silently running over PCIe. Also print a one-line note under GEIST_VK_VERBOSE whenever a scratch buffer lands outside device-local memory.
Why it matters
Larger chunks are what a tensor-core GEMM wants (fewer weight passes, wider N tile), and the same spill hits any model whose chunk × width exceeds the BAR — it currently just shows up as a slowdown.
Acceptance
GEIST_M_MAX=256 and 128 are not slower than 64 on the Bonsai (pp512, RTX 2080 Ti); a test or verbose counter that shows where each scratch buffer lives.
Related: the remainder chunk (m % 16 != 0) still uses the register-tiled GEMM — routing m & ~15 rows through the tensor-core kernel and the rest through the tiled one belongs with this (#487, #467).
Found while tuning the PQ2_0 tensor-core GEMM (#487).
Problem
Prefill throughput collapses with larger chunks: on the RTX 2080 Ti, Ternary-Bonsai-2-27B
pp512is 350 t/s at the defaultm_max64 but 139 t/s atGEIST_M_MAX=256(128 t/s at 128). AGEIST_VK_PROFILE=1run at 256 shows the elementwise/copy ops exploding while the GEMM stays flat:silu_mul12 → 522 µs per call,hadamard11 → 135 µs,copy4970 calls at 98 µs,attention_f160.3 → 3.8 ms. Those buffers are the per-session scratch pool (GEIST_BUFFER_SCRATCH, created host-visible). With a 256 MB BAR heap (no resizable BAR) the pool for m = 256 no longer fits,vk_buffer_createfalls back to plain host-visible system memory, and every dispatch that touches the scratch reads/writes it over PCIe.Proposal
Scratch that only GPU ops touch can live in ordinary device-local VRAM (unmappable): the arch/backend knows which buffers a CPU fallback may map. Options: a
GEIST_BUFFER_SCRATCH_DEVICErole for the activation slabs of a fully GPU-resident model (decided from the fused-op probes at plan build), or an explicit "no host access" flag; a CPU fallback then fails loudly (vk_tensor_hostalready returns nullptr for device-local memory) instead of silently running over PCIe. Also print a one-line note underGEIST_VK_VERBOSEwhenever a scratch buffer lands outside device-local memory.Why it matters
Larger chunks are what a tensor-core GEMM wants (fewer weight passes, wider N tile), and the same spill hits any model whose chunk × width exceeds the BAR — it currently just shows up as a slowdown.
Acceptance
GEIST_M_MAX=256and128are not slower than 64 on the Bonsai (pp512, RTX 2080 Ti); a test or verbose counter that shows where each scratch buffer lives.Related: the remainder chunk (
m % 16 != 0) still uses the register-tiled GEMM — routingm & ~15rows through the tensor-core kernel and the rest through the tiled one belongs with this (#487, #467).