Skip to content

vulkan: failures and a hang above depth 512 / on non-gemma arches (found by #364 campaigns) #409

Description

@geisten

The #364 GPU campaigns on the 2080 Ti exposed four distinct Vulkan-backend defects. Full provenance in the linked run logs; all reproduce under tests/bench_perf_sweep with GEIST_BENCH_BACKEND=vulkan.

  1. gemma4-e4b Q4_K_M: GEIST_E_BACKEND at pp1024 (all repeats; pp128–512 fine) — run 34157240964. Smells like device-memory exhaustion/fragmentation (4.2 GB weights + KV on 11 GB).
  2. bitnet_b1_58-large TQ2_0: GEIST_E_UNSUPPORTED at pp1024 (pp128–512 fine; same model+shape passes on cpu_neon at 198 tok/s) — run 34157633018.
  3. qwen3-0.6b Q8_0: hard hang in the first measured prefill — no output for 6 h until the job timeout killed it; run 34157843292. A plain-attention model, so this is not the DeltaNet gap below.
  4. qwen3.5 arch: state_create fails at model load (decoder state_create failed) — expected in the sense that DeltaNet was only ever built for metal/cpu_neon, but the Vulkan backend should reject the arch with a clear capability error, not a generic load failure — runs 34178688335, 34178881180.

Until these are fixed the matrix carries the affected NVIDIA Vulkan cells as unmeasured/incompatible with this issue as the reference.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions