Skip to content

transformer/vulkan: arena-skip predicate, lazy embed-table resolve and resident mmap after upload (weights_device_copy follow-ups) #468

Description

@geisten

Follow-up to #458 (review finding, altitude/efficiency). These are load-path design debts of caps.weights_device_copy.

1. The arena-skip rule is a proxy predicate

weight_skips_arena (src/archs/transformer/weight_load/tensor_views.c) = "2-D, ≥ 1 MiB, and not the tensor named per_layer_model_proj.weight". It stands in for "a tensor the backend will resolve_weight and upload". It is evaluated in the capacity pre-scan and in the loader (they cannot disagree today, but every new special tensor needs a name clause), it over-matches lookup tables (Gemma's 1.9 GB per_layer_token_embd aliases the mmap and only works because of the lazy resolve below) and misses small matrices. The loader already knows the real signal: load_layer_proj (layer_wiring.c) is the only path with out_weight != nullptr. Pass a storage intent (LINEAR vs BINDABLE) into load_tensor_to_buffer and drop the size threshold + name check; make the arena size come from the same decision (measure pass or chunked growth).

2. Lazy resolve inside the embed op

vk_embedding_lookup_scaled registers an untied token_embd on the first token (repack + upload + sequence flush inside the forward pass; can OOM mid-decode, and re-runs vk_resolve_weight per token when the resolve chooses the CPU fallback). Resolving lookup tables at load time is the clean fix, but a blind resolve_weight on cpu_neon would create a large x8 repack of the table — so it must be opt-in (e.g. a resolve_lookup_table hint or only under weights_device_copy). Alternative for untied tables: dequantize the one row on the host and upload just that row (avoids ~0.7–1 GB of VRAM for a table read once per token).

3. mmap pages stay resident after upload

With aliased weights the GGUF mmap stays open (needed only for the fallbacks) and its pages stay resident, competing with the device copy on unified-memory GPUs (up to the model size on the iGPU). madvise(MADV_DONTNEED) / posix_fadvise the aliased ranges after the resolve loop, or close the mapping when nothing aliases.

Acceptance

  • gemma4-e4b, llama-3.2-3B, qwen3.5-0.8B/4B and the 27B unchanged vs. cpu_scalar; RSS after load reported for the 27B on the iGPU.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions