Follow-up to #458 (review finding, altitude/efficiency). These are load-path design debts of caps.weights_device_copy.
1. The arena-skip rule is a proxy predicate
weight_skips_arena (src/archs/transformer/weight_load/tensor_views.c) = "2-D, ≥ 1 MiB, and not the tensor named per_layer_model_proj.weight". It stands in for "a tensor the backend will resolve_weight and upload". It is evaluated in the capacity pre-scan and in the loader (they cannot disagree today, but every new special tensor needs a name clause), it over-matches lookup tables (Gemma's 1.9 GB per_layer_token_embd aliases the mmap and only works because of the lazy resolve below) and misses small matrices. The loader already knows the real signal: load_layer_proj (layer_wiring.c) is the only path with out_weight != nullptr. Pass a storage intent (LINEAR vs BINDABLE) into load_tensor_to_buffer and drop the size threshold + name check; make the arena size come from the same decision (measure pass or chunked growth).
2. Lazy resolve inside the embed op
vk_embedding_lookup_scaled registers an untied token_embd on the first token (repack + upload + sequence flush inside the forward pass; can OOM mid-decode, and re-runs vk_resolve_weight per token when the resolve chooses the CPU fallback). Resolving lookup tables at load time is the clean fix, but a blind resolve_weight on cpu_neon would create a large x8 repack of the table — so it must be opt-in (e.g. a resolve_lookup_table hint or only under weights_device_copy). Alternative for untied tables: dequantize the one row on the host and upload just that row (avoids ~0.7–1 GB of VRAM for a table read once per token).
3. mmap pages stay resident after upload
With aliased weights the GGUF mmap stays open (needed only for the fallbacks) and its pages stay resident, competing with the device copy on unified-memory GPUs (up to the model size on the iGPU). madvise(MADV_DONTNEED) / posix_fadvise the aliased ranges after the resolve loop, or close the mapping when nothing aliases.
Acceptance
- gemma4-e4b, llama-3.2-3B, qwen3.5-0.8B/4B and the 27B unchanged vs.
cpu_scalar; RSS after load reported for the 27B on the iGPU.
Follow-up to #458 (review finding, altitude/efficiency). These are load-path design debts of
caps.weights_device_copy.1. The arena-skip rule is a proxy predicate
weight_skips_arena(src/archs/transformer/weight_load/tensor_views.c) = "2-D, ≥ 1 MiB, and not the tensor namedper_layer_model_proj.weight". It stands in for "a tensor the backend willresolve_weightand upload". It is evaluated in the capacity pre-scan and in the loader (they cannot disagree today, but every new special tensor needs a name clause), it over-matches lookup tables (Gemma's 1.9 GBper_layer_token_embdaliases the mmap and only works because of the lazy resolve below) and misses small matrices. The loader already knows the real signal:load_layer_proj(layer_wiring.c) is the only path without_weight != nullptr. Pass a storage intent (LINEAR vs BINDABLE) intoload_tensor_to_bufferand drop the size threshold + name check; make the arena size come from the same decision (measure pass or chunked growth).2. Lazy resolve inside the embed op
vk_embedding_lookup_scaledregisters an untiedtoken_embdon the first token (repack + upload + sequence flush inside the forward pass; can OOM mid-decode, and re-runsvk_resolve_weightper token when the resolve chooses the CPU fallback). Resolving lookup tables at load time is the clean fix, but a blindresolve_weighton cpu_neon would create a large x8 repack of the table — so it must be opt-in (e.g. aresolve_lookup_tablehint or only underweights_device_copy). Alternative for untied tables: dequantize the one row on the host and upload just that row (avoids ~0.7–1 GB of VRAM for a table read once per token).3. mmap pages stay resident after upload
With aliased weights the GGUF mmap stays open (needed only for the fallbacks) and its pages stay resident, competing with the device copy on unified-memory GPUs (up to the model size on the iGPU).
madvise(MADV_DONTNEED)/posix_fadvisethe aliased ranges after the resolve loop, or close the mapping when nothing aliases.Acceptance
cpu_scalar; RSS after load reported for the 27B on the iGPU.