Repository navigation
Sync upstream master (2026-10-05, ggml-org 50569eb87) - #17
Merged
Merged
Conversation
* hexagon: add F16 support for activation ops (SILU/GELU/GELU_QUICK/GEGLU/SWIGLU) Widens ggml_hexagon_supported_activations() to accept F16 (src0/dst/src1 must agree on type), and adds F16 per-thread worker functions in act-ops.c mirroring the existing F32 workers, backed by new HVX f16 kernels (hvx_sigmoid_f16_aa, hvx_tanh_f16_aa, hvx_mul_mul_f16_aa, hvx_min_scalar_f16 family). SILU, GELU, GELU_QUICK, GEGLU, and SWIGLU are verified correct on-device (QRD8850) via test-backend-ops CPU-diffed correctness tests. SWIGLU_OAI's F16 path is code-complete and builds clean on host + all 4 DSP arch variants (v73/v75/v79/v81), but has no F16 test-case coverage in test-backend-ops and is therefore unverified on-device in this change. * hex-ops: align macros * hex-ops: minor formatting --------- Co-authored-by: Max Krasnyansky <maxk@qti.qualcomm.com>
cont: fix code style Signed-off-by: Aaron Teo <aaron.teo1@ibm.com>
* Rebase GLM-Next support onto master, and migrate to llama-memory-hybrid-idx * Add initial MTP support * Merge branch optimizations. Reduce allocated compute buffer size, speed up long context decode, fla, and slight MTP improvements. * Review driven changes, remove env vars, protect tensors * Strip MTP for initial PR * Clean up after mtp strip * Clean up after mtp strip * Update speculative.cpp * Update llama-context.h * Clean up after mtp strip * Fix tokenizer ignore merges * Improve quantization protection selection * Refactor mhc helpers, graph base * Lint Fixes * Apply suggestions from code review Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@huggingface.co> * Skip glm5-next in model saver, fix CRLF * Skip glm5-next in sweep * Remove T4 fallback * Review cleanup * Review suggestions * Defer separate MTP gguf handling to MTP PR, drop filter * Repad n_head_kv * kpool init apply * Order by descending score * Drop guard * read kpool from hparams, clarify kpool cache flags, remove kpool_build_state(nullptr) * Add glm5-next support to model saver and add arch test fixture * Review cleanup * Kpool pooled caching clarify * Add multi stream support * Finish Rebase * Sparse FA fir DSA prefill * Const * Update llama-model.cpp to fix rebase error * gguf-py : merge tensor map entries for HC tensors * model : use build_gdn_l2_norm in GLM5_NEXT implementation * chore : remove trailing whitespace * model : use new OP precision setting API in GLM5_NEXT implementation * mtmd : use ggml_swiglu_clamp in GLM5V and apply the image token limit The two clamps around swiglu_split are what ggml_swiglu_clamp already does, so the clamp bounds collapse back to one value. GLM5V also never called set_limit_image_tokens(), so --image-max-tokens had no effect. Assisted-by: Claude Opus 5 (cherry picked from commit 46d18e1) * llama : keep the GLM5-Next k-pool layout across ubatches The layout was rebuilt from a full cell scan on every ubatch. Pools are fixed by the positions relative to the sequence's first one, so the layout now lives on the memory and a ubatch only appends to it. A sequence edit no longer stales every pooled key either, only the ones at or after the edited position, which makes a tail seq_rm free. The pooling subgraph is built unconditionally so the graph shape no longer changes every kpool tokens, and the pool axis is folded into rows before soft_max, which otherwise exceeds the CUDA gridDim.y limit past n_kv 262144. Assisted-by: Claude Opus 5 (cherry picked from commit 5d1c40b) * model : write the GLM5-Next recurrent rollback checkpoints The conv state and the delta net state were only written to the live row, so a rollback restored whatever the checkpoint rows happened to hold. Take the same route as kimi-k3: build_recurrent_attn for the state, and write all K_rs conv groups. That also drops a state view that assumed contiguous rows. Enroll the arch in test-recurrent-state-rollback, which catches this under its garbage-filled cache pass. Assisted-by: Claude Opus 5 (cherry picked from commit 5ace37e) * llama: fix PR ggml-org#27773 test-save-load-state restore failure Clear the attention and indexer cache data after a failed hybrid state restore so restored NaNs cannot affect a later sequence. Assisted-by: Codex * llama: fix PR ggml-org#27773 gpu-rocm graph reallocation Reserve the full GLM5-Next pool capacity and dirty pool count. The gpu-rocm Test step aborts when n_new grows while the graph node count stays fixed; CUDA, Vulkan, Metal, and WebGPU checks report the same error. Assisted-by: Codex * llama : fix GLM5-Next k-pool layout staleness after edits and shared teardown Two defects in the cross-ubatch k-pool layout added by the k-pool commit: 1. Wrong results. An edited sequence only rebuilt its pool layout when its cell count changed, so if the first ubatch after an edit added back exactly as many cells as were removed, the stale position-to-cell list survived. With a unified cache and more than one sequence, where another sequence takes the freed cells, the reused layout points at the wrong cells (CPU: large logit drift, CUDA: NaN). Rebuild whenever the sequence is stale, not only on a size mismatch. 2. Slowdown. "shared" mode was assumed to end only with an edit that forces a rebuild, but sharing also ends when the other sequence is removed. The survivor kept shared = true, pinning cache_safe off and re-pooling every pool on every ubatch (server trigger: n>1 completions with -kvu, via the seq_cp in copy_state_to). In seq_rm, if the layout has shared cells, stale every sequence so one rebuild re-derives sharing and cache_safe returns to 1. Assisted-by: Claude Opus 5 * llama : fix build_attn_mha stream stride for non-contiguous q build_attn_mha split the batch into streams with a stream stride of q->nb[3]/n_stream. That only equals one stream's span, (ne[2]/n_stream)*nb[2], when q is contiguous. GLM5-Next is nope-only, so it does not concat a rope part and passes the permuted q_absorbed straight in, where nb[3] != ne[2]*nb[2]; the stride was then n_head times too large and every stream s >= 1 read another head's queries. Split-KV (-np N without --kv-unified) multi-stream prefill was wrong for every stream past the first. Unified KV and decode were unaffected (n_stream == 1, and decode takes the gather path). Other MLA models concat rope so q is contiguous and the computed value is unchanged for them. Compute the stride from the token dimension, which is identical for a contiguous q. Assisted-by: Claude Opus 5 * llama : re-derive GLM5-Next k-pool sharing on state_read/state_drop The shared-cell teardown added to seq_rm (stale every sequence when the layout has shared cells, so a survivor does not keep shared = true and pin cache_safe off) was missing from the other paths that can free shared cells: state_read and state_drop staled only the one sequence. Apply the same re-derivation there and correct the comment that claimed sharing ends only via an edit or seq_rm. Assisted-by: Claude Opus 5 * quant : drop duplicate GLM5-Next hc_ filter The hc_ name filter was listed twice in the GLM5_NEXT protection block. Assisted-by: Claude Opus 5 * glm5-next: scope K-pool cache access to indexed operations * glm5-next: keep K-pool access in hybrid index memory * glm5-next: keep mHC graph builders model-local * glm5-next: mark only touched pools per ubatch --------- Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@huggingface.co> Co-authored-by: Stanisław Szymczyk <sszymczy@gmail.com> Co-authored-by: Piotr Wilkin <ilintar@gmail.com>
* add models backend check * t4-medium for faster build
The MUSA vendor header never defined __CUDA_ARCH__, so every architecture test in the shared ggml-cuda sources evaluated to 0. Kernel bodies gated on the architecture therefore compiled to nothing, for example the q8_0 -> f16 dequantization kernel in convert.cu, whose NO_DEVICE_CODE fallback expands to an empty body in host code. Report the newest architecture like the HIP backend does and exclude the NVIDIA-only features explicitly, as they are not usable on MUSA. Define it for device passes only: CUB uses defined(__CUDA_ARCH__) to detect device compilation, which is also how nvcc behaves. Drop the now-redundant defined(__CUDA_ARCH__) checks in the architecture comparisons: __CUDA_ARCH__ is undefined in host passes for CUDA and MUSA, and HIP defines it for every pass, so both forms select the same branch.
ggml-org#27773 adds the glm5-next arch without its rows in the Metal fusion baseline, so test-fusion --check fails on it. The rows come from test-fusion --record on an M5 Max, and --check passes 270/270.
…l-org#28381) * openvino: serve GET_ROWS on a weight view from the base Constant Resolve view_src when collecting weight Constants so a view over a quantized weight no longer becomes a dynamic typed Parameter, and fold the row offset of the view into the gather indices instead of slicing the dequantization subgraph. * openvino: lift the quantized GET_ROWS view rejection The supports_op rejection of a quantized src0 view with a nonzero offset keeps the vs0 GET_ROWS cases of ggml-org#28253 away from OpenVINO. The weight view now resolves to the base Constant with the row offset folded into the gather indices, so the rejection goes away.
…re plumbing (ggml-org#29582) * ui : type-safe API types, fetch helpers and download-ready models store plumbing Assisted-by: pi:GLM-5.3-Flash * ui : document the model list index pairing, fix an em-dash Assisted-by: pi:zai-org/GLM-5.3-Flash * Update tools/ui/src/lib/components/app/chat/index.ts Co-authored-by: Pascal <admin@serveurperso.com> --------- Co-authored-by: Pascal <admin@serveurperso.com>
…ml-org#27946) * ui : model id grammar for sidecars, quants and capability parsing Extend the shared model id parser with sidecar tokens (draft variants and auxiliary imatrix/mmproj files), weight-file and custom-quant regexes, and add the tools capability to ModelCapabilities; the selector option row picks it up from the model's declared capabilities. Assisted-by: pi:GLM-5.3-Flash * ui : escape sidecar tokens in the regex alternation Assisted-by: pi:zai-org/GLM-5.3-Flash
* ui : Hugging Face Hub data layer Add HuggingFaceService and its constants/enums/types: GGUF repo search, file tree and model detail fetching, quant/sidecar filename analysis, shard-set collapsing and the llama.app catalog feed, plus an orgOf() helper on the model name utils. Assisted-by: pi:GLM-5.3-Flash * ui : strip provider tilde prefix from hub avatar urls Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash * ui : trim redundant comments in the HF data layer service Per review: drop JSDoc that restates the method name and inline comments that restate the code; keep only comments carrying non-obvious context. Assisted-by: pi:zai-org/GLM-5.3-Flash * ui : harden the HF data layer error typing, cover the helpers in tests Carries the HTTP status on retryable fetch errors instead of matching the message text. Marks expand-dependent catalog fields optional and documents the data/models index pairing. Adds table tests for the pure helpers. Assisted-by: pi:zai-org/GLM-5.3-Flash
* ui : model memory-fit estimation Replace the raw runtime-memory estimate with the app's compatibility check: the smallest real Mac memory tier that fits a model file, budgeted as RAM x 0.75 minus fixed overhead with headroom on the file size. The constants move to lib; the unused runtime-memory estimate is dropped. browser-info's OS detection is exported for reuse. Assisted-by: pi:GLM-5.3-Flash * ui : cover the memory-fit and tool-use heuristics in tests Assisted-by: pi:zai-org/GLM-5.3-Flash
* ui : model download pipeline Track HuggingFace downloads end to end: the server download/cancel endpoints, a status manager fed by the /models/sse download progress events, and a models-discover store holding the catalog and detail state for the discover view. Downloaded and in-flight entries are excluded from the loadable model list. Assisted-by: pi:GLM-5.3-Flash * ui : route sidecar tag lookup through the sidecars util, validate the paused list Assisted-by: pi:zai-org/GLM-5.3-Flash
* ui : shared model display primitives Extract ModelCapabilityIcons (canonical Tools/Reasoning/Vision/Video/Audio order) out of ModelId and reuse it there, add the shared DialogConfirmDownload for destructive download actions, the discover org avatar with dark-mode inversion and the thin download progress bar, and rework ModelId badges to take thinking/tool-use support directly. Assisted-by: pi:GLM-5.3-Flash * ui : remember hub avatars that failed to load Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash * ui : render shared model row hints as native titles Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash * ui : fix badge guard for draft sidecars, keep parameter precision hasBadges now counts draft sidecar badges, so a sidecar-only model still renders. Billions keep one decimal for hub counts and stay bare for whole values. Avatar failures track the org instead of the instance, and the download progress bar no longer pulses while determinate. Assisted-by: pi:zai-org/GLM-5.3-Flash
) This commit adds an optional --add-bos token command line option to the run-org-model.py script. The motivation for this is that there are models, for example Gemma4, that explicitely set the add_bos value to true in llama-vocab.cpp even if the original model does not set this value to True. It would be nice to be able to force the models to agree on the bos token so that logit verification can proceed. Refs: ggml-org#21500
* cpu: accept BF16 in src1 of mul_mat ggml_conv_1d_dw builds its im2col in F32 when the kernel is BF16, then calls ggml_mul_mat(im2col, kernel), which puts F32 in src0 and BF16 in src1. The CPU backend refused that combination, so it was reported as unsupported on every backend and never compared against anything. Widen BF16 into the F32 work buffer, next to the existing packing of F32 into vec_dot_type. This is the arithmetic the Metal mat vec kernel already uses, both operands promoted to float and accumulated in float, so the two agree exactly rather than approximately. Cover it with a conv_1d_dw test over F32, F16 and BF16 kernels, plus three mul_mat cases with BF16 in src1. * vulkan: reject BF16 in src1 of mul_mat unless src0 is BF16 supports_op only checked the src1 type for non contiguous tensors, so a contiguous BF16 src1 was accepted and the pipeline lookup asserted. The only BF16 src1 path is the BF16 x BF16 multiply, every other src0 type now reports the op as unsupported and the scheduler keeps it on the CPU. The BF16 kernel case of the conv_1d_dw test needs the f32 x bf16 mat vec variants of the Metal backend, which land separately.
…g#29675) * ggml : add BF16 unary, GLU, binary and scale ops (CPU, CUDA) * ggml-cpu : use per-op _bf16 functions for BF16 unary and GLU ops Assisted-by: Claude Opus 5.5 * CUDA: use ggml_cuda_cast in binbcast and unary kernels to fix the HIP bf16 build * ggml-openvino : reject BF16 SCALE and mixed-type BF16 ADD/MUL/SUB
* llama: llama_prefetch_rows * llama: support row prefetch on Windows Apply the Windows port contributed by @praneshgo unchanged. Source: ggml-org#29599 (comment) * avoid exposing llama-mmap in model code, route via llama-impl * add windows check, only prefetch in lazy mode * cont : clean-up * cont : fix build * cont : clarify padding token for gemma4 --------- Co-authored-by: Pranesh Gonegandla <pranesh.iitp@gmail.com> Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>
…l-org#29744) The fixture recycles its two blocks over 8 cache slots, so the fp16 error builds up past the 1e-4 NMSE bound on the Vulkan T4 and WebGPU jobs of Models Backend. Two l-cycles keep every branch of the cycle loop and halve the error.
Signed-off-by: Adrien Gallouët <angt@huggingface.co>
* support coerced array attributes * add tests
* convert : update to support dflash * cont : fix Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@huggingface.co> --------- Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@huggingface.co>
…ml-org#29722) * cli: exit on stdin EOF and drop the console wide Ctrl+C broadcast On Windows the simple input reader sends CTRL_C_EVENT to every process attached to the console when stdin reaches EOF, killing unrelated processes such as a supervising agent. The CLI only stopped on EOF because of that self inflicted SIGINT; on POSIX, and with the advanced reader, it spins forever printing prompts. Drop the broadcast so both platforms just return an empty read, and treat an empty read as EOF in the chat loop and the model selection, since a submitted line always ends with a newline. * cli: keep the newline of a trailing "/" and stop mtmd-cli on EOF A lone "/" came back as an empty read and was taken for EOF, and mtmd-cli only stopped on EOF through the removed broadcast.
Co-authored-by: Vishal Singh <numeric-id+vishalMCE@users.noreply.github.com>
* ggml: fix integer overflow guard for zero-element tensors * ggml: validate number of elements in tensor to prevent integer overflow * ggml: fix error print
The sparse indexer mask is built with a set_rows scatter. Padded pools, absent sequences and missing tail cells all pointed to the same n_kv sentinel row, and invisible pools picked by top_k to fill the selection overlap the tail cells of the token, so several CPU threads wrote the same element (ThreadSanitizer data race in the sanitize CI). Allocate the slot mask for both selection paths and route every dead slot to its own dump row n_kv + slot. Live slots address disjoint cells, so the scatter indices of a token are unique.
* migrate the rest * test-thread-safety * rm common_batch_staged
* llama: properly handle KV on training * improve
…org#28324) * convert: fix LoRA conversion crash for Qwen3.5 V-head reorder _reorder_v_heads does reshape+permute+reshape to reorder V heads from grouped to tiled order. LoraTorchTensor.reshape() cannot split its row dimension (A matrix), so converting Qwen3.5 LoRA adapters that target out_proj crashes with NotImplementedError. Fix: detect LoRA tensors and apply the equivalent index permutation directly — column reorder (dim=last) permutes A's columns, row reorder (dim=0) permutes B's rows. This is mathematically identical: (B @ A)[:, perm] == B @ A[:, perm] (B @ A)[perm, :] == B[perm, :] @ A Verified: both paths produce exactly zero diff against the full-tensor reorder on random (rank=32, 4096×4096) matrices. Fixes ggml-org#21125 Signed-off-by: Radu Swigler <radu@swigler.com> * convert: add ty: ignore for hasattr-guarded LoRA call Assisted-By: Claude Opus 4.6 <noreply@anthropic.com> * fix comment * nowrap --------- Signed-off-by: Radu Swigler <radu@swigler.com> Co-authored-by: Radu Swigler <radu@swigler.com> Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@huggingface.co> Assisted-by: Claude Opus 4.6 <noreply@anthropic.com>
…rotation metadata (ggml-org#28498) * kv-cache: save exact KV rotation metadata, reject restoring mismatched rotation * tests : move the state rotation test to test-save-load-state the test is now part of the save/load test matrix and runs against every model under test, like the rest of the suite it probes the KV cache type combinations supported by the model and treats models that do not use attention rotation as passing vacuously Assisted-by: pi:llama.cpp/Qwen3.8-27B * cont : skip unsupported KV caches --------- Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>
…org#29958) * llama : fix unexpected graph reallocation in the k-pool models Both k-pool models built a graph shape that depends on state the full-context reserve cannot know: - qwen4exp branched on inp->cache_safe, which turns false as soon as llama_memory_seq_cp shares cells (e.g. batched-bench -pps): the QSA layers swapped scatter+gather for fill+concat and dropped the new_pool_rep leaf, so the decode graph had 12 fewer nodes than the reserved one - glm5-next branched on gather = n_tokens <= 16 && n_kv > n_sel, so the TG decode built the gather shape (7564 nodes) while the last reserve, the PP one, had the dense shape (7762 nodes) Either mismatch forces a decode-time re-reserve that drops the worst-case sizing and bakes in the current state, so the next state growth (n_pool, n_kv, n_new) needs more room at an unchanged graph size and aborts under GGML_SCHED_DEBUG_REALLOC=1. Reproduce with, e.g.: GGML_SCHED_DEBUG_REALLOC=1 ./bin/llama-batched-bench \ -hf ggml-org/GLM-5.3-Flash-GGUF:Q2_K -npp 2500 -ntg 32 -npl 1,2 \ -c 32768 -pps -kvu Always scatter+gather the pooled keys, and pick gather from context constants only: n_ubatch bounds every ubatch, top_k + kpool - 1 bounds n_sel. Every graph of a context then shares one shape, which the reserve covers, and the dense path measured faster than the gather path at 2.5k and 16k context. Assisted-by: pi:llama.cpp/MiMo-V2.6-Flash-MOPD * llama : drop the unused k-pool cache_safe graph API The k-pool graphs no longer branch on cache_safe, so nothing reads get_kpool_cache_safe() or the conditional new_pool_rep any more: both models always pass the scatter target, which set_input_kpool now requires instead of merely preferring. Also drop the cache_safe copy in kpool_build_sizes(), a sizes-only helper. The layout and state flag itself stays, it still decides which pools a layout with shared cells must re-pool. Assisted-by: pi:llama.cpp/MiMo-V2.6-Flash-MOPD * tests : add a shared-seq graph reserve regression test Decode a prompt into seq 0, share its cells with seq 1 via llama_memory_seq_cp (what llama-batched-bench does for -pps), then keep decoding both sequences. For the k-pool models sharing clears cache_safe, which changes the graph topology while the pools keep growing, so a scheduler that re-reserves with the current state instead of the worst-case one aborts under GGML_SCHED_DEBUG_REALLOC=1. The test registration sets that flag, and the test aborts on both k-pool models before 2220411. kimi-linear and minimax-01 are skipped: they reserve the final pp graph with n_seqs = 1 (see [TAG_RESERVE_DIAG_DECAY] in llama-context.cpp), so every multi-seq graph has a different layout and re-reserves by design. Assisted-by: pi:llama.cpp/MiMo-V2.6-Flash-MOPD * cont : add TODOs * cont : fix comment * cuda: match the moe weighted reduction on empty ubatches ggml_cuda_match_moe_weighted_reduction rejected tensors with zero rows. A ubatch without outputs shrinks the last layer to zero rows through inp_out_ids, so graph_optimize dropped its alloc dep there and the scheduler graph lost one node compared to the reserved one. The scheduler then re-reserved at the size of that ubatch, and the next ubatch with the same node count but larger tensors aborted under GGML_SCHED_DEBUG_REALLOC=1. The compute loop already skips empty nodes before trying any fusion, so the guard only made the alloc deps depend on the row count. * tests: build the rollback test only where internal symbols link The shared-seq case calls llm_arch_from_string, which libllama does not export through LLAMA_API, so linking test-recurrent-state-rollback fails on Windows with shared libraries. Its build now sits in the NOT WIN32 OR NOT BUILD_SHARED_LIBS block, next to test-llama-archs and the test registration it already lives under. * tests: skip archs by name in the shared-seq reserve test The skip of kimi-linear and minimax-01 went through llm_arch_from_string, which libllama does not export through LLAMA_API, so the test could not link on Windows with shared libraries. It now compares the general.architecture string directly, and the test builds on every platform again. --------- Co-authored-by: Pascal <admin@serveurperso.com>
* vulkan: sparse flash attention for quantized K/V Assisted-by: Claude * vulkan: single-scan sparse FA index compaction The compaction ran one workgroup per mask row and walked the row in BLOCK_SIZE chunks, with a workgroup scan per chunk. For decode that is one workgroup doing KV/1024 barrier-bound iterations, so at 128k cells it cost more than the sparse attention it feeds. Split the row into contiguous segments instead: one per subgroup with ballot counting over coalesced loads, or one per thread without subgroups. A single scan over the segment counts then gives each segment its output offset. The index list stays ascending.
* server: reject partial media truncation * server: keep only the keep_first fix Drop the mtmd test helper change, which no longer builds since clip_image_f32_batch stores its entries by value, and drop the vision test: no test fixture reaches a cut between two adjacent media chunks with a reused cache (tinygemma3 uses SWA and wraps images in text tokens, tinyopenjev and small-test are recurrent), so the test passed or failed independently of the fix. --------- Co-authored-by: Pascal <admin@serveurperso.com>
* ggml_cuda: optimize accumulation in mmq_vec_dot_fp4_fp4_mma for better performance * remove whitespace * fix: correct indentation in mma_block_scaled_fp4 loop
* server: support vision input for Clef * move input_attn_causal to private * extend old server_batch::embd * server_batch::token::pos to multi dim * nits * fix abort * fix img tokens cap * fix yield_to_queue mutate data
…tatistics (ggml-org#27990) * webui: Use toLocaleString() format consistently across chat message statistics * Add missing semicolon * ui: respect lint
Assisted-by: pi:llama.cpp/Qwen3.8-27B
* cuda: stage the lightning indexer queries in head passes for MUSA MUSA archs 21 and 22 cap static shared memory at 28 KB, and the tile kernel staged the queries of all four heads next to the key tile for 33 KB. The queries are now staged in passes of LIGHTNING_INDEXER_TILE_HEADS_PER_PASS heads: two on MUSA for 25 KB, four elsewhere where the single pass folds to the previous kernel. * cuda: use the vector lightning indexer kernel on MUSA Address review from am17an: the tile kernel stays off MUSA, whose archs 21 and 22 cap static shared memory at 28 KB, below the 33 KB the tile needs, so MUSA keeps the vector kernel it ran before. This replaces the head passes, CUDA and ROCm run the merged kernel unchanged.
…29936) * Fix Intel prefill regression on MoE models * Revert the n_per_expert change back to nei1
…gml-org#29987) Signed-off-by: Adrien Gallouët <angt@huggingface.co>
* llama.cpp : bump version to 0.6.0 * scripts : update summary prompt (#0)
* hexagon: head-parallel flash_attn partitioning for row-split multicore In row-split mode each core computes its output row shard of every MUL_MAT, but flash_attn was previously partitioning by Q tokens (flat qrow split) instead of by heads. This forced every core to read the full KV cache (all n_kv_heads), negating the memory bandwidth benefit of multicore on flash_attn. Change both HMX and HVX flash_attn kernels to partition by KV heads when n_kv_heads is divisible by n_cores: core i processes heads [i*n_kv_heads/N, (i+1)*n_kv_heads/N) exclusively, reading only its head shard of the KV cache. Falls back to the original token-block split when n_kv_heads % n_cores != 0 (e.g. Gemma-4 with 2 KV heads on 4 cores). Controlled by GGML_HEXAGON_FA_HEAD_SPLIT (default 1 = on). The flag is packed into bit 1 of the existing is_dst_fp32 kparams byte to stay within the 128-byte kernel_params blob limit. Measured gains at 4c row-split (PP t/s, ubatch=1024): Qwen3-0.6B: 6977 -> 11026 (+58%) llama-3.2-3B: 3717 -> 5522 (+49%) Qwen3.5-4B: 2739 -> 2855 (+4%) Gemma-4 MoE: no change (MoE FFN dominates, fallback path) TG is unchanged (flash_attn is a small fraction of decode time relative to the matmul+barrier cost per layer). * hex-fa: cleanup kern_params and head-split selection * hex-fa: add -fa-head-split option to run.py * hex-mdev: update matmul solver to account for reduced work in row-split scenarios * hex-mmid: better work splitting by expers in multi-dev scenarios * hex-fa: update HMX gating based on the model/n-hvx/ctx-len sweep * hex-fa: precompute softcap/scale on the host * hexagon: flatten matmul into 2d to use HMX in multi-sequence * hex-mm: cleanup kparams and use collapse to 3/4D -> 2D mapping * hex-mm: fix typo in collapse fallback * hex-mm: another pass at consistent naming for act tensors * hex-mm: add support for colapsing dims in fused matmuls * hex-build: fix WoS build errors * hex-mm: make sure to enforce dst stride in can_collapse * hex-fa: add a onliner commit for head-split check * hex-fa: remove unused local head_split var * hex-fa: tighten up can_split checks * hex-mm: update unfused paths to use act instead src1 * hex-mm: make sure to check all dsts for splitting * hexagon: fix the second weight chunk address in the batched HMX matmul prologue * hexagon: F16 activation and ragged N in the HMX matmul * hex-mm: tighten the ragged/split checks in mdev cases * hex-mm: enable MM fusion for F16 activations * hex-mm: pass tiled sizes to the solver in fused paths * hex-mmid: remove scalar divs from expert mapping loops * hex-mmid: proper cacheline safety enforcement for mdev splits * hex-mm: improve solver for mdev split scanarios and tail handling * hex-mm: remove redundant checks * hex-mm: fix fused HMX MUL_MAT_NX drops the final partial tile for quantized weights * hex-mm: better handling of ragged shapes (removes scalar memset of vtcm) --------- Co-authored-by: ebateni <ebateni@qti.qualcomm.com> Co-authored-by: Jhen-Jie Hong <iainst0409@gmail.com> Co-authored-by: Yiwei Shao <yiwei@aizip.ai>
…org#30008) Assisted-by: pi:llama.cpp/Qwen3.8-Flash-Next
* hexagon: add pool_2d support * hexagon: add pool_1d support * hex-pool: dma changes * hex-pool: Optimize HTP pooling boundaries and DMA pipelining * hex-pool: code cleanup and correctness fixes * hex-pool: re-write the DMA pipeline * hex-pool: pool chunking support * hex-pool: remove/vectorize all scalar paths * hex-pool: simplify chunk solver (no need for a loop) * hex-pool: remove redundant checks --------- Co-authored-by: Max Krasnyansky <maxk@qti.qualcomm.com>
First of two chunks (a6ea155..436f6f8). Upstream merged ggml-org#27694 (1fb7ef3), a later revision than the one we carried: replay moved into server_accept_replay, the draft params carry temp/seed instead of a sampling pointer, and the candidates are truncated with the draft. Upstream's version replaces ours. Conflicts: - common/sampling.cpp, common/sampling.h: take upstream (the whole fork delta was our older ggml-org#27694 revision with the is_replay parameter) - common/speculative.cpp: take upstream for spec_retune and both call sites (ggml-org#27694); cost routing and ngram-mod changes merged cleanly - common/speculative.h: upstream temp/seed fields, keep local prefer_mtp (cost routing) after them - tools/server/server-context.cpp: take upstream verify/replay path (ggml-org#27694); keep the local .prefer_mtp initializer after .seed - src/models/qwen35.cpp: keep upstream's optional cls_out tensors (ggml-org#29818) and the local d2t guard on the output fallback - ggml/src/ggml-cuda/ggml-cuda.cu: keep both the ggml-org#28702 gate/up SwiGLU matchers and upstream's shared-expert matcher (ggml-org#29184)
Second of two chunks (436f6f8..50569eb). Conflicts: - common/common.h: keep both includes (<memory> for the local ngram-mod table, <cstdio> from upstream ggml-org#29860) - common/speculative.h: keep upstream's common_speculative_are_compatible declaration and the local cost_routing parameter of common_speculative_init
matcher Upstream ggml-org#29633 added a warp_size parameter; the fused Q4_K gate/up SwiGLU matcher still used the old signature and failed to compile.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Merges 158 upstream commits (ggml-org
a6ea155d3..50569eb87) into the fork in two chunks:436f6f89e(graph: gather the recurrent states once, graph: fix CI realloc abort by gathering the recurrent states once ggml-org/llama.cpp#29856)50569eb87(upstream/master, 2026-10-05)Merge with "Create a merge commit" (never squash or rebase). Otherwise the next sync re-conflicts everything.
Dropped carried patches
1fb7ef3e3. Upstream merged a later revision than ours: replay moved intoserver_accept_replay, the draft params carrytemp/seedinstead of a sampling pointer, and the draft candidates are truncated with the draft. The flag is unchanged (--spec-draft-sampling probabilistic, default greedy). No local patch used its internals. Removed from "Carried patches" in AGENTS.md.Conflicts
Chunk 1:
common/sampling.cpp,common/sampling.h: take upstream; the whole fork delta was our older Make the drafter probabilistic and the target verify by rejection sampling for simple draft and MTP ggml-org/llama.cpp#27694 revision.common/speculative.cpp: take upstream forspec_retuneand both call sites (Make the drafter probabilistic and the target verify by rejection sampling for simple draft and MTP ggml-org/llama.cpp#27694); cost routing and ngram-mod merged cleanly.common/speculative.h: upstreamtemp/seedfields; keep localprefer_mtp(cost routing) after them.tools/server/server-context.cpp: take upstream's verify/replay path (Make the drafter probabilistic and the target verify by rejection sampling for simple draft and MTP ggml-org/llama.cpp#27694); keep the local.prefer_mtpinitializer.src/models/qwen35.cpp: keep upstream's optionalcls_outtensors (llama, server: add /v1/systemone API (models: laya, julia-1, lev, openjev, kev) ggml-org/llama.cpp#29818) and the local d2t guard on the output fallback.ggml/src/ggml-cuda/ggml-cuda.cu: keep both the CUDA: fuse FFN gate/up matmuls and GLU in MMQ ggml-org/llama.cpp#28702 gate/up SwiGLU matchers and upstream's shared-expert matcher (CUDA: fuse shared experts into MMVQ ggml-org/llama.cpp#29184).Chunk 2:
common/common.h: keep both includes (<memory>local,<cstdio>upstream).common/speculative.h: keep upstream'scommon_speculative_are_compatibleand the localcost_routingparameter.Checked on macOS (no CUDA)
sync.sh verify: the only files that now match upstream arecommon/sampling.{cpp,h}(Make the drafter probabilistic and the target verify by rejection sampling for simple draft and MTP ggml-org/llama.cpp#27694). The fork diff shrank only in the Make the drafter probabilistic and the target verify by rejection sampling for simple draft and MTP ggml-org/llama.cpp#27694 parts of arg/common/speculative/server-context. No new fork files. Leak check clean.ctest -L main: 53/54 pass.test-tokenizers-ggml-vocabsfails because the localmodels/ggml-vocabscheckout holds Git LFS pointer stubs; this machine's setup, not this merge.Follow-up (not changed in this sync)
Upstream ggml-org#29986 states that
graph_optimizealloc deps must not depend on batch size, or ggml-alloc re-reserves and the scheduler synchronizes at runtime. Our local ggml-org#28702 alloc dep (ggml_cuda_mul_mat_q_gate_up_swiglu_matchesinggml_backend_cuda_graph_optimize) checksy->ne[1]and the mmq/mmvq thresholds, so it is batch dependent. This behaviour predates the sync. If the bench shows sync stalls with varying draft sizes, split the matcher the same way upstream did: type/op checks in graph_optimize, batch checks in try_fuse.Target machine checklist
test-backend-ops -b CUDA0 -o MUL_MAT,MUL_MAT_ID,MUL_MAT_Q_GATE_UP_SWIGLU_DOWN,GATED_DELTA_NET,SSM_CONV,CPY,GLU,FLASH_ATTN_EXT: all OKbench_qwen38.py comparemaster vs this branch: pp/tg and spec acceptance within noisedraft_n_accepted> 0