sync upstream llama.cpp commits 2026 09 15 - #28
Merged
Merged
Conversation
* server: refactor subproc handling * fix Windows build * download: keep concurrent downloads of one blob apart Every process writes the same path + .downloadInProgress, so a second download of the same blob finds that file, takes it for its own partial transfer and asks for the bytes after it, which produces a corrupt result. The in-progress file now carries the pid of the process writing it. std::rename also replaces an existing destination on POSIX but fails on Windows, so a download whose blob appeared in the meantime is dropped after every retry and an etag rewrite silently keeps the old value. std::filesystem::rename has the POSIX behaviour everywhere, and the error now carries the reason reported by the system. * Revert "download: keep concurrent downloads of one blob apart" This reverts commit 917b83f. * tests: serialize the router tests that download the same model Parallel workers share one cache, so the two tests fetch the same blob into the same in-progress file and race to rename it. They now take a file lock around the download, like the session fixture does for the preset models. * Revert "tests: serialize the router tests that download the same model" This reverts commit c368a4a. --------- Co-authored-by: Pascal <admin@serveurperso.com>
* ggml-webgpu: Update to a recent version of Dawn * No module scanning * Accept review suggestion to update comment Co-authored-by: Masashi Yoshimura <yoshimura.masashi.frbs@gmail.com> --------- Co-authored-by: Masashi Yoshimura <yoshimura.masashi.frbs@gmail.com>
…rg#28589) * hex-row-split: add support for multi-device row spliting Co-authored-by: Max Krasnyansky <maxk@qti.qualcomm.com> * hex-mdev: add work splitting to fused kernels * hex-mdev: use mdev_ prefix for all multi-device state * hex-mdev: make device configuration more expressive to support device groups * hex-mdev: fix mdev session init * hex-mdev: fused nx (2x,3x) matmuls must update row counts for each w/o * hex-mdev: fix MUL_MAT work partitioning bugs introduced by mdev * hex-cont: fix crashes with new tests due to wrong striding * hex-mdev: move fences after l2flushes * hex-cont: fix work splitting for mnpu -- align chunks to cachelines * hex-mdev: fix CPY tests with multi-dev * hex-mmid: fix work partitioning with mnpu * hex-mm: fix test failures with mdev * hex-binary: fix work partitioning for mdev * hex-argsort: fix mdev partitioning * hex-mdev: fix work partitioning and general updates for all simple ops * hex-fa: fix mdev work splitting issues * hex-mdev: fixing more failing ops test * hex-mdev: update the rest of the ops * hex-mdev: refactor all mdev splitting logic to be contained within if (mdev_count > 1) {...} * hex-mdev: fix macros * hex-mdev: simplify session flush logic * hex-sync: fix recursion in session flush * hex-mdev: factor out fence buffer and allocator * hex-fence: make fence allocation more robust with reserved slots for mdev * hex-mdev: keep all mdev state in htp_mdev_group * hex-mdev: further cleanup mdev group handling at the host * hex-mdev: update group idx in the opbatch before serializing * hex-batch: remove separate op_pending and use batch_req/rsp_seq * hex-async: workaround another missing tensor_init in ggml-meta * hex-fence: cleanup and robustify fences and error handling in multi-device scenarios * hex-ar: improve ALLREDUCE error handling * hex-async: robust error handling for op_cpy_fence * hex-async: use seq0 from allreduce context to allocate fence_seq * hex-mdev: fix remaining issues with fence and barrier clearing in CPY_FENCE * hex-misc: realign macros and fix misplaces trace events * hex-misc: align macros * hex-mdev: fix unclone buffer re-entrancy * hex-glu: fix mdev partitioning logic * hex-mdev: make buffer uncloning/cleanup work with tensor-split scenarios * hex-mdev: tighten up the can_split check in act-ops * hex-mdev: factor out common bits of the partitioning logic * hex-mm: minor realignment of the macros * hex-bufs: fix incorrectly placed assert for MAX_BUFS * hex-pad: tighten up gating checks for PAD * hex-kparams: make sure all kernels properly use kparams->n_threads * hex-docs: update user and developer docs with new features and detailed guide for ops development * hex-scripts: update run script to properly parse dev groups * hex-misc: formatting * hex-sess: minor cleanup for session init * hex-ar: fix vtcm size calc in allreduce kparams * hex-scripts: fix flake8 warnings * hex-rope: update ROPE to support mdev work split * hex-ops: remove redunant checks and minor reformat * hex-dev-guide: update dev-guide to avoid redundant null checks * hex-async: improve event_wait, event_sync and fence implementations * hex-async: remove synchronous flush from event_sync * hex-async: symplify fence recovery protocol and make sync more robust * hex-async: futher simplify error recovery for fences * hex-err: return status instead of just -1 * hex-async: print all seq nums in hex * hex-async: make sure fences flush dirty ranges * hex-async: add dirty ranges merging to reduce fence flushes * hex-async: properly sync before freeing the event * hex-async: make sure fence owner session is not overriden * hex-async: more fence write order more robust * hex-async: make sure not to fuse ALLREDUCE+ADD if their dsts overlap * hex-fusion: cleanup redundant checks --------- Co-authored-by: Alexander Lu <alexlu@qti.qualcomm.com>
Walk the binding offset back until the distance to the tensor is a whole number of blocks, so block quantized views get a valid element offset in the shader.
…a8_bin` (ggml-org#28677) * opencl: add A8 Q4_K non-MoE binary kernel * opencl: fix layout compatibility * opencl: rename binary kernel selection helpers --------- Co-authored-by: Li He <lih@qti.qualcomm.com>
…g#28747) The child writes its state commands on stdout while the logger writes on stderr, and both share a single pipe. The logger emits the trailing color reset after the newline of a debug, warn or error entry, so that escape sequence has no newline of its own and the router reads it glued in front of the next command. The line prefix check then fails and the command is forwarded as a log line instead of being handled, which leaves a finished download stuck in the downloading state. Writing the command with a leading newline closes the pending line so it always starts at a line boundary.
Signed-off-by: Adrien Gallouët <angt@huggingface.co>
Signed-off-by: Adrien Gallouët <angt@huggingface.co>
…org#28816) Clang stores the modification time of the precompiled header sources inside the header and refuses the header when they differ. A cached header restored from another checkout carries the timestamps of that checkout, so the build fails. The option covers the compilers ccache treats as MSVC while they are clang underneath, clang-cl and the Intel LLVM drivers.
…emas (ggml-org#28736) * common : implement common_schema types * common : implement a json schema optimizer * common : reduce optimizations * common : refactor json-schema-to-grammar to use common_schema * common : use common_trie * common/schema : implement type/kind resolution * cont : cleanup * cont : remove common_chat_tool_parameters * cont : simplify schema resolution * cont : pass common_schema through the json-schema-to-grammar builder * cont : cleanup * cont : move enums under common_schema and add type enum * cont : reduce test cases * cont : clean up * cont : clean up * refactor : rename common_schema_parse to common_schema_from_json * tests : fix gcc dangling-reference warning in test-json-schema * tests : take the schema label as const char * to satisfy gcc dangling-reference * refactor : rename common_schema_builder parse_* methods to build_* * cont : fix may_be_string * cont : properly handle empty tool parameters * cont : add tests for empty $ref * cont : remove dead code * cont : update docs * cont : make "{}" mean any object for json_object as well * cont : restore (min|max)Length to imply string type * cont : rename common_schema to common_chat_schema
* add LOG_JSON macro * fit: add demo LOG_JSON
* chat : improve schema support in qwen3 parser * cont : clean up grammar a bit
There is a driver bug where two queues on the same VkDevice simultaneously submitting can break some internal synchronization. Until it's fixed, add a mutex around queuesubmit.
…gml-org#28833) - Clamp the -j parallelism to min(nproc, 2) so a single-core runner uses -j 1 and multi-core runners use at most -j 2, instead of unconditionally using $(nproc). - Add a 3600s timeout to both test-backend-ops runs (the high-perf CPU path and the default path) so a hung test cannot stall CI indefinitely. - Note a TODO to reduce the timeout to 1800s in the future. Assisted-by: pi:llama.cpp/Qwen3.8-27B
Assisted-by: pi:llama.cpp/Qwen3.8-27B
…28854) Move the EditorConfig Checker and Code Style Checker workflows from the `[self-hosted, fast]` runners to `ubuntu-slim`, which is an established runner label in the repo. Assisted-by: pi:llama.cpp/Qwen3.8-27B
* fix for unsupport zes API * optimize the code * adjust the log level * rm unused head files * Update docs/backend/SYCL.md Co-authored-by: Titaniumtown <titaniumtown@proton.me> * fix the error to detect level zero SDK/dev package, stop build after detect the error * update the message * fix the build error when missed to install level zero dev package * rm GGML_SYCL_DEV_DEBUG, mv read env vars in all entry functions --------- Co-authored-by: Neo Zhang Jianyu <jianyu.zhang@intel.com> Co-authored-by: Titaniumtown <titaniumtown@proton.me> Co-authored-by: Neo Zhang <NA>
…ml-org#28835) Corrects a typo in `tests/test-quant-type-selection` for the Nvidia Nemotron 3 Nano 30B A3B model, which was referred to as *nvidia-nemotron-nano-3-30b-a3b*. The error made the test skip that test case, rather than failing the test. [no release]
…ero divisor (ggml-org#28779) The NextN/MTP tail loop derives the expert FFN size as n_ff/n_expert_used when expert_feed_forward_length gives nothing for the layer. Both values come from per-layer arrays that legitimately hold 0 on layers that are not MoE, so a checkpoint whose predict layers hold 0 in both divides by zero and dies with SIGFPE at load time, with no error message. Report the malformed metadata instead.
* rpc : hash-cache only weights ggml_backend_rpc_buffer_set_tensor and ggml_backend_rpc_set_tensor_async hashed every transfer above HASH_THRESHOLD and let `rpc-server -c` serve it from its file cache. The cache is meant for weights, but the activations ggml_backend_sched copies between backends took the same path: with a two-node split of Qwen3.8-Flash-Next every prefill ubatch above 10 MB was hashed, written to the worker's cache directory (1.4 TB after a day) and later served from there. Use the hash path only for tensors in buffers marked GGML_BACKEND_BUFFER_USAGE_WEIGHTS. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * rpc : save a cache entry only for the tensor that missed the hash check With the client hashing weights only, the server still wrote every SET_TENSOR above HASH_THRESHOLD to the cache directory, so the compute data the scheduler sends kept filling the disk. Remember the hash of the last SET_TENSOR_HASH that missed and save only the SET_TENSOR that follows it with that hash - the weight the client is re-sending. * rpc : signal the cache decision in the SET_TENSOR payload Replace the server-side `pending_cache` state with a `cache_flag` byte in the SET_TENSOR message: the client sets it when SET_TENSOR_HASH reported a miss, the server saves a cache entry only when it is set. Bump RPC_PROTO_MAJOR_VERSION since the wire format changes. --------- Co-authored-by: Patrick Hoffmann <patrickhoffmann@MacBook-Pro-14-HOP.local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
* ci: optimize * keep only the MUSA changes
This was referenced Sep 15, 2026
dzannotti
added a commit
that referenced
this pull request
Sep 15, 2026
dzannotti
added a commit
that referenced
this pull request
Sep 15, 2026
* upstream/master: (72 commits) HIP: Enable AllReduce for ROCm (ggml-org#27825) opencl: choose the MoE expert matmul by batch size for speculative decoding/MTP (ggml-org#27637) ci: build MUSA for only 1 arch (ggml-org#28944) docs: Rule of thumb for AI review time [no ci] (ggml-org#28945) rpc : hash-cache only weights (ggml-org#28789) cuda: support row-contiguous SUM_ROWS (ggml-org#26308) models : move build_arch_graph() after graph() template specialization (ggml-org#28934) vulkan: support sparse Flash Attention (ggml-org#28105) OpenVINO: optimize stateful decode and GPU MoE inference (ggml-org#28638) opencl: add generic ssm_scan (ggml-org#28881) ci: bump kleidiai runners from 22.04 to 24.04 (ggml-org#28885) metal : add FA kernels for HSK=96, HSV=64 (MiniCPM3) (ggml-org#28599) ci: Bump CUDA Windows x64 builds to 13.4.1 (ggml-org#28930) ci : fix android release (ggml-org#28936) cuda : enable i16 and i32 for DUP (ggml-org#28897) cmake : use PROJECT_SOURCE_DIR instead of CMAKE_SOURCE_DIR (ggml-org#28771) webui: stop re-probing disabled /tools endpoint on every message (ggml-org#28646) ci : reuse build tag name when used instead of safe one (ggml-org#28911) CI: hip-quality-check: ignore spill added in bfdc321 (ggml-org#28909) HIP: fattn-mma: use fp32 accumulation on MFMA devices (ggml-org#28576) ...
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
sync upstream llama.cpp commits