Skip to content

llama.cpp : bump version to 0.6.0 - #29997

Merged
ggerganov merged 2 commits into
masterfrom
llama-rc-v0.6.0
Oct 5, 2026
Merged

ggerganov merged 2 commits into
masterfrom
llama-rc-v0.6.0

Conversation

@ggerganov

@ggerganov ggerganov commented Oct 5, 2026 •

Copy link
Copy Markdown
Member

Overview

llama.cpp v0.6.0 introduces the new llama_batch_ext extended batch API (with llama_process) for mixed token/embedding inputs and MTP/deepstack state embeddings, adds support for the GLM-5.3-Flash (GLM5-Next) 320B hybrid model, the Clef decision model (text and vision) and MTP speculative decoding for Qwen4Exp, ships a new /v1/systemone server API for decision models (laya, julia-1, lev, openjev, kev, nimble), overhauls the Web UI with a Hugging Face Hub data layer and model download pipeline, adds a Metal tensor-API flash attention kernel for F16 KV, sparse flash attention for quantized K/V on Vulkan, and updates ggml to v0.26.0.

Highlights

  • New llama_batch_ext extended batch API with llama_process(), supporting mixed token/embedding batches and per-token "state" embeddings for MTP and deepstack models #24669
  • New models: GLM-5.3-Flash (GLM5-Next), a 320B text+vision hybrid model #27773, and the Clef decision model, fully supported with both text and vision #29831 #29969
  • Qwen4Exp: high-quality support is now available, with MTP speculative decoding (~1.5x decode speedup on DGX Spark) and various correctness fixes #29761 #29751
  • llama and server: new /v1/systemone API supporting five decision models - laya, julia-1, lev, openjev (+vision), kev #29818
  • Metal: new tensor API flash attention kernel for F16 KV #29570
  • Metal: new few-row MMA mat-mul kernels for speculative and batched decoding, up to ~3x faster mat-mul on Apple GPUs #29869
  • New llama_prefetch_rows() using MADVISE-based prefetching of PLE tensors in Qwen4Exp and Gemma4 #29599

API changes

  • include/llama.h: new llama_batch_ext batch API with llama_embd, llama_process() and llama_process_type #24669, new llama_get_causal_attn() #28876, session formats bumped to LLAMA_SESSION_VERSION 11 and LLAMA_STATE_SEQ_VERSION 4
  • include/llama-cpp.h: added llama_batch_ext_ptr and deleter for the new extended batch API #24669
  • tools/mtmd/mtmd.h: mtmd_get_memory_usage() now returns an mtmd_memory_usage struct with image_max_tokens and use_non_causal #29773
  • tools/server: new /v1/systemone endpoint for decision models #29818 and /v1/embeddings now accepts typed vision/audio/video content #29556

New models

  • GLM-5.3-Flash (GLM5-Next): 320B KDA/DSA hybrid text+vision model with mHC and MoE #27773
  • Clef decision model, fully supported with both text and vision #29831 #29969
  • Ling 3.0 VL, folded into the BailingMoeV3 architecture #29151
  • Nimble decision model #29844
  • Registered Lfm2BidirectionalForMaskedLM for LFM2.5-Encoder-230M/350M #29862
  • Added classifier_pooling support for rerankers #29627

Core changes

  • Migrated examples, speculative decoding, mtmd and server to the new llama_batch_ext API #29385 #29601; batches now accept both embd and raw tokens #29622
  • Added llama_prec_policy and a model-driven W4A4 (NVFP4/MXFP4) mul_mat path #24364
  • Qwen4Exp: halved indexer score memory #29825, optimized mask constructions #29824, re-enabled the -sm tensor #28569; GLM5-Next: unique scatter rows for dead indexer slots #29745
  • KV cache: fixed restoring mismatched KV cache rotation #28498, fixed K/V and recurrent state cleanup after failed restores #27530 and an invalid assert in recurrent memory #29799
  • Speculative decoding: probabilistic sampling for simple draft and MTP #27694, fixed n-gram drafts rejected at temp > 0 after truncation #29924, preserved original batch order for layer inputs #29019, stop accepting draft tokens at EOG #29638
  • k-pool models: fixed unexpected graph reallocation #29958 and clamped kpool re-pool bound to existing pools #29805
  • Fixed tensor split for fused qkv with uneven K/V head sizes #29294, gather recurrent states once so the reserve covers every split #29856, and properly handle KV on training #28520
  • Context: do not re-reserve the scheduler when toggling causal_attn #28751
  • DFlash drafts: write Gemma embedding scale during conversion #29802 and add dflash support for MiMo #29650

Multi-modality changes

  • Cap max_image to n_ubatch for non-causal models #29773
  • Fixed the mel preprocessor in LFM2 audio #29403
  • Migrated input processing to llama_batch_ext #29385

Server changes

  • New /v1/systemone decision-model API with dedicated decision pipeline #29818, extended to the nimble decision model #29844
  • Support vision input for Clef #29969
  • GET /v1/models and GET /models now report model input/output modalities in a new architecture object #29987
  • /v1/embeddings: accept typed content (vision/audio/video) input #29556 and return HTTP 400 for invalid embedding requests #29060
  • Allow RANK pooling batch splitting for causal LLM rerankers (Qwen3, Qwen3-VL) #28876
  • Reject partial media truncation #24076
  • Logging: self-contained colors and split child commands from logs in router mode #29895, allow preset to set log file #29334, fixed dead LLAMA_ARG_HF_REPO_FILE key in preset allow-list #29938
  • Fixed laya abort by limiting n_batch to n_ubatch #29903
  • Remove the built-in UI's service worker when the UI is not served #29565

UI changes

  • New model download pipeline #27959, Hugging Face Hub data layer #27947, model memory-fit estimation #27957 and model id grammar for sidecars, quants and capability parsing #27946
  • Type-safe API types, fetch helpers and download-ready models store plumbing #29582
  • Shared model display primitives #29644
  • Fixed missing svg use and animation elements in preview and download #28962
  • Use toLocaleString() formatting consistently across chat message statistics #27990

ggml changes

  • ggml updated to v0.26.0: a new alloc_buffer_n/get_alloc_size_n buffer allocation API, sparse flash attention kernels on SYCL, Vulkan and Metal, and major lightning indexer improvements (halved score memory, tiling, MUSA support). The CPU backend gains BF16 ops and a tiled k-quant mul_mat, CUDA gains a model-driven W4A4 (NVFP4/MXFP4) mul_mat path plus MMVQ shared-expert fusion, and the Hexagon backend adds a sampler and more quant types. Model loading is faster with stricter GGUF size validation, Windows ARM64 MSVC builds are enabled, and WebGPU/OpenVINO/OpenCL/SYCL pick up numerous new kernels, ops and fixes.

@github-actions github-actions Bot added the build Compilation issues label Oct 5, 2026
@ggerganov
ggerganov merged commit d812350 into master Oct 5, 2026
22 of 27 checks passed
@ggerganov
ggerganov deleted the llama-rc-v0.6.0 branch October 5, 2026 15:13
@ggerganov

Copy link
Copy Markdown
Member Author

@ckastner The ABI/API check incorrectly find this release to require a major bump: https://github.com/ggml-org/llama.cpp/actions/runs/37328202699/job/111824258081#step:4:1289 . AFAICT the removed functions are not public-facing so it does not warrant a major version bump. WDYT?

@ckastner

ckastner commented Oct 5, 2026 •

Copy link
Copy Markdown
Collaborator

AFAICT the removed functions are not public-facing so it does not warrant a major version bump. WDYT?

Looks to me like the script is checking too many libraries, including all of the non-public ones:

Libraries found in new build: libggml-base libggml-cpu libggml libllama-batched-bench-impl libllama-bench-impl libllama-cli-impl libllama-common libllama-completion-impl libllama-fit-params-impl libllama-perplexity-impl libllama-quantize-impl libllama-server-impl libllama libmtmd

I can reproduce the results from the failed check locally if I use the above list.

The arguments should be reduced to only the public ones, so libggml-base libggml libllama libmtmd (and whatever else I might have missed).

Edit: interestingly it still fails because ggml_backend_meta_alloc_ctx_tensors_from_buft was supposedly removed. I need to look a bit deeper.

@ckastner

ckastner commented Oct 5, 2026

Copy link
Copy Markdown
Collaborator

Edit: interestingly it still fails because ggml_backend_meta_alloc_ctx_tensors_from_buft was supposedly removed. I need to look a bit deeper.

So on the Debian side, we've been seeing ggml_backend_meta_alloc_ctx_tensors_from_buft since ggml v0.10.0.

So removal would technically be an ABI breakage. Though if eg: this symbol was never meant to be exported and just some artifact of wrong visibility, I think it should be OK to remove it without bumping anything.

@ggerganov

Copy link
Copy Markdown
Member Author

Ok, I see we have to narrow down the libs to just libllama + libmtmd. And we'll move the ggml checks to the ggml repo so we make the ABI check during ggml releases.

But I still don't understand why ggml_backend_meta_alloc_ctx_tensors_from_buft is considered to cause a breaking change. It's declared in the private ggml/src/ggml-backend-impl.h.

danbev added a commit to danbev/ggml that referenced this pull request Oct 5, 2026
This commit adds API/ABI release checks similar to what is done for
llama.cpp.

Refs: ggml-org/llama.cpp#29997 (comment)
danbev added a commit to danbev/llama.cpp that referenced this pull request Oct 5, 2026
@ckastner

ckastner commented Oct 5, 2026

Copy link
Copy Markdown
Collaborator

But I still don't understand why ggml_backend_meta_alloc_ctx_tensors_from_buft is considered to cause a breaking change. It's declared in the private ggml/src/ggml-backend-impl.h.

I see the declaration was

GGML_API struct ggml_backend_buffer * ggml_backend_meta_alloc_ctx_tensors_from_buft

and according to ggml.h, on not-Windows with GGML_SHARED, GGML_API is defined as

#define GGML_API __attribute__ ((visibility ("default"))) extern

so it gets default visibility, and is therefore exported.

I'll check if changing the default to hidden breaks anything, and will submit a PR if everything works. That would be the simplest fix.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

build Compilation issues

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants