Repository navigation
llama.cpp : bump version to 0.6.0 - #29997
Conversation
|
@ckastner The ABI/API check incorrectly find this release to require a major bump: https://github.com/ggml-org/llama.cpp/actions/runs/37328202699/job/111824258081#step:4:1289 . AFAICT the removed functions are not public-facing so it does not warrant a major version bump. WDYT? |
Looks to me like the script is checking too many libraries, including all of the non-public ones: I can reproduce the results from the failed check locally if I use the above list. The arguments should be reduced to only the public ones, so Edit: interestingly it still fails because |
So on the Debian side, we've been seeing So removal would technically be an ABI breakage. Though if eg: this symbol was never meant to be exported and just some artifact of wrong visibility, I think it should be OK to remove it without bumping anything. |
|
Ok, I see we have to narrow down the libs to just But I still don't understand why |
This commit adds API/ABI release checks similar to what is done for llama.cpp. Refs: ggml-org/llama.cpp#29997 (comment)
I see the declaration was GGML_API struct ggml_backend_buffer * ggml_backend_meta_alloc_ctx_tensors_from_buftand according to so it gets default visibility, and is therefore exported. I'll check if changing the default to hidden breaks anything, and will submit a PR if everything works. That would be the simplest fix. |
Overview
llama.cpp v0.6.0 introduces the new
llama_batch_extextended batch API (withllama_process) for mixed token/embedding inputs and MTP/deepstack state embeddings, adds support for the GLM-5.3-Flash (GLM5-Next) 320B hybrid model, the Clef decision model (text and vision) and MTP speculative decoding for Qwen4Exp, ships a new/v1/systemoneserver API for decision models (laya, julia-1, lev, openjev, kev, nimble), overhauls the Web UI with a Hugging Face Hub data layer and model download pipeline, adds a Metal tensor-API flash attention kernel for F16 KV, sparse flash attention for quantized K/V on Vulkan, and updates ggml to v0.26.0.Highlights
llama_batch_extextended batch API withllama_process(), supporting mixed token/embedding batches and per-token "state" embeddings for MTP and deepstack models #24669/v1/systemoneAPI supporting five decision models - laya, julia-1, lev, openjev (+vision), kev #29818llama_prefetch_rows()using MADVISE-based prefetching of PLE tensors in Qwen4Exp and Gemma4 #29599API changes
include/llama.h: newllama_batch_extbatch API withllama_embd,llama_process()andllama_process_type#24669, newllama_get_causal_attn()#28876, session formats bumped toLLAMA_SESSION_VERSION11 andLLAMA_STATE_SEQ_VERSION4include/llama-cpp.h: addedllama_batch_ext_ptrand deleter for the new extended batch API #24669tools/mtmd/mtmd.h:mtmd_get_memory_usage()now returns anmtmd_memory_usagestruct withimage_max_tokensanduse_non_causal#29773tools/server: new/v1/systemoneendpoint for decision models #29818 and/v1/embeddingsnow accepts typed vision/audio/video content #29556New models
Lfm2BidirectionalForMaskedLMfor LFM2.5-Encoder-230M/350M #29862classifier_poolingsupport for rerankers #29627Core changes
llama_batch_extAPI #29385 #29601; batches now accept both embd and raw tokens #29622llama_prec_policyand a model-driven W4A4 (NVFP4/MXFP4) mul_mat path #24364-smtensor #28569; GLM5-Next: unique scatter rows for dead indexer slots #29745causal_attn#28751Multi-modality changes
max_imageton_ubatchfor non-causal models #29773llama_batch_ext#29385Server changes
/v1/systemonedecision-model API with dedicated decision pipeline #29818, extended to the nimble decision model #29844GET /v1/modelsandGET /modelsnow report model input/output modalities in a newarchitectureobject #29987/v1/embeddings: accept typed content (vision/audio/video) input #29556 and return HTTP 400 for invalid embedding requests #29060LLAMA_ARG_HF_REPO_FILEkey in preset allow-list #29938n_batchton_ubatch#29903UI changes
svg useand animation elements in preview and download #28962toLocaleString()formatting consistently across chat message statistics #27990ggml changes
alloc_buffer_n/get_alloc_size_nbuffer allocation API, sparse flash attention kernels on SYCL, Vulkan and Metal, and major lightning indexer improvements (halved score memory, tiling, MUSA support). The CPU backend gains BF16 ops and a tiled k-quant mul_mat, CUDA gains a model-driven W4A4 (NVFP4/MXFP4) mul_mat path plus MMVQ shared-expert fusion, and the Hexagon backend adds a sampler and more quant types. Model loading is faster with stricter GGUF size validation, Windows ARM64 MSVC builds are enabled, and WebGPU/OpenVINO/OpenCL/SYCL pick up numerous new kernels, ops and fixes.