Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
31 commits
Select commit Hold shift + click to select a range
cf78650
exclude GPU/NPU failing POOL_2D case
mostafafaheem Aug 22, 2026
f72d986
Fix pool case
mostafafaheem Aug 27, 2026
011271a
ggml-openvino: fix stateful decode for Gemma-4 per-layer-type head sizes
cavusmustafa Jul 17, 2026
70ceed6
ggml-openvino: fix MSVC narrowing error in permute
cavusmustafa Jul 24, 2026
4fb04df
ggml-openvino: classify sliding-window layers structurally on interle…
cavusmustafa Aug 11, 2026
0e9794e
ggml-openvino: add GGML_OPENVINO_REQUANT_KQUANT to select a 4-bit req…
cavusmustafa Aug 11, 2026
c14780c
ggml-openvino: add GGML_OPENVINO_SPILL_DIR to spill weight buffers to…
cavusmustafa Aug 14, 2026
cd327a3
Stateful Performance: Added pass::KVStateSeqAxis to change KV layout
cavusmustafa Aug 20, 2026
80aec82
ggml-openvino: fix stateful decode past the sliding-window size
cavusmustafa Aug 24, 2026
ca08e4f
ggml-openvino: refuse stateful decode that cannot resume from the KV …
cavusmustafa Aug 27, 2026
fecf802
ggml-openvino: use the per-layer KV head count for the stateful KV state
cavusmustafa Aug 27, 2026
6fbe283
ggml-openvino: apply the KV state relayout to any KV head count
cavusmustafa Aug 27, 2026
66d2d07
ggml-openvino : support ggml_rope_set_offset and simplify op support …
mostafafaheem Aug 28, 2026
73ddeb5
add more cpy cases
mostafafaheem Aug 28, 2026
ba4e39d
reject BF16 cpy on NPU
mostafafaheem Sep 1, 2026
f304295
Remove mul_mat_id fallback, gate large mul_mat_id only for mxfp4
wine99 Aug 20, 2026
6645a9e
ggml-openvino: fuse the MoE expert block into MOECompressed on GPU
cavusmustafa Aug 25, 2026
8c24268
ggml-openvino: skip GPU MUL_MAT_ID for unbound expert tensors
cavusmustafa Sep 3, 2026
314afa5
ggml-openvino: requantize grouped 8-bit MoE experts on GPU
cavusmustafa Sep 3, 2026
db83eee
Enable special strided CPY for conv state writeback
wine99 Sep 4, 2026
e871404
openvino: support cacheless encoder models on NPU
zhaixuejun1993 Sep 1, 2026
e13baf3
openvino: optimize norm and RoPE translation
zhaixuejun1993 Sep 1, 2026
a59f311
ggml-openvino : simplify op translators and enable IMROPE/NEOX RoPE f…
mostafafaheem Sep 9, 2026
91903c8
remove unnecessary include and clean up PAD
mostafafaheem Sep 9, 2026
eec425e
fix mulmat bug
mostafafaheem Sep 9, 2026
926ef68
use ov::as_type_ptr instead of std::dynamic_pointer_cast
mostafafaheem Sep 9, 2026
db8913c
ggml-openvino: fix mixed-dtype ADD/SWIGLU_CLAMP, gate unsupported ROP…
ravi9 Sep 15, 2026
88b44d0
openvino: share compiled models with per-context inference state; fix…
wine99 Sep 10, 2026
a81e3f6
ggml-openvino: gate MoE expert-sum ReduceSum shortcut past 8 experts
ravi9 Sep 15, 2026
3200838
ggml-openvino: gate degenerate m=1,n=1 MUL_MAT on GPU
ravi9 Sep 15, 2026
5bcd117
ggml-openvino: make SoftPlus decomposition opt-in native
wine99 Sep 15, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
5 changes: 5 additions & 0 deletions ci/run.sh
Original file line number Diff line number Diff line change
Expand Up @@ -669,6 +669,11 @@ function gg_run_test_backend_ops {
args_extra=""
fi

# TODO: OpenVINO GPU plugin crashes (CL_OUT_OF_RESOURCES) with 2 concurrent workers on GPU.
if [ ! -z "${GG_BUILD_OPENVINO}" ] && [ "${GGML_OPENVINO_DEVICE:-}" = "GPU" ]; then
args_extra=""
fi

# TODO: reduce the test-backend-ops timeout to 1800s
if [ ! -z ${GG_BUILD_HIGH_PERF} ]; then
(time timeout 3600 ./bin/test-backend-ops ${args_extra} -b CPU) 2>&1 | tee -a $OUT/${ci}-test-backend-ops.log
Expand Down
3 changes: 3 additions & 0 deletions docs/backend/OPENVINO.md
Original file line number Diff line number Diff line change
Expand Up @@ -719,10 +719,13 @@ Boolean flags follow a uniform convention: set to a **positive integer** (e.g. `
| `GGML_OPENVINO_STATEFUL_EXECUTION`| Boolean | `0` | Enable stateful KV cache for better performance. Recommended on CPU, GPU. |
| `GGML_OPENVINO_DISABLE_CACHE` | Boolean | `0` | Disable the in-process compiled-model / decoder cache (cache is on by default). Set to `1` to disable. |
| `GGML_OPENVINO_DISABLE_KV_SLICE` | Boolean | `0` | Disable the KV-cache input-tensor slicing optimization (slicing is on by default on CPU/GPU). Set to `1` to disable. |
| `GGML_OPENVINO_DISABLE_KV_STATE_RELAYOUT` | Boolean | `0` | Disable the stateful KV-state sequence-axis relayout (relayout is on by default). It moves the KV state sequence axis from dim 1 to dim 2, so the GPU plugin can append new tokens in place instead of copying the whole state every token, and the reader side no longer transposes the whole accumulated state. Set to `1` to disable. |
| `GGML_OPENVINO_MANUAL_GQA_ATTN` | Boolean | device-based | Tri-state. When **unset**, manual GQA attention is enabled by default on `GPU` and disabled on other devices. Set to a positive integer to force-enable, or `0` to force-disable. |
| `GGML_OPENVINO_MEMORY_OPTIMIZE` | Boolean | `0` | Umbrella switch for compile-time memory reductions. Enables `GGML_OPENVINO_REDUCE_COMPILE_MEM` and, on GPU, `GGML_OPENVINO_RELEASE_WEIGHTS` unless those fine-grained variables are explicitly set. |
| `GGML_OPENVINO_REDUCE_COMPILE_MEM`| Boolean | inherits from `GGML_OPENVINO_MEMORY_OPTIMIZE` | Reduce compile-time host memory use by streaming weight requantization and avoiding extra weight-node materialization where possible. Set explicitly to override the umbrella switch. |
| `GGML_OPENVINO_RELEASE_WEIGHTS` | Boolean | inherits from `GGML_OPENVINO_MEMORY_OPTIMIZE` on GPU | GPU-only. Release host weight buffers after the compiled model cache can reuse the device/plugin copy. Requires stable graph shapes; dynamic workloads that need recompilation should leave this disabled. |
| `GGML_OPENVINO_SPILL_DIR` | String | `not set` | Directory for a disk-backed weight buffer. When set, the repacked weight buffer is mapped from an unlinked file on this path instead of anonymous memory, so its pages are reclaimable under memory pressure instead of staying pinned, cutting the load-time host memory peak. Must point at real storage; a tmpfs mount (e.g. `/tmp` on many systems) backs it with RAM and makes the peak worse. |
| `GGML_OPENVINO_REQUANT_KQUANT` | String | `not set` | Requantize Q6_K/Q5_K weights (and matching MoE expert weights) to a 4-bit target instead of the default Q8_0_C, trading accuracy for less memory traffic. One of `q4_sym128` (Q6_K/Q5_K only), `q4_sym128_all` (Q4_K too, drops its per-group zero point), `q4_asym64_all` (Q6_K/Q5_K/Q4_K, keeps a real zero point at group 64), or `native` (no requantization). |
Comment thread
0cc4m marked this conversation as resolved.
| `GGML_OPENVINO_PROFILING` | Boolean | `0` | Enable execution-time profiling. |
| `GGML_OPENVINO_DUMP_CGRAPH` | Boolean | `0` | Dump the GGML compute graph to `cgraph_ov.txt`. |
| `GGML_OPENVINO_DUMP_IR` | Boolean | `0` | Serialize OpenVINO IR files with timestamps. |
Expand Down
266 changes: 231 additions & 35 deletions ggml/src/ggml-openvino/ggml-decoder.cpp

Large diffs are not rendered by default.

76 changes: 72 additions & 4 deletions ggml/src/ggml-openvino/ggml-decoder.h
Original file line number Diff line number Diff line change
Expand Up @@ -21,18 +21,28 @@ struct ModelParams {
int ctx_per_seq_swa = -1;
int n_seq = 1;
int n_heads_kv = -1;
// Per-layer KV head count. gemma-4 12B interleaves 8 x 256 sliding layers with 1 x 512
// full-attention layers, so no single scalar describes every layer. Keyed by layer, not by
// layer TYPE, because the SWA classification depends on the context size (extents tie at a
// small -c) while the head count does not.
std::map<int, int> n_heads_kv_per_layer;
int head_size = -1;
int state_size = -1; // for SSM molels, eg qwen35
int32_t rope_params[15];
int32_t rope_params[16];
bool mixed_rope_params = false;
bool is_cacheless_attn = false;
std::vector<int> swa_layers;
// The sliding-window mask tensor, identified in compute_llm_params() by grouping attention
// layers on the mask they consume. Only used to tell the two masks apart when naming OV
// parameters -- both carry the same tensor name. Null when the graph has a single mask.
const ggml_tensor * swa_mask = nullptr;

std::vector<std::string> kv_names;
size_t kv_buffer_ctx_id = 0;

bool same_rope_params(const ModelParams & other) const {
return mixed_rope_params == other.mixed_rope_params &&
memcmp(rope_params, other.rope_params, sizeof(int32_t) * 15) == 0;
memcmp(rope_params, other.rope_params, sizeof(int32_t) * 16) == 0;
}

bool can_reuse_dynamically(const ModelParams & other) const { return same_rope_params(other); }
Expand All @@ -48,6 +58,11 @@ struct ComputeParams {
int attention_size = -1;
int attention_size_swa = -1;
int attention_size_static = -1; // encoder/cross-attn KV fill level (whisper)
// Sliding window width, read back from the band of ggml's own SWA mask. ggml never passes
// n_swa down to a backend, but fill_mask() bakes it into the mask contents, so the widest
// unmasked row recovers it. Shorter than n_swa while the sequence is still short, which is
// harmless: every causal pair is inside the window then anyway.
int swa_window = -1;
int input_len = -1;
int token_len_per_seq = -1;
int past_kv_len = -1;
Expand Down Expand Up @@ -96,8 +111,15 @@ struct ComputeParams {
// models use a fixed end-anchored offset in the translator.
};

// defined below; declared here because GgmlOvDecoder uses it inline
std::optional<int> extract_layer_from_name(const std::string & name);

// detects the MoE expert-plane-sum ADD chain (see definition); used by supports_op too
bool is_moe_expert_sum_add(const ggml_tensor * node);

class GgmlOvDecoder : public ov::frontend::ggml::GgmlDecoder {
public:
static std::string get_tensor_name(const ggml_cgraph * cgraph, const ggml_tensor * tensor);
struct NodeInfo {
ggml_tensor * node;
std::string node_name;
Expand Down Expand Up @@ -250,6 +272,21 @@ class GgmlOvDecoder : public ov::frontend::ggml::GgmlDecoder {
m_model_params.swa_layers.end();
}

// KV head count for one layer. Sliding and full layers can differ (gemma-4 12B), so callers
// that reinterpret a KV buffer must use this and not the model-level n_heads_kv.
int get_n_heads_kv_for_layer(int layer) const {
auto it = m_model_params.n_heads_kv_per_layer.find(layer);
return it != m_model_params.n_heads_kv_per_layer.end() ? it->second : m_model_params.n_heads_kv;
}

// Same, for a KV cache tensor: its layer comes from the leaf name (cache_k_l<N>).
int get_n_heads_kv_for_tensor(const ggml_tensor * kv_tensor) const {
if (auto layer = extract_layer_from_name(std::string(kv_tensor->name)); layer.has_value()) {
return get_n_heads_kv_for_layer(layer.value());
}
return m_model_params.n_heads_kv;
}

int get_past_kv_len() const { return m_compute_params.past_kv_len; }

int get_input_len() const { return m_compute_params.input_len; }
Expand Down Expand Up @@ -340,6 +377,12 @@ class GgmlOvDecoder : public ov::frontend::ggml::GgmlDecoder {
(op->op == GGML_OP_SOFT_MAX && tensor == op->src[1]);
}

inline static bool is_inp_mean(const ggml_tensor * tensor, const ggml_tensor * op) {
return op->op == GGML_OP_MUL_MAT && tensor == op->src[1] && tensor->op == GGML_OP_NONE &&
(tensor->flags & GGML_TENSOR_FLAG_INPUT) && tensor->type == GGML_TYPE_F32 &&
op->src[0] != nullptr && op->src[0]->op != GGML_OP_NONE;
}

inline static bool is_rope_freqs_weight(const ggml_tensor * tensor, const ggml_tensor * op) {
return op->op == GGML_OP_ROPE && tensor == op->src[2];
}
Expand All @@ -353,10 +396,21 @@ class GgmlOvDecoder : public ov::frontend::ggml::GgmlDecoder {
(op != nullptr && op->op == GGML_OP_SET_ROWS && op->src[2] == tensor);
}

inline static bool is_conv_state_writeback(const ggml_tensor * node) {
return node->op == GGML_OP_CPY && node->view_src != nullptr && is_kvcache(node->view_src, nullptr) &&
node->src[0] != nullptr && node->src[0]->op == GGML_OP_VIEW && node->src[0]->src[0] != nullptr &&
node->src[0]->src[0]->op == GGML_OP_CONCAT && node->src[1] != nullptr &&
node->src[1]->op == GGML_OP_VIEW && node->src[1]->view_src == node->view_src;
}

inline static bool is_kv_idx(const ggml_tensor * tensor, const ggml_tensor * op) {
return op->op == GGML_OP_SET_ROWS && op->src[1] == tensor;
}

bool is_swa_mask(const ggml_tensor * tensor) const {
return m_model_params.swa_mask != nullptr && tensor == m_model_params.swa_mask;
}

inline static bool is_output_idx(const ggml_tensor * tensor, const ggml_tensor * op) {
return op->op == GGML_OP_GET_ROWS && tensor == op->src[1] && op->src[0]->op != GGML_OP_NONE &&
op->src[1]->op == GGML_OP_NONE;
Expand All @@ -375,8 +429,22 @@ class GgmlOvDecoder : public ov::frontend::ggml::GgmlDecoder {
if (is_inp_emb(tensor, op)) {
return "embd";
}
if (is_stateful() && is_inp_mask(tensor, op)) {
return std::string(tensor->name).find("swa") == std::string::npos ? "self_kq_mask" : "self_kq_mask_swa";
if (is_inp_mask(tensor, op)) {
// Give the two attention masks distinct OV parameter names.
//
// An interleaved-SWA model builds one full-attention mask and one sliding-window mask,
// but build_attn_inp_kq_mask() names them identically, so keying a parameter off
// tensor->name alone makes the second mask OVERWRITE the first in m_model_inputs: both
// attention types then read a single parameter, and the windowed layers silently run
// against an unbanded mask. Disambiguate using the SWA layer set computed in
// compute_llm_params(), which classifies by mask tensor identity rather than by name.
//
// When no SWA layer was found there is only one mask in play, so the plain name is
// correct and no _swa parameter is created.
if (m_model_params.swa_layers.empty()) {
return "self_kq_mask";
}
return is_swa_mask(tensor) ? "self_kq_mask_swa" : "self_kq_mask";
}
return tensor->name;
}
Expand Down
88 changes: 88 additions & 0 deletions ggml/src/ggml-openvino/ggml-openvino-extra.cpp
Original file line number Diff line number Diff line change
Expand Up @@ -31,6 +31,7 @@ void ggml_openvino_device_config::init() {
// String values (use ggml_openvino_getenv_str)
"GGML_OPENVINO_DEVICE",
"GGML_OPENVINO_CACHE_DIR",
"GGML_OPENVINO_SPILL_DIR",
"GGML_OPENVINO_DEBUG_NODE",
"GGML_OPENVINO_COMPILED_MODEL_CACHE_DIR",
"GGML_OPENVINO_NPU_COMPILE_CONFIG",
Expand All @@ -56,6 +57,11 @@ void ggml_openvino_device_config::init() {
"GGML_OPENVINO_RELEASE_WEIGHTS",
"GGML_OPENVINO_REDUCE_COMPILE_MEM",
"GGML_OPENVINO_LOG_UNSUPPORTED_OPS",
"GGML_OPENVINO_LOG_SWA_LAYERS",
"GGML_OPENVINO_NATIVE_SOFTPLUS",
"GGML_OPENVINO_DISABLE_REMOTE_OUTPUTS",
"GGML_OPENVINO_REQUANT_KQUANT",
"GGML_OPENVINO_DISABLE_KV_STATE_RELAYOUT",
};

for (const char * const & env_var : env_var_names) {
Expand Down Expand Up @@ -263,9 +269,81 @@ std::optional<ExtraQuantType> ggml_openvino_get_requant_type(const ggml_tensor *
if (ggml_openvino_is_npu()) {
return ExtraQuantType::Q4_0_128;
}
// By default Q6_K/Q5_K are requantized to Q8_0_C, which *inflates* 6- and 5-bit weights to 8
// while the rest of the model stays at 4 bits, and Q4_K keeps its native group-32 layout
// (an f16 scale plus an f16 zero point per 32 weights = 0.125 B/weight of metadata).
// Decode of a large model is bandwidth-bound, so both cost throughput.
//
// GGML_OPENVINO_REQUANT_KQUANT selects a 4-bit target instead. Names are
// q4_<sym|asym><group>[_all]: <sym|asym> says whether a per-group zero point is kept, <group>
// is the group size, and the _all suffix sends Q4_K down the same path (without it only
// Q6_K/Q5_K are touched):
// q4_sym128 Q6_K/Q5_K -> Q4_0_128 (u4, group 128, symmetric)
// q4_sym128_all and Q4_K too -- drops Q4_K's per-32 zero point, which costs some accuracy
// q4_asym64_all Q6_K/Q5_K and Q4_K -> Q4_1_64 (u4, group 64, asymmetric) -- most of the
// metadata saving while keeping a real zero point
// native no requantization at all (keep Q6_K/Q5_K as they are)
//
// The asymmetric target is only offered in its _all form: leaving Q4_K at its native group 32
// while Q6_K/Q5_K move to group 64 gives the Q/K/V projections different group counts, and the
// GPU plugin's FullyConnectedHorizontalFusion concatenates their scale constants, which then
// fails shape inference. Requantizing all three keeps the group size uniform.
const char * rq = ggml_openvino_getenv_str("GGML_OPENVINO_REQUANT_KQUANT");
auto is_opt = [rq](const char * name) {
return rq && strcmp(rq, name) == 0;
};
const bool sym128 = is_opt("q4_sym128");
const bool sym128_all = is_opt("q4_sym128_all");
const bool asym64_all = is_opt("q4_asym64_all");

if (tensor->type == GGML_TYPE_Q4_K) {
if (sym128_all) {
return ExtraQuantType::Q4_0_128;
}
if (asym64_all) {
return ExtraQuantType::Q4_1_64;
}
}
// MoE expert weights (3D, ne[2] = n_expert) stored as Q5_1/Q8_0 are the expert-side
// equivalent of Q6_K/Q5_K: kept at 8 bits by default while the rest of the model is at 4
// (gemma-4 26B-A4B keeps its down projection there). Send them to 4 bits under the same
// option, at group 64 rather than 128: the down expert has k=704, which 64 divides
// (704/64 = 11) and 128 does not.
if (tensor->ne[2] > 1 && (tensor->type == GGML_TYPE_Q5_1 || tensor->type == GGML_TYPE_Q8_0)) {
if (sym128 || sym128_all) {
return ExtraQuantType::Q4_0_64;
}
if (asym64_all) {
return ExtraQuantType::Q4_1_64;
}
// TODO: temporary workaround for a known OpenVINO GPU-plugin bug -- remove once the
// plugin computes grouped 8-bit GatherMatmulCompressed correctly. This costs accuracy
// (5/8-bit -> 4-bit) on any model it applies to, so it must not outlive the bug.
//
// On GPU these would otherwise stay in their native *grouped 8-bit* layout, which the GPU
// plugin's GatherMatmulCompressed computes incorrectly -- gemma-4 26B-A4B (whose down
// projection is Q5_1) produces garbage, while the same graph is correct on CPU. It is
// specific to grouped 8 bit: the gate/up experts are grouped u4 *with* a zero point and
// are fine, and Qwen3.5 / granite are fine because their Q5_K/Q6_K down projections
// already requantize to per-channel Q8_0_C (grouped=0). Sending these to grouped 4 bit
// avoids the broken layout and restores correct output.
// Opt out with GGML_OPENVINO_REQUANT_KQUANT=native.
if (ggml_openvino_get_device_name() == "GPU" && !is_opt("native")) {
return ExtraQuantType::Q4_0_64;
}
}
switch (tensor->type) {
case GGML_TYPE_Q6_K:
case GGML_TYPE_Q5_K:
if (sym128 || sym128_all) {
return ExtraQuantType::Q4_0_128;
}
if (asym64_all) {
return ExtraQuantType::Q4_1_64;
}
if (is_opt("native")) {
return std::nullopt;
}
return ExtraQuantType::Q8_0_C;
default:
return std::nullopt;
Expand Down Expand Up @@ -331,6 +409,16 @@ ggml_openvino_extracted_layout ggml_openvino_get_extracted_layout(const ggml_ten
layout.weights_per_block = 128;
layout.is_symmetric = true;
break;
case ExtraQuantType::Q4_1_64:
layout.is_u4 = true;
layout.weights_per_block = 64;
layout.is_symmetric = false;
break;
case ExtraQuantType::Q4_0_64:
layout.is_u4 = true;
layout.weights_per_block = 64;
layout.is_symmetric = true;
break;
case ExtraQuantType::Q4_0_C:
layout.is_u4 = true;
layout.weights_per_block = tensor->ne[0];
Expand Down
5 changes: 4 additions & 1 deletion ggml/src/ggml-openvino/ggml-openvino-extra.h
Original file line number Diff line number Diff line change
Expand Up @@ -15,7 +15,10 @@
#include <string>

// ExtraQuantType enum - defines requantization target formats
enum class ExtraQuantType { F16, Q4_0_C, Q8_1_C, Q4_0_128, Q8_0_C, Q8_0_32 };
// Q4_1_64: u4, group 64, *true* asymmetric (per-group scale and zero point). Note that
// Q4_0_128/Q4_0_C are symmetric despite taking the unsigned branch of quantize_q4_0 -- that branch
// pins zp to 8 with d = max/-8, which is algebraically symmetric.
enum class ExtraQuantType { F16, Q4_0_C, Q8_1_C, Q4_0_128, Q4_0_64, Q8_0_C, Q8_0_32, Q4_1_64 };

ov::Core & ov_singleton_core();

Expand Down
Loading
Loading