Commit b92761a51 for llama.cpp
commit b92761a515ea31e852e7fbc1fad5f874b46f3718
Author: Ravi Panchumarthy <ravi.panchumarthy@intel.com>
Date: Sat Oct 3 14:29:25 2026 +0530
ggml-openvino: update to 2026.4.1, optimize performance, expand ops, improve device listing. (#29852)
* ggml-openvino : Qwen3.5 MoE perf (#312)
Squash of ravi9/llama.cpp#312:
- ggml-openvino: add detailed inference profiling (Yu, Zijun)
- ggml-openvino: use remote output tensors by default (Yu, Zijun)
- ggml-openvino: optimize single-sequence recurrent state (Yu, Zijun)
- opt1: remove recurrent reset for single sequence, opt2: direct gdn outputs (break parallel sequence) (Yu, Zijun)
- fix parallel sequences (Yu, Zijun)
- ggml-openvino: simplify graph cache key (ynimmaga)
- enable stateful for qwen35 single sequence (Yu, Zijun)
- Fix after rebasing (Yu, Zijun)
- Add k-requant option q4_asym64 (Yu, Zijun)
- Fix qwen35 llama-bench -p 0 (Yu, Zijun)
- Simplify RESHAPE translation (Yu, Zijun)
- openvino: fuse MoE routing (Yu, Zijun)
- openvino: fuse GDN qk normalization (Yu, Zijun)
- openvino: enable GPU MoE fusion by default (Yu, Zijun)
- ggml-openvino: add cache_only mode to import cached compiled model on disk directly (Yu, Zijun)
- openvino : report the device allocation limit to ggml (Łukasz Ślusarczyk)
- Fix windows build (Yu, Zijun)
Co-authored-by: ynimmaga <ynimmaga@users.noreply.github.com>
Co-authored-by: Łukasz Ślusarczyk <lukasz.slusarczyk@intel.com>
* ggml-openvino: Update doc of compiled model cache
* openvino: implement PRD-compliant device enumeration and memory reporting
* openvino: fix multi-device listing issues from review
- Only the device selected by GGML_OPENVINO_DEVICE reports as GPU; the
other OpenVINO devices report as IGPU so llama.cpp does not offload to
them. Initializing a non-selected device logs a warning.
- Name devices OPENVINO<i> again and show the OpenVINO id in the
description. Raw "CPU" names shadowed the ggml CPU backend.
- Support GPU.N: create the OpenCL queue on OpenVINO's own context for
the selected device, and replace "GPU"/"NPU" string comparisons with
ggml_openvino_is_gpu()/ggml_openvino_is_npu().
- An unavailable GGML_OPENVINO_DEVICE is now an error that lists the
available devices, instead of silently falling back to CPU.
- Memory: cap iGPU/NPU free memory at system available memory, fall back
to system memory instead of 0/0 when the plugin lacks memory
properties, and ignore host USM allocations in GPU usage.
- Initialize the device config once under a lock, even if OpenCL setup
fails.
- Fix supports_op return type for non-selected devices (build error).
* openvino : take USM entry points from the selected device platform
clGetExtensionFunctionAddressForPlatform was called on the first platform
returned by clGetPlatformIDs. The address it returns is only valid for the
platform it was queried on, and the first platform is not always the one that
holds the device OpenVINO selected.
On a host whose first platform comes from another vendor the lookup returns
null, and then every read, write and memset on a GPU buffer fails with
"clEnqueueMemcpyINTEL not available".
Look both entry points up in init(), on the platform of the device OpenVINO
picked, and keep them in the device config next to the command queue.
Assisted-by: Claude Opus 5
* openvino: fuse MoE experts for models with a fused gate_up weight
FuseMoeCompressed only matches models whose gate and up projections are
separate GatherMatmul ops. gemma-4 packs both into one expert weight and
splits the result after the GEMM, so its MoE block stayed unfused and ran
the expert GEMMs as per-token GEMVs.
Add FuseMoeCompressedFusedGateUp, which matches that shape
(one GatherMatmul -> Slice/Slice -> Gelu(ERF) -> Multiply) and folds it into
the same MOECompressed op, using GEMM3_SWIGLU with GEGLU_ERF. The fused
weight, scale and zero point are split into gate/up halves by copying raw
bytes, since a graph Slice would be rewritten to StridedSlice and constant
folded, whose reference evaluator crashes on sub-byte types.
gemma-4 also applies a per-expert output scale to the down projection before
the router weights. MOECompressed takes only one per-expert weight, so that
scale is folded into the routing weights, which is exact.
The op reads the zero point straight off a weight port and needs an integer
Constant there, so the matcher requires one and leaves natively quantized
experts (exact f16 zp) to the unfused path.
gemma-4-26B-A4B on Arc B390, GGML_OPENVINO_REQUANT_KQUANT=q4_asym64_all,
llama-bench -p 512 -n 128 -r 2, against a GGML_OPENVINO_MOE_OP=0 baseline:
pp512 66.16 -> 1608.73 t/s, tg128 25.94 -> 26.46 t/s. Perplexity over 12
chunks is unchanged (1451.3 +/- 177.9 unfused vs 1427.6 +/- 175.1 fused).
No effect without that requant option, on models with separate gate/up
weights, or on CPU. test-backend-ops -b OPENVINO0 is unchanged by this
commit: two MUL_MAT_ID m_v cases fail, the same two on the unmodified base.
* openvino: fix rank-3 axis handling so MoE works under stateful execution
Stateful execution drops the leading size-1 batch dim, so OV tensors are rank
3 while GgmlOvDecoder::get_shape/get_stride still report GGML_MAX_DIMS=4
reversed entries. Several MoE ops derive OV axis indices straight from that
metadata, so they picked the wrong axis. A MoE model with
GGML_OPENVINO_STATEFUL_EXECUTION=1 aborts while building the graph:
Check 'is_axis_valid(axis, r)' failed at src/core/src/validation_util.cpp:336
While validating node 'opset11::TopK ... _ffn_moe_probs ...'
Axis 3 out of the tensor rank range [-3, 2].
Fix idiom throughout: take the axis from the real OV rank, or shift a
metadata-derived axis down by metadata_rank - actual_rank.
argsort.cpp the router top-k axis is 2 on rank 3, not 3. This is the
abort quoted above.
add.cpp the MoE expert-sum bypass collapses the 8-ADD chain into one
ReduceSum on hardcoded axis 2, which on rank 3 reduces n_embd
instead of the expert axis. Now rank-2, with the following
Unsqueeze at rank-3.
get_rows.cpp squeezing a hardcoded {0,1} also strips the batch dim
whenever it is 1, which is every decode step. Squeeze down to
the trailing two dims instead.
mul_mat_id.cpp pick the reshape dims by actual rank, and skip the trailing
Unsqueeze that re-adds the batch dim.
view.cpp the expert-plane slice had the Slice axis, dst_ov_axis, the
ShapeOf+Gather index and the Reshape target all rank-4.
utils.cpp process_view_input_new's "translate_view already resolved
this VIEW, skip re-slicing" shortcut required equal ranks. 4
vs 3 never matched, so every resolved expert plane got
re-sliced. Now compares the common trailing dims. Same axis
shift for the Slice in the view-chain walker.
Stateless is unchanged by construction: every edit is gated on the actual
rank, so axis_shift == 0 reproduces the previous code exactly. Checked on
OV-CPU by diffing greedy output against the unmodified base for dense
gemma-4-E2B, granite-1b-a400m and gemma-4-26B-A4B; all identical.
granite-1b-a400m on OV-CPU aborts with the error above before this change;
after it, it generates and is byte-identical to stateless. Dense gemma-4-E2B
is identical stateless vs stateful both before and after. test-backend-ops
-b OPENVINO0 is unchanged: two pre-existing MUL_MAT_ID m_v cases fail, the
same two on the unmodified base.
gemma-4-26B-A4B is a poor correctness vehicle here. On OV it already drifts
into degenerate repetition a few tokens in, in stateless as much as stateful,
and the two modes diverge somewhere inside that degenerate region instead of
matching token for token. Each mode is self-reproducible across runs.
Known limitation: FuseMoeCompressedFusedGateUp does not match the rank-3
graph, so a MoE model run with GGML_OPENVINO_STATEFUL_EXECUTION=1 loses the
prefill fusion while gaining decode. gemma-4-26B-A4B on Arc B390,
GGML_OPENVINO_REQUANT_KQUANT=q4_asym64_all, llama-bench -p 512 -n 128 -r 2:
unfused (GGML_OPENVINO_MOE_OP=0) pp512 66.16 tg128 25.94
fused, stateless (default) pp512 1608.73 tg128 26.46
fused, stateful pp512 66.18 tg128 29.91
Stateful is opt-in and off by default, and MoE did not run there at all
before this, so nothing that previously worked regresses. Making the pass
match rank 3 is the follow-up.
* OpenVINO Backend: Upgrade graph cache to use node_idx, src_idx, node type
* ggml-openvino : enable more comprehensive conv fusion
* enable conv ops
* Reject kernel size 0 and support IM2COL_3D
* openvino : abort when the GPU remote context cannot be created
init() logged the error and returned, which left the device name a GPU but
remote_context empty. The remote buffer and tensor paths assert only on the
device being a GPU and then dereference that empty optional.
Those paths have no host fallback, and a device that OpenVINO listed should
have a working OpenCL context, so stop instead of continuing. An OpenCL stack
that is broken as a whole is still caught earlier by the device availability
check, which falls back to CPU.
Assisted-by: Claude Opus 5
* openvino : fix build warnings
The single-argument form of the OpenVINO RTTI macros is the intended one, but
their selector macro leaves __VA_ARGS__ empty, which -Wpedantic reports on
every pass and op header. Turn that warning off for this backend only, the
way ggml-cuda and ggml-sycl already do for their own third-party warnings.
Also drop a break and a dead assignment around a GGML_ABORT, which is noreturn.
Assisted-by: Claude Opus 5
* OpenVINO Backend: Support common MTMD ops
* ggml-openvino: give a reshaping view its own ov::Tensor
* ggml-openvino : compute HARDSIGMOID and EXPM1 in f32
HARDSIGMOID used a 1/6 constant in the input type, which is not exact
in bf16, and EXPM1 lost precision for small inputs in f16. Both now
compute in f32 and convert back, except on NPU where the f32 path
gives wrong results.
Fixes the HARDSIGMOID/EXPM1 test-backend-ops failures on GPU.
* ggml-openvino : update device selection and --list-devices
Show the selecting GGML_OPENVINO_DEVICE value and active device in
--list-devices, startup logs, and backend tests.
Clarify OpenVINO selection uses GGML_OPENVINO_DEVICE, not -dev.
* openvino : remove unreachable OpenCL queue checks
A remote buffer exists only on a GPU device, and init() aborts there if the
queue cannot be created, so the queue is never null at these call sites.
Assisted-by: Claude Opus 5
* openvino : update OpenVINO to 2026.4.1 and GPU drivers to 26.35.39758.10
* docs : update OpenVINO validated models and GPU driver version
* ggml-openvino : skip empty views when giving a reshaping view its own tensor
A zero-size view can sit at the end of a GPU USM buffer (Qwen3.5 recurrent cache). Wrapping it as a remote tensor throws "shared USM buffer has smaller size (0)".
Assisted-by: Claude
* ggml-openvino : rebind the cached decoder when llama passes a different graph
llama keeps separate graphs for batches with and without outputs. llama-server splits the prompt into chunks for context checkpoints, so a cached decoder could be reused with a graph built in other memory and bind the previous chunk's input tensors. SWA and recurrent models then lost most of the prompt in llama-cli and llama-server.
Assisted-by: Claude
* docs : update OpenVINO validated models
Smoke test on Lunar Lake (32 GB) with the two fixes above. Re-add the Qwen3.5 and gemma models.
Assisted-by: Claude
---------
Co-authored-by: Yu, Zijun <zijun.yu@intel.com>
Co-authored-by: ynimmaga <ynimmaga@users.noreply.github.com>
Co-authored-by: Łukasz Ślusarczyk <lukasz.slusarczyk@intel.com>
Co-authored-by: haarika-madaka <haarika.madaka@intel.com>
Co-authored-by: Mustafa Cavus <mustafa.cavus@intel.com>
Co-authored-by: Mostafa Faheem <mostafaaafaheem@gmail.com>
diff --git a/.devops/openvino.Dockerfile b/.devops/openvino.Dockerfile
index e301aa8f5..4b5ac734f 100644
--- a/.devops/openvino.Dockerfile
+++ b/.devops/openvino.Dockerfile
@@ -1,12 +1,12 @@
-ARG OPENVINO_VERSION_MAJOR=2026.4
-ARG OPENVINO_VERSION_FULL=2026.4.0.22959.99c81491cc3
+ARG OPENVINO_VERSION_MAJOR=2026.4.1
+ARG OPENVINO_VERSION_FULL=2026.4.1.22982.07f9c262b05
ARG UBUNTU_VERSION=24.04
# Intel GPU driver versions. https://github.com/intel/compute-runtime/releases
-ARG IGC_VERSION=v2.40.13
-ARG IGC_VERSION_FULL=2_2.40.13+22418
-ARG COMPUTE_RUNTIME_VERSION=26.31.39395.13
-ARG COMPUTE_RUNTIME_VERSION_FULL=26.31.39395.13-0
+ARG IGC_VERSION=v2.41.5
+ARG IGC_VERSION_FULL=2_2.41.5+22716
+ARG COMPUTE_RUNTIME_VERSION=26.35.39758.10
+ARG COMPUTE_RUNTIME_VERSION_FULL=26.35.39758.10-0
ARG IGDGMM_VERSION=22.10.0
# Intel NPU driver versions. https://github.com/intel/linux-npu-driver/releases
diff --git a/.github/workflows/build-cache.yml b/.github/workflows/build-cache.yml
index 27512a142..28be179d7 100644
--- a/.github/workflows/build-cache.yml
+++ b/.github/workflows/build-cache.yml
@@ -41,8 +41,8 @@ jobs:
env:
# Sync versions in build-openvino.yml, build-self-hosted.yml, release.yml, build-cache.yml, .devops/openvino.Dockerfile
- OPENVINO_VERSION_MAJOR: "2026.4"
- OPENVINO_VERSION_FULL: "2026.4.0.22959.99c81491cc3"
+ OPENVINO_VERSION_MAJOR: "2026.4.1"
+ OPENVINO_VERSION_FULL: "2026.4.1.22982.07f9c262b05"
steps:
- name: Clone
@@ -69,8 +69,8 @@ jobs:
env:
# Sync versions in build.yml, build-self-hosted.yml, release.yml, build-cache.yml, .devops/openvino.Dockerfile
- OPENVINO_VERSION_MAJOR: "2026.4"
- OPENVINO_VERSION_FULL: "2026.4.0.22959.99c81491cc3"
+ OPENVINO_VERSION_MAJOR: "2026.4.1"
+ OPENVINO_VERSION_FULL: "2026.4.1.22982.07f9c262b05"
steps:
- name: Clone
diff --git a/.github/workflows/build-openvino.yml b/.github/workflows/build-openvino.yml
index daa08b1bf..d325c368d 100644
--- a/.github/workflows/build-openvino.yml
+++ b/.github/workflows/build-openvino.yml
@@ -41,8 +41,8 @@ jobs:
env:
# Sync versions in build-openvino.yml, build-self-hosted.yml, release.yml, build-cache.yml, .devops/openvino.Dockerfile
- OPENVINO_VERSION_MAJOR: "2026.4"
- OPENVINO_VERSION_FULL: "2026.4.0.22959.99c81491cc3"
+ OPENVINO_VERSION_MAJOR: "2026.4.1"
+ OPENVINO_VERSION_FULL: "2026.4.1.22982.07f9c262b05"
steps:
- name: Clone
@@ -96,8 +96,8 @@ jobs:
env:
# Sync versions in build-openvino.yml, build-self-hosted.yml, release.yml, build-cache.yml, .devops/openvino.Dockerfile
- OPENVINO_VERSION_MAJOR: "2026.4"
- OPENVINO_VERSION_FULL: "2026.4.0.22959.99c81491cc3"
+ OPENVINO_VERSION_MAJOR: "2026.4.1"
+ OPENVINO_VERSION_FULL: "2026.4.1.22982.07f9c262b05"
steps:
- name: Clone
diff --git a/.github/workflows/ci-self-hosted-openvino.yml b/.github/workflows/ci-self-hosted-openvino.yml
index e0947c46e..c64c3a1a4 100644
--- a/.github/workflows/ci-self-hosted-openvino.yml
+++ b/.github/workflows/ci-self-hosted-openvino.yml
@@ -47,8 +47,8 @@ jobs:
env:
# Sync versions in build.yml, build-self-hosted.yml, release.yml, build-cache.yml, .devops/openvino.Dockerfile
- OPENVINO_VERSION_MAJOR: "2026.4"
- OPENVINO_VERSION_FULL: "2026.4.0.22959.99c81491cc3"
+ OPENVINO_VERSION_MAJOR: "2026.4.1"
+ OPENVINO_VERSION_FULL: "2026.4.1.22982.07f9c262b05"
steps:
- name: Clone
diff --git a/.github/workflows/release.yml b/.github/workflows/release.yml
index 4784bb718..e88eae3ac 100644
--- a/.github/workflows/release.yml
+++ b/.github/workflows/release.yml
@@ -673,8 +673,8 @@ jobs:
env:
# Sync versions in build-openvino.yml, build-self-hosted.yml, release.yml, build-cache.yml, .devops/openvino.Dockerfile
- OPENVINO_VERSION_MAJOR: "2026.4"
- OPENVINO_VERSION_FULL: "2026.4.0.22959.99c81491cc3"
+ OPENVINO_VERSION_MAJOR: "2026.4.1"
+ OPENVINO_VERSION_FULL: "2026.4.1.22982.07f9c262b05"
steps:
- name: Set OpenVINO version output
@@ -788,8 +788,8 @@ jobs:
env:
# Sync versions in build-openvino.yml, build-self-hosted.yml, release.yml, build-cache.yml, .devops/openvino.Dockerfile
- OPENVINO_VERSION_MAJOR: "2026.4"
- OPENVINO_VERSION_FULL: "2026.4.0.22959.99c81491cc3"
+ OPENVINO_VERSION_MAJOR: "2026.4.1"
+ OPENVINO_VERSION_FULL: "2026.4.1.22982.07f9c262b05"
steps:
- name: Set OpenVINO version output
diff --git a/docs/backend/OPENVINO.md b/docs/backend/OPENVINO.md
index 3d7919775..c9bbc9a9e 100644
--- a/docs/backend/OPENVINO.md
+++ b/docs/backend/OPENVINO.md
@@ -52,8 +52,8 @@ Although OpenVINO supports a wide range of [Intel hardware](https://docs.openvin
- `Q4_1`
- `Q4_K`
- `Q4_K_M`
-- `Q5_K` (converted to `Q8_0_C` at runtime)
-- `Q6_K` (converted to `Q8_0_C` at runtime)
+- `Q5_K` (converted to `Q8_0_C` at runtime by default)
+- `Q6_K` (converted to `Q8_0_C` at runtime by default)
> [!NOTE]
> Accuracy validation and performance optimizations for quantized models are a work in progress.
@@ -93,12 +93,12 @@ Although, the validated models below were tested with `llama-cli` using the `Q4_
> Extensive accuracy validation, performance optimizations, and broader architecture coverage are work in progress.
**Legend & Test Configuration:**
-- **Status:** ✓ = Passed | ✗ = Failed or Unsupported
+- **Status:** ✓ = Passed | ~ = Accuracy issues | ✗ = Failed or Unsupported
- **Execution Modes:**
- **SL** = Stateless (`GGML_OPENVINO_STATEFUL_EXECUTION=0`)
- **SF** = Stateful (`GGML_OPENVINO_STATEFUL_EXECUTION=1`)
- Note: The NPU operates in stateless mode only.
-- **Validation system:** Intel® Core™ Ultra 5 238V (Lunar Lake) | 32 GB RAM | Ubuntu 24.04 | Intel Graphics Compiler 2.41.5 | Intel OpenCL GPU Driver 26.31.39395.13-0 | Intel NPU Driver 1.38.0.
+- **Validation system:** Intel® Core™ Ultra 5 238V (Lunar Lake) | 32 GB RAM | Ubuntu 24.04 | Intel Graphics Compiler 2.41.5 | Intel OpenCL GPU Driver 26.35.39758.10-0 | Intel NPU Driver 1.38.0.
- See [Known Limitations](#known-limitations) for context on observed failures.
| Model | CPU (SL / SF) | GPU (SL / SF) | NPU (SL) |
@@ -113,14 +113,14 @@ Although, the validated models below were tested with `llama-cli` using the `Q4_
| [bartowski/Qwen_Qwen3-1.7B-Q4_K_M](https://huggingface.co/bartowski/Qwen_Qwen3-1.7B-GGUF) | ✓ / ✓ | ✓ / ✓ | ✓ |
| [Qwen/Qwen3-4B-Q4_K_M](https://huggingface.co/Qwen/Qwen3-4B-GGUF) | ✓ / ✓ | ✓ / ✓ | ✓ |
| [lm-kit/Qwen3-8B-Q4_K_M](https://huggingface.co/lm-kit/qwen-3-8b-instruct-gguf) | ✓ / ✓ | ✓ / ✓ | ✓ |
-| [bartowski/Qwen_Qwen3.5-0.8B-Q4_K_M](https://huggingface.co/bartowski/Qwen_Qwen3.5-0.8B-GGUF) | ✓ / ✗ | ✓ / ✗ | ✗ |
-| [bartowski/Qwen_Qwen3.5-2B-Q4_K_M](https://huggingface.co/bartowski/Qwen_Qwen3.5-2B-GGUF) | ✓ / ✗ | ✓ / ✗ | ✗ |
-| [bartowski/Qwen_Qwen3.5-4B-Q4_K_M](https://huggingface.co/bartowski/Qwen_Qwen3.5-4B-GGUF) | ✓ / ✗ | ✓ / ✗ | ✗ |
-| [lmstudio-community/Qwen3.5-9B-Q4_K_M](https://huggingface.co/lmstudio-community/Qwen3.5-9B-GGUF) | ✓ / ✗ | ✓ / ✗ | ✗ |
+| [bartowski/Qwen_Qwen3.5-0.8B-Q4_K_M](https://huggingface.co/bartowski/Qwen_Qwen3.5-0.8B-GGUF) | ✓ / ✓ | ✓ / ~ | ✗ |
+| [bartowski/Qwen_Qwen3.5-2B-Q4_K_M](https://huggingface.co/bartowski/Qwen_Qwen3.5-2B-GGUF) | ✓ / ✓ | ✓ / ~ | ✗ |
+| [bartowski/Qwen_Qwen3.5-4B-Q4_K_M](https://huggingface.co/bartowski/Qwen_Qwen3.5-4B-GGUF) | ✓ / ✓ | ✓ / ~ | ✗ |
+| [lmstudio-community/Qwen3.5-9B-Q4_K_M](https://huggingface.co/lmstudio-community/Qwen3.5-9B-GGUF) | ✓ / ✓ | ✓ / ~ | ✗ |
| | | | |
| [unsloth/gemma-3-4b-it-Q4_K_M](https://huggingface.co/unsloth/gemma-3-4b-it-GGUF) | ✓ / ✓ | ✓ / ✓ | ✓ |
-| [bartowski/google_gemma-4-E2B-it-Q4_K_M](https://huggingface.co/bartowski/google_gemma-4-E2B-it-GGUF) | ✓ / ✓ | ✓ / ✓ | ✗ |
-| [bartowski/google_gemma-4-E4B-it-Q4_K_M](https://huggingface.co/bartowski/google_gemma-4-E4B-it-GGUF) | ✓ / ✓ | ✓ / ✓ | ✓ |
+| [bartowski/google_gemma-4-E2B-it-Q4_K_M](https://huggingface.co/bartowski/google_gemma-4-E2B-it-GGUF) | ✓ / ✓ | ✓ / ~ | ~ |
+| [bartowski/google_gemma-4-E4B-it-Q4_K_M](https://huggingface.co/bartowski/google_gemma-4-E4B-it-GGUF) | ✓ / ✓ | ✗ / ✗ | ✓ |
| [bartowski/gemma-4-12B-it-Q4_K_M](https://huggingface.co/bartowski/gemma-4-12B-it-GGUF) | ✓ / ✓ | ✓ / ✓ | ✗ |
| | | | |
| [bartowski/Phi-3-mini-4k-instruct-Q4_K_M](https://huggingface.co/bartowski/Phi-3-mini-4k-instruct-GGUF) | ✓ / ✓ | ✓ / ✓ | ✓ |
@@ -134,9 +134,9 @@ Although, the validated models below were tested with `llama-cli` using the `Q4_
| [bartowski/DeepSeek-R1-Distill-Llama-8B-Q4_K_M](https://huggingface.co/bartowski/DeepSeek-R1-Distill-Llama-8B-GGUF) | ✓ / ✓ | ✓ / ✓ | ✓ |
| [bartowski/DeepSeek-R1-Distill-Qwen-7B-Q4_K_M](https://huggingface.co/bartowski/DeepSeek-R1-Distill-Qwen-7B-GGUF) | ✓ / ✓ | ✓ / ✓ | ✓ |
| | | | |
-| [ibm-granite/granite-4.0-350m-Q4_K_M](https://huggingface.co/ibm-granite/granite-4.0-350m-GGUF) | ✓ / ✓ | ✓ / ✓ | ✓ |
+| [ibm-granite/granite-4.0-350m-Q4_K_M](https://huggingface.co/ibm-granite/granite-4.0-350m-GGUF) | ✓ / ✓ | ~ / ~ | ✓ |
| [ibm-granite/granite-4.0-micro-Q4_K_M](https://huggingface.co/ibm-granite/granite-4.0-micro-GGUF) | ✓ / ✓ | ✓ / ✓ | ✓ |
-| [ibm-granite/granite-4.0-1b-Q4_K_M](https://huggingface.co/ibm-granite/granite-4.0-1b-GGUF) | ✓ / ✓ | ✓ / ✓ | ✗ |
+| [ibm-granite/granite-4.0-1b-Q4_K_M](https://huggingface.co/ibm-granite/granite-4.0-1b-GGUF) | ✓ / ✓ | ~ / ~ | ~ |
| [ibm-research/granite-3.2-8b-instruct-Q4_K_M](https://huggingface.co/ibm-research/granite-3.2-8b-instruct-GGUF) | ✓ / ✓ | ✓ / ✓ | ✓ |
| | | | |
| [HuggingFaceTB/smollm2-1.7b-instruct-q4_k_m](https://huggingface.co/HuggingFaceTB/SmolLM2-1.7B-Instruct-GGUF) | ✓ / ✓ | ✓ / ✓ | ✓ |
@@ -244,8 +244,8 @@ chmod +x build-llamacpp-ov.sh
# ============================================
set -euo pipefail
-OPENVINO_VERSION_MAJOR="2026.4"
-OPENVINO_VERSION_FULL="2026.4.0.22959.99c81491cc3"
+OPENVINO_VERSION_MAJOR="2026.4.1"
+OPENVINO_VERSION_FULL="2026.4.1.22982.07f9c262b05"
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
OPENVINO_INSTALL_DIR="/opt/intel/openvino_${OPENVINO_VERSION_MAJOR}"
@@ -342,7 +342,7 @@ echo " ./build/ReleaseOV/bin/llama-cli -m model.gguf"
```
> [!NOTE]
-> The script pins OpenVINO `2026.4` via the `OPENVINO_VERSION_MAJOR` / `OPENVINO_VERSION_FULL` variables at the top — edit them to track a different release.
+> The script pins OpenVINO `2026.4.1` via the `OPENVINO_VERSION_MAJOR` / `OPENVINO_VERSION_FULL` variables at the top — edit them to track a different release.
</details>
@@ -372,8 +372,8 @@ REM ============================================
REM llama.cpp OpenVINO Build Script (Ninja)
REM ============================================
-set "OPENVINO_VERSION_MAJOR=2026.4"
-set "OPENVINO_VERSION_FULL=2026.4.0.22959.99c81491cc3"
+set "OPENVINO_VERSION_MAJOR=2026.4.1"
+set "OPENVINO_VERSION_FULL=2026.4.1.22982.07f9c262b05"
set "SCRIPT_DIR=%~dp0"
set "VCPKG_DIR=C:\vcpkg"
@@ -552,7 +552,7 @@ endlocal
```
> [!NOTE]
-> The script pins OpenVINO `2026.4` via the `OPENVINO_VERSION_MAJOR` / `OPENVINO_VERSION_FULL` variables at the top — edit them to track a different release. From any new shell, source the matching `setupvars` script via the junction — `call "C:\Intel\openvino\setupvars.bat"` from `cmd`, or `& "C:\Intel\openvino\setupvars.ps1"` from PowerShell. If `winget` cannot register Visual Studio Build Tools on first run, install them once manually and re-run the script from an elevated **Developer Command Prompt for VS 2022**.
+> The script pins OpenVINO `2026.4.1` via the `OPENVINO_VERSION_MAJOR` / `OPENVINO_VERSION_FULL` variables at the top — edit them to track a different release. From any new shell, source the matching `setupvars` script via the junction — `call "C:\Intel\openvino\setupvars.bat"` from `cmd`, or `& "C:\Intel\openvino\setupvars.ps1"` from PowerShell. If `winget` cannot register Visual Studio Build Tools on first run, install them once manually and re-run the script from an elevated **Developer Command Prompt for VS 2022**.
</details>
@@ -625,7 +625,7 @@ $env:GGML_OPENVINO_DEVICE = "NPU"
build\ReleaseOV\bin\llama-cli.exe -m "C:\models\Llama-3.2-1B-Instruct-Q4_K_M.gguf" -c 512
```
> [!NOTE]
-> On systems with multiple GPUs, use `GPU.0` or `GPU.1` to explicitly target specific GPU. See [OpenVINO GPU Device](https://docs.openvino.ai/2026/openvino-workflow/running-inference/inference-devices-and-modes/gpu-device.html) for more details.
+> On systems with multiple GPUs, use `GPU.0` or `GPU.1` to explicitly target specific GPU. A device that is not available is an error (no fallback to CPU), and the error message lists the available OpenVINO devices with their names. Run `llama-cli --list-devices` to see the valid values: each OpenVINO device shows the `GGML_OPENVINO_DEVICE=<value>` to set, and `(selected)` marks the active one. Select the OpenVINO device with this variable, not with `-dev`. See [OpenVINO GPU Device](https://docs.openvino.ai/2026/openvino-workflow/running-inference/inference-devices-and-modes/gpu-device.html) for more details.
### 5. Docker Build
@@ -713,12 +713,13 @@ Boolean flags follow a uniform convention: set to a **positive integer** (e.g. `
| Variable | Type | Default | Description |
|-----------------------------------|-----------|------------|-------------------------------------------------------------------------------------------------------------|
-| `GGML_OPENVINO_DEVICE` | String | `CPU` | Specify the target device (CPU, GPU, NPU). On systems with multiple GPUs, use `GPU.0` or `GPU.1` to explicitly target specific GPU. See [OpenVINO GPU Device](https://docs.openvino.ai/2026/openvino-workflow/running-inference/inference-devices-and-modes/gpu-device.html). When set to **NPU**, static compilation mode is enabled for optimal performance. |
-| `GGML_OPENVINO_CACHE_DIR` | String | `not set` | Directory for OpenVINO model caching (recommended: `/tmp/ov_cache`). Enables model caching when set. **Not supported on NPU devices.** |
-| `GGML_OPENVINO_COMPILED_MODEL_CACHE_DIR` | String | `not set` | Directory for the frontend compiled-model cache. When set, OpenVINO compiled models are exported as blobs and imported on later runs to skip weight requantization, graph conversion, and compilation for matching single-graph models. |
+| `GGML_OPENVINO_DEVICE` | String | `CPU` | Specify the target device (CPU, GPU, NPU). On systems with multiple GPUs, use `GPU.0` or `GPU.1` to explicitly target specific GPU. A device that is not available is an error (no fallback to CPU), and the error message lists the available OpenVINO devices with their names. See [OpenVINO GPU Device](https://docs.openvino.ai/2026/openvino-workflow/running-inference/inference-devices-and-modes/gpu-device.html). When set to **NPU**, static compilation mode is enabled for optimal performance. |
+| `GGML_OPENVINO_CACHE_DIR` | String | `not set` | Directory for OpenVINO's separate plugin cache. On NPU, this sets `NPUW_CACHE_DIR`. |
+| `GGML_OPENVINO_COMPILED_MODEL_CACHE_DIR` | String | `not set` | Directory for standalone compiled blobs with weights. Dynamic CPU/GPU graphs can import matching blobs on later runs. |
+| `GGML_OPENVINO_COMPILED_MODEL_CACHE_ONLY` | Boolean | `0` | Require an existing compiled blob and skip weight uploads and compilation. Requires Linux or Windows mmap loading and a full dynamic CPU/GPU graph on OpenVINO. |
| `GGML_OPENVINO_PREFILL_CHUNK_SIZE`| Integer | `256` | Token chunk size for **NPU** prefill (NPU-only; ignored on CPU/GPU). Must be a positive integer; otherwise the default is used. |
| `GGML_OPENVINO_NPU_COMPILE_CONFIG` | String | `not set` | NPU-only compiler mode parameters forwarded to OpenVINO as `NPU_COMPILATION_MODE_PARAMS`, for example `optimization-level=3`. |
-| `GGML_OPENVINO_STATEFUL_EXECUTION`| Boolean | `0` | Enable stateful KV cache for better performance. Recommended on CPU, GPU. |
+| `GGML_OPENVINO_STATEFUL_EXECUTION`| Boolean | `0` | Keep KV and supported recurrent caches inside the model. Single-slot CPU/GPU execution only. |
| `GGML_OPENVINO_DISABLE_CACHE` | Boolean | `0` | Disable the in-process compiled-model / decoder cache (cache is on by default). Set to `1` to disable. |
| `GGML_OPENVINO_DISABLE_KV_SLICE` | Boolean | `0` | Disable the KV-cache input-tensor slicing optimization (slicing is on by default on CPU/GPU). Set to `1` to disable. |
| `GGML_OPENVINO_DISABLE_KV_STATE_RELAYOUT` | Boolean | `0` | Disable the stateful KV-state sequence-axis relayout (relayout is on by default). It moves the KV state sequence axis from dim 1 to dim 2, so the GPU plugin can append new tokens in place instead of copying the whole state every token, and the reader side no longer transposes the whole accumulated state. Set to `1` to disable. |
@@ -727,8 +728,10 @@ Boolean flags follow a uniform convention: set to a **positive integer** (e.g. `
| `GGML_OPENVINO_REDUCE_COMPILE_MEM`| Boolean | inherits from `GGML_OPENVINO_MEMORY_OPTIMIZE` | Reduce compile-time host memory use by streaming weight requantization and avoiding extra weight-node materialization where possible. Set explicitly to override the umbrella switch. |
| `GGML_OPENVINO_RELEASE_WEIGHTS` | Boolean | inherits from `GGML_OPENVINO_MEMORY_OPTIMIZE` on GPU | GPU-only. Release host weight buffers after the compiled model cache can reuse the device/plugin copy. Requires stable graph shapes; dynamic workloads that need recompilation should leave this disabled. |
| `GGML_OPENVINO_SPILL_DIR` | String | `not set` | Directory for a disk-backed weight buffer. When set, the repacked weight buffer is mapped from an unlinked file on this path instead of anonymous memory, so its pages are reclaimable under memory pressure instead of staying pinned, cutting the load-time host memory peak. Must point at real storage; a tmpfs mount (e.g. `/tmp` on many systems) backs it with RAM and makes the peak worse. |
-| `GGML_OPENVINO_REQUANT_KQUANT` | String | `not set` | Requantize Q6_K/Q5_K weights (and matching MoE expert weights) to a 4-bit target instead of the default Q8_0_C, trading accuracy for less memory traffic. One of `q4_sym128` (Q6_K/Q5_K only), `q4_sym128_all` (Q4_K too, drops its per-group zero point), `q4_asym64_all` (Q6_K/Q5_K/Q4_K, keeps a real zero point at group 64), or `native` (no requantization). |
-| `GGML_OPENVINO_PROFILING` | Boolean | `0` | Enable execution-time profiling. |
+| `GGML_OPENVINO_REQUANT_KQUANT` | String | `not set` | Requantize Q6_K/Q5_K weights (and matching MoE expert weights) to a 4-bit target instead of the default Q8_0_C, trading accuracy for less memory traffic. One of `q4_asym64` (Q6_K/Q5_K only, keeps a real zero point at group 64), `q4_asym64_all` (also requantizes Q4_K), `q4_sym128` (Q6_K/Q5_K only), `q4_sym128_all` (Q4_K too, drops its per-group zero point), or `native` (no requantization). |
+| `GGML_OPENVINO_PROFILING` | Integer | `0` | `1` logs execution timing; `2` or higher also enables OpenVINO and OpenCL profiling. |
+| `GGML_OPENVINO_DEBUG_NODE` | String | `not set` | Add the named graph nodes as compiled outputs for debugging. Separate multiple names with commas. |
+| `GGML_OPENVINO_MOE_OP` | Boolean | `1` | On GPU, set to `0` to keep the unfused GatherMatmul path. |
| `GGML_OPENVINO_DUMP_CGRAPH` | Boolean | `0` | Dump the GGML compute graph to `cgraph_ov.txt`. |
| `GGML_OPENVINO_DUMP_IR` | Boolean | `0` | Serialize OpenVINO IR files with timestamps. |
| `GGML_OPENVINO_DEBUG_INPUT` | Boolean | `0` | Enable input debugging and print input tensor info. |
@@ -737,8 +740,9 @@ Boolean flags follow a uniform convention: set to a **positive integer** (e.g. `
| `GGML_OPENVINO_LOG_UNSUPPORTED_OPS`| Boolean | `0` | Log warning messages with tensor details and rejection reasons for any ops not supported by the OpenVINO backend. Emits at `WARN` level (requires `--log-verbosity >= 2`, enabled by default). |
> [!NOTE]
-> - `GGML_OPENVINO_STATEFUL_EXECUTION` is an **Experimental** feature to allow stateful execution for managing the KV cache internally inside the OpenVINO model, improving performance on CPUs and GPUs. Stateful execution is not effective on NPUs, and not all models currently support this feature. This feature is experimental and has been validated only with the llama-simple, llama-cli, llama-bench, and llama-run applications and is recommended to enable for the best performance. Other applications, such as llama-server and llama-perplexity, are not yet supported.
+> - `GGML_OPENVINO_STATEFUL_EXECUTION` is an **Experimental** feature for managing caches internally inside the OpenVINO model on CPUs and GPUs. Use a single slot (`-np 1`). KV caches retain the append-based state layout and sequence-axis optimization. Qwen3.5 adds recurrent cache states in their GGML layouts. Qwen3.5 requires an unsplit graph with model caching enabled and no recurrent rollback. A prompt starting at position 0 resets all states. State save/restore, sequence rewind, context shift, and mid-sequence graph replacement are unsupported. Stateful execution is not effective on NPUs.
> - `GGML_OPENVINO_LOG_UNSUPPORTED_OPS` emits logs at `WARN` level (`GGML_LOG_WARN`), which requires application log verbosity `--log-verbosity >= 2` (or `-lv 2`).
+> - With `GGML_OPENVINO_COMPILED_MODEL_CACHE_ONLY=1`, use the same compilation settings as the export run. One directory can hold blobs for different models and settings; `GGML_OPENVINO_SPILL_DIR` does not affect the cache key and is ignored in cache-only mode. See [Compiled model cache](../../ggml/src/ggml-openvino/README.md) for the workflow and restrictions.
### Example Usage
diff --git a/ggml/src/ggml-openvino/CMakeLists.txt b/ggml/src/ggml-openvino/CMakeLists.txt
index af3e0758c..347c24f4f 100644
--- a/ggml/src/ggml-openvino/CMakeLists.txt
+++ b/ggml/src/ggml-openvino/CMakeLists.txt
@@ -13,6 +13,17 @@ ggml_add_backend_library(ggml-openvino
target_link_libraries(ggml-openvino PRIVATE openvino::runtime openvino::threading OpenCL::OpenCL)
+# the OpenVINO RTTI macros take one argument and leave __VA_ARGS__ empty, which -Wpedantic reports
+if (CMAKE_CXX_COMPILER_ID MATCHES "Clang" OR CMAKE_CXX_COMPILER_ID STREQUAL "IntelLLVM")
+ target_compile_options(ggml-openvino PRIVATE -Wno-gnu-zero-variadic-macro-arguments)
+elseif (CMAKE_CXX_COMPILER_ID STREQUAL "GNU")
+ target_compile_options(ggml-openvino PRIVATE -Wno-pedantic)
+endif()
+
+if (WIN32)
+ target_link_libraries(ggml-openvino PRIVATE psapi)
+endif()
+
if (GGML_OPENVINO)
if (CMAKE_SYSTEM_PROCESSOR STREQUAL "aarch64")
elseif (CMAKE_SYSTEM_PROCESSOR STREQUAL "x86_64" OR CMAKE_SYSTEM_PROCESSOR STREQUAL "amd64" OR CMAKE_SYSTEM_PROCESSOR STREQUAL "AMD64")
diff --git a/ggml/src/ggml-openvino/README.md b/ggml/src/ggml-openvino/README.md
new file mode 100644
index 000000000..6a23f62bf
--- /dev/null
+++ b/ggml/src/ggml-openvino/README.md
@@ -0,0 +1,44 @@
+# Compiled model cache
+
+`GGML_OPENVINO_COMPILED_MODEL_CACHE_DIR` exports compiled CPU/GPU graphs with their weights. It bypasses the plugin-level `GGML_OPENVINO_CACHE_DIR` and uses `OPTIMIZE_SPEED`, so weightless caching is disabled.
+
+One directory can hold blobs for different models and compilation settings. Run each intended workload once to export its dynamic graph:
+
+```sh
+GGML_OPENVINO_DEVICE=GPU \
+GGML_OPENVINO_NATIVE_SOFTPLUS=1 \
+GGML_OPENVINO_DISABLE_KV_SLICE=1 \
+GGML_OPENVINO_REQUANT_KQUANT=q4_asym64_all \
+GGML_OPENVINO_COMPILED_MODEL_CACHE_DIR=/path/to/qwen-cache \
+./build/ReleaseOV/bin/llama-bench -m /path/to/model.gguf -r 1
+```
+
+`GGML_OPENVINO_SPILL_DIR` remains optional for this first run. Wait for the `model cache WROTE` message and completion of the workload before stopping it. Compatible prefill and decode graphs share one blob and manifest. A graph with different ports or incompatible shapes gets an exact entry instead; interrupted exports are not cache hits.
+
+On later runs, supply the same compilation settings and enable `GGML_OPENVINO_COMPILED_MODEL_CACHE_ONLY=1`:
+
+```sh
+GGML_OPENVINO_DEVICE=GPU \
+GGML_OPENVINO_NATIVE_SOFTPLUS=1 \
+GGML_OPENVINO_DISABLE_KV_SLICE=1 \
+GGML_OPENVINO_REQUANT_KQUANT=q4_asym64_all \
+GGML_OPENVINO_COMPILED_MODEL_CACHE_DIR=/path/to/qwen-cache \
+GGML_OPENVINO_COMPILED_MODEL_CACHE_ONLY=1 \
+./build/ReleaseOV/bin/llama-bench -m /path/to/model.gguf -r 1
+```
+
+Cache-only mode allocates backend address space without filling weight pages. On Windows, this also uses system commit capacity. The model-buffer size in the loader log is this virtual size. Weight uploads only record source identity; they do not read or requantize the weights. Graph conversion and compilation are skipped. Runtime buffers are still allocated and populated normally.
+
+Cache-only mode uses the settings provided by the current process. Keep these values exactly the same, including set versus unset: `GGML_OPENVINO_REQUANT_KQUANT`, `GGML_OPENVINO_NATIVE_SOFTPLUS`, `GGML_OPENVINO_DISABLE_KV_SLICE`, `GGML_OPENVINO_MANUAL_GQA_ATTN`, `GGML_OPENVINO_STATEFUL_EXECUTION`, `GGML_OPENVINO_DISABLE_KV_STATE_RELAYOUT`, `GGML_OPENVINO_DISABLE_REMOTE_OUTPUTS`, `GGML_OPENVINO_REDUCE_COMPILE_MEM`, `GGML_OPENVINO_MEMORY_OPTIMIZE`, `GGML_OPENVINO_PROFILING`, and `GGML_OPENVINO_DEBUG_NODE`. On GPU, also repeat `GGML_OPENVINO_MOE_OP=0` if used. `GGML_OPENVINO_SPILL_DIR` is optional on the first run and ignored in cache-only mode; host-weight release is disabled in cache-only mode.
+
+A missing or incompatible graph fails with an error instead of compiling with absent weights. The fingerprint uses the dynamic graph's topology, ports, model parameters, weights, settings, and OpenVINO version; changing only dynamic token or KV sizes does not require a new entry. A different workload can still require another graph; populate it first without cache-only mode.
+
+## Restrictions
+
+- Cache-only mode requires Linux or Windows and mmap loading (`--load-mode mmap`, or the default when all selected devices support mmap). Do not use tensor validation or mlock when trying to avoid weight reads.
+- On Windows, the backend commits virtual memory for its buffers without touching weight pages. Large models can still reach the system commit limit.
+- The model must execute entirely on OpenVINO, with dynamic CPU/GPU graphs and in-process caching enabled. Static/NPU execution and CPU fallback are unsupported in cache-only mode.
+- GGUF metadata, tokenizer data, tensor descriptors, and graph construction are still needed. The llama.cpp loader is unchanged: depending on its prefetch settings, it may request pages with `MAP_POPULATE` or read-ahead on Linux, or `PrefetchVirtualMemory` on Windows. Non-mmap loading also reads the payload before the backend sees it.
+- File identity, size, modification/change timestamps, tensor offsets, graph structure, settings, and OpenVINO version identify cache entries. Linux uses device/inode and Windows uses volume serial/file index. Replacing, copying, or modifying a GGUF invalidates its entries. This avoids reading weight bytes and ties the cache to the local source files. Keep those files unchanged throughout loading and inference.
+- Use the same target device and compatible OpenVINO/plugin installation. Import support depends on the plugin; the tested CPU plugin cannot import MoE graphs containing `GatherMatmulCompressed`. GPU MoE and CPU dense graph imports were tested.
+- Blobs contain weights and can approach model size for each compiled graph. Import still reads those blobs and initializes the device.
diff --git a/ggml/src/ggml-openvino/ggml-decoder.cpp b/ggml/src/ggml-openvino/ggml-decoder.cpp
index cec32f6df..ae36a0959 100644
--- a/ggml/src/ggml-openvino/ggml-decoder.cpp
+++ b/ggml/src/ggml-openvino/ggml-decoder.cpp
@@ -87,6 +87,14 @@ void GgmlOvDecoder::update_io(ggml_cgraph * cgraph) {
compute_model_outputs();
}
+// llama keeps separate graphs for batches with and without outputs, so a cache hit can come from a
+// graph built in other memory. The decoder then still points at the old graph's tensors.
+bool GgmlOvDecoder::is_bound_to(const ggml_cgraph * cgraph) const {
+ return m_cgraph == cgraph && cgraph->n_nodes > 0 && m_node_info_list.size() == (size_t) cgraph->n_nodes &&
+ m_node_info_list.front().node == cgraph->nodes[0] &&
+ m_node_info_list.back().node == cgraph->nodes[cgraph->n_nodes - 1];
+}
+
GgmlOvDecoder::GgmlOvDecoder(ggml_cgraph * cgraph, std::map<std::string, std::shared_ptr<ov::Node>> & model_weights) {
m_cgraph = cgraph;
m_model_weights = model_weights;
@@ -117,6 +125,12 @@ bool is_same_shape(const ggml_tensor * a, const ggml_tensor * b) {
bool is_conv_states_all_tensor(const ggml_tensor * tensor) {
return tensor != nullptr && strncmp(tensor->name, "conv_states_all", strlen("conv_states_all")) == 0;
}
+
+bool is_full_single_slot_writeback(const ggml_tensor * node) {
+ return node->view_src != nullptr && node->view_src->ne[1] == 1 && node->src[1] != nullptr &&
+ node->src[1]->op == GGML_OP_VIEW && node->src[1]->view_src == node->view_src &&
+ node->src[1]->view_offs == 0 && ggml_nbytes(node->src[1]) == ggml_nbytes(node->view_src);
+}
} // namespace
// MoE expert aggregation (build_moe_ffn in llama-graph.cpp): each expert plane is
@@ -274,8 +288,34 @@ int GgmlOvDecoder::compute_op_case(const ggml_tensor * node) const {
int op_case = 0;
switch (node->op) {
case GGML_OP_RESHAPE: {
+ if (m_naive) {
+ break;
+ }
auto name = std::string(node->name);
auto * src = node->src[0];
+ // Identify recurrent sequence reshapes before size checks, which are ambiguous for one token.
+ bool recurrent_sequence = false;
+ for (int i = 0; i < m_cgraph->n_nodes && !recurrent_sequence; ++i) {
+ const auto * consumer = m_cgraph->nodes[i];
+ if (consumer->op == GGML_OP_MUL_MAT_ID && consumer->src[1] == node) {
+ return 1;
+ } else if (consumer->op == GGML_OP_SSM_CONV) {
+ const auto * concat = consumer->src[0];
+ if (concat->op == GGML_OP_CONCAT) {
+ const auto * transposed = concat->src[1];
+ recurrent_sequence = transposed->op == GGML_OP_TRANSPOSE && transposed->src[0] == node;
+ }
+ } else if (consumer->op == GGML_OP_UNARY && ggml_get_unary_op(consumer) == GGML_UNARY_OP_SOFTPLUS) {
+ const auto * biased = consumer->src[0];
+ recurrent_sequence = biased->op == GGML_OP_ADD && biased->src[0] == node;
+ }
+ }
+ if (recurrent_sequence && node->ne[0] == src->ne[0] && node->ne[3] == 1) {
+ return 6;
+ }
+ if (node->ne[0] == src->ne[0] && node->ne[2] == 1 && node->ne[3] == 1) {
+ return 5;
+ }
if (src->op == GGML_OP_RESHAPE && src->src[0]->ne[0] == node->ne[0] && src->src[0]->ne[1] == node->ne[1]) {
op_case = 4;
} else if (node->ne[0] * node->ne[1] == src->ne[0]) {
@@ -285,7 +325,7 @@ int GgmlOvDecoder::compute_op_case(const ggml_tensor * node) const {
if (src->ne[2] * src->ne[3] == node->ne[1]) {
op_case = 5;
}
- } else if (src->ne[0] * src->ne[1] * src->ne[2] == node->ne[1]) {
+ } else if (node->ne[0] == 1 && src->ne[0] * src->ne[1] * src->ne[2] == node->ne[1]) {
op_case = 3;
} else if (name.find("linear_attn_qkv_mixed") == 0 || name.find("alpha") == 0) {
op_case = 6;
@@ -294,6 +334,38 @@ int GgmlOvDecoder::compute_op_case(const ggml_tensor * node) const {
} else if (name.find("state_predelta") == 0) {
op_case = 8;
}
+ if (op_case == 1 && m_is_stateful) {
+ // Recurrent convolution and GDN gates retain their rank-4 layout.
+ bool recurrent = src->op == GGML_OP_GET_ROWS && is_recurrent_cache(src->src[0]);
+ for (int i = 0; i < m_cgraph->n_nodes && !recurrent; ++i) {
+ const auto * consumer = m_cgraph->nodes[i];
+ if (consumer->op == GGML_OP_GATED_DELTA_NET) {
+ for (int j : {3, 4}) {
+ const auto * gate = consumer->src[j];
+ if (gate->op == GGML_OP_UNARY) {
+ gate = gate->src[0];
+ }
+ recurrent = recurrent || gate == node;
+ }
+ } else if (consumer->op == GGML_OP_MUL) {
+ for (int j = 0; j < 2; ++j) {
+ const auto * gate = consumer->src[j];
+ const auto * norm = consumer->src[1 - j];
+ if (gate->op != GGML_OP_UNARY || gate->src[0] != node) {
+ continue;
+ }
+ if (norm->op == GGML_OP_MUL) {
+ norm = norm->src[0];
+ }
+ recurrent = recurrent || (norm->op == GGML_OP_RMS_NORM && norm->src[0]->op == GGML_OP_VIEW &&
+ norm->src[0]->src[0]->op == GGML_OP_GATED_DELTA_NET);
+ }
+ }
+ }
+ if (recurrent) {
+ op_case = 9;
+ }
+ }
break;
}
case GGML_OP_PERMUTE: {
@@ -342,11 +414,12 @@ int GgmlOvDecoder::compute_op_case(const ggml_tensor * node) const {
if (node->src[1]->op == GGML_OP_VIEW) {
// GET_ROWS gathering recurrent state cache rows via the inp->s_copy index list:
// src[0] is a reshape of cache_r/cache_s, src[1] is a view of the s_copy leaf.
- // op_case 3: main view (active sequences, view offset 0)
- // op_case 4: extra view (defrag remainder, nonzero view offset)
+ // op_case 1/2: active/extra rows of a multi-slot cache
+ // op_case 3/4: active/extra rows of a single-slot cache
if (node->src[0]->op == GGML_OP_RESHAPE && node->src[0]->src[0] != nullptr &&
- is_kvcache(node->src[0]->src[0], nullptr)) {
- op_case = node->src[1]->view_offs == 0 ? 1 : 2;
+ is_recurrent_cache(node->src[0]->src[0])) {
+ const bool single_slot = node->src[0]->src[0]->ne[1] == 1;
+ op_case = (node->src[1]->view_offs == 0 ? 1 : 2) + (single_slot ? 2 : 0);
}
}
break;
@@ -362,6 +435,14 @@ int GgmlOvDecoder::compute_op_case(const ggml_tensor * node) const {
op_case = 2;
break;
}
+ case GGML_ROPE_TYPE_VISION: {
+ op_case = 3;
+ break;
+ }
+ case GGML_ROPE_TYPE_MROPE: {
+ op_case = 4;
+ break;
+ }
default:
op_case = 0;
break;
@@ -369,6 +450,12 @@ int GgmlOvDecoder::compute_op_case(const ggml_tensor * node) const {
break;
}
case GGML_OP_VIEW: {
+ if (!m_model_params.has_rs_rollback && node->src[0] != nullptr &&
+ node->src[0]->op == GGML_OP_GATED_DELTA_NET) {
+ // The GDN translator publishes native attention/state outputs under these VIEW names.
+ op_case = 2;
+ break;
+ }
if (m_is_static && node->src[0] != nullptr &&
(node->src[0]->op == GGML_OP_GATED_DELTA_NET || node->src[0]->op == GGML_OP_CONCAT)) {
// VIEW slicing a GATED_DELTA_NET combined [attn|state] output, or the conv_input
@@ -426,6 +513,10 @@ int GgmlOvDecoder::compute_op_case(const ggml_tensor * node) const {
if (node->src[0]->op == GGML_OP_VIEW) {
if (is_same_shape(node->src[0]->src[0], node->src[0])) {
op_case = 1;
+ } else if (!m_model_params.has_rs_rollback &&
+ node->src[0]->src[0]->op == GGML_OP_GATED_DELTA_NET) {
+ // GDN attention is routed directly to this VIEW by get_output_names().
+ op_case = 3;
} else if (node->src[0]->src[0]->op == GGML_OP_GATED_DELTA_NET) {
op_case = 2;
}
@@ -449,12 +540,40 @@ int GgmlOvDecoder::compute_op_case(const ggml_tensor * node) const {
}
break;
}
+ case GGML_OP_UPSCALE: {
+ const int32_t mode_flags = node->op_params[0];
+ const ggml_scale_mode scale_mode = static_cast<ggml_scale_mode>(mode_flags & 0xFF);
+ switch (scale_mode) {
+ case GGML_SCALE_MODE_NEAREST: {
+ op_case = 1;
+ break;
+ }
+ case GGML_SCALE_MODE_BILINEAR: {
+ op_case = 2;
+ break;
+ }
+ case GGML_SCALE_MODE_BICUBIC: {
+ op_case = 3;
+ break;
+ }
+ default:
+ op_case = 0;
+ break;
+ }
+ break;
+ }
case GGML_OP_CPY: {
if (node->src[0]->op == GGML_OP_VIEW) {
if (node->src[0]->src[0]->op == GGML_OP_GATED_DELTA_NET) {
- op_case = 1;
+ if (!m_model_params.has_rs_rollback) {
+ // op_case 7 replaces a single-slot cache; op_case 10 writes native GDN state
+ // into an active range of a larger non-rollback cache.
+ op_case = is_full_single_slot_writeback(node) ? 7 : 10;
+ } else {
+ op_case = 1;
+ }
} else if (GgmlOvDecoder::is_conv_state_writeback(node)) {
- op_case = 2;
+ op_case = is_full_single_slot_writeback(node) ? 8 : 2;
break;
} else if (is_conv_states_all_tensor(node->view_src) && node->src[1] != nullptr &&
node->src[1]->op == GGML_OP_VIEW && node->src[1]->view_src == node->view_src) {
@@ -463,9 +582,9 @@ int GgmlOvDecoder::compute_op_case(const ggml_tensor * node) const {
}
} else if (node->src[0]->op == GGML_OP_GET_ROWS && node->src[1] != nullptr &&
node->src[1]->op == GGML_OP_VIEW && node->src[1]->view_src != nullptr &&
- is_kvcache(node->src[1]->view_src, nullptr)) {
+ is_recurrent_cache(node->src[1]->view_src)) {
// s_copy defrag remainder writeback: gathered extra state rows copied back into the cache
- op_case = 3;
+ op_case = node->src[1]->view_src->ne[1] == 1 ? 9 : 3;
} else if (node->src[1] != nullptr && node->src[1]->op == GGML_OP_VIEW && node->src[1]->view_src != nullptr) {
// op_case 5: KV write for decoder self-attention (dynamic write offset)
// op_case 6: KV write for encoder self-attn or cross-attn (static offset)
@@ -504,7 +623,7 @@ int GgmlOvDecoder::compute_op_case(const ggml_tensor * node) const {
}
case GGML_OP_SCALE: {
if (node->view_src && node->buffer->usage == GGML_BACKEND_BUFFER_USAGE_ANY) {
- op_case = 1;
+ op_case = node->view_src->ne[1] == 1 ? 2 : 1;
}
break;
}
@@ -858,35 +977,48 @@ std::pair<ModelParams, ComputeParams> GgmlOvDecoder::compute_llm_params(ggml_cgr
if (node->op == GGML_OP_GATED_DELTA_NET) {
model_params.state_size = node->src[0]->ne[0];
}
- if (node->op == GGML_OP_SCALE && node->view_src != nullptr && is_kvcache(node->view_src, nullptr)) {
+ if (node->op == GGML_OP_SCALE && node->view_src != nullptr && is_recurrent_cache(node->view_src)) {
+ if (model_params.n_rs_slots == -1) {
+ model_params.n_rs_slots = node->view_src->ne[1];
+ } else {
+ GGML_ASSERT(model_params.n_rs_slots == node->view_src->ne[1]);
+ }
compute_params.cache_rs_reset_len = ggml_nelements(node) / node->view_src->ne[0];
compute_params.cache_rs_reset_idx = node->src[0]->view_offs / node->view_src->ne[0];
}
// Capture the destination slot block of every recurrent state cache writeback, plus the
- // conv_input window the conv state writeback copies. The active sequences occupy a
- // contiguous slot block [begin, begin + n_seqs) of the cache; the block and the window move
+ // source window needed by conv state and packed GDN rollback writes. The active sequences
+ // occupy a contiguous slot block [begin, begin + n_seqs) of the cache; these offsets move
// with the batch, so they are fed to the cached model as runtime inputs.
- if (node->op == GGML_OP_CPY && node->view_src != nullptr && is_kvcache(node->view_src, nullptr) &&
+ if (node->op == GGML_OP_CPY && node->view_src != nullptr && is_recurrent_cache(node->view_src) &&
node->src[1] != nullptr && node->src[1]->op == GGML_OP_VIEW && node->src[1]->view_src == node->view_src) {
const bool is_conv = is_conv_state_writeback(node);
const bool is_gdn = node->src[0]->op == GGML_OP_VIEW && node->src[0]->src[0] != nullptr &&
node->src[0]->src[0]->op == GGML_OP_GATED_DELTA_NET;
const bool is_extra = node->src[0]->op == GGML_OP_GET_ROWS;
+ const bool is_gdn_rollback = is_gdn && is_same_shape(node->src[0], node->src[1]);
const ggml_tensor * dest_view = node->src[1];
const ggml_tensor * cache = node->view_src;
const size_t row_bytes = cache->ne[0] * ggml_type_size(cache->type);
- if (row_bytes > 0 && (is_conv || is_gdn || is_extra)) {
+ if (is_gdn_rollback) {
+ // Rollback GDN exposes an already-flattened [state, seq, snapshot] VIEW and copies
+ // it to an identically-shaped cache VIEW. Non-rollback copies native 4-D state
+ // [value, key, head, seq] into flattened cache rows, so the shapes differ. This
+ // signature is local to the CPY and still works when fallback splits the graph.
+ model_params.has_rs_rollback = true;
+ }
+ if (row_bytes > 0 && (is_conv || is_gdn || is_extra) && !is_full_single_slot_writeback(node)) {
ComputeParams::RsWriteback writeback;
writeback.slot_begin = (int) (dest_view->view_offs / row_bytes);
if (is_conv) {
writeback.src_begin = (int) (node->src[0]->view_offs / node->src[0]->view_src->nb[0]);
- } else if (is_gdn) {
+ } else if (is_gdn_rollback) {
writeback.src_begin = (int) (node->src[0]->view_offs / node->src[0]->view_src->nb[1]);
}
compute_params.rs_writebacks[get_tensor_ov_name(cgraph, node)] = writeback;
}
- if (is_conv || is_gdn) {
+ if ((is_conv || is_gdn) && !is_full_single_slot_writeback(node)) {
compute_params.s_copy_active_slot_len = (int) dest_view->ne[1];
}
}
@@ -975,6 +1107,12 @@ ov::PartialShape GgmlOvDecoder::get_graph_input_shape(const ggml_tensor * op,
input_shape = ov::PartialShape{-1, 1, -1, -1};
}
+ } else if (is_recurrent_cache(input)) {
+ input_shape = ov::PartialShape{get_shape(input)};
+ if (!m_is_static && !m_is_stateful && input->ne[1] > 1) {
+ input_shape[2] = -1;
+ }
+
} else if (is_kvcache(input, op)) {
// kvcache
input_shape = ov::PartialShape{get_shape(input)};
@@ -1057,7 +1195,7 @@ bool GgmlOvDecoder::is_s_copy_leaf(const ggml_tensor * tensor) const {
while (data != nullptr && (data->op == GGML_OP_VIEW || data->op == GGML_OP_RESHAPE)) {
data = data->src[0];
}
- if (data != nullptr && is_kvcache(data, nullptr)) {
+ if (data != nullptr && is_recurrent_cache(data)) {
return true;
}
}
@@ -1095,7 +1233,7 @@ void GgmlOvDecoder::add_extra_inputs() {
}
// create_1d_input("token_len", m_compute_params.token_len_per_seq * m_compute_params.n_seq_active);
- if (m_compute_params.cache_rs_reset_idx != -1) {
+ if (m_compute_params.cache_rs_reset_idx != -1 && m_model_params.n_rs_slots != 1) {
// Whether/which cache slot to reset varies per compute call (e.g. a new sequence starting
// vs. continued decoding). can_reuse_statically() does not invalidate the cached static
// model on ComputeParams changes, so these must stay runtime Parameters even when static
@@ -1119,7 +1257,7 @@ void GgmlOvDecoder::add_extra_inputs() {
for (const auto & [node_name, writeback] : m_compute_params.rs_writebacks) {
create_1d_input("rs_slot_begin_" + node_name, writeback.slot_begin);
- if (!m_is_static) {
+ if (!m_is_static && writeback.src_begin >= 0) {
create_1d_input("rs_src_begin_" + node_name, writeback.src_begin);
}
}
@@ -1216,6 +1354,9 @@ void GgmlOvDecoder::compute_model_outputs() {
if (cur_node->op == GGML_OP_NONE || cur_node->op == GGML_OP_VIEW || cur_node->op == GGML_OP_RESHAPE) {
continue;
}
+ if (::is_inplace_op(cur_node) && ggml_nbytes(cur_node) == 0) {
+ continue;
+ }
auto cur_node_use_count = m_cgraph->use_counts[ggml_hash_find(&m_cgraph->visited_hash_set, cur_node)];
if (cur_node_use_count == 0) {
// The output of in-place ops is the view_src tensor, which is updated in place. We should use the view_src name as the output name to make sure it can be correctly matched with the later ops that use the view_src.
@@ -1822,6 +1963,27 @@ std::vector<size_t> GgmlOvDecoder::get_output_stride(int node_idx) const {
}
std::vector<std::string> GgmlOvDecoder::get_output_names(int node_idx) const {
+ auto * node = m_node_info_list[node_idx].node;
+ if (node->op == GGML_OP_GATED_DELTA_NET && !m_model_params.has_rs_rollback) {
+ std::string attn_name;
+ std::string state_name;
+ for (int i = node_idx + 1; i < m_cgraph->n_nodes; i++) {
+ auto * consumer = m_cgraph->nodes[i];
+ if (consumer->op != GGML_OP_VIEW || consumer->src[0] != node) {
+ continue;
+ }
+ // GGML packs [attention | state]. The attention VIEW starts at offset 0 and the
+ // state VIEW starts after the token-dependent attention segment.
+ auto & name = consumer->view_offs == 0 ? attn_name : state_name;
+ if (!name.empty()) {
+ return {m_node_info_list[node_idx].node_name};
+ }
+ name = get_tensor_ov_name(m_cgraph, consumer);
+ }
+ if (!attn_name.empty() && !state_name.empty()) {
+ return {attn_name, state_name};
+ }
+ }
return {m_node_info_list[node_idx].node_name};
}
@@ -2154,6 +2316,11 @@ void GgmlOvDecoder::compute_node_dynamic_dims() {
case GGML_OP_DIV:
case GGML_OP_CLAMP:
case GGML_OP_PAD:
+ case GGML_OP_UPSCALE:
+ case GGML_OP_SIN:
+ case GGML_OP_COS:
+ case GGML_OP_LOG:
+ case GGML_OP_ROLL:
m_node_dynamic_dims[node] = m_node_dynamic_dims[node->src[0]];
break;
case GGML_OP_SUM_ROWS:
@@ -2168,6 +2335,8 @@ void GgmlOvDecoder::compute_node_dynamic_dims() {
break;
case GGML_OP_CPY:
case GGML_OP_SET_ROWS:
+ case GGML_OP_SUM:
+ case GGML_OP_MEAN:
m_node_dynamic_dims[node] = -1;
break;
case GGML_OP_IM2COL: {
@@ -2198,6 +2367,25 @@ void GgmlOvDecoder::compute_node_dynamic_dims() {
}
break;
}
+ case GGML_OP_IM2COL_3D: {
+ m_node_dynamic_dims[node] = -1;
+ if (m_node_dynamic_dims[node->src[1]] != -1) {
+ const int src_dyn = m_node_dynamic_dims[node->src[1]];
+ if (src_dyn == 0) {
+ m_node_dynamic_dims[node] = 1; // IW -> OW
+ } else if (src_dyn == 1) {
+ m_node_dynamic_dims[node] = 2; // IH -> OH
+ } else if (src_dyn == 3) {
+ m_node_dynamic_dims[node] = 3; // N -> N
+ }
+ if (m_node_dynamic_dims[node] != -1) {
+ OPENVINO_ASSERT(node->src[1]->ne[src_dyn] == node->ne[m_node_dynamic_dims[node]],
+ "Dynamic dim value mismatch for IM2COL_3D node: " + std::string(node->name) +
+ " and its src[1]: " + std::string(node->src[1]->name));
+ }
+ }
+ break;
+ }
default:
GGML_LOG_DEBUG("ggml-openvino: compute_node_dynamic_dims: unhandled op %s for node '%s'\n",
ggml_op_name(node->op), node->name);
diff --git a/ggml/src/ggml-openvino/ggml-decoder.h b/ggml/src/ggml-openvino/ggml-decoder.h
index 056e39e87..33b95340a 100644
--- a/ggml/src/ggml-openvino/ggml-decoder.h
+++ b/ggml/src/ggml-openvino/ggml-decoder.h
@@ -28,7 +28,9 @@ struct ModelParams {
std::map<int, int> n_heads_kv_per_layer;
int head_size = -1;
int state_size = -1; // for SSM molels, eg qwen35
- int32_t rope_params[16];
+ int32_t rope_params[16] = {};
+ int n_rs_slots = -1;
+ bool has_rs_rollback = false;
bool mixed_rope_params = false;
bool is_cacheless_attn = false;
std::vector<int> swa_layers;
@@ -45,9 +47,15 @@ struct ModelParams {
memcmp(rope_params, other.rope_params, sizeof(int32_t) * 16) == 0;
}
- bool can_reuse_dynamically(const ModelParams & other) const { return same_rope_params(other); }
+ bool can_reuse_dynamically(const ModelParams & other) const {
+ return same_rope_params(other) && n_rs_slots == other.n_rs_slots &&
+ has_rs_rollback == other.has_rs_rollback;
+ }
- bool can_reuse_statically(const ModelParams & other) const { return same_rope_params(other) && ctx == other.ctx; }
+ bool can_reuse_statically(const ModelParams & other) const {
+ return same_rope_params(other) && ctx == other.ctx && n_rs_slots == other.n_rs_slots &&
+ has_rs_rollback == other.has_rs_rollback;
+ }
bool kv_buffer_changed(const ModelParams & other) const { return kv_buffer_ctx_id != other.kv_buffer_ctx_id; }
};
@@ -100,7 +108,7 @@ struct ComputeParams {
struct RsWriteback {
int slot_begin = 0; // first cache slot written by the CPY
- int src_begin = 0; // first source row or column copied by the CPY
+ int src_begin = -1; // first source column copied by a conv-state CPY
};
std::map<std::string, RsWriteback> rs_writebacks;
@@ -353,6 +361,7 @@ public:
void add_extra_inputs();
void update_io(ggml_cgraph * cgraph);
+ bool is_bound_to(const ggml_cgraph * cgraph) const;
static bool is_inp_tok(const ggml_tensor * tensor, const ggml_tensor * op) {
return op->op == GGML_OP_GET_ROWS && tensor == op->src[1] && op->src[0]->op == GGML_OP_NONE;
@@ -362,10 +371,11 @@ public:
return op->op == GGML_OP_ROPE && tensor == op->src[1];
}
- // IMROPE packs 4 stacked position planes (t/h/w/e) into inp_pos, each of length
+ // IMROPE and VISION pack 4 stacked position planes (t/h/w/e) into inp_pos, each of length
// n_tokens; other modes carry a single position per token.
static int get_inp_pos_n_planes(const ggml_tensor * op) {
- return op->op_params[2] == GGML_ROPE_TYPE_IMROPE ? 4 : 1;
+ const int mode = op->op_params[2];
+ return (mode == GGML_ROPE_TYPE_IMROPE || mode == GGML_ROPE_TYPE_VISION || (mode & GGML_ROPE_TYPE_MROPE)) ? 4 : 1;
}
static bool is_inp_emb(const ggml_tensor * tensor, const ggml_tensor * op) {
@@ -387,17 +397,26 @@ public:
return op->op == GGML_OP_ROPE && tensor == op->src[2];
}
- // also returns true for cache_s and cache_r in SSM/DeltaNet models
- static bool is_kvcache(const ggml_tensor * tensor, const ggml_tensor * op) {
- if (tensor == nullptr) {
+ inline static bool is_recurrent_cache(const ggml_tensor * tensor) {
+ return tensor != nullptr && (strncmp(tensor->name, "cache_r_l", strlen("cache_r_l")) == 0 ||
+ strncmp(tensor->name, "cache_s_l", strlen("cache_s_l")) == 0 ||
+ strncmp(tensor->name, "cache_ple_r_l", strlen("cache_ple_r_l")) == 0);
+ }
+
+ inline static bool is_cache(const ggml_tensor * tensor, const ggml_tensor * op) {
+ return is_recurrent_cache(tensor) || is_kvcache(tensor, op);
+ }
+
+ inline static bool is_kvcache(const ggml_tensor * tensor, const ggml_tensor * op) {
+ if (tensor == nullptr || is_recurrent_cache(tensor)) {
return false;
}
return (tensor->buffer != nullptr && tensor->buffer->usage == GGML_BACKEND_BUFFER_USAGE_ANY) ||
(op != nullptr && op->op == GGML_OP_SET_ROWS && op->src[2] == tensor);
}
- static bool is_conv_state_writeback(const ggml_tensor * node) {
- return node->op == GGML_OP_CPY && node->view_src != nullptr && is_kvcache(node->view_src, nullptr) &&
+ inline static bool is_conv_state_writeback(const ggml_tensor * node) {
+ return node->op == GGML_OP_CPY && node->view_src != nullptr && is_recurrent_cache(node->view_src) &&
node->src[0] != nullptr && node->src[0]->op == GGML_OP_VIEW && node->src[0]->src[0] != nullptr &&
node->src[0]->src[0]->op == GGML_OP_CONCAT && node->src[1] != nullptr &&
node->src[1]->op == GGML_OP_VIEW && node->src[1]->view_src == node->view_src;
diff --git a/ggml/src/ggml-openvino/ggml-openvino-extra.cpp b/ggml/src/ggml-openvino/ggml-openvino-extra.cpp
index 216e3b8a6..0257e23db 100644
--- a/ggml/src/ggml-openvino/ggml-openvino-extra.cpp
+++ b/ggml/src/ggml-openvino/ggml-openvino-extra.cpp
@@ -2,12 +2,15 @@
#include "ggml-impl.h"
#include "ggml.h"
+#include "model-cache.h"
+#include <algorithm>
#include <cstdlib>
#include <cstring>
#include <openvino/runtime/intel_gpu/ocl/ocl.hpp>
#include <openvino/runtime/intel_npu/level_zero/level_zero.hpp>
#include <openvino/runtime/properties.hpp>
+#include <mutex>
#include <optional>
ov::Core & ov_singleton_core() {
@@ -15,14 +18,88 @@ ov::Core & ov_singleton_core() {
return core;
}
+static bool has_prefix(const std::string & s, const std::string & prefix) {
+ return s.size() >= prefix.size() && std::equal(prefix.begin(), prefix.end(), s.begin());
+}
+
+static bool is_virtual_routing_device(const std::string & device_name) {
+ return has_prefix(device_name, "AUTO") || has_prefix(device_name, "MULTI") || has_prefix(device_name, "HETERO");
+}
+
+static std::vector<std::string> ov_enumerate_devices() {
+ std::vector<std::string> result;
+
+ for (const auto & device : ov_singleton_core().get_available_devices()) {
+ if (!is_virtual_routing_device(device)) {
+ result.push_back(device);
+ }
+ }
+
+ if (result.empty()) {
+ result.push_back("CPU");
+ }
+
+ std::sort(result.begin(), result.end());
+ result.erase(std::unique(result.begin(), result.end()), result.end());
+ return result;
+}
+
+std::string ggml_openvino_get_device_description(const std::string & device_name) {
+ std::string description = device_name;
+ try {
+ description = ov_singleton_core().get_property(device_name, ov::device::full_name);
+ } catch (...) {
+ return device_name;
+ }
+
+ if (has_prefix(device_name, "NPU")) {
+ try {
+ const std::string arch = ov_singleton_core().get_property(device_name, "DEVICE_ARCHITECTURE").as<std::string>();
+ if (!arch.empty()) {
+ description += " (NPU " + arch + ")";
+ }
+ } catch (...) {
+ }
+ }
+
+ return description;
+}
+
+// requested: GGML_OPENVINO_DEVICE, nullptr if unset. available_devices is never empty (see ov_enumerate_devices)
+static std::string resolve_openvino_device_name(const std::vector<std::string> & available_devices,
+ const char * requested) {
+ auto available = [&](const std::string & name) {
+ return std::find(available_devices.begin(), available_devices.end(), name) != available_devices.end();
+ };
+ if (requested == nullptr) {
+ return available("CPU") ? "CPU" : available_devices.front();
+ }
+ if (!available(requested)) {
+ // No fallback to CPU (easy to miss) and no GPU -> GPU.0 alias (with iGPU + dGPU, GPU.0 is often the
+ // wrong one). List the devices here: --list-devices initializes this backend and would abort too.
+ std::string list;
+ for (const std::string & name : available_devices) {
+ list += "\n " + name + ": " + ggml_openvino_get_device_description(name);
+ }
+ GGML_ABORT("GGML OpenVINO Backend: GGML_OPENVINO_DEVICE=%s is not available. "
+ "Set it to one of the available OpenVINO devices:%s",
+ requested, list.c_str());
+ }
+ return requested;
+}
+
// =====================================================
// Device Configuration Implementations
// =====================================================
void ggml_openvino_device_config::init() {
+ static std::mutex mutex;
+ std::lock_guard<std::mutex> lock(mutex);
if (initialized) {
return;
}
+ // Set up front: a failed OpenCL setup below is not retried on every call
+ initialized = true;
// All recognized GGML_OPENVINO_* env vars. Their values are cached here
// once at backend init time and read back via ggml_openvino_getenv_str()
@@ -34,6 +111,7 @@ void ggml_openvino_device_config::init() {
"GGML_OPENVINO_SPILL_DIR",
"GGML_OPENVINO_DEBUG_NODE",
"GGML_OPENVINO_COMPILED_MODEL_CACHE_DIR",
+ "GGML_OPENVINO_COMPILED_MODEL_CACHE_ONLY",
"GGML_OPENVINO_NPU_COMPILE_CONFIG",
// Integer values (use ggml_openvino_getenv_int)
"GGML_OPENVINO_PREFILL_CHUNK_SIZE",
@@ -53,6 +131,7 @@ void ggml_openvino_device_config::init() {
"GGML_OPENVINO_DISABLE_KV_SLICE",
"GGML_OPENVINO_ENABLE_FALLBACK",
"GGML_OPENVINO_MANUAL_GQA_ATTN",
+ "GGML_OPENVINO_MOE_OP",
"GGML_OPENVINO_MEMORY_OPTIMIZE",
"GGML_OPENVINO_RELEASE_WEIGHTS",
"GGML_OPENVINO_REDUCE_COMPILE_MEM",
@@ -62,6 +141,8 @@ void ggml_openvino_device_config::init() {
"GGML_OPENVINO_DISABLE_REMOTE_OUTPUTS",
"GGML_OPENVINO_REQUANT_KQUANT",
"GGML_OPENVINO_DISABLE_KV_STATE_RELAYOUT",
+ // Build the precise (but O(n_nodes)) graph cache key. Needed by op tests.
+ "GGML_OPENVINO_FULL_GRAPH_KEY",
};
for (const char * const & env_var : env_var_names) {
@@ -71,16 +152,14 @@ void ggml_openvino_device_config::init() {
}
}
- device_name = ggml_openvino_getenv_str("GGML_OPENVINO_DEVICE", "CPU");
- auto available_devices = ov_singleton_core().get_available_devices();
- if (std::find(available_devices.begin(), available_devices.end(), device_name) == available_devices.end()) {
- GGML_LOG_WARN("GGML OpenVINO Backend: device %s is not available, fallback to CPU\n", device_name.c_str());
- device_name = "CPU";
- }
- is_npu = (device_name == "NPU");
+ available_devices = ov_enumerate_devices();
+ device_name = resolve_openvino_device_name(available_devices, ggml_openvino_getenv_str("GGML_OPENVINO_DEVICE"));
+ is_npu = has_prefix(device_name, "NPU");
+
+ ggml_openvino_model_cache_init();
const char * cache_dir = ggml_openvino_getenv_str("GGML_OPENVINO_CACHE_DIR");
- if (device_name == "NPU") {
+ if (has_prefix(device_name, "NPU")) {
compile_config = {
{"NPU_COMPILER_DYNAMIC_QUANTIZATION", "YES" },
{"NPU_USE_NPUW", "YES" },
@@ -106,48 +185,69 @@ void ggml_openvino_device_config::init() {
compile_config.insert(ov::cache_mode(ov::CacheMode::OPTIMIZE_SIZE));
}
+ if (ggml_openvino_getenv_int("GGML_OPENVINO_PROFILING") >= 2) {
+ compile_config.insert(ov::enable_profiling(true));
+ }
+
// Initialize remote context with queue sharing for GPU
- if (device_name == "GPU") {
- // Create OpenCL context and queue
- cl_int err;
- cl_platform_id platform;
- err = clGetPlatformIDs(1, &platform, nullptr);
- if (err != CL_SUCCESS) {
- GGML_LOG_ERROR("Failed to get OpenCL platform: %d\n", err);
- return;
+ if (has_prefix(device_name, "GPU")) {
+ // Use the OpenCL context OpenVINO created for this device, so GPU.N gets its own device
+ cl_context cl_ctx;
+ try {
+ auto ov_ctx = ov_singleton_core().get_default_context(device_name).as<ov::intel_gpu::ocl::ClContext>();
+ cl_ctx = ov_ctx.get();
+ } catch (const std::exception & e) {
+ // The consumers of the remote context have no host fallback, and OpenVINO
+ // already reported the device as present.
+ GGML_ABORT("ggml-openvino: failed to get the OpenCL context for %s: %s", device_name.c_str(), e.what());
}
+ cl_int err;
cl_device_id cl_device;
- err = clGetDeviceIDs(platform, CL_DEVICE_TYPE_GPU, 1, &cl_device, nullptr);
+ err = clGetContextInfo(cl_ctx, CL_CONTEXT_DEVICES, sizeof(cl_device), &cl_device, nullptr);
if (err != CL_SUCCESS) {
- GGML_LOG_ERROR("Failed to get OpenCL device: %d\n", err);
- return;
+ GGML_ABORT("ggml-openvino: failed to get the OpenCL device for %s: %d", device_name.c_str(), err);
}
- cl_context cl_ctx = clCreateContext(nullptr, 1, &cl_device, nullptr, nullptr, &err);
+ cl_platform_id cl_platform;
+ err = clGetDeviceInfo(cl_device, CL_DEVICE_PLATFORM, sizeof(cl_platform), &cl_platform, nullptr);
if (err != CL_SUCCESS) {
- GGML_LOG_ERROR("Failed to create OpenCL context: %d\n", err);
- return;
+ GGML_ABORT("ggml-openvino: failed to get the OpenCL platform for %s: %d", device_name.c_str(), err);
}
- cl_queue = clCreateCommandQueueWithProperties(cl_ctx, cl_device, nullptr, &err);
+ cl_mem_fill_fn =
+ (clEnqueueMemFillINTEL_fn) clGetExtensionFunctionAddressForPlatform(cl_platform, "clEnqueueMemFillINTEL");
+ cl_mem_cpy_fn =
+ (clEnqueueMemcpyINTEL_fn) clGetExtensionFunctionAddressForPlatform(cl_platform, "clEnqueueMemcpyINTEL");
+
+ cl_ulong device_max_alloc = 0;
+ err = clGetDeviceInfo(cl_device, CL_DEVICE_MAX_MEM_ALLOC_SIZE, sizeof(device_max_alloc), &device_max_alloc,
+ nullptr);
+ if (err == CL_SUCCESS) {
+ max_alloc_size = device_max_alloc;
+ } else {
+ // not fatal, ggml then allocates one buffer
+ GGML_LOG_WARN("Failed to get OpenCL max allocation size: %d\n", err);
+ }
+
+ const cl_queue_properties profiling_properties[] = {
+ CL_QUEUE_PROPERTIES,
+ CL_QUEUE_PROFILING_ENABLE,
+ 0,
+ };
+ const cl_queue_properties * queue_properties =
+ ggml_openvino_getenv_int("GGML_OPENVINO_PROFILING") >= 2 ? profiling_properties : nullptr;
+ cl_queue = clCreateCommandQueueWithProperties(cl_ctx, cl_device, queue_properties, &err);
if (err != CL_SUCCESS) {
- GGML_LOG_ERROR("Failed to create OpenCL command queue: %d\n", err);
- clReleaseContext(cl_ctx);
- return;
+ GGML_ABORT("ggml-openvino: failed to create the OpenCL queue for %s: %d", device_name.c_str(), err);
}
// Create OpenVINO remote context with queue sharing
remote_context = ov::intel_gpu::ocl::ClContext(ov_singleton_core(), cl_queue);
-
- // Release the context (queue keeps a reference)
- clReleaseContext(cl_ctx);
- } else if (device_name == "NPU") {
+ } else if (has_prefix(device_name, "NPU")) {
// remote tensor is not used for NPU yet
// remote_context = ov_singleton_core().get_default_context(device_name);
}
-
- initialized = true;
}
ggml_openvino_device_config::~ggml_openvino_device_config() {
@@ -173,6 +273,12 @@ const std::string & ggml_openvino_get_device_name() {
return ggml_openvino_get_device_config().device_name;
}
+std::vector<std::string> ggml_openvino_get_available_devices() {
+ auto & config = ggml_openvino_get_device_config();
+ config.init();
+ return config.available_devices;
+}
+
// Get the value of a GGML_OPENVINO_* env var as a string. Returns
// default_value when the var is unset or set to an empty string.
const char * ggml_openvino_getenv_str(const char * var, const char * default_value) {
@@ -198,12 +304,12 @@ bool ggml_openvino_reduce_compile_mem_enabled() {
return ggml_openvino_getenv_int("GGML_OPENVINO_MEMORY_OPTIMIZE") != 0;
}
-bool ggml_openvino_release_weights_enabled(const std::string & device) {
+bool ggml_openvino_release_weights_enabled() {
const char * release_weights = ggml_openvino_getenv_str("GGML_OPENVINO_RELEASE_WEIGHTS");
if (release_weights != nullptr) {
- return device == "GPU" && ggml_openvino_getenv_int("GGML_OPENVINO_RELEASE_WEIGHTS") != 0;
+ return ggml_openvino_is_gpu() && ggml_openvino_getenv_int("GGML_OPENVINO_RELEASE_WEIGHTS") != 0;
}
- return device == "GPU" && ggml_openvino_getenv_int("GGML_OPENVINO_MEMORY_OPTIMIZE") != 0;
+ return ggml_openvino_is_gpu() && ggml_openvino_getenv_int("GGML_OPENVINO_MEMORY_OPTIMIZE") != 0;
}
// Check if running on NPU
@@ -211,6 +317,14 @@ bool ggml_openvino_is_npu() {
return ggml_openvino_get_device_config().is_npu;
}
+bool ggml_openvino_is_gpu() {
+ return has_prefix(ggml_openvino_get_device_name(), "GPU");
+}
+
+size_t ggml_openvino_max_alloc_size() {
+ return ggml_openvino_get_device_config().max_alloc_size;
+}
+
// Get the remote context for the current device (returns empty optional for CPU)
std::optional<ov::RemoteContext> ggml_openvino_get_remote_context() {
return ggml_openvino_get_device_config().remote_context;
@@ -226,32 +340,14 @@ cl_command_queue ggml_openvino_get_cl_queue() {
return ggml_openvino_get_device_config().cl_queue;
}
-// Get the clEnqueueMemFillINTEL function pointer (lazy load)
+// Get the clEnqueueMemFillINTEL function pointer
clEnqueueMemFillINTEL_fn ggml_openvino_get_clEnqueueMemFillINTEL() {
- static clEnqueueMemFillINTEL_fn fn = nullptr;
- static bool loaded = false;
- if (!loaded) {
- loaded = true;
- cl_platform_id platform;
- if (clGetPlatformIDs(1, &platform, nullptr) == CL_SUCCESS) {
- fn = (clEnqueueMemFillINTEL_fn) clGetExtensionFunctionAddressForPlatform(platform, "clEnqueueMemFillINTEL");
- }
- }
- return fn;
+ return ggml_openvino_get_device_config().cl_mem_fill_fn;
}
-// Get the clEnqueueMemcpyINTEL function pointer (lazy load)
+// Get the clEnqueueMemcpyINTEL function pointer
clEnqueueMemcpyINTEL_fn ggml_openvino_get_clEnqueueMemcpyINTEL() {
- static clEnqueueMemcpyINTEL_fn fn = nullptr;
- static bool loaded = false;
- if (!loaded) {
- loaded = true;
- cl_platform_id platform;
- if (clGetPlatformIDs(1, &platform, nullptr) == CL_SUCCESS) {
- fn = (clEnqueueMemcpyINTEL_fn) clGetExtensionFunctionAddressForPlatform(platform, "clEnqueueMemcpyINTEL");
- }
- }
- return fn;
+ return ggml_openvino_get_device_config().cl_mem_cpy_fn;
}
// Get requantization type for a tensor type (returns nullopt if no requant needed)
@@ -280,14 +376,11 @@ std::optional<ExtraQuantType> ggml_openvino_get_requant_type(const ggml_tensor *
// Q6_K/Q5_K are touched):
// q4_sym128 Q6_K/Q5_K -> Q4_0_128 (u4, group 128, symmetric)
// q4_sym128_all and Q4_K too -- drops Q4_K's per-32 zero point, which costs some accuracy
- // q4_asym64_all Q6_K/Q5_K and Q4_K -> Q4_1_64 (u4, group 64, asymmetric) -- most of the
- // metadata saving while keeping a real zero point
+ // q4_asym64 Q6_K/Q5_K -> Q4_1_64 (u4, group 64, asymmetric)
+ // q4_asym64_all Q6_K/Q5_K and Q4_K -> Q4_1_64 (u4, group 64, asymmetric)
// native no requantization at all (keep Q6_K/Q5_K as they are)
//
- // The asymmetric target is only offered in its _all form: leaving Q4_K at its native group 32
- // while Q6_K/Q5_K move to group 64 gives the Q/K/V projections different group counts, and the
- // GPU plugin's FullyConnectedHorizontalFusion concatenates their scale constants, which then
- // fails shape inference. Requantizing all three keeps the group size uniform.
+ // q4_asym64 leaves Q4_K at its native group 32. Use q4_asym64_all to keep the group size uniform.
const char * rq = ggml_openvino_getenv_str("GGML_OPENVINO_REQUANT_KQUANT");
auto is_opt = [rq](const char * name) {
return rq && strcmp(rq, name) == 0;
@@ -295,6 +388,7 @@ std::optional<ExtraQuantType> ggml_openvino_get_requant_type(const ggml_tensor *
const bool sym128 = is_opt("q4_sym128");
const bool sym128_all = is_opt("q4_sym128_all");
const bool asym64_all = is_opt("q4_asym64_all");
+ const bool asym64 = is_opt("q4_asym64");
if (tensor->type == GGML_TYPE_Q4_K) {
if (sym128_all) {
@@ -313,7 +407,7 @@ std::optional<ExtraQuantType> ggml_openvino_get_requant_type(const ggml_tensor *
if (sym128 || sym128_all) {
return ExtraQuantType::Q4_0_64;
}
- if (asym64_all) {
+ if (asym64 || asym64_all) {
return ExtraQuantType::Q4_1_64;
}
// TODO: temporary workaround for a known OpenVINO GPU-plugin bug -- remove once the
@@ -328,7 +422,7 @@ std::optional<ExtraQuantType> ggml_openvino_get_requant_type(const ggml_tensor *
// already requantize to per-channel Q8_0_C (grouped=0). Sending these to grouped 4 bit
// avoids the broken layout and restores correct output.
// Opt out with GGML_OPENVINO_REQUANT_KQUANT=native.
- if (ggml_openvino_get_device_name() == "GPU" && !is_opt("native")) {
+ if (ggml_openvino_is_gpu() && !is_opt("native")) {
return ExtraQuantType::Q4_0_64;
}
}
@@ -338,7 +432,7 @@ std::optional<ExtraQuantType> ggml_openvino_get_requant_type(const ggml_tensor *
if (sym128 || sym128_all) {
return ExtraQuantType::Q4_0_128;
}
- if (asym64_all) {
+ if (asym64 || asym64_all) {
return ExtraQuantType::Q4_1_64;
}
if (is_opt("native")) {
@@ -439,9 +533,7 @@ ggml_openvino_extracted_layout ggml_openvino_get_extracted_layout(const ggml_ten
layout.weights_per_block = tensor->ne[0];
break;
default:
- layout.weights_per_block = -1;
GGML_ABORT("Code of re-quantizing to channel-wise is not updated");
- break;
}
if (layout.is_requant) {
@@ -560,12 +652,11 @@ ggml_openvino_tensor_extra * ggml_openvino_create_tensor_extra(const ggml_tensor
return nullptr;
}
- const auto & device_name = ggml_openvino_get_device_name();
auto remote_context = ggml_openvino_get_remote_context();
std::shared_ptr<ov::Tensor> ov_tensor;
if (is_remote) {
- GGML_ASSERT(device_name == "GPU");
+ GGML_ASSERT(ggml_openvino_is_gpu());
auto gpu_context = remote_context->as<ov::intel_gpu::ocl::ClContext>();
auto usm_tensor = gpu_context.create_tensor(element_type, shape, tensor->data);
ov_tensor = std::make_shared<ov::intel_gpu::ocl::USMTensor>(std::move(usm_tensor));
diff --git a/ggml/src/ggml-openvino/ggml-openvino-extra.h b/ggml/src/ggml-openvino/ggml-openvino-extra.h
index 9d827d969..3a58b2d02 100644
--- a/ggml/src/ggml-openvino/ggml-openvino-extra.h
+++ b/ggml/src/ggml-openvino/ggml-openvino-extra.h
@@ -63,12 +63,16 @@ clEnqueueMemcpyINTEL_fn ggml_openvino_get_clEnqueueMemcpyINTEL();
struct ggml_openvino_device_config {
std::string device_name = "CPU";
+ std::vector<std::string> available_devices;
bool is_npu = false;
bool initialized = false;
std::optional<ov::RemoteContext> remote_context;
+ size_t max_alloc_size = SIZE_MAX;
ov::AnyMap compile_config;
std::unordered_map<std::string, std::string> environment_variables;
cl_command_queue cl_queue = nullptr;
+ clEnqueueMemFillINTEL_fn cl_mem_fill_fn = nullptr;
+ clEnqueueMemcpyINTEL_fn cl_mem_cpy_fn = nullptr;
void init();
~ggml_openvino_device_config();
@@ -83,6 +87,12 @@ void ggml_openvino_init_device_config();
// Get the device name
const std::string & ggml_openvino_get_device_name();
+// Get all available physical OpenVINO devices
+std::vector<std::string> ggml_openvino_get_available_devices();
+
+// Human-readable device name, e.g. "Intel(R) AI Boost (NPU 4000)"; the device id if unavailable
+std::string ggml_openvino_get_device_description(const std::string & device_name);
+
// Environment variable accessors. All GGML_OPENVINO_* env vars are read once
// during backend init and cached on the device config; consumers must go
// through these helpers (never call ::getenv directly) so behavior stays
@@ -102,11 +112,17 @@ int ggml_openvino_getenv_int(const char * var, int default_value = 0);
// Memory optimization toggles. GGML_OPENVINO_MEMORY_OPTIMIZE is an umbrella
// switch; the fine-grained env vars still override it when explicitly set.
bool ggml_openvino_reduce_compile_mem_enabled();
-bool ggml_openvino_release_weights_enabled(const std::string & device);
+bool ggml_openvino_release_weights_enabled();
// Check if running on NPU
bool ggml_openvino_is_npu();
+// Check if running on a GPU (GPU, GPU.0, GPU.1, ...)
+bool ggml_openvino_is_gpu();
+
+// Largest single memory object the device can allocate, SIZE_MAX when there is no known limit
+size_t ggml_openvino_max_alloc_size();
+
// Host weight-buffer release (GGML_OPENVINO_RELEASE_WEIGHTS, GPU only).
// register: record a host weight buffer (idempotent per data pointer).
// release: madvise(MADV_DONTNEED) all registered buffers, dropping their RSS.
diff --git a/ggml/src/ggml-openvino/ggml-openvino.cpp b/ggml/src/ggml-openvino/ggml-openvino.cpp
index 336c793a0..87ee096fa 100644
--- a/ggml/src/ggml-openvino/ggml-openvino.cpp
+++ b/ggml/src/ggml-openvino/ggml-openvino.cpp
@@ -8,6 +8,7 @@
#include "ggml-openvino/utils.h"
#include "ggml-quants.h"
#include "ggml.h"
+#include "model-cache.h"
#include <algorithm>
#include <atomic>
@@ -24,7 +25,10 @@
#include <openvino/runtime/allocator.hpp>
#include <openvino/runtime/intel_gpu/ocl/ocl.hpp>
#include <openvino/runtime/intel_npu/level_zero/level_zero.hpp>
+#include <openvino/runtime/properties.hpp>
#include <openvino/runtime/tensor.hpp>
+#include <algorithm>
+#include <map>
#include <set>
#include <string>
#include <vector>
@@ -69,8 +73,7 @@ struct ggml_backend_openvino_buffer_context {
size_t size;
bool is_remote;
- // Set when the buffer is a file-backed spill mapping (GGML_OPENVINO_SPILL_DIR); it must be
- // munmap'd rather than freed.
+ // File-backed spill or cache-only virtual memory.
void * spill_mapping = nullptr;
size_t spill_size = 0;
@@ -79,6 +82,8 @@ struct ggml_backend_openvino_buffer_context {
// Track all extras for cleanup
std::map<ggml_tensor *, ggml_openvino_extra_base *> tensor_extras;
+ std::map<const void *, uint64_t> weight_fingerprints;
+ std::vector<ggml_openvino_source_mapping> source_mappings;
// Used for re-allocation on device for kvcache
void * data_prev;
@@ -100,7 +105,7 @@ struct ggml_backend_openvino_buffer_context {
const auto & device_name = ggml_openvino_get_device_name();
if (is_remote) {
- GGML_ASSERT(device_name == "GPU");
+ GGML_ASSERT(ggml_openvino_is_gpu());
auto remote_context = ggml_openvino_get_remote_context();
auto gpu_context = remote_context->as<ov::intel_gpu::ocl::ClContext>();
ov::intel_gpu::ocl::USMTensor usm_tensor =
@@ -108,8 +113,25 @@ struct ggml_backend_openvino_buffer_context {
data = usm_tensor.get();
ov_buffer = std::make_shared<ov::intel_gpu::ocl::USMTensor>(std::move(usm_tensor));
} else {
-#ifndef _WIN32
- if (const char * spill_dir = ggml_openvino_getenv_str("GGML_OPENVINO_SPILL_DIR")) {
+#ifdef _WIN32
+ if (ggml_openvino_model_cache_only()) {
+ data = spill_mapping = VirtualAlloc(nullptr, size, MEM_RESERVE | MEM_COMMIT, PAGE_READWRITE);
+ if (data == nullptr) {
+ return;
+ }
+ spill_size = size;
+ ov_buffer = std::make_shared<ov::Tensor>(ov::element::u8, ov::Shape{size}, data);
+ } else
+#else
+ if (ggml_openvino_model_cache_only()) {
+ void * m = mmap(nullptr, size, PROT_READ | PROT_WRITE, MAP_PRIVATE | MAP_ANONYMOUS, -1, 0);
+ if (m == MAP_FAILED) {
+ return;
+ }
+ data = spill_mapping = m;
+ spill_size = size;
+ ov_buffer = std::make_shared<ov::Tensor>(ov::element::u8, ov::Shape{size}, data);
+ } else if (const char * spill_dir = ggml_openvino_getenv_str("GGML_OPENVINO_SPILL_DIR")) {
// Disk-backed weight buffer: back the repacked weights with a temp file via MAP_SHARED
// instead of anonymous memory. Anonymous pages can only be evicted to swap, so the
// repacked buffer stays pinned alongside the mmap'd source and both are resident at once
@@ -180,7 +202,11 @@ struct ggml_backend_openvino_buffer_context {
delete pair.second;
}
tensor_extras.clear();
-#ifndef _WIN32
+#ifdef _WIN32
+ if (spill_mapping != nullptr) {
+ VirtualFree(spill_mapping, 0, MEM_RELEASE);
+ } else
+#else
if (spill_mapping != nullptr) {
munmap(spill_mapping, spill_size);
} else
@@ -295,7 +321,7 @@ static enum ggml_status ggml_backend_openvino_buffer_init_tensor(ggml_backend_bu
ggml_backend_openvino_buffer_context * ctx = (ggml_backend_openvino_buffer_context *) buffer->context;
// Put kvcache on device memory for GPU (NPU memory is too small even for kvcache)
- if (strncmp(tensor->name, "cache_", 6) == 0 && !ctx->is_remote && ggml_openvino_get_device_name() == "GPU" &&
+ if (strncmp(tensor->name, "cache_", 6) == 0 && !ctx->is_remote && ggml_openvino_is_gpu() &&
!is_stateful_enabled()) {
GGML_ASSERT(ctx->tensor_extras.empty());
auto device = ctx->device;
@@ -311,6 +337,26 @@ static enum ggml_status ggml_backend_openvino_buffer_init_tensor(ggml_backend_bu
if (tensor->view_src != nullptr) {
GGML_ASSERT(tensor->view_src->buffer->buft == buffer->buft);
if (tensor->view_src->extra != nullptr) {
+ // The cached ov::Tensor carries the shape it was built with, so sharing view_src's
+ // extra hands out the wrong shape for a reshaping view (e.g. Vcur reshaped from
+ // [n_embd, n_tokens] to [head_size, n_heads_kv, n_tokens]). When such a view is a
+ // graph input, binding it fails the shape check. Give it its own extra instead;
+ // ggml_openvino_create_tensor_extra reads ne and data off the view, so the offset is
+ // handled too. Only safe for a contiguous view - the ov::Tensor assumes dense strides.
+ // Skip empty views: they have no data, and on GPU one can sit at the end of the USM buffer.
+ if (!ggml_are_same_shape(tensor, tensor->view_src) && ggml_is_contiguous(tensor) &&
+ !ggml_is_quantized(tensor->type) && tensor->data != nullptr && ggml_nbytes(tensor) > 0) {
+ if (ggml_openvino_tensor_extra * extra =
+ ggml_openvino_create_tensor_extra(tensor, ctx->is_remote)) {
+ auto it = ctx->tensor_extras.find(tensor);
+ if (it != ctx->tensor_extras.end()) {
+ delete it->second;
+ }
+ ctx->tensor_extras[tensor] = extra;
+ tensor->extra = extra;
+ return GGML_STATUS_SUCCESS;
+ }
+ }
tensor->extra = tensor->view_src->extra;
}
return GGML_STATUS_SUCCESS;
@@ -346,7 +392,7 @@ static void ggml_backend_openvino_buffer_memset_tensor(ggml_backend_buffer_t buf
// For remote (device) buffers, use OpenCL USM memfill
cl_command_queue queue = ggml_openvino_get_cl_queue();
auto mem_fill_fn = ggml_openvino_get_clEnqueueMemFillINTEL();
- if (queue != nullptr && mem_fill_fn != nullptr) {
+ if (mem_fill_fn != nullptr) {
uint8_t pattern = value;
cl_int err = mem_fill_fn(queue, (char *) tensor->data + offset, &pattern, sizeof(pattern), size, 0, nullptr,
nullptr);
@@ -355,7 +401,7 @@ static void ggml_backend_openvino_buffer_memset_tensor(ggml_backend_buffer_t buf
}
clFinish(queue);
} else {
- GGML_LOG_ERROR("%s: no OpenCL queue or clEnqueueMemFillINTEL not available for GPU buffer\n", __func__);
+ GGML_LOG_ERROR("%s: clEnqueueMemFillINTEL not available for GPU buffer\n", __func__);
}
} else {
memset((char *) tensor->data + offset, value, size);
@@ -375,6 +421,17 @@ static void ggml_backend_openvino_buffer_set_tensor(ggml_backend_buffer_t buffer
bool is_weight_buffer = (buffer->usage == GGML_BACKEND_BUFFER_USAGE_WEIGHTS);
// Full tensor set: offset=0, full size, not a view
bool is_full_tensor_set = (offset == 0 && size == ggml_nbytes(tensor) && tensor->view_src == nullptr);
+ if (is_weight_buffer && ggml_openvino_getenv_str("GGML_OPENVINO_COMPILED_MODEL_CACHE_DIR")) {
+ if (is_full_tensor_set) {
+ ctx->weight_fingerprints[tensor->data] = ggml_openvino_source_fingerprint(data, size, ctx->source_mappings);
+ }
+ if (ggml_openvino_model_cache_only()) {
+ if (!is_full_tensor_set) {
+ GGML_ABORT("ggml-openvino: cache-only mode requires whole mmap weight uploads");
+ }
+ return;
+ }
+ }
// 2D tensor (typical weight shape), or a 3D quantized MoE expert weight (MUL_MAT_ID). Dense 3D
// expert weights are handled later in create_weight_node instead.
bool is_2d = (tensor->ne[2] == 1 && tensor->ne[3] == 1);
@@ -441,14 +498,14 @@ static void ggml_backend_openvino_buffer_set_tensor(ggml_backend_buffer_t buffer
if (ctx->is_remote) {
cl_command_queue queue = ggml_openvino_get_cl_queue();
auto mem_cpy_fn = ggml_openvino_get_clEnqueueMemcpyINTEL();
- if (queue != nullptr && mem_cpy_fn != nullptr) {
+ if (mem_cpy_fn != nullptr) {
cl_int err =
mem_cpy_fn(queue, CL_TRUE, (char *) tensor->data + offset, data, size, 0, nullptr, nullptr);
if (err != CL_SUCCESS) {
GGML_LOG_ERROR("%s: clEnqueueMemcpyINTEL failed with error %d\n", __func__, err);
}
} else {
- GGML_LOG_ERROR("%s: no OpenCL queue or clEnqueueMemcpyINTEL not available for GPU buffer\n", __func__);
+ GGML_LOG_ERROR("%s: clEnqueueMemcpyINTEL not available for GPU buffer\n", __func__);
}
} else {
memcpy((char *) tensor->data + offset, data, size);
@@ -478,18 +535,22 @@ static void ggml_backend_openvino_buffer_get_tensor(ggml_backend_buffer_t buffer
GGML_ASSERT(tensor != nullptr && tensor->data != nullptr);
ggml_backend_openvino_buffer_context * ctx = (ggml_backend_openvino_buffer_context *) buffer->context;
+ if (ggml_openvino_model_cache_only() && buffer->usage == GGML_BACKEND_BUFFER_USAGE_WEIGHTS) {
+ GGML_ABORT("ggml-openvino: cannot read unloaded weights in cache-only mode");
+ }
+
if (ctx->is_remote) {
// For remote (device) buffers, use OpenCL USM memcpy (device-to-host)
cl_command_queue queue = ggml_openvino_get_cl_queue();
auto mem_cpy_fn = ggml_openvino_get_clEnqueueMemcpyINTEL();
- if (queue != nullptr && mem_cpy_fn != nullptr) {
+ if (mem_cpy_fn != nullptr) {
cl_int err =
mem_cpy_fn(queue, CL_TRUE, data, (const char *) tensor->data + offset, size, 0, nullptr, nullptr);
if (err != CL_SUCCESS) {
GGML_LOG_ERROR("%s: clEnqueueMemcpyINTEL failed with error %d\n", __func__, err);
}
} else {
- GGML_LOG_ERROR("%s: no OpenCL queue or clEnqueueMemcpyINTEL not available for GPU buffer\n", __func__);
+ GGML_LOG_ERROR("%s: clEnqueueMemcpyINTEL not available for GPU buffer\n", __func__);
}
} else {
memcpy(data, (const char *) tensor->data + offset, size);
@@ -507,8 +568,8 @@ static bool ggml_backend_openvino_buffer_cpy_tensor(ggml_backend_buffer_t buffer
// For remote (device) buffers, use OpenCL USM memcpy
cl_command_queue queue = ggml_openvino_get_cl_queue();
auto mem_cpy_fn = ggml_openvino_get_clEnqueueMemcpyINTEL();
- if (queue == nullptr || mem_cpy_fn == nullptr) {
- GGML_LOG_ERROR("%s: no OpenCL queue or clEnqueueMemcpyINTEL not available for GPU buffer\n", __func__);
+ if (mem_cpy_fn == nullptr) {
+ GGML_LOG_ERROR("%s: clEnqueueMemcpyINTEL not available for GPU buffer\n", __func__);
return false;
}
// Can copy from host to device
@@ -550,7 +611,7 @@ static void ggml_backend_openvino_buffer_clear(ggml_backend_buffer_t buffer, uin
if (ctx->is_remote) {
cl_command_queue queue = ggml_openvino_get_cl_queue();
auto mem_fill_fn = ggml_openvino_get_clEnqueueMemFillINTEL();
- if (queue != nullptr && mem_fill_fn != nullptr) {
+ if (mem_fill_fn != nullptr) {
uint8_t pattern = value;
cl_int err = mem_fill_fn(queue, ctx->data, &pattern, sizeof(pattern), ctx->size, 0, nullptr, nullptr);
if (err != CL_SUCCESS) {
@@ -558,8 +619,7 @@ static void ggml_backend_openvino_buffer_clear(ggml_backend_buffer_t buffer, uin
}
clFinish(queue);
} else {
- GGML_LOG_WARN("%s: no OpenCL queue or clEnqueueMemFillINTEL not available for GPU buffer clear\n",
- __func__);
+ GGML_LOG_WARN("%s: clEnqueueMemFillINTEL not available for GPU buffer clear\n", __func__);
}
} else {
memset(ctx->data, value, ctx->size);
@@ -609,7 +669,8 @@ static size_t ggml_backend_openvino_buffer_type_get_alignment(ggml_backend_buffe
static size_t ggml_backend_openvino_buffer_type_get_max_size(ggml_backend_buffer_type_t buft) {
GGML_UNUSED(buft);
- return SIZE_MAX;
+ // A GPU caps a single memory object, so let ggml split a large buffer into parts that fit
+ return ggml_openvino_max_alloc_size();
}
static size_t ggml_backend_openvino_buffer_type_get_alloc_size(ggml_backend_buffer_type_t buft,
@@ -617,7 +678,7 @@ static size_t ggml_backend_openvino_buffer_type_get_alloc_size(ggml_backend_buff
GGML_UNUSED(buft);
// For quantized weight tensors, we need extra space for extracted data.
- if (ggml_is_quantized(tensor->type) && tensor->ne[3] == 1) {
+ if (!ggml_openvino_model_cache_only() && ggml_is_quantized(tensor->type) && tensor->ne[3] == 1) {
ggml_openvino_extracted_layout layout = ggml_openvino_get_extracted_layout(tensor);
if (layout.total_size > 0) {
// GGML_LOG_DEBUG("%s: tensor %s needs %zu bytes (original %zu, extracted: weights=%zu scales=%zu zp=%zu)\n",
@@ -772,6 +833,19 @@ bool ggml_backend_buft_is_openvino_host(ggml_backend_buffer_type_t buft) {
return buft->iface.get_name == ggml_backend_openvino_host_buffer_type_get_name;
}
+uint64_t ggml_backend_openvino_weight_fingerprint(const ggml_tensor * tensor) {
+ if (ggml_backend_buffer_is_openvino(tensor->buffer)) {
+ auto * ctx = static_cast<ggml_backend_openvino_buffer_context *>(tensor->buffer->context);
+ auto it = ctx->weight_fingerprints.find(tensor->data);
+ if (it != ctx->weight_fingerprints.end()) {
+ return it->second;
+ }
+ GGML_ABORT("ggml-openvino: missing source identity for weight %s", tensor->name);
+ }
+ std::vector<ggml_openvino_source_mapping> mappings;
+ return ggml_openvino_source_fingerprint(tensor->data, ggml_nbytes(tensor), mappings);
+}
+
static void ggml_backend_openvino_free(ggml_backend_t backend) {
ggml_backend_openvino_context * ctx = (ggml_backend_openvino_context *) backend->context;
@@ -825,7 +899,7 @@ static const ggml_backend_i ggml_backend_openvino_interface = {
};
int ggml_backend_openvino_get_device_count() {
- return 1;
+ return (int) ggml_openvino_get_available_devices().size();
}
static ggml_guid_t ggml_backend_openvino_guid(void) {
@@ -884,10 +958,122 @@ namespace {
struct ggml_backend_openvino_device_context {
int device;
std::string name;
+ std::string ov_name; // OpenVINO device id: CPU, GPU, GPU.1, NPU, ...
std::string description;
+ size_t total_memory;
};
}
+static bool ov_device_has_prefix(const std::string & s, const std::string & prefix) {
+ return s.size() >= prefix.size() && std::equal(prefix.begin(), prefix.end(), s.begin());
+}
+
+static bool ov_try_get_size_t_property(const std::string & device, const std::string & property, size_t & out) {
+ try {
+ const ov::Any value = ov_singleton_core().get_property(device, property);
+ if (value.is<size_t>()) {
+ out = value.as<size_t>();
+ return true;
+ }
+ if (value.is<uint64_t>()) {
+ out = (size_t) value.as<uint64_t>();
+ return true;
+ }
+ if (value.is<unsigned long long>()) {
+ out = (size_t) value.as<unsigned long long>();
+ return true;
+ }
+ if (value.is<int64_t>()) {
+ const int64_t v = value.as<int64_t>();
+ if (v >= 0) {
+ out = (size_t) v;
+ return true;
+ }
+ }
+ } catch (...) {
+ }
+ return false;
+}
+
+// System memory available to new allocations (MemAvailable on Linux), SIZE_MAX if unknown
+static size_t ov_system_available_memory() {
+#ifdef _WIN32
+ MEMORYSTATUSEX status;
+ status.dwLength = sizeof(status);
+ if (GlobalMemoryStatusEx(&status)) {
+ return (size_t) status.ullAvailPhys;
+ }
+#else
+ if (FILE * f = fopen("/proc/meminfo", "r")) {
+ char line[256];
+ unsigned long long kb = 0;
+ bool found = false;
+ while (!found && fgets(line, sizeof(line), f)) {
+ found = sscanf(line, "MemAvailable: %llu kB", &kb) == 1;
+ }
+ fclose(f);
+ if (found) {
+ return (size_t) std::min<unsigned long long>(kb * 1024, SIZE_MAX);
+ }
+ }
+#endif
+ return SIZE_MAX;
+}
+
+// iGPU and NPU allocate from system RAM, so their free memory can't exceed what the OS has available
+static bool ov_device_shares_system_memory(const std::string & device) {
+ if (ov_device_has_prefix(device, "NPU")) {
+ return true;
+ }
+ if (!ov_device_has_prefix(device, "GPU")) {
+ return false;
+ }
+ try {
+ return ov_singleton_core().get_property(device, ov::device::type) == ov::device::Type::INTEGRATED;
+ } catch (...) {
+ return false;
+ }
+}
+
+// usm_host / usm_shared allocations live in system RAM on a discrete GPU
+static bool ov_gpu_stat_is_host_memory(const std::string & key) {
+ return key == "usm_host" || key == "usm_shared";
+}
+
+static bool ov_try_get_gpu_used_memory(const std::string & device, size_t & out) {
+ out = 0;
+ try {
+ const ov::Any stats_any = ov_singleton_core().get_property(device, "GPU_MEMORY_STATISTICS");
+ if (stats_any.is<std::map<std::string, uint64_t>>()) {
+ const auto stats = stats_any.as<std::map<std::string, uint64_t>>();
+ for (const auto & kv : stats) {
+ if (!ov_gpu_stat_is_host_memory(kv.first)) {
+ out += (size_t) kv.second;
+ }
+ }
+ return true;
+ }
+ if (stats_any.is<ov::AnyMap>()) {
+ const auto stats = stats_any.as<ov::AnyMap>();
+ for (const auto & kv : stats) {
+ if (ov_gpu_stat_is_host_memory(kv.first)) {
+ continue;
+ }
+ if (kv.second.is<size_t>()) {
+ out += kv.second.as<size_t>();
+ } else if (kv.second.is<uint64_t>()) {
+ out += (size_t) kv.second.as<uint64_t>();
+ } else if (kv.second.is<unsigned long long>()) {
+ out += (size_t) kv.second.as<unsigned long long>();
+ }
+ }
+ return true;
+ }
+ } catch (...) {
+ }
+ return false;
+}
+
static const char * ggml_backend_openvino_device_get_name(ggml_backend_dev_t dev) {
ggml_backend_openvino_device_context * ctx = (ggml_backend_openvino_device_context *) dev->context;
return ctx->name.c_str();
@@ -899,27 +1085,45 @@ static const char * ggml_backend_openvino_device_get_description(ggml_backend_de
}
static void ggml_backend_openvino_device_get_memory(ggml_backend_dev_t dev, size_t * free, size_t * total) {
+ ggml_backend_openvino_device_context * ctx = (ggml_backend_openvino_device_context *) dev->context;
+
+ // total_memory is only set for GPU/NPU; used = this process's OpenVINO allocations on the device
+ size_t used = 0;
+ const bool known = ctx->total_memory > 0 &&
+ (ov_device_has_prefix(ctx->ov_name, "GPU") ?
+ ov_try_get_gpu_used_memory(ctx->ov_name, used) :
+ ov_try_get_size_t_property(ctx->ov_name, "NPU_DEVICE_ALLOC_MEM_SIZE", used));
+ if (known) {
+ *total = ctx->total_memory;
+ *free = (used >= *total) ? 0 : (*total - used);
+ } else {
+ // CPU, or a plugin without memory properties: report system memory
#ifdef _WIN32
- MEMORYSTATUSEX status;
- status.dwLength = sizeof(status);
- GlobalMemoryStatusEx(&status);
- *total = status.ullTotalPhys;
- *free = status.ullAvailPhys;
+ MEMORYSTATUSEX status;
+ status.dwLength = sizeof(status);
+ GlobalMemoryStatusEx(&status);
+ *total = status.ullTotalPhys;
+ *free = status.ullAvailPhys;
#else
- long pages = sysconf(_SC_PHYS_PAGES);
- long page_size = sysconf(_SC_PAGE_SIZE);
- *total = pages * page_size;
+ long pages = sysconf(_SC_PHYS_PAGES);
+ long page_size = sysconf(_SC_PAGE_SIZE);
+ *total = pages * page_size;
- // "free" system memory is ill-defined, for practical purposes assume that all of it is free:
- *free = *total;
+ // "free" system memory is ill-defined, for practical purposes assume that all of it is free:
+ *free = *total;
#endif // _WIN32
+ }
- GGML_UNUSED(dev);
+ if (ov_device_shares_system_memory(ctx->ov_name)) {
+ *free = std::min(*free, ov_system_available_memory());
+ }
}
static enum ggml_backend_dev_type ggml_backend_openvino_device_get_type(ggml_backend_dev_t dev) {
- GGML_UNUSED(dev);
- return GGML_BACKEND_DEVICE_TYPE_GPU;
+ ggml_backend_openvino_device_context * ctx = (ggml_backend_openvino_device_context *) dev->context;
+ // Only the device selected by GGML_OPENVINO_DEVICE is offered for offload. The others are
+ // registered for discovery (--list-devices) only; llama.cpp skips IGPU devices when a GPU exists.
+ return ctx->ov_name == ggml_openvino_get_device_name() ? GGML_BACKEND_DEVICE_TYPE_GPU : GGML_BACKEND_DEVICE_TYPE_IGPU;
}
static void ggml_backend_openvino_device_get_props(ggml_backend_dev_t dev, ggml_backend_dev_props * props) {
@@ -940,6 +1144,12 @@ static void ggml_backend_openvino_device_get_props(ggml_backend_dev_t dev, ggml_
static ggml_backend_t ggml_backend_openvino_device_init(ggml_backend_dev_t dev, const char * params) {
GGML_UNUSED(params);
ggml_backend_openvino_device_context * ctx = (ggml_backend_openvino_device_context *) dev->context;
+ if (ctx->ov_name != ggml_openvino_get_device_name()) {
+ // Not an error: test-backend-ops initializes every device
+ GGML_LOG_WARN("%s: %s (OpenVINO %s) is not the selected device, no ops will run on it; "
+ "set GGML_OPENVINO_DEVICE=%s to use it\n",
+ __func__, ctx->name.c_str(), ctx->ov_name.c_str(), ctx->ov_name.c_str());
+ }
return ggml_backend_openvino_init(ctx->device);
}
@@ -1159,7 +1369,7 @@ static ggml_openvino_op_support is_op_supported_case(const ggml_tensor * op) {
if (op->type == GGML_TYPE_I64) {
return {false, "CONCAT with I64 type is not supported"};
}
- if (ggml_openvino_get_device_name() == "GPU" && op->type == GGML_TYPE_BF16 && has_view_op_input(op)) {
+ if (ggml_openvino_is_gpu() && op->type == GGML_TYPE_BF16 && has_view_op_input(op)) {
return {false, "CONCAT with BF16 type and VIEW input is not supported on GPU"};
}
break;
@@ -1183,7 +1393,7 @@ static ggml_openvino_op_support is_op_supported_case(const ggml_tensor * op) {
if (op->ne[3] != 1) {
return {false, "GET_ROWS/SET_ROWS with ne[3] != 1 (ne[3]=" + std::to_string(op->ne[3]) + ") is not supported"};
}
- if (op->op == GGML_OP_GET_ROWS && ggml_openvino_get_device_name() == "GPU" &&
+ if (op->op == GGML_OP_GET_ROWS && ggml_openvino_is_gpu() &&
op->src[0]->type == GGML_TYPE_BF16) {
return {false, "GET_ROWS with BF16 src0 is not supported on GPU"};
}
@@ -1246,26 +1456,37 @@ static ggml_openvino_op_support is_op_supported_case(const ggml_tensor * op) {
// The GPU plugin can fuse broadcast DIV into the preceding FFN GEMM path
// and produce infs for per-channel scale vectors. Keep those DIVs on CPU
// until the fused GPU kernel is reliable. (falied case llama-arch-test mpt)
- if (ggml_openvino_get_device_name() == "GPU" && op->src[1]->ne[0] == op->ne[0] &&
+ if (ggml_openvino_is_gpu() && op->src[1]->ne[0] == op->ne[0] &&
op->src[1]->ne[1] == 1 && op->src[1]->ne[2] == 1 && op->src[1]->ne[3] == 1) {
return {false, "DIV per-channel scale broadcast is not supported on GPU"};
}
break;
}
case GGML_OP_POOL_2D: {
- const auto& name = ggml_openvino_get_device_name();
- if (name == "GPU") {
+ if (ggml_openvino_is_gpu()) {
const int32_t * params = op->op_params;
const int k0 = params[1];
const int k1 = params[2];
const int p0 = params[5];
const int p1 = params[6];
if ((p0 > 0 || p1 > 0) && (k0 < 3 || k1 < 3)) {
- return {false, "POOL_2D with padding and kernel size < 3 is not supported on " + name};
+ return {false, "POOL_2D with padding and kernel size < 3 is not supported on " + ggml_openvino_get_device_name()};
}
}
break;
}
+ case GGML_OP_SUM: {
+ if (op->src[0]->op == GGML_OP_PERMUTE) {
+ return {false, "SUM with PERMUTE input is not supported"};
+ }
+ break;
+ }
+ case GGML_OP_MEAN: {
+ if (op->src[0]->op == GGML_OP_PERMUTE && op->src[0]->src[0] != nullptr && op->src[0]->src[0]->op == GGML_OP_VIEW) {
+ return {false, "MEAN with PERMUTE of VIEW input is not supported"};
+ }
+ break;
+ }
case GGML_OP_SUM_ROWS: {
if (op->src[0]->op == GGML_OP_PERMUTE) {
return {false, "SUM_ROWS with PERMUTE input is not supported"};
@@ -1303,7 +1524,7 @@ static ggml_openvino_op_support is_op_supported_case(const ggml_tensor * op) {
break;
}
case GGML_OP_PERMUTE: {
- if (op->type == GGML_TYPE_BF16 && ggml_openvino_get_device_name() == "GPU") {
+ if (op->type == GGML_TYPE_BF16 && ggml_openvino_is_gpu()) {
return {false, "PERMUTE with BF16 type is not supported on GPU"};
}
break;
@@ -1312,7 +1533,7 @@ static ggml_openvino_op_support is_op_supported_case(const ggml_tensor * op) {
if (op->src[0]->type != GGML_TYPE_BF16 && op->src[1]->type == GGML_TYPE_BF16) {
return {false, "CPY with BF16 src[1] type is not supported"};
}
- if (ggml_openvino_get_device_name() == "NPU" && (op->src[0]->type == GGML_TYPE_BF16 || op->src[1]->type == GGML_TYPE_BF16)) {
+ if (ggml_openvino_is_npu() && (op->src[0]->type == GGML_TYPE_BF16 || op->src[1]->type == GGML_TYPE_BF16)) {
return {false, "CPY with BF16 is not supported is not supported on NPU"};
}
// CPY to a quantized destination (e.g. f32 -> q4_0) is numerically unstable with OpenVINO backend.
@@ -1337,13 +1558,13 @@ static ggml_openvino_op_support is_op_supported_case(const ggml_tensor * op) {
break;
}
case GGML_OP_MUL_MAT: {
- if (ggml_openvino_get_device_name() == "GPU" && op->src[0] != nullptr && op->src[1] != nullptr &&
+ if (ggml_openvino_is_gpu() && op->src[0] != nullptr && op->src[1] != nullptr &&
ggml_is_quantized(op->src[0]->type) && strcmp(op->src[0]->name, "a") == 0 &&
strcmp(op->src[1]->name, "b") == 0 && op->src[0]->ne[1] == 1 && op->src[1]->ne[1] == 64 &&
op->src[0]->ne[0] == 256 && op->src[1]->ne[0] == 256) {
return {false, "MUL_MAT quantized benchmark test case on GPU is not supported"};
}
- if (ggml_openvino_get_device_name() == "GPU" && op->type == GGML_TYPE_F32 && op->ne[0] == 1 && op->ne[1] == 1 &&
+ if (ggml_openvino_is_gpu() && op->type == GGML_TYPE_F32 && op->ne[0] == 1 && op->ne[1] == 1 &&
(op->src[0]->buffer == nullptr || op->src[0]->buffer->usage != GGML_BACKEND_BUFFER_USAGE_WEIGHTS)) {
return {false, "MUL_MAT scalar dot product with non-weight src[0] on GPU is not supported"};
}
@@ -1363,7 +1584,7 @@ static ggml_openvino_op_support is_op_supported_case(const ggml_tensor * op) {
return {false, "MUL_MAT_ID with single-expert or empty ne[2] <= 1 (ne[2]=" +
std::to_string(op->src[0]->ne[2]) + ") is not supported"};
}
- if (ggml_openvino_get_device_name() == "GPU" && op->src[0] != nullptr && !ggml_is_quantized(op->src[0]->type)) {
+ if (ggml_openvino_is_gpu() && op->src[0] != nullptr && !ggml_is_quantized(op->src[0]->type)) {
return {false, "MUL_MAT_ID with non-quantized weights on GPU is not supported"};
}
// The GPU plugin's GatherMatmul returns wrong values for the layouts test-backend-ops
@@ -1372,55 +1593,21 @@ static ggml_openvino_op_support is_op_supported_case(const ggml_tensor * op) {
// The same graph is correct on the CPU plugin, and correct on GPU for every real model,
// which always feeds experts from a bound tensor buffer. Standalone op-test tensors have
// no buffer at all, so use that to exclude them and let the scheduler run them on CPU.
- if (ggml_openvino_get_device_name() == "GPU" && op->src[0] != nullptr && op->src[0]->buffer == nullptr) {
+ if (ggml_openvino_is_gpu() && op->src[0] != nullptr && op->src[0]->buffer == nullptr) {
return {false, "MUL_MAT_ID with unbound expert tensors on GPU is not supported"};
}
// Only MXFP4 still needs the large-temporary guard; every other quantized type goes
// through GatherMatmul, which never materializes the selected expert weights.
- if (ggml_openvino_get_device_name() == "GPU" && op->src[0] != nullptr && op->src[0]->type == GGML_TYPE_MXFP4 &&
+ if (ggml_openvino_is_gpu() && op->src[0] != nullptr && op->src[0]->type == GGML_TYPE_MXFP4 &&
mul_mat_id_requires_large_tmp(op)) {
return {false, "MUL_MAT_ID with MXFP4 weights requires large temporary on GPU"};
}
break;
}
case GGML_OP_ROPE: {
- const int32_t * op_params = op->op_params;
- const int n_dims = op_params[1];
- const int mode = op_params[2];
- const int64_t n_offs = op_params[15];
- if (mode != GGML_ROPE_TYPE_NORMAL && mode != GGML_ROPE_TYPE_NEOX && mode != GGML_ROPE_TYPE_IMROPE) {
- return {false, "ROPE with mode " + std::to_string(mode) + " is not supported"};
- }
- if (n_offs < 0 || (n_offs % 2) != 0) {
- return {false, "ROPE with invalid n_offs=" + std::to_string(n_offs)};
- }
- const int64_t head_dim = op->src[0]->ne[0];
- const int64_t rope_dims = n_dims == 0 ? head_dim : n_dims;
- if (rope_dims <= 0 || rope_dims + n_offs > head_dim || (rope_dims % 2) != 0) {
- return {false, "ROPE with n_dims=" + std::to_string(n_dims) + ", n_offs=" + std::to_string(n_offs) +
- ", head_dim=" + std::to_string(head_dim) + " is not supported"};
- }
- if (op->type != GGML_TYPE_F32 && op->type != GGML_TYPE_F16) {
- return {false, "ROPE with type " + std::string(ggml_type_name(op->type)) + " is not supported"};
- }
if (op->view_src != nullptr && !ggml_is_contiguous(op->src[0])) {
return {false, "ROPE on VIEW / non-contiguous input is not supported"};
}
- if (op->src[0]->ne[3] > 1) {
- // translate_rope's cos/sin tables cover one sequence only; ne[3] > 1 fails to broadcast.
- return {false, "ROPE with multiple sequences (ne[3]=" + std::to_string(op->src[0]->ne[3]) +
- ") is not supported"};
- }
- float freq_scale;
- float ext_factor;
- float attn_factor;
- memcpy(&freq_scale, op_params + 6, sizeof(float));
- memcpy(&ext_factor, op_params + 7, sizeof(float));
- memcpy(&attn_factor, op_params + 8, sizeof(float));
- if (mode == GGML_ROPE_TYPE_IMROPE &&
- (op->src[2] != nullptr || freq_scale != 1.0f || ext_factor != 0.0f || attn_factor != 1.0f)) {
- return {false, "IMROPE with freq_factors, freq_scale, ext_factor, or attn_factor is not supported"};
- }
break;
}
case GGML_OP_TRANSPOSE: {
@@ -1430,7 +1617,7 @@ static ggml_openvino_op_support is_op_supported_case(const ggml_tensor * op) {
break;
}
case GGML_OP_REPEAT: {
- if (ggml_openvino_get_device_name() == "GPU" && op->type == GGML_TYPE_BF16) {
+ if (ggml_openvino_is_gpu() && op->type == GGML_TYPE_BF16) {
return {false, "REPEAT with BF16 type is not supported on GPU"};
}
break;
@@ -1438,7 +1625,7 @@ static ggml_openvino_op_support is_op_supported_case(const ggml_tensor * op) {
case GGML_OP_GATED_DELTA_NET: {
// enable after https://github.com/openvinotoolkit/openvino/pull/35917 is included in OV release
// return true;
- // if (ggml_openvino_get_device_name() == "GPU" && op->src[0]->ne[2] > 1) {
+ // if (ggml_openvino_is_gpu() && op->src[0]->ne[2] > 1) {
// // CVS-186471
// return true;
// }
@@ -1469,6 +1656,84 @@ static ggml_openvino_op_support is_op_supported_case(const ggml_tensor * op) {
}
break;
}
+ case GGML_OP_CONV_2D:
+ case GGML_OP_CONV_2D_DW: {
+ if (op->src[0]->ne[0] <= 0 || op->src[0]->ne[1] <= 0) {
+ return {false, "CONV_2D kernel size must be positive"};
+ }
+ if (op->src[0]->op == GGML_OP_PERMUTE || op->src[1]->op == GGML_OP_PERMUTE) {
+ return {false, "CONV_2D with PERMUTE input is not supported"};
+ }
+ if (has_non_contiguous_view_input(op)) {
+ return {false, "CONV_2D with non-contiguous view input is not supported"};
+ }
+ const int32_t * params = op->op_params;
+ const int p0 = params[2];
+ const int p1 = params[3];
+ const int d0 = params[4];
+ const int d1 = params[5];
+ const int64_t dilated_kw = (int64_t) d0 * (op->src[0]->ne[0] - 1) + 1;
+ const int64_t dilated_kh = (int64_t) d1 * (op->src[0]->ne[1] - 1) + 1;
+ const int64_t padded_w = op->src[1]->ne[0] + 2 * p0;
+ const int64_t padded_h = op->src[1]->ne[1] + 2 * p1;
+ if (padded_w < dilated_kw || padded_h < dilated_kh) {
+ return {false, "CONV_2D padded input is smaller than kernel"};
+ }
+ break;
+ }
+ case GGML_OP_CONV_3D: {
+ if (op->src[0]->ne[0] <= 0 || op->src[0]->ne[1] <= 0 || op->src[0]->ne[2] <= 0) {
+ return {false, "CONV_3D kernel size must be positive"};
+ }
+ if (op->src[0]->op == GGML_OP_PERMUTE || op->src[1]->op == GGML_OP_PERMUTE) {
+ return {false, "CONV_3D with PERMUTE input is not supported"};
+ }
+ if (has_non_contiguous_view_input(op)) {
+ return {false, "CONV_3D with non-contiguous view input is not supported"};
+ }
+ const int32_t * params = op->op_params;
+ const int p0 = params[3];
+ const int p1 = params[4];
+ const int p2 = params[5];
+ const int d0 = params[6];
+ const int d1 = params[7];
+ const int d2 = params[8];
+ const int64_t dilated_kw = (int64_t) d0 * (op->src[0]->ne[0] - 1) + 1;
+ const int64_t dilated_kh = (int64_t) d1 * (op->src[0]->ne[1] - 1) + 1;
+ const int64_t dilated_kd = (int64_t) d2 * (op->src[0]->ne[2] - 1) + 1;
+ const int64_t padded_w = op->src[1]->ne[0] + 2 * p0;
+ const int64_t padded_h = op->src[1]->ne[1] + 2 * p1;
+ const int64_t padded_d = op->src[1]->ne[2] + 2 * p2;
+ if (padded_w < dilated_kw || padded_h < dilated_kh || padded_d < dilated_kd) {
+ return {false, "CONV_3D padded input is smaller than kernel"};
+ }
+ break;
+ }
+ case GGML_OP_CONV_TRANSPOSE_1D:
+ case GGML_OP_CONV_TRANSPOSE_2D: {
+ if (op->src[0]->ne[0] <= 0 || op->src[0]->ne[1] <= 0) {
+ return {false, "CONV_TRANSPOSE kernel size must be positive"};
+ }
+ if (op->src[0]->op == GGML_OP_PERMUTE || op->src[1]->op == GGML_OP_PERMUTE) {
+ return {false, "CONV_TRANSPOSE with PERMUTE input is not supported"};
+ }
+ if (has_non_contiguous_view_input(op)) {
+ return {false, "CONV_TRANSPOSE with non-contiguous view input is not supported"};
+ }
+ break;
+ }
+ case GGML_OP_IM2COL: {
+ if (op->src[0]->ne[0] <= 0 || op->src[0]->ne[1] <= 0) {
+ return {false, "IM2COL kernel size must be positive"};
+ }
+ break;
+ }
+ case GGML_OP_IM2COL_3D: {
+ if (op->src[0]->ne[0] <= 0 || op->src[0]->ne[1] <= 0 || op->src[0]->ne[2] <= 0) {
+ return {false, "IM2COL_3D kernel size must be positive"};
+ }
+ break;
+ }
default:
break;
}
@@ -1478,6 +1743,24 @@ static ggml_openvino_op_support is_op_supported_case(const ggml_tensor * op) {
static ggml_openvino_op_support ggml_backend_openvino_device_supports_op_impl(ggml_backend_dev_t dev, const ggml_tensor * op) {
GGML_ASSERT(dev->reg != nullptr);
+ ggml_backend_openvino_device_context * dev_ctx = (ggml_backend_openvino_device_context *) dev->context;
+ if (dev_ctx->ov_name != ggml_openvino_get_device_name()) {
+ // Data placed on a non-selected device (e.g. with -dev) can never run here; stop with a hint
+ // instead of the generic scheduler abort. Unallocated tensors (test-backend-ops) pass through.
+ for (int i = -1; i < GGML_MAX_SRC; i++) {
+ const ggml_tensor * t = i < 0 ? op : op->src[i];
+ ggml_backend_buffer_t buf = t == nullptr ? nullptr : (t->view_src ? t->view_src->buffer : t->buffer);
+ if (buf != nullptr &&
+ (ggml_backend_buft_is_openvino(buf->buft) || ggml_backend_buft_is_openvino_host(buf->buft)) &&
+ ((ggml_backend_openvino_buffer_type_context *) buf->buft->context)->device == dev_ctx->device) {
+ GGML_ABORT("%s is not the selected OpenVINO device (%s). The OpenVINO device is chosen with the "
+ "GGML_OPENVINO_DEVICE environment variable, not -dev: set GGML_OPENVINO_DEVICE=%s",
+ dev_ctx->name.c_str(), ggml_openvino_get_device_name().c_str(), dev_ctx->ov_name.c_str());
+ }
+ }
+ return {false, "device is not the selected OpenVINO device"};
+ }
+
static std::unordered_set<ggml_type> supported_types{
GGML_TYPE_F32, GGML_TYPE_F16, GGML_TYPE_BF16, GGML_TYPE_I64, GGML_TYPE_I32, GGML_TYPE_Q4_0,
GGML_TYPE_Q4_1, GGML_TYPE_Q4_K, GGML_TYPE_Q5_1, GGML_TYPE_Q5_K, GGML_TYPE_Q8_0, GGML_TYPE_Q6_K,
@@ -1527,8 +1810,9 @@ static ggml_openvino_op_support ggml_backend_openvino_device_supports_op_impl(gg
if (!supported) {
return {false, "unary op " + std::string(ggml_unary_op_name(ggml_get_unary_op(op))) + " has no op translator"};
}
- if (ggml_get_unary_op(op) == GGML_UNARY_OP_EXP && op->type == GGML_TYPE_F32) {
- return {false, "UNARY_EXP with F32 type is not supported"};
+ if (op->type == GGML_TYPE_F32 && (ggml_get_unary_op(op) == GGML_UNARY_OP_EXP ||
+ ggml_get_unary_op(op) == GGML_UNARY_OP_EXPM1)) {
+ return {false, "UNARY_EXP / UNARY_EXPM1 with F32 type is not supported"};
}
break;
}
@@ -1665,15 +1949,26 @@ GGML_BACKEND_API ggml_backend_reg_t ggml_backend_openvino_reg(void) {
std::lock_guard<std::mutex> lock(mutex);
if (!initialized) {
ggml_openvino_init();
+ const std::vector<std::string> openvino_devices = ggml_openvino_get_available_devices();
ggml_backend_openvino_reg_context * ctx = new ggml_backend_openvino_reg_context;
for (int i = 0; i < ggml_backend_openvino_get_device_count(); i++) {
ggml_backend_openvino_device_context * dev_ctx = new ggml_backend_openvino_device_context;
dev_ctx->device = i;
+ // Not the raw OpenVINO id: "CPU" would shadow the ggml CPU backend in ggml_backend_dev_by_name
dev_ctx->name = GGML_OPENVINO_NAME + std::to_string(i);
-
- dev_ctx->description = ov::get_openvino_version().description;
+ dev_ctx->ov_name = openvino_devices[i];
+ // The device is chosen with GGML_OPENVINO_DEVICE, not -dev, so show the value to set
+ dev_ctx->description = "GGML_OPENVINO_DEVICE=" + dev_ctx->ov_name +
+ (dev_ctx->ov_name == ggml_openvino_get_device_name() ? " (selected)" : "") +
+ " - " + ggml_openvino_get_device_description(dev_ctx->ov_name);
+ dev_ctx->total_memory = 0;
+ if (ov_device_has_prefix(dev_ctx->ov_name, "GPU")) {
+ ov_try_get_size_t_property(dev_ctx->ov_name, "GPU_DEVICE_TOTAL_MEM_SIZE", dev_ctx->total_memory);
+ } else if (ov_device_has_prefix(dev_ctx->ov_name, "NPU")) {
+ ov_try_get_size_t_property(dev_ctx->ov_name, "NPU_DEVICE_TOTAL_MEM_SIZE", dev_ctx->total_memory);
+ }
ggml_backend_dev_t dev =
new ggml_backend_device{/* .interface = */ ggml_backend_openvino_device_interface,
diff --git a/ggml/src/ggml-openvino/model-cache.cpp b/ggml/src/ggml-openvino/model-cache.cpp
index 3725fbd22..c8c5fb5ad 100644
--- a/ggml/src/ggml-openvino/model-cache.cpp
+++ b/ggml/src/ggml-openvino/model-cache.cpp
@@ -6,9 +6,13 @@
#include "ggml-openvino-extra.h"
#include <cerrno>
+#include <algorithm>
#include <cstdio>
#include <cstring>
#include <fstream>
+#include <iomanip>
+#include <limits>
+#include <sstream>
#include <openvino/core/version.hpp>
#include <string>
#include <sys/stat.h>
@@ -16,7 +20,19 @@
#include <vector>
#if defined(_WIN32)
+# define WIN32_LEAN_AND_MEAN
+# ifndef NOMINMAX
+# define NOMINMAX
+# endif
+# include <windows.h>
+# include <psapi.h>
# include <direct.h>
+# include <process.h>
+#else
+# include <unistd.h>
+#endif
+#ifdef __linux__
+# include <sys/sysmacros.h>
#endif
namespace {
@@ -37,10 +53,7 @@ inline uint64_t fnv1a_u64(uint64_t h, uint64_t v) {
constexpr uint64_t FNV_OFFSET = 0xcbf29ce484222325ull;
-// Bytes sampled from each end of a weight tensor for the sampled hash. The whole
-// model is never hashed (that would cost seconds every run); instead we sample a
-// bounded window from the head and tail of each weight's bytes. The manifest
-// re-verify (same sample) guards the residual collision risk.
+// Fallback when source-file identity is unavailable outside cache-only mode.
constexpr size_t WEIGHT_SAMPLE_BYTES = 4096;
// Is this src a model weight, mirroring create_weight_nodes()'s selection:
@@ -52,8 +65,7 @@ bool is_weight_src(const ggml_tensor * src) {
return src->buffer->usage == GGML_BACKEND_BUFFER_USAGE_WEIGHTS || ggml_is_quantized(src->type);
}
-// Per-weight sampled fingerprint: identity (name/shape/type) + a bounded byte
-// sample. Returns FNV offset basis if data is unavailable (kept deterministic).
+// Weight metadata and source identity; do not read repacked or unloaded buffers.
uint64_t weight_fingerprint(const ggml_tensor * t) {
uint64_t h = FNV_OFFSET;
h = fnv1a(h, t->name, strlen(t->name));
@@ -63,15 +75,7 @@ uint64_t weight_fingerprint(const ggml_tensor * t) {
h = fnv1a_u64(h, static_cast<uint64_t>(t->type));
const size_t nbytes = ggml_nbytes(t);
h = fnv1a_u64(h, nbytes);
- if (t->data != nullptr && nbytes > 0) {
- const size_t head = nbytes < WEIGHT_SAMPLE_BYTES ? nbytes : WEIGHT_SAMPLE_BYTES;
- h = fnv1a(h, t->data, head);
- if (nbytes > WEIGHT_SAMPLE_BYTES) {
- const size_t tail = nbytes < 2 * WEIGHT_SAMPLE_BYTES ? nbytes - WEIGHT_SAMPLE_BYTES : WEIGHT_SAMPLE_BYTES;
- h = fnv1a(h, static_cast<const uint8_t *>(t->data) + (nbytes - tail), tail);
- }
- }
- return h;
+ return fnv1a_u64(h, ggml_backend_openvino_weight_fingerprint(t));
}
// Walk the cgraph and invoke fn(weight_tensor) for each distinct weight, in node
@@ -156,12 +160,162 @@ bool make_dirs(const std::string & path) {
} // namespace
+bool ggml_openvino_model_cache_only() {
+ return ggml_openvino_getenv_int("GGML_OPENVINO_COMPILED_MODEL_CACHE_ONLY") != 0;
+}
+
+static const char * cache_settings[] = {
+ "GGML_OPENVINO_REQUANT_KQUANT",
+ "GGML_OPENVINO_NATIVE_SOFTPLUS",
+ "GGML_OPENVINO_DISABLE_KV_SLICE",
+ "GGML_OPENVINO_MANUAL_GQA_ATTN",
+ "GGML_OPENVINO_STATEFUL_EXECUTION",
+ "GGML_OPENVINO_DISABLE_KV_STATE_RELAYOUT",
+ "GGML_OPENVINO_DISABLE_REMOTE_OUTPUTS",
+ "GGML_OPENVINO_REDUCE_COMPILE_MEM",
+ "GGML_OPENVINO_MEMORY_OPTIMIZE",
+ "GGML_OPENVINO_PROFILING",
+};
+
+void ggml_openvino_model_cache_init() {
+ const bool cache_only = ggml_openvino_model_cache_only();
+ const std::string dir = ggml_openvino_model_cache_dir();
+ if (dir.empty()) {
+ if (cache_only) {
+ GGML_ABORT("ggml-openvino: cache-only mode requires GGML_OPENVINO_COMPILED_MODEL_CACHE_DIR");
+ }
+ return;
+ }
+ if (cache_only && (ggml_openvino_is_npu() || ggml_openvino_getenv_int("GGML_OPENVINO_FORCE_STATIC") ||
+ ggml_openvino_getenv_int("GGML_OPENVINO_DISABLE_CACHE") ||
+ ggml_openvino_getenv_int("GGML_OPENVINO_ENABLE_FALLBACK"))) {
+ GGML_ABORT("ggml-openvino: cache-only mode requires dynamic CPU/GPU execution with caching and without fallback");
+ }
+#if !defined(__linux__) && !defined(_WIN32)
+ if (cache_only) {
+ GGML_ABORT("ggml-openvino: cache-only mmap identification requires Linux or Windows");
+ }
+#endif
+ if (cache_only) {
+ auto & config = ggml_openvino_get_device_config();
+ config.environment_variables.erase("GGML_OPENVINO_SPILL_DIR");
+ config.environment_variables["GGML_OPENVINO_RELEASE_WEIGHTS"] = "0";
+ }
+}
+
+uint64_t ggml_openvino_source_fingerprint(const void * data, size_t size, std::vector<ggml_openvino_source_mapping> & mappings) {
+ const uintptr_t address = reinterpret_cast<uintptr_t>(data);
+ auto contains = [&](const ggml_openvino_source_mapping & m) {
+ return address >= m.begin && address < m.end && size <= m.end - address;
+ };
+ auto fingerprint = [&](const ggml_openvino_source_mapping & m) {
+ return fnv1a_u64(m.identity, m.offset + address - m.begin);
+ };
+ for (const auto & m : mappings) {
+ if (contains(m)) {
+ return fingerprint(m);
+ }
+ }
+#ifdef __linux__
+ std::ifstream maps("/proc/self/maps");
+ std::string line;
+ while (std::getline(maps, line)) {
+ unsigned long long begin, end, offset, inode;
+ unsigned int dev_major, dev_minor;
+ char permissions[5];
+ int path_start = 0;
+ if (sscanf(line.c_str(), "%llx-%llx %4s %llx %x:%x %llu %n", &begin, &end, permissions,
+ &offset, &dev_major, &dev_minor, &inode, &path_start) != 7 || inode == 0) {
+ continue;
+ }
+ ggml_openvino_source_mapping m{uintptr_t(begin), uintptr_t(end), offset, FNV_OFFSET};
+ if (!contains(m)) {
+ continue;
+ }
+ struct stat st;
+ const std::string path = line.substr(path_start);
+ if (stat(path.c_str(), &st) != 0 || !S_ISREG(st.st_mode) || uint64_t(st.st_ino) != inode ||
+ major(st.st_dev) != dev_major || minor(st.st_dev) != dev_minor) {
+ break;
+ }
+ m.identity = fnv1a_u64(m.identity, st.st_dev);
+ m.identity = fnv1a_u64(m.identity, st.st_ino);
+ m.identity = fnv1a_u64(m.identity, st.st_size);
+ m.identity = fnv1a_u64(m.identity, st.st_mtim.tv_sec);
+ m.identity = fnv1a_u64(m.identity, st.st_mtim.tv_nsec);
+ m.identity = fnv1a_u64(m.identity, st.st_ctim.tv_sec);
+ m.identity = fnv1a_u64(m.identity, st.st_ctim.tv_nsec);
+ mappings.push_back(m);
+ return fingerprint(m);
+ }
+#elif defined(_WIN32)
+ MEMORY_BASIC_INFORMATION memory;
+ if (VirtualQuery(data, &memory, sizeof(memory)) == sizeof(memory) && memory.Type == MEM_MAPPED) {
+ std::wstring name(MAX_PATH, L'\0');
+ DWORD length = 0;
+ while (name.size() <= 32768) {
+ length = GetMappedFileNameW(GetCurrentProcess(), const_cast<void *>(data), name.data(), static_cast<DWORD>(name.size()));
+ if (length == 0 || length < name.size() - 1) {
+ break;
+ }
+ name.resize(name.size() * 2);
+ }
+ if (length > 0 && length < name.size() - 1) {
+ name.resize(length);
+ const std::wstring path = L"\\\\?\\GLOBALROOT" + name;
+ HANDLE file = CreateFileW(path.c_str(), FILE_READ_ATTRIBUTES,
+ FILE_SHARE_READ | FILE_SHARE_WRITE | FILE_SHARE_DELETE, nullptr,
+ OPEN_EXISTING, FILE_ATTRIBUTE_NORMAL, nullptr);
+ if (file != INVALID_HANDLE_VALUE) {
+ BY_HANDLE_FILE_INFORMATION info;
+ FILE_BASIC_INFO basic;
+ const bool valid = GetFileInformationByHandle(file, &info) &&
+ GetFileInformationByHandleEx(file, FileBasicInfo, &basic, sizeof(basic));
+ CloseHandle(file);
+ if (valid) {
+ const uint64_t file_size = (uint64_t(info.nFileSizeHigh) << 32) | info.nFileSizeLow;
+ const uintptr_t begin = reinterpret_cast<uintptr_t>(memory.AllocationBase);
+ if (file_size <= std::numeric_limits<uintptr_t>::max() - begin) {
+ ggml_openvino_source_mapping m{begin, begin + static_cast<uintptr_t>(file_size), 0,
+ fnv1a(FNV_OFFSET, "win32", 5)};
+ if (contains(m)) {
+ m.identity = fnv1a_u64(m.identity, info.dwVolumeSerialNumber);
+ m.identity = fnv1a_u64(m.identity, (uint64_t(info.nFileIndexHigh) << 32) | info.nFileIndexLow);
+ m.identity = fnv1a_u64(m.identity, file_size);
+ m.identity = fnv1a_u64(m.identity, (uint64_t(info.ftLastWriteTime.dwHighDateTime) << 32) |
+ info.ftLastWriteTime.dwLowDateTime);
+ m.identity = fnv1a_u64(m.identity, static_cast<uint64_t>(basic.ChangeTime.QuadPart));
+ mappings.push_back(m);
+ return fingerprint(m);
+ }
+ }
+ }
+ }
+ }
+ }
+#endif
+ if (ggml_openvino_model_cache_only()) {
+ GGML_ABORT("ggml-openvino: could not identify mapped GGUF weight; use --load-mode mmap");
+ }
+ uint64_t h = FNV_OFFSET;
+ const size_t head = std::min(size, WEIGHT_SAMPLE_BYTES);
+ h = fnv1a(h, data, head);
+ if (size > head) {
+ const size_t tail = std::min(size - head, WEIGHT_SAMPLE_BYTES);
+ h = fnv1a(h, static_cast<const uint8_t *>(data) + size - tail, tail);
+ }
+ return h;
+}
+
std::string ggml_openvino_model_cache_dir() {
const char * dir = ggml_openvino_getenv_str("GGML_OPENVINO_COMPILED_MODEL_CACHE_DIR");
if (!dir || strlen(dir) == 0) {
return std::string();
}
std::string path(dir);
+ if (ggml_openvino_model_cache_only()) {
+ return path;
+ }
// Create the cache directory (and parents) on first use so callers don't
// have to pre-create it; a missing dir would otherwise silently disable the
// cache (manifest/blob writes fail with no directory to write into).
@@ -173,13 +327,36 @@ std::string ggml_openvino_model_cache_dir() {
return path;
}
+std::string ggml_openvino_model_cache_temp_path(const std::string & path) {
+#ifdef _WIN32
+ const int pid = _getpid();
+#else
+ const int pid = getpid();
+#endif
+ return path + ".tmp." + std::to_string(pid) + "." + std::to_string(ggml_time_us());
+}
+
uint64_t ggml_openvino_model_fingerprint(const ggml_cgraph * cgraph,
const std::string & device,
bool fa,
const int32_t * rope_params,
int rope_len,
- uint64_t extra_cfg) {
+ uint64_t extra_cfg,
+ const std::string & graph_signature) {
uint64_t h = FNV_OFFSET;
+ h = fnv1a_u64(h, 2);
+ h = fnv1a(h, graph_signature.data(), graph_signature.size());
+ for (const char * name : cache_settings) {
+ const char * value = ggml_openvino_getenv_str(name, "");
+ h = fnv1a(h, value, strlen(value) + 1);
+ }
+ if (const char * debug_nodes = ggml_openvino_getenv_str("GGML_OPENVINO_DEBUG_NODE")) {
+ h = fnv1a(h, "GGML_OPENVINO_DEBUG_NODE", sizeof("GGML_OPENVINO_DEBUG_NODE"));
+ h = fnv1a(h, debug_nodes, strlen(debug_nodes) + 1);
+ }
+ if (ggml_openvino_is_gpu() && ggml_openvino_getenv_int("GGML_OPENVINO_MOE_OP", 1) == 0) {
+ h = fnv1a(h, "GGML_OPENVINO_MOE_OP=0", sizeof("GGML_OPENVINO_MOE_OP=0"));
+ }
// Topology: node count + each node's op and name (cheap, and distinguishes
// graphs that share weights but differ structurally).
@@ -193,7 +370,7 @@ uint64_t ggml_openvino_model_fingerprint(const ggml_cgraph * cgraph,
// Weights: the model identity.
for_each_weight(cgraph, [&](const ggml_tensor * t) { h = fnv1a_u64(h, weight_fingerprint(t)); });
- // Config that changes the produced blob.
+ // Device, model parameters, and backend configuration.
h = fnv1a(h, device.data(), device.size());
h = fnv1a_u64(h, fa ? 1u : 0u);
if (rope_params && rope_len > 0) {
@@ -216,7 +393,9 @@ std::string ggml_openvino_model_cache_manifest_path(const std::string & dir, uin
bool ggml_openvino_model_cache_write_manifest(const std::string & path,
const ggml_cgraph * cgraph,
- uint64_t fingerprint) {
+ uint64_t fingerprint,
+ const std::vector<std::string> & inputs,
+ const std::vector<std::string> & outputs) {
std::ofstream f(path, std::ios::trunc);
if (!f.is_open()) {
return false;
@@ -227,12 +406,21 @@ bool ggml_openvino_model_cache_write_manifest(const std::string & path,
f << t->name << " " << t->ne[0] << " " << t->ne[1] << " " << t->ne[2] << " " << t->ne[3] << " "
<< static_cast<int>(t->type) << " " << hex64(weight_fingerprint(t)) << "\n";
});
+ f << "ports\n";
+ for (const auto * names : { &inputs, &outputs }) {
+ f << names->size() << '\n';
+ for (const auto & name : *names) {
+ f << std::quoted(name) << '\n';
+ }
+ }
return f.good();
}
bool ggml_openvino_model_cache_verify_manifest(const std::string & path,
const ggml_cgraph * cgraph,
- uint64_t fingerprint) {
+ uint64_t fingerprint,
+ std::vector<std::string> & inputs,
+ std::vector<std::string> & outputs) {
std::ifstream f(path);
if (!f.is_open()) {
return false;
@@ -260,7 +448,7 @@ bool ggml_openvino_model_cache_verify_manifest(const std::string & path,
size_t idx = 0;
std::string line;
std::getline(f, line); // consume rest of ov_version line
- while (std::getline(f, line)) {
+ while (idx < expected.size() && std::getline(f, line)) {
if (line.empty()) {
continue;
}
@@ -269,5 +457,28 @@ bool ggml_openvino_model_cache_verify_manifest(const std::string & path,
}
++idx;
}
- return idx == expected.size();
+ if (idx != expected.size()) {
+ return false;
+ }
+ if (!std::getline(f, line)) {
+ return true;
+ }
+ if (line != "ports") {
+ return false;
+ }
+ for (auto * names : { &inputs, &outputs }) {
+ size_t count;
+ if (!(f >> count) || count > 100000) {
+ return false;
+ }
+ for (size_t i = 0; i < count; ++i) {
+ std::string name;
+ if (!(f >> std::quoted(name))) {
+ return false;
+ }
+ names->push_back(name);
+ }
+ }
+ f >> std::ws;
+ return f.eof();
}
diff --git a/ggml/src/ggml-openvino/model-cache.h b/ggml/src/ggml-openvino/model-cache.h
index 15967b962..ee69fca83 100644
--- a/ggml/src/ggml-openvino/model-cache.h
+++ b/ggml/src/ggml-openvino/model-cache.h
@@ -1,38 +1,39 @@
#pragma once
-// Frontend-level compiled-model cache (GGML_OPENVINO_COMPILED_MODEL_CACHE_DIR).
-//
-// The OpenVINO plugin's own ov::cache_dir caches the compiled blob keyed by the
-// *OV model*, but producing that model still runs the full frontend every time:
-// weight requantization (incl. the large token_embd F32 transient) and the
-// ggml->OV graph conversion. This cache keys off a fingerprint computed directly
-// from the ggml cgraph, so a hit skips requant + convert + compile entirely and
-// instead imports a previously exported CompiledModel blob.
-//
-// Opt-in and independent from GGML_OPENVINO_CACHE_DIR. Default off.
+// Compiled blobs include weights. Cache-only execution skips weight uploads and graph compilation.
#include "ggml.h"
#include <cstdint>
#include <string>
+#include <vector>
-// Returns the compiled-model cache directory from GGML_OPENVINO_COMPILED_MODEL_CACHE_DIR,
-// or empty if unset/disabled. When empty, callers must not use the cache.
+bool ggml_openvino_model_cache_only();
+void ggml_openvino_model_cache_init();
+
+struct ggml_openvino_source_mapping {
+ uintptr_t begin;
+ uintptr_t end;
+ uint64_t offset;
+ uint64_t identity;
+};
+
+// Identify mmap weights without reading their pages. Cache mappings for one buffer lifetime.
+uint64_t ggml_openvino_source_fingerprint(const void * data, size_t size, std::vector<ggml_openvino_source_mapping> & mappings);
+uint64_t ggml_backend_openvino_weight_fingerprint(const ggml_tensor * tensor);
+
+// Returns the compiled-model cache directory, or empty if unset.
std::string ggml_openvino_model_cache_dir();
+std::string ggml_openvino_model_cache_temp_path(const std::string & path);
-// Compute a stable 64-bit fingerprint identifying the model+config that a cgraph
-// would compile to. Combines graph topology, a sampled hash of every weight
-// tensor (name/shape/dtype + bounded byte sample), and the config that changes
-// the produced blob (device, flash-attention, rope params, the compile-memory
-// flags, stateful, and the OpenVINO version). `device` is the resolved device
-// string; `fa` is the flash-attention flag; `rope_params`/`rope_len` cover the
-// model's rope configuration; `extra_cfg` folds in any other blob-affecting bits.
+// Hash graph structure, source weight identities, configuration, and OpenVINO version.
uint64_t ggml_openvino_model_fingerprint(const ggml_cgraph * cgraph,
const std::string & device,
bool fa,
const int32_t * rope_params,
int rope_len,
- uint64_t extra_cfg);
+ uint64_t extra_cfg,
+ const std::string & graph_signature);
// Path to the compiled-blob file for a fingerprint (<dir>/<hex>.blob).
std::string ggml_openvino_model_cache_blob_path(const std::string & dir, uint64_t fingerprint);
@@ -41,16 +42,16 @@ std::string ggml_openvino_model_cache_blob_path(const std::string & dir, uint64_
// fingerprints, used to re-verify a hit before trusting the blob.
std::string ggml_openvino_model_cache_manifest_path(const std::string & dir, uint64_t fingerprint);
-// Write/read the manifest. The manifest is a newline-separated list of
-// "name ne0 ne1 ne2 ne3 type sample_hash" lines plus a header line with the
-// fingerprint and OV version. Returns false on I/O error.
+// Record weight metadata and source identities. Returns false on I/O error.
bool ggml_openvino_model_cache_write_manifest(const std::string & path,
const ggml_cgraph * cgraph,
- uint64_t fingerprint);
+ uint64_t fingerprint,
+ const std::vector<std::string> & inputs,
+ const std::vector<std::string> & outputs);
-// Verify that the cgraph's weights still match the stored manifest (guards the
-// sampled-hash collision risk: a blob is only trusted if every weight's
-// name/shape/type/sample-hash matches what was cached). Returns true on match.
+// Require all weight metadata and source identities to match the manifest.
bool ggml_openvino_model_cache_verify_manifest(const std::string & path,
const ggml_cgraph * cgraph,
- uint64_t fingerprint);
+ uint64_t fingerprint,
+ std::vector<std::string> & inputs,
+ std::vector<std::string> & outputs);
diff --git a/ggml/src/ggml-openvino/openvino/node_context.h b/ggml/src/ggml-openvino/openvino/node_context.h
index f1ea0e4f0..c653bc0e6 100644
--- a/ggml/src/ggml-openvino/openvino/node_context.h
+++ b/ggml/src/ggml-openvino/openvino/node_context.h
@@ -33,6 +33,8 @@ public:
const std::vector<std::string> & get_input_names() const { return m_input_names; }
+ const std::vector<std::string> & get_output_names() const { return m_output_names; }
+
size_t get_input_size() const override { return m_decoder->get_input_size(m_node_idx); }
ov::element::Type get_input_type(size_t index) const {
@@ -120,7 +122,10 @@ public:
auto view_it = m_tensor_map->find(m_input_names[idx]);
if (!base_name.empty() && view_it != m_tensor_map->end()) {
auto base_it = m_tensor_map->find(base_name);
- if (base_it != m_tensor_map->end() &&
+ // A multi-output translator can publish a VIEW directly without materializing
+ // its packed parent (GatedDeltaNet attention/state). In that case the VIEW is the
+ // authoritative value. The node comparison retains the existing resolved-VIEW path.
+ if (base_it == m_tensor_map->end() ||
view_it->second.get_node_shared_ptr() != base_it->second.get_node_shared_ptr()) {
return view_it->second;
}
diff --git a/ggml/src/ggml-openvino/openvino/op/add.cpp b/ggml/src/ggml-openvino/openvino/op/add.cpp
index a45520d92..84bc85618 100644
--- a/ggml/src/ggml-openvino/openvino/op/add.cpp
+++ b/ggml/src/ggml-openvino/openvino/op/add.cpp
@@ -27,10 +27,18 @@ OutputVector translate_add(const NodeContext & context) {
auto base_name = context.get_view_input_src_name(1, view_size - 1);
auto base = context.get_input(base_name);
+ // Stateful models drop the leading batch dim, so the base is rank 3 and both axes
+ // below shift down by one. Take them from the actual rank: the expert axis is always
+ // second from last, and the token axis is re-added just before it.
+ const auto base_rank = base.get_partial_shape().rank();
+ FRONT_END_OP_CONVERSION_CHECK(base_rank.is_static() && base_rank.get_length() >= 3,
+ "MoE expert sum needs a static rank of at least 3");
+ const int64_t rank = base_rank.get_length();
+
auto reduced = std::make_shared<ov::op::v1::ReduceSum>(
- base, ov::op::v0::Constant::create(ov::element::i64, ov::Shape{1}, {2}), false);
- auto res =
- std::make_shared<ov::op::v0::Unsqueeze>(reduced, ov::op::v0::Constant::create(ov::element::i64, {1}, {1}));
+ base, ov::op::v0::Constant::create(ov::element::i64, ov::Shape{1}, {rank - 2}), false);
+ auto res = std::make_shared<ov::op::v0::Unsqueeze>(
+ reduced, ov::op::v0::Constant::create(ov::element::i64, {1}, {rank - 3}));
return rename_outputs_with_suffix({res}, context.get_name());
}
diff --git a/ggml/src/ggml-openvino/openvino/op/argsort.cpp b/ggml/src/ggml-openvino/openvino/op/argsort.cpp
index bb8344af8..dd0872516 100644
--- a/ggml/src/ggml-openvino/openvino/op/argsort.cpp
+++ b/ggml/src/ggml-openvino/openvino/op/argsort.cpp
@@ -32,10 +32,14 @@ OutputVector translate_argsort(const NodeContext & context) {
FRONT_END_OP_CONVERSION_CHECK(false, "Unsupported GGML_OP_ARGSORT order: ", order);
}
- auto k = std::make_shared<ov::op::v0::Squeeze>(get_dimensions(input.get_node_shared_ptr(), {3}),
+ // Stateful models drop the leading size-1 batch dim, so the expert axis is 2 there
+ // instead of 3 (same rank-3-vs-rank-4 split as get_rows.cpp / process_view_input).
+ const int axis = (context.is_stateful() && input.get_partial_shape().rank() == 3) ? 2 : 3;
+
+ auto k = std::make_shared<ov::op::v0::Squeeze>(get_dimensions(input.get_node_shared_ptr(), {axis}),
ov::op::v0::Constant::create(ov::element::i64, {1}, {0}));
- auto topk = std::make_shared<ov::op::v11::TopK>(input, k, 3, mode, ov::op::v11::TopK::SortType::SORT_VALUES,
+ auto topk = std::make_shared<ov::op::v11::TopK>(input, k, axis, mode, ov::op::v11::TopK::SortType::SORT_VALUES,
context.get_output_type(), false);
return rename_outputs_with_suffix({topk->output(1)}, context.get_name());
diff --git a/ggml/src/ggml-openvino/openvino/op/conv.cpp b/ggml/src/ggml-openvino/openvino/op/conv.cpp
new file mode 100644
index 000000000..94173dc4f
--- /dev/null
+++ b/ggml/src/ggml-openvino/openvino/op/conv.cpp
@@ -0,0 +1,233 @@
+#include "../node_context.h"
+#include "../op_table.h"
+#include "../utils.h"
+
+#include <openvino/op/constant.hpp>
+#include <openvino/op/convert.hpp>
+#include <openvino/op/convolution.hpp>
+#include <openvino/op/group_conv.hpp>
+#include <openvino/op/reshape.hpp>
+#include <openvino/op/unsqueeze.hpp>
+
+namespace ov {
+namespace frontend {
+namespace ggml {
+namespace op {
+
+OutputVector translate_conv_2d(const NodeContext & context) {
+ num_inputs_check(context, 2, 2);
+
+ ov::Output<Node> kernel = process_view_input_new(context, 0);
+ ov::Output<Node> input = process_view_input_new(context, 1);
+
+ if (kernel.get_element_type() != input.get_element_type()) {
+ kernel = std::make_shared<ov::op::v0::Convert>(kernel, input.get_element_type());
+ }
+
+ const int32_t * params = context.get_output_op_params();
+ const int s0 = params[0];
+ const int s1 = params[1];
+ const int p0 = params[2];
+ const int p1 = params[3];
+ const int d0 = params[4];
+ const int d1 = params[5];
+
+ ov::Strides strides{static_cast<size_t>(s1), static_cast<size_t>(s0)};
+ ov::CoordinateDiff pads_begin{static_cast<ptrdiff_t>(p1), static_cast<ptrdiff_t>(p0)};
+ ov::CoordinateDiff pads_end{static_cast<ptrdiff_t>(p1), static_cast<ptrdiff_t>(p0)};
+ ov::Strides dilations{static_cast<size_t>(d1), static_cast<size_t>(d0)};
+
+ ov::Output<Node> res = std::make_shared<ov::op::v1::Convolution>(
+ input, kernel, strides, pads_begin, pads_end, dilations, ov::op::PadType::EXPLICIT);
+
+ const auto output_type = context.get_output_type();
+ if (res.get_element_type() != output_type) {
+ res = std::make_shared<ov::op::v0::Convert>(res, output_type);
+ }
+
+ return rename_outputs_with_suffix({res}, context.get_name());
+}
+
+OutputVector translate_conv_2d_dw(const NodeContext & context) {
+ num_inputs_check(context, 2, 2);
+
+ ov::Output<Node> kernel = process_view_input_new(context, 0);
+ ov::Output<Node> input = process_view_input_new(context, 1);
+
+ if (kernel.get_element_type() != input.get_element_type()) {
+ kernel = std::make_shared<ov::op::v0::Convert>(kernel, input.get_element_type());
+ }
+
+ const int32_t * params = context.get_output_op_params();
+ const int s0 = params[0];
+ const int s1 = params[1];
+ const int p0 = params[2];
+ const int p1 = params[3];
+ const int d0 = params[4];
+ const int d1 = params[5];
+
+ // Reshape kernel from [C, 1, KH, KW] to [C, 1, 1, KH, KW] for 2D GroupConvolution
+ auto unsqueeze_axis = ov::op::v0::Constant::create(ov::element::i64, ov::Shape{1}, {1});
+ auto kernel_5d = std::make_shared<ov::op::v0::Unsqueeze>(kernel, unsqueeze_axis);
+
+ ov::Strides strides{static_cast<size_t>(s1), static_cast<size_t>(s0)};
+ ov::CoordinateDiff pads_begin{static_cast<ptrdiff_t>(p1), static_cast<ptrdiff_t>(p0)};
+ ov::CoordinateDiff pads_end{static_cast<ptrdiff_t>(p1), static_cast<ptrdiff_t>(p0)};
+ ov::Strides dilations{static_cast<size_t>(d1), static_cast<size_t>(d0)};
+
+ ov::Output<Node> res = std::make_shared<ov::op::v1::GroupConvolution>(
+ input, kernel_5d, strides, pads_begin, pads_end, dilations, ov::op::PadType::EXPLICIT);
+
+ const auto output_type = context.get_output_type();
+ if (res.get_element_type() != output_type) {
+ res = std::make_shared<ov::op::v0::Convert>(res, output_type);
+ }
+
+ return rename_outputs_with_suffix({res}, context.get_name());
+}
+
+OutputVector translate_conv_transpose_1d(const NodeContext & context) {
+ num_inputs_check(context, 2, 2);
+
+ ov::Output<Node> kernel = process_view_input_new(context, 0);
+ ov::Output<Node> input = process_view_input_new(context, 1);
+
+ if (kernel.get_element_type() != input.get_element_type()) {
+ kernel = std::make_shared<ov::op::v0::Convert>(kernel, input.get_element_type());
+ }
+
+ const int32_t * params = context.get_output_op_params();
+ const int s0 = params[0];
+ const int p0 = params[1];
+ const int d0 = params[2];
+
+ const auto kernel_shape = context.get_input_shape(0).to_shape(); // [1, Cin, Cout, K]
+ const int64_t Cin = kernel_shape[1];
+ const int64_t Cout = kernel_shape[2];
+ const int64_t K = kernel_shape[3];
+
+ const auto input_shape = context.get_input_shape(1).to_shape(); // [1, N, Cin, L]
+ const int64_t N = input_shape[0] * input_shape[1];
+ const int64_t L = input_shape[3];
+
+ auto kernel_3d = std::make_shared<ov::op::v1::Reshape>(
+ kernel, ov::op::v0::Constant::create(ov::element::i64, {3}, {Cin, Cout, K}), false);
+ auto input_3d = std::make_shared<ov::op::v1::Reshape>(
+ input, ov::op::v0::Constant::create(ov::element::i64, {3}, {N, Cin, L}), false);
+
+ ov::Strides strides{static_cast<size_t>(s0)};
+ ov::CoordinateDiff pads_begin{static_cast<ptrdiff_t>(p0)};
+ ov::CoordinateDiff pads_end{static_cast<ptrdiff_t>(p0)};
+ ov::Strides dilations{static_cast<size_t>(d0)};
+
+ auto conv_tr = std::make_shared<ov::op::v1::ConvolutionBackpropData>(
+ input_3d, kernel_3d, strides, pads_begin, pads_end, dilations);
+
+ const auto out_shape = context.get_output_shape().to_shape();
+ auto out_shape_const = ov::op::v0::Constant::create(
+ ov::element::i64, {4}, {static_cast<int64_t>(out_shape[0]), static_cast<int64_t>(out_shape[1]),
+ static_cast<int64_t>(out_shape[2]), static_cast<int64_t>(out_shape[3])});
+ ov::Output<Node> res = std::make_shared<ov::op::v1::Reshape>(conv_tr, out_shape_const, false);
+
+ const auto output_type = context.get_output_type();
+ if (res.get_element_type() != output_type) {
+ res = std::make_shared<ov::op::v0::Convert>(res, output_type);
+ }
+
+ return rename_outputs_with_suffix({res}, context.get_name());
+}
+
+OutputVector translate_conv_transpose_2d(const NodeContext & context) {
+ num_inputs_check(context, 2, 2);
+
+ ov::Output<Node> kernel = process_view_input_new(context, 0);
+ ov::Output<Node> input = process_view_input_new(context, 1);
+
+ if (kernel.get_element_type() != input.get_element_type()) {
+ kernel = std::make_shared<ov::op::v0::Convert>(kernel, input.get_element_type());
+ }
+
+ const int32_t * params = context.get_output_op_params();
+ const int stride = params[0];
+
+ ov::Strides strides{static_cast<size_t>(stride), static_cast<size_t>(stride)};
+ ov::CoordinateDiff pads_begin{0, 0};
+ ov::CoordinateDiff pads_end{0, 0};
+ ov::Strides dilations{1, 1};
+
+ ov::Output<Node> res = std::make_shared<ov::op::v1::ConvolutionBackpropData>(
+ input, kernel, strides, pads_begin, pads_end, dilations);
+
+ const auto output_type = context.get_output_type();
+ if (res.get_element_type() != output_type) {
+ res = std::make_shared<ov::op::v0::Convert>(res, output_type);
+ }
+
+ return rename_outputs_with_suffix({res}, context.get_name());
+}
+
+OutputVector translate_conv_3d(const NodeContext & context) {
+ num_inputs_check(context, 2, 2);
+
+ ov::Output<Node> kernel = process_view_input_new(context, 0);
+ ov::Output<Node> input = process_view_input_new(context, 1);
+
+ if (kernel.get_element_type() != input.get_element_type()) {
+ kernel = std::make_shared<ov::op::v0::Convert>(kernel, input.get_element_type());
+ }
+
+ const int32_t * params = context.get_output_op_params();
+ const int s0 = params[0];
+ const int s1 = params[1];
+ const int s2 = params[2];
+ const int p0 = params[3];
+ const int p1 = params[4];
+ const int p2 = params[5];
+ const int d0 = params[6];
+ const int d1 = params[7];
+ const int d2 = params[8];
+ const int c = params[9];
+ const int n = params[10];
+ const int oc = params[11];
+
+ const auto kshape = context.get_input_shape(0).to_shape(); // [c*oc, KD, KH, KW]
+ const int64_t KD = kshape[1];
+ const int64_t KH = kshape[2];
+ const int64_t KW = kshape[3];
+
+ const auto inshape = context.get_input_shape(1).to_shape(); // [c*n, ID, IH, IW]
+ const int64_t ID = inshape[1];
+ const int64_t IH = inshape[2];
+ const int64_t IW = inshape[3];
+
+ auto kernel_5d = std::make_shared<ov::op::v1::Reshape>(
+ kernel, ov::op::v0::Constant::create(ov::element::i64, {5}, {static_cast<int64_t>(oc), static_cast<int64_t>(c), KD, KH, KW}), false);
+ auto input_5d = std::make_shared<ov::op::v1::Reshape>(
+ input, ov::op::v0::Constant::create(ov::element::i64, {5}, {static_cast<int64_t>(n), static_cast<int64_t>(c), ID, IH, IW}), false);
+
+ ov::Strides strides{static_cast<size_t>(s2), static_cast<size_t>(s1), static_cast<size_t>(s0)};
+ ov::CoordinateDiff pads_begin{static_cast<ptrdiff_t>(p2), static_cast<ptrdiff_t>(p1), static_cast<ptrdiff_t>(p0)};
+ ov::CoordinateDiff pads_end{static_cast<ptrdiff_t>(p2), static_cast<ptrdiff_t>(p1), static_cast<ptrdiff_t>(p0)};
+ ov::Strides dilations{static_cast<size_t>(d2), static_cast<size_t>(d1), static_cast<size_t>(d0)};
+
+ auto conv = std::make_shared<ov::op::v1::Convolution>(
+ input_5d, kernel_5d, strides, pads_begin, pads_end, dilations, ov::op::PadType::EXPLICIT);
+
+ const auto out_shape = context.get_output_shape().to_shape(); // [oc*n, OD, OH, OW]
+ auto out_shape_const = ov::op::v0::Constant::create(
+ ov::element::i64, {4}, {static_cast<int64_t>(out_shape[0]), static_cast<int64_t>(out_shape[1]),
+ static_cast<int64_t>(out_shape[2]), static_cast<int64_t>(out_shape[3])});
+ ov::Output<Node> res = std::make_shared<ov::op::v1::Reshape>(conv, out_shape_const, false);
+
+ const auto output_type = context.get_output_type();
+ if (res.get_element_type() != output_type) {
+ res = std::make_shared<ov::op::v0::Convert>(res, output_type);
+ }
+
+ return rename_outputs_with_suffix({res}, context.get_name());
+}
+
+} // namespace op
+} // namespace ggml
+} // namespace frontend
+} // namespace ov
diff --git a/ggml/src/ggml-openvino/openvino/op/cpy.cpp b/ggml/src/ggml-openvino/openvino/op/cpy.cpp
index 6f1e34779..2c7299a89 100644
--- a/ggml/src/ggml-openvino/openvino/op/cpy.cpp
+++ b/ggml/src/ggml-openvino/openvino/op/cpy.cpp
@@ -90,10 +90,24 @@ OutputVector translate_cpy(const NodeContext & context) {
return {context.get_input(1)};
}
}
+ // op_case 7/8/9 are the single-slot variants; op_case 10 writes native GDN state into a
+ // multi-slot cache without rollback snapshots.
+ const bool single_slot_assign = op_case >= 7 && op_case <= 9;
+ const bool direct_gdn_state = op_case == 7 || op_case == 10;
+ int writeback_case = op_case;
+ if (op_case == 10) {
+ writeback_case = 1;
+ } else if (single_slot_assign) {
+ writeback_case = op_case - 6;
+ }
const std::string slot_begin_name = "rs_slot_begin_" + context.get_name();
- const bool slice_assign =
- context.has_input(slot_begin_name) && !context.is_stateful() && (op_case >= 1 && op_case <= 3);
+ const bool slice_assign = writeback_case >= 1 && writeback_case <= 3 &&
+ (single_slot_assign || context.has_input(slot_begin_name));
if (slice_assign) {
+ if (single_slot_assign && writeback_case == 3) {
+ return {context.get_input(1)};
+ }
+
const int64_t slot_axis = 2;
auto zero = ov::op::v0::Constant::create(ov::element::i64, {1}, {0});
auto one = ov::op::v0::Constant::create(ov::element::i64, {1}, {1});
@@ -103,28 +117,25 @@ OutputVector translate_cpy(const NodeContext & context) {
std::vector<int64_t>{1, 1, -1, output_shape[3].get_length()});
ov::Output<ov::Node> src;
- ov::Output<ov::Node> begin = context.get_input(slot_begin_name);
+ ov::Output<ov::Node> begin;
+ if (!single_slot_assign) {
+ begin = context.get_input(slot_begin_name);
+ }
auto base = context.get_input(1);
- if (op_case == 1) {
- ov::Output<ov::Node> state_begin;
- const std::string src_begin_name = "rs_src_begin_" + context.get_name();
- if (context.has_input(src_begin_name)) {
- state_begin = context.get_input(src_begin_name);
+ if (writeback_case == 1) {
+ if (direct_gdn_state) {
+ // Non-rollback GDN publishes state directly as [active_slots, heads, value_dim,
+ // key_dim]. Flatten each active slot before replacing or updating the cache.
+ src = std::make_shared<ov::op::v1::Reshape>(context.get_input(0), feature, false);
} else {
- auto ssm_state_size = context.get_ssm_state_size();
- if (context.has_input("s_copy_active_slot_len")) {
- auto len = context.get_input("s_copy_active_slot_len");
- auto state_rows = std::make_shared<ov::op::v1::Multiply>(
- ov::op::v0::Constant::create(ov::element::i64, {1}, {ssm_state_size}), len);
- state_begin = std::make_shared<ov::op::v0::Negative>(state_rows);
- } else {
- state_begin = ov::op::v0::Constant::create(ov::element::i64, {1}, {-ssm_state_size});
- }
+ // Multi-slot rollback still consumes GGML's packed [attention | state snapshots]
+ // layout. Slice the state block using the runtime source offset.
+ auto src_begin = context.get_input("rs_src_begin_" + context.get_name());
+ auto state_part =
+ std::make_shared<ov::op::v8::Slice>(context.get_input(0), src_begin, int_max, one, axis);
+ src = std::make_shared<ov::op::v1::Reshape>(state_part, feature, false);
}
- auto state_part =
- std::make_shared<ov::op::v8::Slice>(context.get_input(0), state_begin, int_max, one, axis);
- src = std::make_shared<ov::op::v1::Reshape>(state_part, feature, false);
- } else if (op_case == 2) {
+ } else if (writeback_case == 2) {
// conv_input is [previous conv state | new tokens]; the snapshot is the conv_kernel_size - 1
// columns ending at the last *valid* token. Gather (rather than Slice) keeps the output
// shape static even though the window start is a runtime value.
@@ -177,6 +188,10 @@ OutputVector translate_cpy(const NodeContext & context) {
src = std::make_shared<ov::op::v0::Convert>(src, context.get_output_type());
}
+ if (single_slot_assign) {
+ return rename_outputs_with_suffix({src}, context.get_name());
+ }
+
auto src_len = std::make_shared<ov::op::v8::Gather>(
std::make_shared<ov::op::v3::ShapeOf>(src, ov::element::i64), axis,
ov::op::v0::Constant::create(ov::element::i64, {}, {0}));
@@ -200,6 +215,10 @@ OutputVector translate_cpy(const NodeContext & context) {
src = std::make_shared<ov::op::v0::Convert>(src, context.get_output_type());
}
+ if (single_slot_assign) {
+ return rename_outputs_with_suffix({src}, context.get_name());
+ }
+
auto src_len =
std::make_shared<ov::op::v8::Gather>(std::make_shared<ov::op::v3::ShapeOf>(src, ov::element::i64), axis,
ov::op::v0::Constant::create(ov::element::i64, {}, {0}));
diff --git a/ggml/src/ggml-openvino/openvino/op/flash_attn_ext.cpp b/ggml/src/ggml-openvino/openvino/op/flash_attn_ext.cpp
index b06d01dca..44c6dbc03 100644
--- a/ggml/src/ggml-openvino/openvino/op/flash_attn_ext.cpp
+++ b/ggml/src/ggml-openvino/openvino/op/flash_attn_ext.cpp
@@ -126,8 +126,7 @@ OutputVector translate_flash_attn_ext(const NodeContext & context) {
if (env != nullptr) {
return ggml_openvino_getenv_int("GGML_OPENVINO_MANUAL_GQA_ATTN") > 0;
}
- const char * dev = ggml_openvino_getenv_str("GGML_OPENVINO_DEVICE");
- return dev != nullptr && std::string(dev) == "GPU";
+ return ggml_openvino_is_gpu();
}();
const bool use_manual_gqa_attention =
manual_gqa_enabled && factor > 1 && num_heads_kv > 1 && !context.is_stateful();
diff --git a/ggml/src/ggml-openvino/openvino/op/gated_delta_net.cpp b/ggml/src/ggml-openvino/openvino/op/gated_delta_net.cpp
index 8d07c90bf..9cedf6811 100644
--- a/ggml/src/ggml-openvino/openvino/op/gated_delta_net.cpp
+++ b/ggml/src/ggml-openvino/openvino/op/gated_delta_net.cpp
@@ -3,6 +3,7 @@
#include "../node_context.h"
#include "../op_table.h"
#include "../utils.h"
+#include "ggml-openvino/ggml-openvino-extra.h"
#include <cmath>
#include <cstdint>
@@ -13,14 +14,17 @@
#include <openvino/op/concat.hpp>
#include <openvino/op/constant.hpp>
#include <openvino/op/convert.hpp>
+#include <openvino/op/divide.hpp>
#include <openvino/op/exp.hpp>
#include <openvino/op/gather.hpp>
#include <openvino/op/less.hpp>
#include <openvino/op/loop.hpp>
#include <openvino/op/matmul.hpp>
#include <openvino/op/multiply.hpp>
+#include <openvino/op/reduce_mean.hpp>
#include <openvino/op/reshape.hpp>
#include <openvino/op/squeeze.hpp>
+#include <openvino/op/sqrt.hpp>
#include <openvino/op/subtract.hpp>
#include <openvino/op/tile.hpp>
#include <openvino/op/transpose.hpp>
@@ -34,6 +38,60 @@ namespace op {
static OutputVector translate_gated_delta_net_ref(const NodeContext & context);
+static bool match_gdn_l2_norm(const Output<Node> & normalized, Output<Node> & input, float & eps) {
+ // Match the RMSNorm decomposition emitted by translate_rms_norm, followed by GGML SCALE.
+ const auto scale = ov::as_type_ptr<ov::op::v1::Multiply>(normalized.get_node_shared_ptr());
+ if (!scale) {
+ return false;
+ }
+ const auto factor = ov::as_type_ptr<ov::op::v0::Constant>(scale->get_input_node_shared_ptr(1));
+ const auto rms = ov::as_type_ptr<ov::op::v1::Multiply>(scale->get_input_node_shared_ptr(0));
+ if (!factor || ov::shape_size(factor->get_shape()) != 1 || !rms) {
+ return false;
+ }
+ const auto x = rms->input_value(0);
+ const auto & shape = x.get_partial_shape();
+ if (x.get_element_type() != ov::element::f32 || shape.rank() != 4 || shape[3].is_dynamic() ||
+ shape[3].get_length() <= 0 || normalized.get_partial_shape() != shape) {
+ return false;
+ }
+ const float dim = static_cast<float>(shape[3].get_length());
+ if (factor->cast_vector<float>()[0] != 1.0f / std::sqrt(dim)) {
+ return false;
+ }
+ const auto reciprocal = ov::as_type_ptr<ov::op::v1::Divide>(rms->get_input_node_shared_ptr(1));
+ if (!reciprocal) {
+ return false;
+ }
+ const auto one = ov::as_type_ptr<ov::op::v0::Constant>(reciprocal->get_input_node_shared_ptr(0));
+ const auto root = ov::as_type_ptr<ov::op::v0::Sqrt>(reciprocal->get_input_node_shared_ptr(1));
+ if (!one || ov::shape_size(one->get_shape()) != 1 || one->cast_vector<float>()[0] != 1.0f || !root) {
+ return false;
+ }
+ const auto add = ov::as_type_ptr<ov::op::v1::Add>(root->get_input_node_shared_ptr(0));
+ if (!add) {
+ return false;
+ }
+ const auto mean = ov::as_type_ptr<ov::op::v1::ReduceMean>(add->get_input_node_shared_ptr(0));
+ const auto rms_eps = ov::as_type_ptr<ov::op::v0::Constant>(add->get_input_node_shared_ptr(1));
+ if (!mean || !mean->get_keep_dims() || !rms_eps || ov::shape_size(rms_eps->get_shape()) != 1) {
+ return false;
+ }
+ const auto axes = ov::as_type_ptr<ov::op::v0::Constant>(mean->get_input_node_shared_ptr(1));
+ const auto square = ov::as_type_ptr<ov::op::v1::Multiply>(mean->get_input_node_shared_ptr(0));
+ if (!axes || axes->cast_vector<int64_t>() != std::vector<int64_t>{-1} || !square ||
+ square->input_value(0) != x || square->input_value(1) != x) {
+ return false;
+ }
+ // RMSNorm(x, rms_eps) / sqrt(D) = x / sqrt(sum(x*x) + D*rms_eps).
+ eps = dim * rms_eps->cast_vector<float>()[0];
+ if (!std::isfinite(eps) || eps <= 0.0f) {
+ return false;
+ }
+ input = x;
+ return true;
+}
+
OutputVector translate_gated_delta_net(const NodeContext & context) {
auto v_shape = context.get_input_shape(2).to_shape(); // [B, T, H_v, S_v]
auto q_shape = context.get_input_shape(0).to_shape(); // [B, T, H_k, S_k]
@@ -53,11 +111,21 @@ OutputVector translate_gated_delta_net(const NodeContext & context) {
auto q = context.get_input(0);
auto k = context.get_input(1);
- auto v = process_view_input(context, 2, H_v * S_v);
+ auto v = process_view_input(context, 2, H_v * S_v, 3);
auto g = context.get_input(3);
auto beta = context.get_input(4);
auto state = context.get_input(5);
+ Output<Node> raw_q, raw_k;
+ float q_eps = 1e-6f, k_eps = 1e-6f;
+ const bool fuse_qk_l2norm = ggml_openvino_is_gpu() &&
+ match_gdn_l2_norm(q, raw_q, q_eps) && match_gdn_l2_norm(k, raw_k, k_eps);
+ if (fuse_qk_l2norm) {
+ // Keep head tiling below; GDN applies normalization per head and keeps its attention scale.
+ q = raw_q;
+ k = raw_k;
+ }
+
// ggml maps GQA heads in tiled order, while OV GDN maps repeated heads in grouped order.
if (H_v != H_k) {
const int64_t repeat = H_v / H_k;
@@ -109,7 +177,7 @@ OutputVector translate_gated_delta_net(const NodeContext & context) {
// << ", v=" << v.get_partial_shape() << ", g=" << g.get_partial_shape()
// << ", beta=" << beta.get_partial_shape() << ", state=" << state.get_partial_shape() << std::endl;
- auto gdn = std::make_shared<ov::op::internal::GatedDeltaNet>(q, k, v, state, g, beta);
+ auto gdn = std::make_shared<ov::op::internal::GatedDeltaNet>(q, k, v, state, g, beta, fuse_qk_l2norm, q_eps, k_eps);
auto attn_4d = gdn->output(0);
auto state_4d = gdn->output(1); // [B, H_v, key_dim, value_dim]
@@ -118,6 +186,13 @@ OutputVector translate_gated_delta_net(const NodeContext & context) {
// Transpose output state back to ggml layout [B, H_v, value_dim, key_dim]
auto state_transposed = std::make_shared<ov::op::v1::Transpose>(state_4d, state_perm);
+ if (context.get_output_names().size() == 2) {
+ // The canonical graph consumes the packed GGML result only through separate attention
+ // and state VIEWs. Publish the native outputs under those VIEW names to avoid
+ // flatten -> concat -> reshape -> slice -> reshape chains. This also works for B > 1.
+ return rename_outputs_with_suffix({attn_4d, state_transposed}, context.get_name());
+ }
+
auto flat_shape_1d = ov::op::v0::Constant::create(ov::element::i64, {1}, {-1});
auto attn = std::make_shared<ov::op::v1::Reshape>(attn_4d, flat_shape_1d, false);
auto new_state = std::make_shared<ov::op::v1::Reshape>(state_transposed, flat_shape_1d, false);
@@ -310,6 +385,11 @@ static OutputVector translate_gated_delta_net_ref(const NodeContext & context) {
// state: [B*H_v, S_v, S_v] -> [B, H_v, S_v, S_v] -> flatten
auto state_4d_shape = ov::op::v0::Constant::create(ov::element::i64, {4}, std::vector<int64_t>{B, H_v, S_v, S_v});
auto state_4d = std::make_shared<ov::op::v1::Reshape>(final_state_out, state_4d_shape, false);
+ if (context.get_output_names().size() == 2) {
+ // Match the fused translator's direct attention/state contract.
+ return rename_outputs_with_suffix({attn_perm, state_4d}, context.get_name());
+ }
+
auto state_1d = std::make_shared<ov::op::v1::Reshape>(state_4d, flat_shape_1d, false);
// Concat [attn | state] and reshape to final output
diff --git a/ggml/src/ggml-openvino/openvino/op/get_rows.cpp b/ggml/src/ggml-openvino/openvino/op/get_rows.cpp
index d122722b7..200821d97 100644
--- a/ggml/src/ggml-openvino/openvino/op/get_rows.cpp
+++ b/ggml/src/ggml-openvino/openvino/op/get_rows.cpp
@@ -44,6 +44,10 @@ OutputVector translate_get_rows(const NodeContext & context) {
}
auto op_case = context.get_op_case();
+ if (op_case == 3 || op_case == 4) {
+ return {data};
+ }
+
ov::Output<ov::Node> indices;
if ((op_case == 1 || op_case == 2) && context.has_input("s_copy_active_slot_len")) {
// Recurrent state reorder (inp->s_copy): slice the active (op_case 1) or extra (op_case 2)
@@ -66,8 +70,19 @@ OutputVector translate_get_rows(const NodeContext & context) {
// data[1,b,x,y] ind[1,1,b,x'] test-backend-ops case
// data[x,y] ind[1,1,1,x'] normal case
- indices =
- std::make_shared<ov::op::v0::Squeeze>(indices, ov::op::v0::Constant::create(ov::element::i64, {2}, {0, 1}));
+ // Squeeze the leading dims down to [b,x']. Stateful models drop one rank, so a hardcoded
+ // {0,1} would also strip the batch dim whenever b == 1 (every decode step).
+ const auto indices_rank = indices.get_partial_shape().rank();
+ FRONT_END_OP_CONVERSION_CHECK(indices_rank.is_static(), "Expected static rank for GET_ROWS indices");
+ std::vector<int64_t> indices_squeeze_axes;
+ for (int64_t i = 0; i + 2 < indices_rank.get_length(); ++i) {
+ indices_squeeze_axes.push_back(i);
+ }
+ if (!indices_squeeze_axes.empty()) {
+ indices = std::make_shared<ov::op::v0::Squeeze>(
+ indices, ov::op::v0::Constant::create(ov::element::i64, {indices_squeeze_axes.size()},
+ indices_squeeze_axes));
+ }
if (row_offset != 0) {
indices = std::make_shared<ov::op::v1::Add>(
indices, ov::op::v0::Constant::create(indices.get_element_type(), {}, {row_offset}));
diff --git a/ggml/src/ggml-openvino/openvino/op/im2col.cpp b/ggml/src/ggml-openvino/openvino/op/im2col.cpp
index 08b53f260..469d1de1f 100644
--- a/ggml/src/ggml-openvino/openvino/op/im2col.cpp
+++ b/ggml/src/ggml-openvino/openvino/op/im2col.cpp
@@ -6,11 +6,13 @@
#include <memory>
#include <openvino/core/shape.hpp>
#include <openvino/core/strides.hpp>
+#include <openvino/op/concat.hpp>
#include <openvino/op/constant.hpp>
#include <openvino/op/convert.hpp>
#include <openvino/op/extractimagepatches.hpp>
#include <openvino/op/pad.hpp>
#include <openvino/op/reshape.hpp>
+#include <openvino/op/slice.hpp>
#include <openvino/op/transpose.hpp>
#include <openvino/op/util/attr_types.hpp>
@@ -113,6 +115,124 @@ OutputVector translate_im2col(const NodeContext & context) {
return rename_outputs_with_suffix({res}, context.get_name());
}
+OutputVector translate_im2col_3d(const NodeContext & context) {
+ num_inputs_check(context, 2, 2);
+ const int32_t * params = context.get_output_op_params();
+ int32_t s0 = params[0];
+ int32_t s1 = params[1];
+ int32_t s2 = params[2];
+ int32_t p0 = params[3];
+ int32_t p1 = params[4];
+ int32_t p2 = params[5];
+ int32_t d0 = params[6];
+ int32_t d1 = params[7];
+ int32_t d2 = params[8];
+ int32_t IC = params[9];
+
+ ov::Output<Node> image = process_view_input_new(context, 1);
+ const ov::Shape kernel_shape = context.get_input(0).get_shape();
+ const ov::Shape image_shape = image.get_shape();
+ const ov::Shape out_shape = context.get_output_shape().to_shape();
+
+ const size_t KD = kernel_shape[1];
+ const size_t KH = kernel_shape[2];
+ const size_t KW = kernel_shape[3];
+
+ const size_t N = image_shape[0] / static_cast<size_t>(IC);
+ const size_t ID = image_shape[1];
+ const size_t IH = image_shape[2];
+ const size_t IW = image_shape[3];
+
+ const size_t OD = (ID + 2 * p2 - d2 * (KD - 1) - 1) / s2 + 1;
+ const size_t OH = (IH + 2 * p1 - d1 * (KH - 1) - 1) / s1 + 1;
+ const size_t OW = (IW + 2 * p0 - d0 * (KW - 1) - 1) / s0 + 1;
+
+ if (N == 0 || OD == 0 || OH == 0 || OW == 0) {
+ auto output_type = context.get_output_type();
+ ov::Output<Node> res = ov::op::v0::Constant::create(
+ output_type, ov::Shape{N * OD, OH, OW, static_cast<size_t>(IC * KD * KH * KW)}, {});
+ return rename_outputs_with_suffix({res}, context.get_name());
+ }
+
+ const size_t IH_pad = IH + 2 * p1;
+ const size_t IW_pad = IW + 2 * p0;
+
+ auto image_5d_shape = ov::op::v0::Constant::create(
+ ov::element::i64, ov::Shape{5},
+ std::vector<int64_t>{static_cast<int64_t>(N), static_cast<int64_t>(IC), static_cast<int64_t>(ID),
+ static_cast<int64_t>(IH), static_cast<int64_t>(IW)});
+ auto image_5d = std::make_shared<ov::op::v1::Reshape>(image, image_5d_shape, false);
+
+ auto pads_begin = ov::op::v0::Constant::create(
+ ov::element::i64, ov::Shape{5}, std::vector<int64_t>{0, 0, p2, p1, p0});
+ auto pads_end = ov::op::v0::Constant::create(
+ ov::element::i64, ov::Shape{5}, std::vector<int64_t>{0, 0, p2, p1, p0});
+ auto pad_3d = std::make_shared<ov::op::v1::Pad>(image_5d, pads_begin, pads_end, ov::op::PadMode::CONSTANT);
+
+ const ov::Shape patch_sizes = {KH, KW};
+ const ov::Strides strides = {static_cast<size_t>(s1), static_cast<size_t>(s0)};
+ const ov::Shape rates = {static_cast<size_t>(d1), static_cast<size_t>(d0)};
+
+ ov::OutputVector kd_slices;
+ kd_slices.reserve(KD);
+
+ auto perm_nod = ov::op::v0::Constant::create(ov::element::i64, ov::Shape{5}, {0, 2, 1, 3, 4});
+ auto reshape_4d_shape = ov::op::v0::Constant::create(
+ ov::element::i64, ov::Shape{4},
+ std::vector<int64_t>{static_cast<int64_t>(N * OD), static_cast<int64_t>(IC),
+ static_cast<int64_t>(IH_pad), static_cast<int64_t>(IW_pad)});
+ auto perm1 = ov::op::v0::Constant::create(ov::element::i64, ov::Shape{4}, {0, 2, 3, 1});
+ auto r1_shape = ov::op::v0::Constant::create(
+ ov::element::i64, ov::Shape{5},
+ std::vector<int64_t>{static_cast<int64_t>(N * OD), static_cast<int64_t>(OH), static_cast<int64_t>(OW),
+ static_cast<int64_t>(KH * KW), static_cast<int64_t>(IC)});
+ auto perm2 = ov::op::v0::Constant::create(ov::element::i64, ov::Shape{5}, {0, 1, 2, 4, 3});
+ auto r2_shape = ov::op::v0::Constant::create(
+ ov::element::i64, ov::Shape{6},
+ std::vector<int64_t>{static_cast<int64_t>(N * OD), static_cast<int64_t>(OH), static_cast<int64_t>(OW),
+ static_cast<int64_t>(IC), 1, static_cast<int64_t>(KH * KW)});
+ auto step_c = ov::op::v0::Constant::create(ov::element::i64, ov::Shape{1}, {static_cast<int64_t>(s2)});
+ auto axes_c = ov::op::v0::Constant::create(ov::element::i64, ov::Shape{1}, {2});
+
+ for (size_t ikd = 0; ikd < KD; ++ikd) {
+ auto start_c = ov::op::v0::Constant::create(
+ ov::element::i64, ov::Shape{1}, {static_cast<int64_t>(ikd * d2)});
+ auto stop_c = ov::op::v0::Constant::create(
+ ov::element::i64, ov::Shape{1}, {static_cast<int64_t>(ikd * d2 + OD * s2)});
+ auto depth_slice = std::make_shared<ov::op::v8::Slice>(pad_3d, start_c, stop_c, step_c, axes_c);
+ auto depth_slice_trans = std::make_shared<ov::op::v1::Transpose>(depth_slice, perm_nod);
+ auto depth_slice_4d = std::make_shared<ov::op::v1::Reshape>(depth_slice_trans, reshape_4d_shape, false);
+
+ auto patches = std::make_shared<ov::op::v3::ExtractImagePatches>(
+ depth_slice_4d, patch_sizes, strides, rates, ov::op::PadType::VALID);
+ auto t1 = std::make_shared<ov::op::v1::Transpose>(patches, perm1);
+ auto r1 = std::make_shared<ov::op::v1::Reshape>(t1, r1_shape, false);
+ auto t2 = std::make_shared<ov::op::v1::Transpose>(r1, perm2);
+ auto r2 = std::make_shared<ov::op::v1::Reshape>(t2, r2_shape, false);
+ kd_slices.push_back(r2);
+ }
+
+ ov::Output<Node> res;
+ if (KD == 1) {
+ res = kd_slices[0];
+ } else {
+ res = std::make_shared<ov::op::v0::Concat>(kd_slices, 4);
+ }
+
+ auto final_shape = ov::op::v0::Constant::create(
+ ov::element::i64, ov::Shape{4},
+ std::vector<int64_t>{static_cast<int64_t>(N * OD), static_cast<int64_t>(OH), static_cast<int64_t>(OW),
+ static_cast<int64_t>(IC * KD * KH * KW)});
+ res = std::make_shared<ov::op::v1::Reshape>(res, final_shape, false);
+
+ auto output_type = context.get_output_type();
+ if (res.get_element_type() != output_type) {
+ res = std::make_shared<ov::op::v0::Convert>(res, output_type);
+ }
+
+ return rename_outputs_with_suffix({res}, context.get_name());
+}
+
} // namespace op
} // namespace ggml
} // namespace frontend
diff --git a/ggml/src/ggml-openvino/openvino/op/l2_norm.cpp b/ggml/src/ggml-openvino/openvino/op/l2_norm.cpp
index 4c9bc06c9..6e0485da2 100644
--- a/ggml/src/ggml-openvino/openvino/op/l2_norm.cpp
+++ b/ggml/src/ggml-openvino/openvino/op/l2_norm.cpp
@@ -28,7 +28,7 @@ OutputVector translate_l2_norm(const NodeContext & context) {
// 93: [ 128, 16, 1, 2] L2_NORM q_conv_predelta-1
// [ 128, 16, 1, 2] 0: VIEW q_conv-1
auto output_shape = context.get_output_shape().to_shape();
- input_node = process_view_input(context, 0, output_shape[2] * output_shape[3]);
+ input_node = process_view_input(context, 0, output_shape[2] * output_shape[3], 3);
input_node =
std::make_shared<ov::op::v0::Squeeze>(input_node, ov::op::v0::Constant::create(ov::element::i64, {1}, {0}));
diff --git a/ggml/src/ggml-openvino/openvino/op/mean.cpp b/ggml/src/ggml-openvino/openvino/op/mean.cpp
new file mode 100644
index 000000000..1e9b10bbd
--- /dev/null
+++ b/ggml/src/ggml-openvino/openvino/op/mean.cpp
@@ -0,0 +1,27 @@
+#include "../node_context.h"
+#include "../op_table.h"
+#include "../utils.h"
+
+#include <memory>
+#include <openvino/op/constant.hpp>
+#include <openvino/op/reduce_mean.hpp>
+
+namespace ov {
+namespace frontend {
+namespace ggml {
+namespace op {
+
+OutputVector translate_mean(const NodeContext & context) {
+ num_inputs_check(context, 1, 1);
+
+ auto input = process_view_input_new(context, 0);
+ auto axis = ov::op::v0::Constant::create(ov::element::i64, ov::Shape{1}, {-1});
+ auto res = std::make_shared<ov::op::v1::ReduceMean>(input, axis, true);
+
+ return rename_outputs_with_suffix({res}, context.get_name());
+}
+
+} // namespace op
+} // namespace ggml
+} // namespace frontend
+} // namespace ov
diff --git a/ggml/src/ggml-openvino/openvino/op/mul_mat_id.cpp b/ggml/src/ggml-openvino/openvino/op/mul_mat_id.cpp
index a336924e1..52c3f5eda 100644
--- a/ggml/src/ggml-openvino/openvino/op/mul_mat_id.cpp
+++ b/ggml/src/ggml-openvino/openvino/op/mul_mat_id.cpp
@@ -40,6 +40,21 @@ ov::Output<ov::Node> slice_axis(const ov::Output<ov::Node> & input, int64_t axis
const_i64({axis}));
}
+// GGML tensors are rank 4, but stateful models drop the leading size-1 batch dim, so
+// activations and ids arrive one rank lower. Pick the trailing dims by actual rank.
+std::vector<int> trailing_dims(const ov::Output<ov::Node> & input, int count) {
+ const auto rank = input.get_partial_shape().rank();
+ FRONT_END_OP_CONVERSION_CHECK(rank.is_static(), "Expected static rank for MUL_MAT_ID input");
+ const int rank_len = static_cast<int>(rank.get_length());
+ FRONT_END_OP_CONVERSION_CHECK(rank_len >= count, "MUL_MAT_ID input rank is too low");
+
+ std::vector<int> dims;
+ for (int i = rank_len - count; i < rank_len; ++i) {
+ dims.push_back(i);
+ }
+ return dims;
+}
+
ov::Output<ov::Node> static_shape_dims_or_shapeof(const ov::Output<ov::Node> & input,
const std::vector<int> & dims) {
const auto & partial_shape = input.get_partial_shape();
@@ -157,6 +172,13 @@ OutputVector translate_mul_mat_id(const NodeContext & context) {
auto activations = process_view_input_new(context, 1);
auto ids = process_view_input_new(context, 2);
+ if (activations.get_partial_shape().rank() == 3) {
+ activations = std::make_shared<ov::op::v0::Unsqueeze>(activations, const_i64({0}));
+ }
+ if (ids.get_partial_shape().rank() == 3) {
+ ids = std::make_shared<ov::op::v0::Unsqueeze>(ids, const_i64({0}));
+ }
+
if (expert_weights.get_element_type() == ov::element::u8 && expert_weights.get_partial_shape().rank().is_static() &&
expert_weights.get_partial_shape().rank().get_length() == 5) {
return rename_outputs_with_suffix({translate_mul_mat_id_mxfp4_packed(context, expert_weights, activations, ids)},
@@ -186,8 +208,8 @@ OutputVector translate_mul_mat_id(const NodeContext & context) {
expert_weights = std::make_shared<ov::op::v1::Reshape>(expert_weights, expert_weights_shape_3d, false);
}
- auto activations_shape_3d = static_shape_dims_or_shapeof(activations, {1, 2, 3});
- auto ids_shape_2d = static_shape_dims_or_shapeof(ids, {2, 3});
+ auto activations_shape_3d = static_shape_dims_or_shapeof(activations, trailing_dims(activations, 3));
+ auto ids_shape_2d = static_shape_dims_or_shapeof(ids, trailing_dims(ids, 2));
activations = std::make_shared<ov::op::v1::Reshape>(activations, activations_shape_3d, false);
ids = std::make_shared<ov::op::v1::Reshape>(ids, ids_shape_2d, false);
@@ -197,7 +219,7 @@ OutputVector translate_mul_mat_id(const NodeContext & context) {
}
const auto output_type = context.get_output_type();
- const auto activations_type = ggml_openvino_get_device_name() == "GPU" ? ov::element::f16 : ov::element::f32;
+ const auto activations_type = ggml_openvino_is_gpu() ? ov::element::f16 : ov::element::f32;
if (activations.get_element_type() != activations_type) {
activations = std::make_shared<ov::op::v0::Convert>(activations, activations_type);
}
@@ -210,11 +232,14 @@ OutputVector translate_mul_mat_id(const NodeContext & context) {
ov::Output<ov::Node> result = std::make_shared<ov::op::internal::GatherMatmul>(activations_for_gather, expert_weights, ids);
- // result is [n_used, n_tokens, m]; GGML expects [1, n_tokens, n_used, m].
+ // result is [n_used, n_tokens, m]; GGML expects [1, n_tokens, n_used, m], except on the
+ // stateful path where the leading batch dim is dropped.
auto result_transpose_order = const_i64({1, 0, 2});
result = std::make_shared<ov::op::v1::Transpose>(result, result_transpose_order);
- auto unsqueeze_axes = ov::op::v0::Constant::create(ov::element::i64, {1}, {0});
- result = std::make_shared<ov::op::v0::Unsqueeze>(result, unsqueeze_axes);
+ if (!context.is_stateful()) {
+ auto unsqueeze_axes = ov::op::v0::Constant::create(ov::element::i64, {1}, {0});
+ result = std::make_shared<ov::op::v0::Unsqueeze>(result, unsqueeze_axes);
+ }
if (result.get_element_type() != output_type) {
result = std::make_shared<ov::op::v0::Convert>(result, output_type);
diff --git a/ggml/src/ggml-openvino/openvino/op/reshape.cpp b/ggml/src/ggml-openvino/openvino/op/reshape.cpp
index 272001814..5ed27b1a7 100644
--- a/ggml/src/ggml-openvino/openvino/op/reshape.cpp
+++ b/ggml/src/ggml-openvino/openvino/op/reshape.cpp
@@ -24,84 +24,57 @@ OutputVector translate_reshape(const NodeContext & context) {
return {context.get_input(0)};
}
- int op_case = context.get_op_case();
-
- auto output_shape = context.get_output_shape().to_shape();
+ const int op_case = context.get_op_case();
+ const auto output_shape = context.get_output_shape().to_shape();
+ std::vector<int64_t> shape(output_shape.begin(), output_shape.end());
std::shared_ptr<ov::Node> new_shape_node;
- if (op_case == 0) {
- new_shape_node = ov::op::v0::Constant::create(ov::element::i64, {4}, context.get_output_shape().to_shape());
- } else if (op_case == 1) {
- if (context.is_stateful()) {
- new_shape_node = ov::op::v0::Constant::create(
- ov::element::i64, {3}, std::vector<int64_t>{-1, (int64_t) output_shape[2], (int64_t) output_shape[3]});
- } else {
- new_shape_node = ov::op::v0::Constant::create(
- ov::element::i64, {4},
- std::vector<int64_t>{(int64_t) output_shape[0], -1, (int64_t) output_shape[2],
- (int64_t) output_shape[3]});
+ switch (op_case) {
+ case 0:
+ break;
+ case 1:
+ case 9:
+ shape[1] = -1;
+ if (context.is_stateful() && op_case == 1) {
+ shape.erase(shape.begin());
}
- } else if (op_case == 2) {
- new_shape_node = ov::op::v0::Constant::create(
- ov::element::i64, {4},
- std::vector<int64_t>{(int64_t) output_shape[0], (int64_t) output_shape[1], -1, (int64_t) output_shape[3]});
-
- } else if (op_case == 3) {
- // - 14: [ 1, 1024, 1, 1] RESHAPE Vcur-0 (reshaped) (reshaped)
- // [ 512, 2, 1, 1] 0: RESHAPE Vcur-0 (reshaped)
- // - 15: [ 1, 524288, 1, 1] RESHAPE cache_v_l0 (reshaped)
- // [ 512, 1024, 1, 1] 0: NONE cache_v_l0
- // - 16: [ 1, 524288, 1, 1] SET_ROWS cache_v_l0 (reshaped) (view)
- // [ 1, 1024, 1, 1] 0: RESHAPE Vcur-0 (reshaped) (reshaped)
- // [ 1024, 1, 1, 1] 1: NONE leaf_11
- // [ 1, 524288, 1, 1] 2: RESHAPE cache_v_l0 (reshaped)
- new_shape_node = ov::op::v0::Constant::create(
- ov::element::i64, {4}, std::vector<int64_t>{(int64_t) output_shape[0], (int64_t) output_shape[1], -1, 1});
-
- } else if (op_case == 4) {
+ break;
+ case 2:
+ case 3:
+ shape[2] = -1;
+ if (op_case == 3) {
+ shape[3] = 1;
+ }
+ break;
+ case 4:
return {context.get_input(0).get_node_shared_ptr()->input_value(0)};
-
- } else if (op_case == 5) {
- if (context.is_stateful()) {
- std::vector<int64_t> shape_vec = {1, -1, (int64_t) context.get_output_shape().to_shape()[3]};
- new_shape_node = ov::op::v0::Constant::create(ov::element::i64, {3}, shape_vec);
- } else {
- std::vector<int64_t> shape_vec = {1, 1, -1, (int64_t) context.get_output_shape().to_shape()[3]};
- new_shape_node = ov::op::v0::Constant::create(ov::element::i64, {4}, shape_vec);
+ case 5:
+ case 7:
+ shape = {1, 1, -1, shape[3]};
+ if (context.is_stateful() && op_case == 5) {
+ shape.erase(shape.begin());
}
-
- // // Alternative
- // auto token_len = context.get_input("token_len");
- // auto emb_size =
- // ov::op::v0::Constant::create(ov::element::i64, {1}, {(int64_t) context.get_output_shape().to_shape()[3]});
- // auto one = ov::op::v0::Constant::create(ov::element::i64, {1}, {1});
- // new_shape_node = std::make_shared<ov::op::v0::Concat>(ov::OutputVector{one, one, token_len, emb_size}, 0);
- } else if (op_case == 6) {
- // 14: [ 6144, 1, 2, 1] RESHAPE linear_attn_qkv_mixed-0
- // [ 6144, 2, 1, 1] 0: MUL_MAT node_13
- // reshape to [1, n_slot_active_len, -1, 6144]
+ break;
+ case 6:
+ // Recurrent inputs keep the active sequence count separate from the token count.
if (context.has_input("s_copy_active_slot_len")) {
auto n_slot_active_len = context.get_input("s_copy_active_slot_len");
- auto emb_size = ov::op::v0::Constant::create(ov::element::i64, {1},
- {(int64_t) context.get_output_shape().to_shape()[3]});
+ auto emb_size = ov::op::v0::Constant::create(ov::element::i64, {1}, {shape[3]});
auto one = ov::op::v0::Constant::create(ov::element::i64, {1}, {1});
auto neg_one = ov::op::v0::Constant::create(ov::element::i64, {1}, {-1});
new_shape_node =
std::make_shared<ov::op::v0::Concat>(ov::OutputVector{one, n_slot_active_len, neg_one, emb_size}, 0);
} else {
- new_shape_node = ov::op::v0::Constant::create(ov::element::i64, {4}, context.get_output_shape().to_shape());
+ shape = {1, 1, -1, shape[3]};
}
- } else if (op_case == 7) {
- // 57: [ 2048, 2, 1, 1] RESHAPE linear_attn_out-0 (reshaped)
- // [ 2048, 1, 2, 1] 0: MUL_MAT linear_attn_out-0
- std::vector<int64_t> shape_vec = {1, 1, -1, (int64_t) context.get_output_shape().to_shape()[3]};
- new_shape_node = ov::op::v0::Constant::create(ov::element::i64, {4}, shape_vec);
- } else if (op_case == 8) {
- // 106: [ 128, 128, 16, 2] RESHAPE state_predelta-1
- // [ 262144, 2, 1, 1] 0: GET_ROWS node_86
- auto output_shape = context.get_output_shape().to_shape();
- std::vector<int64_t> shape_vec = {-1, (int64_t) output_shape[1], (int64_t) output_shape[2],
- (int64_t) output_shape[3]};
- new_shape_node = ov::op::v0::Constant::create(ov::element::i64, {4}, shape_vec);
+ break;
+ case 8:
+ shape[0] = -1;
+ break;
+ default:
+ FRONT_END_OP_CONVERSION_CHECK(false, "Unsupported RESHAPE case: ", op_case);
+ }
+ if (!new_shape_node) {
+ new_shape_node = ov::op::v0::Constant::create(ov::element::i64, {shape.size()}, shape);
}
auto res = std::make_shared<ov::op::v1::Reshape>(context.get_input(0), new_shape_node, false);
return rename_outputs_with_suffix({res}, context.get_name());
diff --git a/ggml/src/ggml-openvino/openvino/op/rms_norm.cpp b/ggml/src/ggml-openvino/openvino/op/rms_norm.cpp
index 25c953545..980fa3cee 100644
--- a/ggml/src/ggml-openvino/openvino/op/rms_norm.cpp
+++ b/ggml/src/ggml-openvino/openvino/op/rms_norm.cpp
@@ -25,7 +25,17 @@ OutputVector translate_rms_norm(const NodeContext & context) {
auto op_case = context.get_op_case();
ov::Output<ov::Node> input_node;
- if (op_case == 2) {
+ if (op_case == 3) {
+ // Flatten sequence and token dimensions to match the gate layout.
+ auto input_shape = context.get_input_shape(0).to_shape();
+ input_node = std::make_shared<ov::op::v1::Reshape>(
+ context.get_input(0),
+ ov::op::v0::Constant::create(
+ ov::element::i64, {4}, std::vector<int64_t>{1, -1, (int64_t) input_shape[2], (int64_t) input_shape[3]}),
+ false);
+ } else if (op_case == 1) {
+ input_node = process_view_input_new(context, 0);
+ } else if (op_case == 2) {
auto ssm_state_size = context.get_ssm_state_size();
// The GDN op packs [attn | new_state] along the row axis; the state occupies the last
// ssm_state_size * n_seqs rows. Slice it off (scaling by the active sequence count) to keep
diff --git a/ggml/src/ggml-openvino/openvino/op/rope.cpp b/ggml/src/ggml-openvino/openvino/op/rope.cpp
index a3da7d1fb..2c7fe6256 100644
--- a/ggml/src/ggml-openvino/openvino/op/rope.cpp
+++ b/ggml/src/ggml-openvino/openvino/op/rope.cpp
@@ -44,6 +44,8 @@ OutputVector translate_rope(const NodeContext & context) {
constexpr int TYPE_NORMAL = 0;
constexpr int TYPE_NEOX = 1;
constexpr int TYPE_IMROPE = 2;
+ constexpr int TYPE_VISION = 3;
+ constexpr int TYPE_MROPE = 4;
Output<Node> cos_theta_node;
Output<Node> sin_theta_node;
@@ -67,7 +69,7 @@ OutputVector translate_rope(const NodeContext & context) {
if (context.get_input_size() == 3) {
rope_freqs_weight = context.get_input(2).get_node_shared_ptr();
}
- auto sin_cos = make_sin_cos(op_params, inp_pos, rope_freqs_weight, mode == TYPE_IMROPE, false);
+ auto sin_cos = make_sin_cos(op_params, inp_pos, rope_freqs_weight, mode, false, head_dim);
sin_theta_node = sin_cos.first;
cos_theta_node = sin_cos.second;
context.put_shared(cache_key + "_cos", cos_theta_node);
@@ -80,10 +82,11 @@ OutputVector translate_rope(const NodeContext & context) {
data_node = std::make_shared<ov::op::v0::Convert>(data_node, ov::element::f32);
}
+ const int64_t total_rope_dims = (mode == TYPE_VISION) ? (2 * n_dims) : n_dims;
FRONT_END_OP_CONVERSION_CHECK(n_offs >= 0 && (n_offs % 2 == 0),
"ROPE expects non-negative even n_offs");
- FRONT_END_OP_CONVERSION_CHECK(n_dims > 0 && n_dims + n_offs <= head_dim && (n_dims % 2 == 0),
- "ROPE expects even n_dims in [1, head_dim - n_offs]");
+ FRONT_END_OP_CONVERSION_CHECK(n_dims > 0 && total_rope_dims + n_offs <= head_dim && (n_dims % 2 == 0),
+ "ROPE expects even n_dims with total_rope_dims + n_offs <= head_dim");
// RoPEFusionFlux requires rank_equals(4) on x, t_cos and t_sin. The cos/sin
// tables are already built rank-4 ([1, S, 1, head_size/2]) for both modes. In
@@ -91,9 +94,10 @@ OutputVector translate_rope(const NodeContext & context) {
// to rank-4 ([1, S, n_heads, head_size]) here. Stateful RoPE already produced
// rank-4 output, so downstream attention is unaffected.
if (context.is_stateful()) {
+ const int64_t batch = static_cast<int64_t>(output_shape[0]);
auto r4_shape = ov::op::v0::Constant::create(
ov::element::i64, {4},
- std::vector<int64_t>{1, -1, (int64_t) output_shape[2], (int64_t) output_shape[3]});
+ std::vector<int64_t>{batch, -1, (int64_t) output_shape[2], (int64_t) output_shape[3]});
data_node = std::make_shared<ov::op::v1::Reshape>(data_node, r4_shape, false);
}
// For TYPE_NORMAL rope (both stateful and stateless) we emit the Flux-style
@@ -103,6 +107,7 @@ OutputVector translate_rope(const NodeContext & context) {
auto axis_last = ov::op::v0::Constant::create(ov::element::i64, {1}, {-1});
auto step_one = ov::op::v0::Constant::create(ov::element::i64, {1}, {1});
+ const int64_t batch = static_cast<int64_t>(output_shape[0]);
const int64_t n_heads = static_cast<int64_t>(output_shape[2]);
const int64_t half = n_dims / 2;
auto rot_start = ov::op::v0::Constant::create(ov::element::i64, {1}, {n_offs});
@@ -112,7 +117,7 @@ OutputVector translate_rope(const NodeContext & context) {
auto neg_one_f = ov::op::v0::Constant::create(data_node->get_element_type(), ov::Shape{}, {-1.0f});
auto paired_shape = ov::op::v0::Constant::create(
- ov::element::i64, {5}, std::vector<int64_t>{1, -1, n_heads, half, 2});
+ ov::element::i64, {5}, std::vector<int64_t>{batch, -1, n_heads, half, 2});
auto x_paired = std::make_shared<ov::op::v1::Reshape>(rot_data, paired_shape, false);
auto split_axis = ov::op::v0::Constant::create(ov::element::i64, ov::Shape{}, {-1});
@@ -124,7 +129,7 @@ OutputVector translate_rope(const NodeContext & context) {
auto x_rotated_paired = std::make_shared<ov::op::v0::Concat>(ov::OutputVector{x1_neg, x0}, -1);
auto flat_shape =
- ov::op::v0::Constant::create(ov::element::i64, {4}, std::vector<int64_t>{1, -1, n_heads, n_dims});
+ ov::op::v0::Constant::create(ov::element::i64, {4}, std::vector<int64_t>{batch, -1, n_heads, n_dims});
auto x_rotated =
std::make_shared<ov::op::v1::Reshape>(x_rotated_paired, flat_shape, false);
@@ -167,10 +172,13 @@ OutputVector translate_rope(const NodeContext & context) {
} else {
res = std::make_shared<ov::op::v0::Concat>(concat_parts, -1);
}
- } else if (mode == TYPE_NEOX || mode == TYPE_IMROPE) {
- if (mode == TYPE_IMROPE) {
+ } else if (mode == TYPE_NEOX || mode == TYPE_IMROPE || mode == TYPE_MROPE || mode == TYPE_VISION) {
+ const int64_t half = (mode == TYPE_VISION) ? n_dims : (n_dims / 2);
+ const int64_t rot_dims = 2 * half;
+
+ if (mode != TYPE_NEOX) {
auto cos_sin_shape = std::make_shared<ov::op::v0::Constant>(ov::element::i64, ov::Shape{4},
- std::vector<int64_t>{1, -1, 1, (n_dims >> 1)});
+ std::vector<int64_t>{1, -1, 1, half});
cos_theta_node = std::make_shared<ov::op::v1::Reshape>(cos_theta_node, cos_sin_shape, true);
sin_theta_node = std::make_shared<ov::op::v1::Reshape>(sin_theta_node, cos_sin_shape, true);
}
@@ -179,13 +187,12 @@ OutputVector translate_rope(const NodeContext & context) {
auto step_one = ov::op::v0::Constant::create(ov::element::i64, {1}, {1});
Output<Node> rot_data = data_node;
- if (n_offs > 0 || n_offs + n_dims < head_dim) {
+ if (n_offs > 0 || n_offs + rot_dims < head_dim) {
auto rot_start = ov::op::v0::Constant::create(ov::element::i64, {1}, {n_offs});
- auto rot_end = ov::op::v0::Constant::create(ov::element::i64, {1}, {n_offs + n_dims});
+ auto rot_end = ov::op::v0::Constant::create(ov::element::i64, {1}, {n_offs + rot_dims});
rot_data = std::make_shared<ov::op::v8::Slice>(data_node, rot_start, rot_end, step_one, axis_last);
}
- const int64_t half = n_dims / 2;
auto neg_one_f = ov::op::v0::Constant::create(data_node->get_element_type(), ov::Shape{}, {-1.0f});
auto split_axis = ov::op::v0::Constant::create(ov::element::i64, ov::Shape{}, {3});
@@ -212,8 +219,8 @@ OutputVector translate_rope(const NodeContext & context) {
concat_parts.push_back(head);
}
concat_parts.push_back(rotated);
- if (n_offs + n_dims < head_dim) {
- auto tail_start = ov::op::v0::Constant::create(ov::element::i64, {1}, {n_offs + n_dims});
+ if (n_offs + rot_dims < head_dim) {
+ auto tail_start = ov::op::v0::Constant::create(ov::element::i64, {1}, {n_offs + rot_dims});
auto tail_end = ov::op::v0::Constant::create(ov::element::i64, {1}, {head_dim});
auto tail = std::make_shared<ov::op::v8::Slice>(data_node, tail_start, tail_end, step_one, axis_last);
concat_parts.push_back(tail);
diff --git a/ggml/src/ggml-openvino/openvino/op/scale.cpp b/ggml/src/ggml-openvino/openvino/op/scale.cpp
index 1d5ef4ffa..2947ee03c 100644
--- a/ggml/src/ggml-openvino/openvino/op/scale.cpp
+++ b/ggml/src/ggml-openvino/openvino/op/scale.cpp
@@ -37,6 +37,10 @@ OutputVector translate_scale(const NodeContext & context) {
auto scale_node = std::make_shared<ov::op::v0::Constant>(ov::element::f32, ov::Shape{}, std::vector<float>{scale});
+ if (context.get_op_case() == 2) {
+ return {context.get_input(0)};
+ }
+
if (context.get_op_case() == 1 && context.has_input("cache_rs_reset_len")) {
auto cache_rs_reset_idx = context.get_input("cache_rs_reset_idx");
auto cache_rs_reset_len = context.get_input("cache_rs_reset_len");
diff --git a/ggml/src/ggml-openvino/openvino/op/sum.cpp b/ggml/src/ggml-openvino/openvino/op/sum.cpp
new file mode 100644
index 000000000..c480ee456
--- /dev/null
+++ b/ggml/src/ggml-openvino/openvino/op/sum.cpp
@@ -0,0 +1,30 @@
+#include "../node_context.h"
+#include "../op_table.h"
+#include "../utils.h"
+
+#include <memory>
+#include <openvino/op/constant.hpp>
+#include <openvino/op/range.hpp>
+#include <openvino/op/reduce_sum.hpp>
+#include <openvino/op/reshape.hpp>
+#include <openvino/op/shape_of.hpp>
+
+namespace ov {
+namespace frontend {
+namespace ggml {
+namespace op {
+
+OutputVector translate_sum(const NodeContext & context) {
+ num_inputs_check(context, 1, 1);
+
+ auto input = process_view_input_new(context, 0);
+ auto axes = ov::op::v0::Constant::create(ov::element::i64, ov::Shape{4}, {0, 1, 2, 3});
+ auto res = std::make_shared<ov::op::v1::ReduceSum>(input, axes, true);
+
+ return rename_outputs_with_suffix({res}, context.get_name());
+}
+
+} // namespace op
+} // namespace ggml
+} // namespace frontend
+} // namespace ov
diff --git a/ggml/src/ggml-openvino/openvino/op/unary.cpp b/ggml/src/ggml-openvino/openvino/op/unary.cpp
new file mode 100644
index 000000000..1bb7aa2c8
--- /dev/null
+++ b/ggml/src/ggml-openvino/openvino/op/unary.cpp
@@ -0,0 +1,138 @@
+#include "../node_context.h"
+#include "../op_table.h"
+#include "../utils.h"
+#include "ggml-openvino/ggml-openvino-extra.h"
+
+#include <memory>
+#include <openvino/op/abs.hpp>
+#include <openvino/op/add.hpp>
+#include <openvino/op/clamp.hpp>
+#include <openvino/op/constant.hpp>
+#include <openvino/op/convert.hpp>
+#include <openvino/op/elu.hpp>
+#include <openvino/op/exp.hpp>
+#include <openvino/op/gelu.hpp>
+#include <openvino/op/greater.hpp>
+#include <openvino/op/hard_sigmoid.hpp>
+#include <openvino/op/log.hpp>
+#include <openvino/op/multiply.hpp>
+#include <openvino/op/negative.hpp>
+#include <openvino/op/relu.hpp>
+#include <openvino/op/round.hpp>
+#include <openvino/op/sigmoid.hpp>
+#include <openvino/op/softplus.hpp>
+#include <openvino/op/subtract.hpp>
+
+namespace ov {
+namespace frontend {
+namespace ggml {
+namespace op {
+
+OutputVector translate_unary_gelu(const NodeContext & context) {
+ num_inputs_check(context, 1, 1);
+ auto input = process_view_input_new(context, 0);
+ auto res = std::make_shared<ov::op::v7::Gelu>(input, ov::op::GeluApproximationMode::TANH);
+ return rename_outputs_with_suffix({res}, context.get_name());
+}
+
+OutputVector translate_unary_gelu_erf(const NodeContext & context) {
+ num_inputs_check(context, 1, 1);
+ auto input = process_view_input_new(context, 0);
+ auto res = std::make_shared<ov::op::v7::Gelu>(input, ov::op::GeluApproximationMode::ERF);
+ return rename_outputs_with_suffix({res}, context.get_name());
+}
+
+OutputVector translate_unary_gelu_quick(const NodeContext & context) {
+ num_inputs_check(context, 1, 1);
+ auto input = process_view_input_new(context, 0);
+ auto scale = ov::op::v0::Constant::create(input.get_element_type(), ov::Shape{}, {1.702f});
+ auto mul = std::make_shared<ov::op::v1::Multiply>(input, scale);
+ auto sig = std::make_shared<ov::op::v0::Sigmoid>(mul);
+ auto res = std::make_shared<ov::op::v1::Multiply>(input, sig);
+ return rename_outputs_with_suffix({res}, context.get_name());
+}
+
+OutputVector translate_unary_elu(const NodeContext & context) {
+ num_inputs_check(context, 1, 1);
+ auto input = process_view_input_new(context, 0);
+ auto res = std::make_shared<ov::op::v0::Elu>(input, 1.0);
+ return rename_outputs_with_suffix({res}, context.get_name());
+}
+
+OutputVector translate_unary_hardsigmoid(const NodeContext & context) {
+ num_inputs_check(context, 1, 1);
+ // compute in f32 like the ggml reference: 1/6 is not exact in f16/bf16 (NPU cannot take the f32 path)
+ auto input = process_view_input_new(context, 0);
+ const auto type = ggml_openvino_is_npu() ? input.get_element_type() : ov::element::f32;
+ ov::Output<ov::Node> x = input;
+ if (type != input.get_element_type()) {
+ x = std::make_shared<ov::op::v0::Convert>(input, type);
+ }
+ auto alpha = ov::op::v0::Constant::create(type, ov::Shape{}, {1.0f / 6.0f});
+ auto beta = ov::op::v0::Constant::create(type, ov::Shape{}, {0.5f});
+ ov::Output<ov::Node> res = std::make_shared<ov::op::v0::HardSigmoid>(x, alpha, beta);
+ if (type != input.get_element_type()) {
+ res = std::make_shared<ov::op::v0::Convert>(res, input.get_element_type());
+ }
+ return rename_outputs_with_suffix({res}, context.get_name());
+}
+
+OutputVector translate_unary_step(const NodeContext & context) {
+ num_inputs_check(context, 1, 1);
+ auto input = process_view_input_new(context, 0);
+ auto zero = ov::op::v0::Constant::create(input.get_element_type(), ov::Shape{}, {0.0f});
+ auto cond = std::make_shared<ov::op::v1::Greater>(input, zero);
+ auto res = std::make_shared<ov::op::v0::Convert>(cond, input.get_element_type());
+ return rename_outputs_with_suffix({res}, context.get_name());
+}
+
+OutputVector translate_unary_round(const NodeContext & context) {
+ num_inputs_check(context, 1, 1);
+ auto input = process_view_input_new(context, 0);
+ auto res = std::make_shared<ov::op::v5::Round>(input, ov::op::v5::Round::RoundMode::HALF_AWAY_FROM_ZERO);
+ return rename_outputs_with_suffix({res}, context.get_name());
+}
+
+OutputVector translate_unary_expm1(const NodeContext & context) {
+ num_inputs_check(context, 1, 1);
+ // compute in f32 like the ggml reference: exp(x) - 1 in f16 loses the small-x digits (NPU cannot take the f32 path)
+ auto input = process_view_input_new(context, 0);
+ const auto type = ggml_openvino_is_npu() ? input.get_element_type() : ov::element::f32;
+ ov::Output<ov::Node> x = input;
+ if (type != input.get_element_type()) {
+ x = std::make_shared<ov::op::v0::Convert>(input, type);
+ }
+ auto exp = std::make_shared<ov::op::v0::Exp>(x);
+ auto one = ov::op::v0::Constant::create(type, ov::Shape{}, {1.0f});
+ ov::Output<ov::Node> res = std::make_shared<ov::op::v1::Subtract>(exp, one);
+ if (type != input.get_element_type()) {
+ res = std::make_shared<ov::op::v0::Convert>(res, input.get_element_type());
+ }
+ return rename_outputs_with_suffix({res}, context.get_name());
+}
+
+OutputVector translate_unary_softplus(const NodeContext & context) {
+ num_inputs_check(context, 1, 1);
+
+ if (ggml_openvino_getenv_int("GGML_OPENVINO_NATIVE_SOFTPLUS") != 0) {
+ return translate_1to1_match_1_input<ov::op::v4::SoftPlus>(context);
+ }
+
+ auto input = process_view_input_new(context, 0);
+ const auto element_type = input.get_element_type();
+ auto one = ov::op::v0::Constant::create(element_type, ov::Shape{}, {1.0f});
+
+ auto positive = std::make_shared<ov::op::v0::Relu>(input);
+ auto abs = std::make_shared<ov::op::v0::Abs>(input);
+ auto neg_abs = std::make_shared<ov::op::v0::Negative>(abs);
+ auto exp_neg_abs = std::make_shared<ov::op::v0::Exp>(neg_abs);
+ auto log_term = std::make_shared<ov::op::v0::Log>(std::make_shared<ov::op::v1::Add>(one, exp_neg_abs));
+ auto res = std::make_shared<ov::op::v1::Add>(positive, log_term);
+
+ return rename_outputs_with_suffix({res}, context.get_name());
+}
+
+} // namespace op
+} // namespace ggml
+} // namespace frontend
+} // namespace ov
diff --git a/ggml/src/ggml-openvino/openvino/op/unary_softplus.cpp b/ggml/src/ggml-openvino/openvino/op/unary_softplus.cpp
deleted file mode 100644
index a9e495c37..000000000
--- a/ggml/src/ggml-openvino/openvino/op/unary_softplus.cpp
+++ /dev/null
@@ -1,44 +0,0 @@
-#include "../node_context.h"
-#include "../op_table.h"
-#include "../utils.h"
-#include "ggml-openvino/ggml-openvino-extra.h"
-
-#include <openvino/op/abs.hpp>
-#include <openvino/op/add.hpp>
-#include <openvino/op/constant.hpp>
-#include <openvino/op/exp.hpp>
-#include <openvino/op/log.hpp>
-#include <openvino/op/negative.hpp>
-#include <openvino/op/relu.hpp>
-#include <openvino/op/softplus.hpp>
-
-namespace ov {
-namespace frontend {
-namespace ggml {
-namespace op {
-
-OutputVector translate_unary_softplus(const NodeContext & context) {
- num_inputs_check(context, 1, 1);
-
- if (ggml_openvino_getenv_int("GGML_OPENVINO_NATIVE_SOFTPLUS") != 0) {
- return translate_1to1_match_1_input<ov::op::v4::SoftPlus>(context);
- }
-
- auto input = process_view_input_new(context, 0);
- const auto element_type = input.get_element_type();
- auto one = ov::op::v0::Constant::create(element_type, ov::Shape{}, {1.0f});
-
- auto positive = std::make_shared<ov::op::v0::Relu>(input);
- auto abs = std::make_shared<ov::op::v0::Abs>(input);
- auto neg_abs = std::make_shared<ov::op::v0::Negative>(abs);
- auto exp_neg_abs = std::make_shared<ov::op::v0::Exp>(neg_abs);
- auto log_term = std::make_shared<ov::op::v0::Log>(std::make_shared<ov::op::v1::Add>(one, exp_neg_abs));
- auto res = std::make_shared<ov::op::v1::Add>(positive, log_term);
-
- return rename_outputs_with_suffix({res}, context.get_name());
-}
-
-} // namespace op
-} // namespace ggml
-} // namespace frontend
-} // namespace ov
diff --git a/ggml/src/ggml-openvino/openvino/op/upscale.cpp b/ggml/src/ggml-openvino/openvino/op/upscale.cpp
new file mode 100644
index 000000000..03f3a91b2
--- /dev/null
+++ b/ggml/src/ggml-openvino/openvino/op/upscale.cpp
@@ -0,0 +1,181 @@
+#include "../node_context.h"
+#include "../op_table.h"
+#include "../utils.h"
+#include "ggml.h"
+
+#include <cmath>
+#include <cstddef>
+#include <memory>
+#include <openvino/op/constant.hpp>
+#include <openvino/op/gather.hpp>
+#include <openvino/op/interpolate.hpp>
+#include <openvino/op/matmul.hpp>
+#include <vector>
+
+namespace ov {
+namespace frontend {
+namespace ggml {
+namespace op {
+
+OutputVector translate_upscale(const NodeContext & context) {
+ num_inputs_check(context, 1, 1);
+
+ using Interpolate = ov::op::v4::Interpolate;
+ auto input = process_view_input_new(context, 0);
+
+ const auto input_shape = context.get_input_shape(0).to_shape();
+ const auto output_shape = context.get_output_shape().to_shape();
+
+ if (input_shape == output_shape) {
+ return rename_outputs_with_suffix({input}, context.get_name());
+ }
+
+ ov::Output<ov::Node> res = input;
+
+ // Resample batch / channel dimensions (ne[3] and ne[2], corresponding to OV axes 0 and 1)
+ // using nearest-neighbor index mapping: i0_d = floor(i_d * in_d / out_d).
+ for (size_t axis = 0; axis < 2; ++axis) {
+ const size_t in_dim = input_shape[axis];
+ const size_t out_dim = output_shape[axis];
+ if (in_dim != out_dim) {
+ std::vector<int64_t> indices(out_dim);
+ for (size_t i = 0; i < out_dim; ++i) {
+ indices[i] = static_cast<int64_t>((i * in_dim) / out_dim);
+ }
+ auto indices_node = ov::op::v0::Constant::create(ov::element::i64, {indices.size()}, indices);
+ auto axis_node = ov::op::v0::Constant::create(ov::element::i64, ov::Shape{}, {axis});
+ res = std::make_shared<ov::op::v8::Gather>(res, indices_node, axis_node);
+ }
+ }
+
+ // Spatial interpolation for ne[1] and ne[0] (OV axes 2 and 3)
+ if (input_shape[2] != output_shape[2] || input_shape[3] != output_shape[3]) {
+ const int32_t * op_params = context.get_output_op_params();
+ const int32_t mode_flags = op_params != nullptr ? op_params[0] : 0;
+
+ const int op_case = context.get_op_case();
+ const bool align_corners = (mode_flags & GGML_SCALE_FLAG_ALIGN_CORNERS) != 0;
+ const bool antialias = (mode_flags & GGML_SCALE_FLAG_ANTIALIAS) != 0;
+
+ if (op_case == 2 && antialias) { // GGML_SCALE_MODE_BILINEAR with antialias
+ const size_t H_in = input_shape[2];
+ const size_t W_in = input_shape[3];
+ const size_t H_out = output_shape[2];
+ const size_t W_out = output_shape[3];
+
+ const float pixel_offset = 0.5f;
+
+ // Height projection: [H_out, H_in]
+ std::vector<float> Wy_data(H_out * H_in, 0.0f);
+ const float sf1 = static_cast<float>(H_out) / H_in;
+ const float support1 = std::max(1.0f, 1.0f / sf1);
+ const float invscale1 = 1.0f / support1;
+
+ for (size_t i1 = 0; i1 < H_out; ++i1) {
+ const float y = (static_cast<float>(i1) + pixel_offset) / sf1;
+ const auto y_start = static_cast<int64_t>(y - support1 + pixel_offset);
+ const size_t y_min = y_start > 0 ? static_cast<size_t>(y_start) : 0;
+ const auto y_end = static_cast<int64_t>(y + support1 + pixel_offset);
+ const size_t y_max = y_end > 0 ? std::min<size_t>(static_cast<size_t>(y_end), H_in) : 0;
+
+ float total_weight = 0.0f;
+ for (size_t sy = y_min; sy < y_max; ++sy) {
+ float diff = std::abs((static_cast<float>(sy) - y + pixel_offset) * invscale1);
+ float weight = std::max(1.0f - diff, 0.0f);
+ Wy_data[i1 * H_in + sy] = weight;
+ total_weight += weight;
+ }
+ if (total_weight > 0.0f) {
+ for (size_t sy = y_min; sy < y_max; ++sy) {
+ Wy_data[i1 * H_in + sy] /= total_weight;
+ }
+ }
+ }
+
+ // Width projection: [W_in, W_out]
+ std::vector<float> Wx_data(W_in * W_out, 0.0f);
+ const float sf0 = static_cast<float>(W_out) / W_in;
+ const float support0 = std::max(1.0f, 1.0f / sf0);
+ const float invscale0 = 1.0f / support0;
+
+ for (size_t i0 = 0; i0 < W_out; ++i0) {
+ const float x = (static_cast<float>(i0) + pixel_offset) / sf0;
+ const auto x_start = static_cast<int64_t>(x - support0 + pixel_offset);
+ const size_t x_min = x_start > 0 ? static_cast<size_t>(x_start) : 0;
+ const auto x_end = static_cast<int64_t>(x + support0 + pixel_offset);
+ const size_t x_max = x_end > 0 ? std::min<size_t>(static_cast<size_t>(x_end), W_in) : 0;
+
+ float total_weight = 0.0f;
+ for (size_t sx = x_min; sx < x_max; ++sx) {
+ float diff = std::abs((static_cast<float>(sx) - x + pixel_offset) * invscale0);
+ float weight = std::max(1.0f - diff, 0.0f);
+ Wx_data[sx * W_out + i0] = weight;
+ total_weight += weight;
+ }
+ if (total_weight > 0.0f) {
+ for (size_t sx = x_min; sx < x_max; ++sx) {
+ Wx_data[sx * W_out + i0] /= total_weight;
+ }
+ }
+ }
+
+ auto Wy_node = ov::op::v0::Constant::create(ov::element::f32, {H_out, H_in}, Wy_data);
+ auto Wx_node = ov::op::v0::Constant::create(ov::element::f32, {W_in, W_out}, Wx_data);
+
+ auto res_y = std::make_shared<ov::op::v0::MatMul>(Wy_node, res);
+ res = std::make_shared<ov::op::v0::MatMul>(res_y, Wx_node);
+ } else {
+ Interpolate::InterpolateAttrs attrs;
+ attrs.shape_calculation_mode = Interpolate::ShapeCalcMode::SIZES;
+ attrs.antialias = antialias;
+ attrs.pads_begin = {0, 0, 0, 0};
+ attrs.pads_end = {0, 0, 0, 0};
+
+ switch (op_case) {
+ case 1: // GGML_SCALE_MODE_NEAREST
+ attrs.mode = Interpolate::InterpolateMode::NEAREST;
+ attrs.nearest_mode = Interpolate::NearestMode::FLOOR;
+ attrs.coordinate_transformation_mode = Interpolate::CoordinateTransformMode::ASYMMETRIC;
+ break;
+ case 2: // GGML_SCALE_MODE_BILINEAR
+ attrs.mode = Interpolate::InterpolateMode::LINEAR;
+ attrs.coordinate_transformation_mode = align_corners
+ ? Interpolate::CoordinateTransformMode::ALIGN_CORNERS
+ : Interpolate::CoordinateTransformMode::HALF_PIXEL;
+ break;
+ case 3: // GGML_SCALE_MODE_BICUBIC
+ attrs.mode = Interpolate::InterpolateMode::CUBIC;
+ attrs.cube_coeff = -0.75;
+ attrs.coordinate_transformation_mode = align_corners
+ ? Interpolate::CoordinateTransformMode::ALIGN_CORNERS
+ : Interpolate::CoordinateTransformMode::HALF_PIXEL;
+ break;
+ default:
+ FRONT_END_OP_CONVERSION_CHECK(false, "Unsupported upscale op_case: ", op_case);
+ }
+
+ std::vector<int64_t> target_shape_vec = {
+ static_cast<int64_t>(output_shape[2]),
+ static_cast<int64_t>(output_shape[3]),
+ };
+ std::vector<float> scales_vec = {
+ static_cast<float>(output_shape[2]) / static_cast<float>(input_shape[2]),
+ static_cast<float>(output_shape[3]) / static_cast<float>(input_shape[3]),
+ };
+
+ auto target_shape_node = ov::op::v0::Constant::create(ov::element::i64, {2}, target_shape_vec);
+ auto scales_node = ov::op::v0::Constant::create(ov::element::f32, {2}, scales_vec);
+ auto axes_node = ov::op::v0::Constant::create(ov::element::i64, {2}, {2, 3});
+
+ res = std::make_shared<Interpolate>(
+ res, target_shape_node, scales_node, axes_node, attrs);
+ }
+ }
+
+ return rename_outputs_with_suffix({res}, context.get_name());
+}
+
+} // namespace op
+} // namespace ggml
+} // namespace frontend
+} // namespace ov
diff --git a/ggml/src/ggml-openvino/openvino/op/view.cpp b/ggml/src/ggml-openvino/openvino/op/view.cpp
index ca2d2dc08..486c2ff28 100644
--- a/ggml/src/ggml-openvino/openvino/op/view.cpp
+++ b/ggml/src/ggml-openvino/openvino/op/view.cpp
@@ -44,7 +44,18 @@ OutputVector translate_view(const NodeContext & context) {
auto ss = src_ps.to_shape();
auto dd = dst_ps.to_shape();
const size_t nd = ss.size();
- if (sst.size() == nd && dst.size() == nd) {
+ // Stateful models drop the leading size-1 batch dim, so the real OV tensor can
+ // be one rank lower than the ggml shape metadata above. Axis indices derived
+ // from that metadata must be shifted down by the difference before they are
+ // used as OV axes. Bail out if the dims we would drop are not all 1.
+ const auto in_rank = context.get_input(0).get_partial_shape().rank();
+ const int axis_shift =
+ in_rank.is_static() ? (int) nd - (int) in_rank.get_length() : 0;
+ bool shift_ok = axis_shift >= 0 && (size_t) axis_shift < nd;
+ for (int a = 0; a < axis_shift && shift_ok; ++a) {
+ shift_ok = (ss[a] == 1 && dd[a] == 1);
+ }
+ if (shift_ok && sst.size() == nd && dst.size() == nd) {
// Map each dst axis of size>1 to a src axis with equal (size,stride);
// the unmatched src axis of size>1 is the indexed expert axis.
// dst_to_src[d] records which src axis each dst axis came from, so we can
@@ -93,7 +104,8 @@ OutputVector translate_view(const NodeContext & context) {
ov::op::v0::Constant::create(ov::element::i64, {1}, {sel}),
ov::op::v0::Constant::create(ov::element::i64, {1}, {sel + 1}),
ov::op::v0::Constant::create(ov::element::i64, {1}, {1}),
- ov::op::v0::Constant::create(ov::element::i64, {1}, {dropped}));
+ ov::op::v0::Constant::create(ov::element::i64, {1},
+ {dropped - axis_shift}));
// Build the reshape target from the (concrete) dst shape, but
// keep the dynamic token axis dynamic instead of freezing it
// to the captured n_tokens. Without this the constant dst
@@ -106,20 +118,22 @@ OutputVector translate_view(const NodeContext & context) {
// dynamic dim from the correct SOURCE axis via ShapeOf+Gather
// and place it at the dst token position.
const int32_t dyn = context.get_op_dynamic_dim(); // output ggml axis, -1 if none
- int dst_ov_axis = (dyn != -1) ? (3 - (int) dyn) : -1; // get_shape() reverses ggml order
- int src_ov_axis = (dst_ov_axis >= 0 && dst_ov_axis < (int) nd)
+ // still in ggml metadata axis space; get_shape() reverses ggml order
+ int dst_ov_axis = (dyn != -1) ? ((int) nd - 1 - (int) dyn) : -1;
+ int src_ov_axis = (dst_ov_axis >= axis_shift && dst_ov_axis < (int) nd)
? dst_to_src[dst_ov_axis]
: -1;
- if (dst_ov_axis >= 0 && src_ov_axis >= 0) {
+ if (dst_ov_axis >= 0 && src_ov_axis >= axis_shift) {
// target = concat of per-axis scalars; the token axis is a
// runtime Gather of the slice's shape, the rest are constants.
auto sl_shape = std::make_shared<ov::op::v3::ShapeOf>(sl, ov::element::i64);
auto tok_dim = std::make_shared<ov::op::v8::Gather>(
sl_shape,
- ov::op::v0::Constant::create(ov::element::i64, {1}, {src_ov_axis}),
+ ov::op::v0::Constant::create(ov::element::i64, {1},
+ {src_ov_axis - axis_shift}),
ov::op::v0::Constant::create(ov::element::i64, {}, {0}));
ov::OutputVector parts;
- for (int a = 0; a < (int) nd; ++a) {
+ for (int a = axis_shift; a < (int) nd; ++a) {
if (a == dst_ov_axis) {
parts.push_back(tok_dim);
} else {
@@ -131,8 +145,8 @@ OutputVector translate_view(const NodeContext & context) {
auto rs = std::make_shared<ov::op::v1::Reshape>(sl, dc, false);
return rename_outputs_with_suffix({rs}, context.get_name());
}
- auto dc = ov::op::v0::Constant::create(
- ov::element::i64, {nd}, std::vector<int64_t>(dd.begin(), dd.end()));
+ std::vector<int64_t> dd_ov(dd.begin() + axis_shift, dd.end());
+ auto dc = ov::op::v0::Constant::create(ov::element::i64, {dd_ov.size()}, dd_ov);
auto rs = std::make_shared<ov::op::v1::Reshape>(sl, dc, false);
return rename_outputs_with_suffix({rs}, context.get_name());
}
diff --git a/ggml/src/ggml-openvino/openvino/op_table.cpp b/ggml/src/ggml-openvino/openvino/op_table.cpp
index f249a06bb..12a0953e3 100644
--- a/ggml/src/ggml-openvino/openvino/op_table.cpp
+++ b/ggml/src/ggml-openvino/openvino/op_table.cpp
@@ -2,16 +2,25 @@
#include "utils.h"
+#include <openvino/op/abs.hpp>
#include <openvino/op/add.hpp>
+#include <openvino/op/ceiling.hpp>
+#include <openvino/op/cos.hpp>
#include <openvino/op/divide.hpp>
#include <openvino/op/exp.hpp>
+#include <openvino/op/floor.hpp>
#include <openvino/op/gather.hpp>
#include <openvino/op/gelu.hpp>
+#include <openvino/op/hswish.hpp>
+#include <openvino/op/log.hpp>
#include <openvino/op/matmul.hpp>
#include <openvino/op/multiply.hpp>
#include <openvino/op/negative.hpp>
#include <openvino/op/relu.hpp>
#include <openvino/op/sigmoid.hpp>
+#include <openvino/op/sign.hpp>
+#include <openvino/op/sin.hpp>
+#include <openvino/op/softplus.hpp>
#include <openvino/op/subtract.hpp>
#include <openvino/op/swish.hpp>
#include <openvino/op/tanh.hpp>
@@ -32,6 +41,7 @@ std::unordered_map<std::string, CreatorFunction> get_supported_ops() {
{"GGML_OP_FILL", op::translate_fill },
{"GGML_OP_GET_ROWS", op::translate_get_rows },
{"GGML_OP_IM2COL", op::translate_im2col },
+ {"GGML_OP_IM2COL_3D", op::translate_im2col_3d },
{"GGML_OP_MUL", op::translate_1to1_match_2_inputs<v1::Multiply>},
{"GGML_OP_MUL_MAT", op::translate_mulmat },
{"GGML_OP_MUL_MAT_ID", op::translate_mul_mat_id },
@@ -49,7 +59,24 @@ std::unordered_map<std::string, CreatorFunction> get_supported_ops() {
{"GGML_OP_ARGSORT", op::translate_argsort },
{"GGML_OP_SUB", op::translate_1to1_match_2_inputs<v1::Subtract>},
{"GGML_OP_TRANSPOSE", op::translate_transpose },
- {"GGML_UNARY_OP_GELU", op::translate_1to1_match_1_input<v7::Gelu> },
+ {"GGML_OP_SIN", op::translate_1to1_match_1_input<v0::Sin> },
+ {"GGML_OP_COS", op::translate_1to1_match_1_input<v0::Cos> },
+ {"GGML_OP_LOG", op::translate_1to1_match_1_input<v0::Log> },
+ {"GGML_OP_MEAN", op::translate_mean },
+ {"GGML_OP_SUM", op::translate_sum },
+ {"GGML_UNARY_OP_GELU", op::translate_unary_gelu },
+ {"GGML_UNARY_OP_GELU_ERF", op::translate_unary_gelu_erf },
+ {"GGML_UNARY_OP_GELU_QUICK", op::translate_unary_gelu_quick },
+ {"GGML_UNARY_OP_ELU", op::translate_unary_elu },
+ {"GGML_UNARY_OP_HARDSWISH", op::translate_1to1_match_1_input<v4::HSwish> },
+ {"GGML_UNARY_OP_HARDSIGMOID", op::translate_unary_hardsigmoid },
+ {"GGML_UNARY_OP_STEP", op::translate_unary_step },
+ {"GGML_UNARY_OP_ABS", op::translate_1to1_match_1_input<v0::Abs> },
+ {"GGML_UNARY_OP_SGN", op::translate_1to1_match_1_input<v0::Sign> },
+ {"GGML_UNARY_OP_FLOOR", op::translate_1to1_match_1_input<v0::Floor> },
+ {"GGML_UNARY_OP_CEIL", op::translate_1to1_match_1_input<v0::Ceiling> },
+ {"GGML_UNARY_OP_ROUND", op::translate_unary_round },
+ {"GGML_UNARY_OP_EXPM1", op::translate_unary_expm1 },
{"GGML_UNARY_OP_SIGMOID", op::translate_1to1_match_1_input<v0::Sigmoid> },
{"GGML_UNARY_OP_SILU", op::translate_1to1_match_1_input<v4::Swish> },
{"GGML_UNARY_OP_SOFTPLUS", op::translate_unary_softplus },
@@ -60,7 +87,7 @@ std::unordered_map<std::string, CreatorFunction> get_supported_ops() {
{"GGML_OP_VIEW", op::translate_view },
{"GGML_GLU_OP_SWIGLU", op::translate_glu_swiglu },
{"GGML_GLU_OP_SWIGLU_OAI", op::translate_glu_swiglu_oai },
- {"GGML_GLU_OP_SWIGLU_CLAMP", op::translate_glu_swiglu_clamp },
+ {"GGML_GLU_OP_SWIGLU_CLAMP", op::translate_glu_swiglu_clamp },
{"GGML_GLU_OP_GEGLU", op::translate_glu_geglu },
{"GGML_GLU_OP_GEGLU_QUICK", op::translate_glu_geglu_quick },
{"GGML_OP_SET_ROWS", op::translate_set_rows },
@@ -78,6 +105,12 @@ std::unordered_map<std::string, CreatorFunction> get_supported_ops() {
{"GGML_OP_SET", op::translate_set },
{"GGML_OP_POOL_2D", op::translate_pool_2d },
{"GGML_OP_ROLL", op::translate_roll },
+ {"GGML_OP_UPSCALE", op::translate_upscale },
+ {"GGML_OP_CONV_2D", op::translate_conv_2d },
+ {"GGML_OP_CONV_2D_DW", op::translate_conv_2d_dw },
+ {"GGML_OP_CONV_TRANSPOSE_1D", op::translate_conv_transpose_1d },
+ {"GGML_OP_CONV_TRANSPOSE_2D", op::translate_conv_transpose_2d },
+ {"GGML_OP_CONV_3D", op::translate_conv_3d },
// solve_tri has accuracy issues on GPU
// {"GGML_OP_SOLVE_TRI", op::translate_solve_tri },
};
diff --git a/ggml/src/ggml-openvino/openvino/op_table.h b/ggml/src/ggml-openvino/openvino/op_table.h
index 3dc98bd96..6ed7549a2 100644
--- a/ggml/src/ggml-openvino/openvino/op_table.h
+++ b/ggml/src/ggml-openvino/openvino/op_table.h
@@ -18,6 +18,7 @@ GGML_OP_CONVERTER(translate_div);
GGML_OP_CONVERTER(translate_fill);
GGML_OP_CONVERTER(translate_get_rows);
GGML_OP_CONVERTER(translate_im2col);
+GGML_OP_CONVERTER(translate_im2col_3d);
GGML_OP_CONVERTER(translate_mulmat);
GGML_OP_CONVERTER(translate_mul_mat_id);
GGML_OP_CONVERTER(translate_permute);
@@ -56,6 +57,22 @@ GGML_OP_CONVERTER(translate_tri);
GGML_OP_CONVERTER(translate_solve_tri);
GGML_OP_CONVERTER(translate_pool_2d);
GGML_OP_CONVERTER(translate_roll);
+GGML_OP_CONVERTER(translate_upscale);
+GGML_OP_CONVERTER(translate_mean);
+GGML_OP_CONVERTER(translate_sum);
+GGML_OP_CONVERTER(translate_unary_gelu);
+GGML_OP_CONVERTER(translate_unary_gelu_erf);
+GGML_OP_CONVERTER(translate_unary_gelu_quick);
+GGML_OP_CONVERTER(translate_unary_elu);
+GGML_OP_CONVERTER(translate_unary_hardsigmoid);
+GGML_OP_CONVERTER(translate_unary_step);
+GGML_OP_CONVERTER(translate_unary_round);
+GGML_OP_CONVERTER(translate_unary_expm1);
+GGML_OP_CONVERTER(translate_conv_2d);
+GGML_OP_CONVERTER(translate_conv_2d_dw);
+GGML_OP_CONVERTER(translate_conv_transpose_1d);
+GGML_OP_CONVERTER(translate_conv_transpose_2d);
+GGML_OP_CONVERTER(translate_conv_3d);
} // namespace op
diff --git a/ggml/src/ggml-openvino/openvino/pass/fuse_argsort_topk.cpp b/ggml/src/ggml-openvino/openvino/pass/fuse_argsort_topk.cpp
new file mode 100644
index 000000000..9cf541656
--- /dev/null
+++ b/ggml/src/ggml-openvino/openvino/pass/fuse_argsort_topk.cpp
@@ -0,0 +1,70 @@
+#include "fuse_argsort_topk.h"
+
+#include <openvino/op/constant.hpp>
+#include <openvino/op/slice.hpp>
+#include <openvino/op/topk.hpp>
+#include <openvino/pass/pattern/op/wrap_type.hpp>
+
+namespace ov {
+namespace frontend {
+namespace ggml {
+namespace pass {
+
+FuseArgsortTopK::FuseArgsortTopK() {
+ // GGML argsort materializes the full ordering before a VIEW keeps the prefix.
+ // Let TopK produce only that prefix when all index consumers are such views.
+ auto pattern = ov::pass::pattern::wrap_type<ov::op::v11::TopK>();
+
+ const auto callback = [](ov::pass::pattern::Matcher & m) {
+ auto topk = ov::as_type_ptr<ov::op::v11::TopK>(m.get_match_root());
+ if (!topk->output(0).get_target_inputs().empty() || topk->output(1).get_target_inputs().empty() ||
+ topk->get_sort_type() != ov::op::v11::TopK::SortType::SORT_VALUES) {
+ return false;
+ }
+
+ const auto shape = topk->get_output_partial_shape(1);
+ if (shape.rank().is_dynamic()) {
+ return false;
+ }
+ const int64_t rank = shape.rank().get_length();
+ const int64_t axis = topk->get_axis();
+ if (shape[axis].is_dynamic()) {
+ return false;
+ }
+
+ int64_t prefix = -1;
+ for (const auto & input : topk->output(1).get_target_inputs()) {
+ const auto * slice = ov::as_type<ov::op::v8::Slice>(input.get_node());
+ if (!slice || input.get_index() != 0 || slice->get_input_size() != 5) {
+ return false;
+ }
+ int64_t values[4];
+ for (size_t i = 0; i < 4; ++i) {
+ auto value = ov::as_type_ptr<ov::op::v0::Constant>(slice->get_input_node_shared_ptr(i + 1));
+ if (!value || ov::shape_size(value->get_shape()) != 1) {
+ return false;
+ }
+ values[i] = value->cast_vector<int64_t>()[0];
+ }
+ const int64_t slice_axis = values[3] < 0 ? values[3] + rank : values[3];
+ if (values[0] != 0 || values[2] != 1 || slice_axis != axis || values[1] <= 0 ||
+ values[1] >= shape[axis].get_length() || (prefix != -1 && prefix != values[1])) {
+ return false;
+ }
+ prefix = values[1];
+ }
+
+ // ggml_argsort_top_k sorts the full row, then exposes only its first k indices.
+ // Keep other TopK forms unchanged because their ordering can be observable.
+ topk->set_argument(1, ov::op::v0::Constant::create(ov::element::i64, ov::Shape{}, {prefix}));
+ topk->validate_and_infer_types();
+ return true;
+ };
+
+ register_matcher(std::make_shared<ov::pass::pattern::Matcher>(pattern, "ov::frontend::ggml::pass::FuseArgsortTopK"), callback);
+}
+
+} // namespace pass
+} // namespace ggml
+} // namespace frontend
+} // namespace ov
diff --git a/ggml/src/ggml-openvino/openvino/pass/fuse_argsort_topk.h b/ggml/src/ggml-openvino/openvino/pass/fuse_argsort_topk.h
new file mode 100644
index 000000000..c248ebed3
--- /dev/null
+++ b/ggml/src/ggml-openvino/openvino/pass/fuse_argsort_topk.h
@@ -0,0 +1,19 @@
+#pragma once
+
+#include <openvino/pass/matcher_pass.hpp>
+
+namespace ov {
+namespace frontend {
+namespace ggml {
+namespace pass {
+
+class FuseArgsortTopK : public ov::pass::MatcherPass {
+public:
+ OPENVINO_MATCHER_PASS_RTTI("ov::frontend::ggml::pass::FuseArgsortTopK")
+ FuseArgsortTopK();
+};
+
+} // namespace pass
+} // namespace ggml
+} // namespace frontend
+} // namespace ov
diff --git a/ggml/src/ggml-openvino/openvino/pass/fuse_moe_compressed.cpp b/ggml/src/ggml-openvino/openvino/pass/fuse_moe_compressed.cpp
index c4872ac2e..db8fc6133 100644
--- a/ggml/src/ggml-openvino/openvino/pass/fuse_moe_compressed.cpp
+++ b/ggml/src/ggml-openvino/openvino/pass/fuse_moe_compressed.cpp
@@ -1,5 +1,6 @@
#include "fuse_moe_compressed.h"
+#include <cstring>
#include <limits>
#include <set>
#include <memory>
@@ -7,9 +8,11 @@
#include <openvino/core/rt_info.hpp>
#include <openvino/op/constant.hpp>
#include <openvino/op/convert.hpp>
+#include <openvino/op/gelu.hpp>
#include <openvino/op/multiply.hpp>
#include <openvino/op/reduce_sum.hpp>
#include <openvino/op/reshape.hpp>
+#include <openvino/op/slice.hpp>
#include <openvino/op/squeeze.hpp>
#include <openvino/op/subtract.hpp>
#include <openvino/op/swish.hpp>
@@ -84,6 +87,44 @@ size_t logical_k(const ov::Shape & shape) {
return shape.size() == 4 ? shape[2] * shape[3] : shape.back();
}
+// Slice a [n_expert, m, ...] compressed weight/scale/zp Constant along axis 1 (m), [begin, end).
+// Built by copying raw bytes directly instead of a graph Slice op: CommonOptimizations rewrites
+// v8::Slice into v1::StridedSlice for constant folding regardless of disable_constant_folding
+// (the rewrite doesn't carry the marking over), and that reference evaluator crashes on
+// sub-byte (u4/i4) element types. Each axis-1 "row" (the trailing dims) is confirmed
+// byte-aligned here -- weight/scale/zp trailing sizes are always whole groups -- so a
+// byte-range memcpy per row is exact for both regular and sub-byte element types.
+ov::Output<ov::Node> slice_experts_dim1(const ov::Output<ov::Node> & tensor, int64_t begin, int64_t end) {
+ auto constant = ov::as_type_ptr<ov::op::v0::Constant>(tensor.get_node_shared_ptr());
+ OPENVINO_ASSERT(constant, "slice_experts_dim1 expects a Constant input");
+
+ const auto & shape = constant->get_shape();
+ OPENVINO_ASSERT(shape.size() >= 2, "slice_experts_dim1 expects rank >= 2");
+ const auto & type = constant->get_element_type();
+
+ size_t row_elems = 1;
+ for (size_t i = 2; i < shape.size(); ++i) {
+ row_elems *= shape[i];
+ }
+ const size_t row_bits = row_elems * type.bitwidth();
+ OPENVINO_ASSERT(row_bits % 8 == 0, "slice_experts_dim1: row is not byte-aligned");
+ const size_t row_bytes = row_bits / 8;
+
+ ov::Shape out_shape = shape;
+ out_shape[1] = static_cast<size_t>(end - begin);
+
+ std::vector<uint8_t> out_data(shape[0] * static_cast<size_t>(end - begin) * row_bytes);
+ const auto * src = static_cast<const uint8_t *>(constant->get_data_ptr());
+ const size_t src_row_bytes = shape[1] * row_bytes;
+ for (size_t e = 0; e < shape[0]; ++e) {
+ std::memcpy(out_data.data() + e * static_cast<size_t>(end - begin) * row_bytes,
+ src + e * src_row_bytes + static_cast<size_t>(begin) * row_bytes,
+ static_cast<size_t>(end - begin) * row_bytes);
+ }
+
+ return std::make_shared<ov::op::v0::Constant>(type, out_shape, out_data.data());
+}
+
} // namespace
FuseMoeCompressed::FuseMoeCompressed() {
@@ -267,6 +308,208 @@ FuseMoeCompressed::FuseMoeCompressed() {
register_matcher(std::make_shared<Matcher>(root_m, "ov::frontend::ggml::pass::FuseMoeCompressed"), callback);
}
+FuseMoeCompressedFusedGateUp::FuseMoeCompressedFusedGateUp() {
+ using namespace ov::pass::pattern;
+
+ // A single GatherMatmul computes the fused gate+up projection; the gate/up split happens
+ // AFTER the GEMM via two Slice ops on the last (feature) axis (see get_glu_inputs in
+ // op/glu_geglu.cpp), unlike FuseMoeCompressed's two-separate-GatherMatmul models.
+ auto hidden_m = any_input();
+ auto a_reshape_m = wrap_type<ov::op::v1::Reshape>({ hidden_m, any_input() });
+ auto a_m = wrap_type<ov::op::v1::Transpose>({ optional<ov::op::v0::Convert>({ a_reshape_m }), any_input() });
+
+ auto gate_up_w_m = any_input();
+ auto ids_gate_up_m = any_input();
+ auto bgm_fused_m = wrap_type<ov::op::internal::GatherMatmul>({ a_m, gate_up_w_m, ids_gate_up_m, any_input() });
+ auto gu_u_m = optional<ov::op::v0::Convert>({ wrap_type<ov::op::v0::Unsqueeze>(
+ { wrap_type<ov::op::v1::Transpose>({ bgm_fused_m, any_input() }), any_input() }) });
+
+ auto gate_slice_m = wrap_type<ov::op::v8::Slice>({ gu_u_m, any_input(), any_input(), any_input(), any_input() });
+ auto up_slice_m = wrap_type<ov::op::v8::Slice>({ gu_u_m, any_input(), any_input(), any_input(), any_input() });
+ auto gelu_m = wrap_type<ov::op::v7::Gelu>({ gate_slice_m });
+ auto geglu_m = wrap_type<ov::op::v1::Multiply>({ gelu_m, up_slice_m });
+
+ auto d_t_m = wrap_type<ov::op::v1::Transpose>(
+ { optional<ov::op::v0::Convert>({ wrap_type<ov::op::v1::Reshape>({ geglu_m, any_input() }) }),
+ any_input() });
+ auto down_w_m = any_input();
+ auto ids_down_m = any_input();
+ auto bgm_down_m = wrap_type<ov::op::internal::GatherMatmul>({ d_t_m, down_w_m, ids_down_m, any_input() });
+ auto down_u_m = optional<ov::op::v0::Convert>({ wrap_type<ov::op::v0::Unsqueeze>(
+ { wrap_type<ov::op::v1::Transpose>({ bgm_down_m, any_input() }), any_input() }) });
+
+ // gemma-4 applies an extra per-expert output scale to the down projection before the
+ // router-weight multiply (llama-graph.cpp's ffn_down_exps.scale); FuseMoeCompressed's
+ // Qwen-shaped models have no such scale, so this node is specific to this pattern.
+ auto down_scale_m = any_input();
+ auto down_scaled_m = wrap_type<ov::op::v1::Multiply>({ down_u_m, down_scale_m });
+
+ auto routing_m = any_input();
+ auto weighted_m = wrap_type<ov::op::v1::Multiply>({ down_scaled_m, routing_m });
+ auto root_m = wrap_type<ov::op::v1::ReduceSum>({ weighted_m, any_input() });
+
+ const auto callback = [=](Matcher & m) {
+ auto & pm = m.get_pattern_value_map();
+
+ const auto gate_up = unwrap_dequant(pm.at(gate_up_w_m));
+ const auto down = unwrap_dequant(pm.at(down_w_m));
+ if (!gate_up.ok || !down.ok) {
+ return false;
+ }
+
+ // oneDNN's weight-decompression GEMM needs a real zero point, so a symmetric (no zp)
+ // projection cannot be fused here -- e.g. the GPU-only Q4_0 fallback in
+ // ggml_openvino_get_extracted_layout for grouped 8-bit expert weights.
+ if (!gate_up.has_zp || !down.has_zp) {
+ return false;
+ }
+
+ const auto gate_up_shape = gate_up.weight.get_shape();
+ const auto down_shape = down.weight.get_shape();
+ if (gate_up_shape.size() < 3 || down_shape.size() < 3 || gate_up_shape.size() != down_shape.size()) {
+ return false;
+ }
+
+ // The op reads the zero point straight off a weight port, so it must already be an
+ // integer Constant. Requantized experts (GGML_OPENVINO_REQUANT_KQUANT=q4_asym64_all)
+ // are; a native Q4_K expert keeps an exact f16 zp and is left to the unfused path
+ // rather than rounded here -- rounding would put a live node on that port and the
+ // expert GEMM would read the wrong zero point.
+ static const std::set<ov::element::Type> int_types = { ov::element::u4, ov::element::i4,
+ ov::element::u8, ov::element::i8 };
+ if (int_types.count(gate_up.weight.get_element_type()) == 0 ||
+ int_types.count(down.weight.get_element_type()) == 0 ||
+ int_types.count(gate_up.zp.get_element_type()) == 0 ||
+ int_types.count(down.zp.get_element_type()) == 0) {
+ return false;
+ }
+
+ const auto group_of = [](const dequant_inputs & w) {
+ const auto s = w.weight.get_shape();
+ return s.size() == 4 ? s[3] : logical_k(s);
+ };
+ if (group_of(gate_up) != group_of(down)) {
+ return false;
+ }
+
+ auto gelu_node = ov::as_type_ptr<ov::op::v7::Gelu>(pm.at(gelu_m).get_node_shared_ptr());
+ if (!gelu_node || gelu_node->get_approximation_mode() != ov::op::GeluApproximationMode::ERF) {
+ return false;
+ }
+
+ // Read the actual gate/up split point from the matched Slice nodes instead of assuming
+ // inter_size/2, so this stays correct if the model's ffn_dim convention ever changes.
+ auto gate_slice = ov::as_type_ptr<ov::op::v8::Slice>(pm.at(gate_slice_m).get_node_shared_ptr());
+ auto up_slice = ov::as_type_ptr<ov::op::v8::Slice>(pm.at(up_slice_m).get_node_shared_ptr());
+ auto gate_end_c = ov::as_type_ptr<ov::op::v0::Constant>(gate_slice->input_value(2).get_node_shared_ptr());
+ auto up_begin_c = ov::as_type_ptr<ov::op::v0::Constant>(up_slice->input_value(1).get_node_shared_ptr());
+ if (!gate_end_c || !up_begin_c) {
+ return false;
+ }
+ const auto gate_end = gate_end_c->cast_vector<int64_t>().at(0);
+ const auto up_begin = up_begin_c->cast_vector<int64_t>().at(0);
+ if (gate_end != up_begin || gate_end <= 0 || gate_end >= static_cast<int64_t>(gate_up_shape[1])) {
+ return false;
+ }
+ const int64_t mid = gate_end;
+ const int64_t full = static_cast<int64_t>(gate_up_shape[1]);
+
+ auto ids = pm.at(ids_down_m);
+ const auto ids_pshape = ids.get_partial_shape();
+ if (ids_pshape.rank().is_dynamic() || ids_pshape[ids_pshape.rank().get_length() - 1].is_dynamic()) {
+ return false;
+ }
+ const size_t top_k = ids_pshape[ids_pshape.rank().get_length() - 1].get_length();
+
+ // routing weights arrive as [1, n_tokens, top_k, 1]; the op wants [..., top_k]
+ auto routing = pm.at(routing_m);
+ const auto routing_pshape = routing.get_partial_shape();
+ if (routing_pshape.rank().is_dynamic() || routing_pshape.rank().get_length() != 4 ||
+ routing_pshape[3] != 1) {
+ return false;
+ }
+ // Fold gemma-4's per-expert output scale into the routing weights: the reduction is
+ // sum_e(routing[e] * scale[e] * down_out[e]), and MOECompressed only takes one
+ // per-expert weight, so pre-multiply it into routing here (same [.., top_k, 1] shape).
+ auto down_scale = pm.at(down_scale_m);
+ if (down_scale.get_partial_shape() != routing_pshape) {
+ return false;
+ }
+ routing = std::make_shared<ov::op::v1::Multiply>(routing, down_scale);
+ routing = std::make_shared<ov::op::v0::Squeeze>(
+ routing, ov::op::v0::Constant::create(ov::element::i64, ov::Shape{ 1 }, { 3 }));
+ if (ids_pshape.rank().get_length() == 2) {
+ ids = std::make_shared<ov::op::v0::Unsqueeze>(
+ ids, ov::op::v0::Constant::create(ov::element::i64, ov::Shape{ 1 }, { 0 }));
+ }
+ if (routing.get_partial_shape() != ids.get_partial_shape()) {
+ return false;
+ }
+
+ const size_t down_k = logical_k(down_shape);
+ const auto down_scale_shape = down.scale.get_shape();
+ const size_t down_groups = down_scale_shape.size() >= 3 ? down_scale_shape[2] : 1;
+
+ // Split the fused gate_up weight/scale/zp on the output axis. slice_experts_dim1 copies
+ // raw bytes out of the Constant, so each half stays a Constant -- which is what the op
+ // needs on its weight ports.
+ const ov::Output<ov::Node> gate_weight = slice_experts_dim1(gate_up.weight, 0, mid);
+ const ov::Output<ov::Node> up_weight = slice_experts_dim1(gate_up.weight, mid, full);
+ const ov::Output<ov::Node> gate_scale = slice_experts_dim1(gate_up.scale, 0, mid);
+ const ov::Output<ov::Node> up_scale = slice_experts_dim1(gate_up.scale, mid, full);
+ const ov::Output<ov::Node> gate_zp = slice_experts_dim1(gate_up.zp, 0, mid);
+ const ov::Output<ov::Node> up_zp = slice_experts_dim1(gate_up.zp, mid, full);
+
+ ov::op::internal::MOECompressed::Config config;
+ config.expert_type = ov::op::internal::MOE::Expert_type::GEMM3_SWIGLU;
+ config.activation_type = ov::op::internal::MOE::Activation_type::GEGLU_ERF;
+ config.expert_alpha = 0.0f;
+ config.expert_beta = 1.0f;
+ config.gate_idx = 0;
+ config.hidden_size = logical_k(gate_up_shape);
+ config.inter_size = static_cast<size_t>(mid);
+ config.num_expert = gate_up_shape[0];
+ config.num_shared_expert = 0;
+ config.top_k = top_k;
+ config.group_size = down_groups <= 1 ? std::numeric_limits<size_t>::max() : down_k / down_groups;
+ config.has_batch_dim = true;
+ config.has_zp = true;
+ config.out_type = ov::element::dynamic;
+
+ // Rebuild the activations Transpose before mul_mat_id's f16 GPU Convert, so the op
+ // stays f32 like the block it replaces (mirrors FuseMoeCompressed).
+ const auto a_transpose = pm.at(a_m).get_node_shared_ptr();
+ ov::Output<ov::Node> hidden =
+ std::make_shared<ov::op::v1::Transpose>(pm.at(a_reshape_m), a_transpose->input_value(1));
+
+ // w0 is the activated (gate) lane and w1 the multiplied (up) lane -- see
+ // moe_3gemm_swiglu_opt.cpp: "scratch.up = up(x) * silu(gate(x))".
+ const ov::OutputVector args = {
+ hidden, routing, ids,
+ gate_weight, gate_scale, gate_zp,
+ up_weight, up_scale, up_zp,
+ down.weight, down.scale, down.zp,
+ };
+
+ auto moe = std::make_shared<ov::op::internal::MOECompressed>(args, config);
+
+ ov::Output<ov::Node> result = moe->output(0);
+ const auto root_type = m.get_match_root()->get_output_element_type(0);
+ if (result.get_element_type() != root_type) {
+ result = std::make_shared<ov::op::v0::Convert>(result, root_type);
+ }
+
+ result.get_node_shared_ptr()->set_friendly_name(m.get_match_root()->get_friendly_name());
+ ov::copy_runtime_info(m.get_matched_nodes(), result.get_node_shared_ptr());
+ ov::replace_node(m.get_match_root(), result.get_node_shared_ptr());
+ register_new_node(moe);
+ return true;
+ };
+
+ register_matcher(
+ std::make_shared<Matcher>(root_m, "ov::frontend::ggml::pass::FuseMoeCompressedFusedGateUp"), callback);
+}
+
} // namespace pass
} // namespace ggml
} // namespace frontend
diff --git a/ggml/src/ggml-openvino/openvino/pass/fuse_moe_compressed.h b/ggml/src/ggml-openvino/openvino/pass/fuse_moe_compressed.h
index 5500bed68..5c574ebad 100644
--- a/ggml/src/ggml-openvino/openvino/pass/fuse_moe_compressed.h
+++ b/ggml/src/ggml-openvino/openvino/pass/fuse_moe_compressed.h
@@ -13,6 +13,17 @@ public:
FuseMoeCompressed();
};
+// Folds the MoE expert block emitted for a model whose gate and up projections share one
+// fused MUL_MAT_ID weight (gemma-4: one GatherMatmul + Slice/Slice split, GEGLU activation)
+// into a single ov::op::internal::MOECompressed, GEMM3_SWIGLU/GEGLU_ERF. Splits the fused
+// weight/scale/zp into gate/up halves so it lands on the same 3-GEMM expert-grouped kernel
+// FuseMoeCompressed uses for models with separate gate/up weights.
+class FuseMoeCompressedFusedGateUp : public ov::pass::MatcherPass {
+public:
+ OPENVINO_MATCHER_PASS_RTTI("ov::frontend::ggml::pass::FuseMoeCompressedFusedGateUp")
+ FuseMoeCompressedFusedGateUp();
+};
+
} // namespace pass
} // namespace ggml
} // namespace frontend
diff --git a/ggml/src/ggml-openvino/openvino/pass/fuse_moe_router.cpp b/ggml/src/ggml-openvino/openvino/pass/fuse_moe_router.cpp
new file mode 100644
index 000000000..d6ed7f551
--- /dev/null
+++ b/ggml/src/ggml-openvino/openvino/pass/fuse_moe_router.cpp
@@ -0,0 +1,153 @@
+#include "fuse_moe_router.h"
+
+#include <openvino/core/graph_util.hpp>
+#include <openvino/core/rt_info.hpp>
+#include <openvino/op/broadcast.hpp>
+#include <openvino/op/clamp.hpp>
+#include <openvino/op/concat.hpp>
+#include <openvino/op/constant.hpp>
+#include <openvino/op/divide.hpp>
+#include <openvino/op/gather.hpp>
+#include <openvino/op/matmul.hpp>
+#include <openvino/op/reduce_sum.hpp>
+#include <openvino/op/reshape.hpp>
+#include <openvino/op/shape_of.hpp>
+#include <openvino/op/slice.hpp>
+#include <openvino/op/softmax.hpp>
+#include <openvino/op/squeeze.hpp>
+#include <openvino/op/tile.hpp>
+#include <openvino/op/topk.hpp>
+#include <openvino/op/unsqueeze.hpp>
+#include <openvino/pass/pattern/op/wrap_type.hpp>
+
+namespace ov {
+namespace frontend {
+namespace ggml {
+namespace pass {
+
+namespace {
+
+bool constant_is(const ov::Output<ov::Node> & output, const std::vector<int64_t> & values) {
+ auto node = ov::as_type_ptr<ov::op::v0::Constant>(output.get_node_shared_ptr());
+ return node && node->get_element_type().is_integral_number() && node->cast_vector<int64_t>() == values;
+}
+
+} // namespace
+
+FuseMoeRouter::FuseMoeRouter() {
+ using namespace ov::pass::pattern;
+ using namespace ov::op;
+
+ // Match the GGML rank-4 routing chain. The GPU plugin recognizes the resulting
+ // Softmax -> TopK -> ReduceSum -> Divide subgraph as MoERouterFused.
+ auto logits = wrap_type<v0::MatMul>();
+ auto softmax = wrap_type<v8::Softmax>({logits});
+ auto topk = wrap_type<v11::TopK>({softmax, any_input()});
+ topk->set_output_size(2);
+ auto ids = wrap_type<v8::Slice>({topk->output(1), any_input(), any_input(), any_input(), any_input()});
+ auto ids_2d = wrap_type<v0::Squeeze>({ids, any_input()});
+ auto probs = wrap_type<v1::Reshape>({softmax, any_input()});
+ auto data = wrap_type<v0::Squeeze>({probs, any_input()});
+ auto data_batch = wrap_type<v8::Gather>({wrap_type<v3::ShapeOf>({data}), any_input(), any_input()});
+ auto ids_count = wrap_type<v8::Gather>({wrap_type<v3::ShapeOf>({ids_2d}), any_input(), any_input()});
+ auto target = wrap_type<v0::Concat>({data_batch, ids_count});
+ auto broadcast = wrap_type<v3::Broadcast>({ids_2d, target});
+ auto gather = wrap_type<v8::Gather>({data, broadcast, any_input()});
+ auto weights_4d = wrap_type<v0::Unsqueeze>({gather, any_input()});
+ auto weights = wrap_type<v1::Reshape>({weights_4d, any_input()});
+ auto sum = wrap_type<v1::ReduceSum>({weights, any_input()});
+ auto clamp = wrap_type<v0::Clamp>({sum});
+ auto tile = wrap_type<v0::Tile>({clamp, any_input()});
+ auto norm = wrap_type<v1::Divide>({weights, tile});
+
+ const auto callback = [=](Matcher & m) {
+ const auto & pm = m.get_pattern_value_map();
+ const auto node = [&](const std::shared_ptr<ov::Node> & p) { return pm.at(p).get_node_shared_ptr(); };
+ const auto input_is = [&](const std::shared_ptr<ov::Node> & p, size_t i, const std::vector<int64_t> & v) {
+ return constant_is(node(p)->input_value(i), v);
+ };
+ auto mm = ov::as_type_ptr<v0::MatMul>(node(logits));
+ auto sm = ov::as_type_ptr<v8::Softmax>(node(softmax));
+ auto tk = ov::as_type_ptr<v11::TopK>(node(topk));
+ auto reduce = ov::as_type_ptr<v1::ReduceSum>(node(sum));
+ auto limit = ov::as_type_ptr<v0::Clamp>(node(clamp));
+ const auto shape = mm->get_output_partial_shape(0);
+ if (shape.rank() != 4 || shape[0] != 1 || shape[1] != 1 || shape[3].is_dynamic() ||
+ mm->get_transpose_a() || mm->get_input_partial_shape(0).rank() != 4 ||
+ mm->get_input_partial_shape(1).rank() != 2 || (sm->get_axis() != -1 && sm->get_axis() != 3) ||
+ tk->get_axis() != 3 || tk->get_mode() != v11::TopK::Mode::MAX ||
+ tk->get_sort_type() != v11::TopK::SortType::SORT_VALUES || tk->get_stable() ||
+ tk->get_index_element_type() != ov::element::i32 ||
+ !tk->output(0).get_target_inputs().empty()) {
+ return false;
+ }
+ const int64_t experts = shape[3].get_length();
+ auto end = ov::as_type_ptr<v0::Constant>(node(ids)->get_input_node_shared_ptr(2));
+ if (!end || !end->get_element_type().is_integral_number() || ov::shape_size(end->get_shape()) != 1) {
+ return false;
+ }
+ const int64_t k = end->cast_vector<int64_t>()[0];
+ const auto sorted_shape = tk->get_output_partial_shape(1);
+ if (k <= 0 || k > experts || sorted_shape[3].is_dynamic() || sorted_shape[3].get_length() < k ||
+ !input_is(probs, 1, {1, -1, experts, 1}) ||
+ !input_is(data, 1, {0}) || !input_is(ids_2d, 1, {0, 1}) ||
+ !input_is(weights_4d, 1, {0}) || !input_is(weights, 1, {1, 1, -1, k}) ||
+ ov::as_type_ptr<v1::Reshape>(node(probs))->get_special_zero() ||
+ ov::as_type_ptr<v1::Reshape>(node(weights))->get_special_zero() ||
+ !input_is(gather, 2, {1}) || ov::as_type_ptr<v8::Gather>(node(gather))->get_batch_dims() != 1 ||
+ !input_is(data_batch, 1, {0}) || !input_is(data_batch, 2, {0}) ||
+ !input_is(ids_count, 1, {1}) || !input_is(ids_count, 2, {0}) ||
+ ov::as_type_ptr<v8::Gather>(node(data_batch))->get_batch_dims() != 0 ||
+ ov::as_type_ptr<v8::Gather>(node(ids_count))->get_batch_dims() != 0 ||
+ ov::as_type_ptr<v0::Concat>(node(target))->get_axis() != 0 ||
+ ov::as_type_ptr<v3::Broadcast>(node(broadcast))->get_broadcast_spec().m_type != ov::op::BroadcastType::BIDIRECTIONAL ||
+ !reduce->get_keep_dims() || (!input_is(sum, 1, {-1}) && !input_is(sum, 1, {3})) ||
+ !input_is(tile, 1, {1, 1, 1, k}) ||
+ ov::as_type_ptr<v1::Divide>(node(norm))->get_autob().m_type != ov::op::AutoBroadcastType::NUMPY) {
+ return false;
+ }
+ // The top-k softmax sum is at least k / experts. Leave margin for rounding.
+ // This guard is only valid for unbiased softmax routing.
+ if (!(limit->get_min() <= 0.5 * double(k) / experts && limit->get_max() >= 2.0)) {
+ return false;
+ }
+ std::vector<std::shared_ptr<ov::Node>> slices;
+ for (const auto & input : tk->output(1).get_target_inputs()) {
+ auto slice = ov::as_type_ptr<v8::Slice>(input.get_node()->shared_from_this());
+ if (!slice || input.get_index() != 0 || slice->get_input_size() != 5 ||
+ !constant_is(slice->input_value(1), {0}) || !constant_is(slice->input_value(2), {k}) ||
+ !constant_is(slice->input_value(3), {1}) ||
+ (!constant_is(slice->input_value(4), {3}) && !constant_is(slice->input_value(4), {-1}))) {
+ return false;
+ }
+ slices.push_back(slice);
+ }
+
+ // Remove GGML's leading singleton dimension while building the plugin pattern,
+ // then restore it on both outputs for the following GGML nodes.
+ auto axis0 = v0::Constant::create(ov::element::i64, ov::Shape{1}, {0});
+ auto hidden = std::make_shared<v0::Squeeze>(mm->input_value(0), axis0);
+ auto routing = std::make_shared<v0::MatMul>(hidden, mm->input_value(1), false, mm->get_transpose_b());
+ auto probabilities = std::make_shared<v8::Softmax>(routing, -1);
+ auto selected = std::make_shared<v11::TopK>(probabilities,
+ v0::Constant::create(ov::element::i64, ov::Shape{}, {k}), 2,
+ v11::TopK::Mode::MAX, v11::TopK::SortType::SORT_VALUES, tk->get_index_element_type(), false);
+ auto total = std::make_shared<v1::ReduceSum>(selected->output(0),
+ v0::Constant::create(ov::element::i64, ov::Shape{1}, {-1}), true);
+ auto normalized = std::make_shared<v1::Divide>(selected->output(0), total);
+ auto weights_out = std::make_shared<v0::Unsqueeze>(normalized, axis0);
+ auto ids_out = std::make_shared<v0::Unsqueeze>(selected->output(1), axis0);
+ ov::copy_runtime_info({mm, sm, tk, node(norm)}, {hidden, routing, probabilities, selected, total, normalized, weights_out, ids_out});
+ ov::replace_node(node(norm), weights_out);
+ for (const auto & slice : slices) {
+ ov::replace_node(slice, ids_out);
+ }
+ return true;
+ };
+ register_matcher(std::make_shared<Matcher>(norm, "ov::frontend::ggml::pass::FuseMoeRouter"), callback);
+}
+
+} // namespace pass
+} // namespace ggml
+} // namespace frontend
+} // namespace ov
diff --git a/ggml/src/ggml-openvino/openvino/pass/fuse_moe_router.h b/ggml/src/ggml-openvino/openvino/pass/fuse_moe_router.h
new file mode 100644
index 000000000..49e01f16d
--- /dev/null
+++ b/ggml/src/ggml-openvino/openvino/pass/fuse_moe_router.h
@@ -0,0 +1,19 @@
+#pragma once
+
+#include <openvino/pass/matcher_pass.hpp>
+
+namespace ov {
+namespace frontend {
+namespace ggml {
+namespace pass {
+
+class FuseMoeRouter : public ov::pass::MatcherPass {
+public:
+ OPENVINO_MATCHER_PASS_RTTI("ov::frontend::ggml::pass::FuseMoeRouter")
+ FuseMoeRouter();
+};
+
+} // namespace pass
+} // namespace ggml
+} // namespace frontend
+} // namespace ov
diff --git a/ggml/src/ggml-openvino/openvino/pass/fuse_to_conv.cpp b/ggml/src/ggml-openvino/openvino/pass/fuse_to_conv.cpp
index 21801c0f3..6fffd6718 100644
--- a/ggml/src/ggml-openvino/openvino/pass/fuse_to_conv.cpp
+++ b/ggml/src/ggml-openvino/openvino/pass/fuse_to_conv.cpp
@@ -7,13 +7,14 @@
#include <openvino/op/convert.hpp>
#include <openvino/op/convolution.hpp>
#include <openvino/op/extractimagepatches.hpp>
+#include <openvino/op/group_conv.hpp>
#include <openvino/op/matmul.hpp>
#include <openvino/op/pad.hpp>
#include <openvino/op/reshape.hpp>
#include <openvino/op/transpose.hpp>
-#include <openvino/pass/pattern/op/label.hpp>
#include <openvino/pass/pattern/op/pattern.hpp>
#include <openvino/pass/pattern/op/wrap_type.hpp>
+#include <utility>
namespace opp = ov::pass::pattern;
@@ -22,180 +23,250 @@ namespace frontend {
namespace ggml {
namespace pass {
-// This pass fuses an IM2COL + MatMul convolution into OpenVINO's Convolution op for performance gains.
-// Reference the im2col.cpp translator for reference on the pattern being matched.
-
FuseToConv::FuseToConv() {
- const auto m_wei = opp::any_input();
- const auto m_act = opp::any_input();
- const auto m_matmul = opp::wrap_type<ov::op::v0::MatMul>({m_wei, m_act});
-
- const auto callback = [=](ov::pass::pattern::Matcher & m) {
- const auto & pm = m.get_pattern_value_map();
+ const auto m_in0 = opp::any_input();
+ const auto m_in1 = opp::any_input();
+ const auto m_matmul = opp::wrap_type<ov::op::v0::MatMul>({m_in0, m_in1});
- auto matmul_node = ov::as_type_ptr<ov::op::v0::MatMul>(pm.at(m_matmul).get_node_shared_ptr());
- if (!matmul_node || matmul_node->get_transpose_a() || !matmul_node->get_transpose_b()) {
+ const auto callback = [=](opp::Matcher & m) {
+ auto matmul_node = ov::as_type_ptr<ov::op::v0::MatMul>(m.get_match_root());
+ if (!matmul_node) {
return false;
}
- auto trace = matmul_node->input_value(1);
-
- // Optional Convert
- if (auto n = ov::as_type_ptr<ov::op::v0::Convert>(trace.get_node_shared_ptr())) {
- trace = n->input_value(0);
- }
-
- for (int i = 0; i < 2; ++i) {
- auto n = ov::as_type_ptr<ov::op::v1::Reshape>(trace.get_node_shared_ptr());
- if (!n) {
+ auto unwrap = [](ov::Output<Node> n) {
+ while (ov::is_type<ov::op::v0::Convert>(n.get_node_shared_ptr()) ||
+ ov::is_type<ov::op::v1::Reshape>(n.get_node_shared_ptr())) {
+ n = n.get_node_shared_ptr()->input_value(0);
+ }
+ return n;
+ };
+
+ auto get_im2col = [&](ov::Output<Node> trace)
+ -> std::pair<std::shared_ptr<ov::op::v3::ExtractImagePatches>, std::shared_ptr<ov::op::v1::Pad>> {
+ auto t2 = ov::as_type_ptr<ov::op::v1::Transpose>(unwrap(std::move(trace)).get_node_shared_ptr());
+ auto r1 = t2 ? ov::as_type_ptr<ov::op::v1::Reshape>(t2->get_input_node_shared_ptr(0)) : nullptr;
+ auto t1 = r1 ? ov::as_type_ptr<ov::op::v1::Transpose>(r1->get_input_node_shared_ptr(0)) : nullptr;
+ auto eip =
+ t1 ? ov::as_type_ptr<ov::op::v3::ExtractImagePatches>(t1->get_input_node_shared_ptr(0)) : nullptr;
+ auto pad = eip ? ov::as_type_ptr<ov::op::v1::Pad>(eip->get_input_node_shared_ptr(0)) : nullptr;
+ return {eip, pad};
+ };
+
+ bool weight_is_in0 = true;
+ auto [eip, pad] = get_im2col(matmul_node->input_value(1));
+ ov::Output<Node> w_trace = matmul_node->input_value(0);
+ if (!pad) {
+ std::tie(eip, pad) = get_im2col(matmul_node->input_value(0));
+ w_trace = matmul_node->input_value(1);
+ weight_is_in0 = false;
+ if (!pad) {
return false;
}
- trace = n->input_value(0);
}
- if (auto n = ov::as_type_ptr<ov::op::v1::Transpose>(trace.get_node_shared_ptr())) {
- trace = n->input_value(0);
- } else {
+ auto pb_const = ov::as_type_ptr<ov::op::v0::Constant>(pad->get_input_node_shared_ptr(1));
+ auto pe_const = ov::as_type_ptr<ov::op::v0::Constant>(pad->get_input_node_shared_ptr(2));
+ if (!pb_const || !pe_const) {
return false;
}
- if (auto n = ov::as_type_ptr<ov::op::v1::Reshape>(trace.get_node_shared_ptr())) {
- trace = n->input_value(0);
- } else {
+ const auto pb = pb_const->cast_vector<std::ptrdiff_t>();
+ const auto pe = pe_const->cast_vector<std::ptrdiff_t>();
+ if (pb.size() < 4 || pe.size() < 4) {
return false;
}
- if (auto n = ov::as_type_ptr<ov::op::v1::Transpose>(trace.get_node_shared_ptr())) {
- trace = n->input_value(0);
- } else {
+ auto image_input = pad->input_value(0);
+ auto image_shape = image_input.get_partial_shape();
+ if (image_shape.rank() != 4 || image_shape[1].is_dynamic()) {
return false;
}
+ const size_t IC = static_cast<size_t>(image_shape[1].get_length());
- auto eip = ov::as_type_ptr<ov::op::v3::ExtractImagePatches>(trace.get_node_shared_ptr());
- if (!eip) {
+ w_trace = unwrap(w_trace);
+ auto weight_pshape = w_trace.get_partial_shape();
+ if (!weight_pshape.is_static()) {
return false;
}
- const auto eip_strides = eip->get_strides(); // {stride_h, stride_w}
- const auto eip_rates = eip->get_rates(); // {dil_h, dil_w}
- auto pad = ov::as_type_ptr<ov::op::v1::Pad>(eip->input_value(0).get_node_shared_ptr());
- if (!pad) {
- return false;
+ const auto & ws = weight_pshape.to_shape();
+ const size_t KH = eip->get_sizes()[0];
+ const size_t KW = eip->get_sizes()[1];
+ const size_t kernel_spatial_ic = IC * KH * KW;
+
+ size_t groups = 0;
+ if (IC == 1 && image_shape[0].is_static() && image_shape[2].is_static() && image_shape[3].is_static()) {
+ if ((ws.size() == 4 && ws[1] == 1 && ws[2] == KH && ws[3] == KW && ws[0] > 1) ||
+ (ws.size() == 2 && ws[1] == KH * KW && ws[0] > 1) ||
+ (ws.size() == 3 && ws[1] == 1 && ws[2] == KH * KW && ws[0] > 1)) {
+ groups = ws[0];
+ } else if (ws.size() == 4 && ws[0] == 1 && ws[2] == 1 && ws[3] == KH * KW && ws[1] > 1) {
+ groups = ws[1];
+ }
}
- auto pads_begin_const =
- ov::as_type_ptr<ov::op::v0::Constant>(pad->input_value(1).get_node_shared_ptr());
- const auto pads_begin_vals = pads_begin_const->cast_vector<int64_t>(); // {0, 0, pad_h, pad_w}
- const std::ptrdiff_t pad_h = static_cast<std::ptrdiff_t>(pads_begin_vals[2]);
- const std::ptrdiff_t pad_w = static_cast<std::ptrdiff_t>(pads_begin_vals[3]);
+ const bool is_depthwise = groups > 1 && (image_shape[0].get_length() % groups == 0);
- auto image_input = pad->input_value(0); // [N, IC, 1, IW] NCHW
+ ov::Output<Node> conv_out;
+ size_t OC = 0;
- auto w_trace = matmul_node->input_value(0);
- if (auto n = ov::as_type_ptr<ov::op::v0::Convert>(w_trace.get_node_shared_ptr())) {
- w_trace = n->input_value(0);
- }
- for (int i = 0; i < 2; ++i) {
- auto n = ov::as_type_ptr<ov::op::v1::Reshape>(w_trace.get_node_shared_ptr());
- if (!n) {
- break;
+ if (is_depthwise) {
+ const size_t N = static_cast<size_t>(image_shape[0].get_length()) / groups;
+ const size_t IH = static_cast<size_t>(image_shape[2].get_length());
+ const size_t IW = static_cast<size_t>(image_shape[3].get_length());
+ OC = groups;
+
+ auto img_shape_const = register_new_node<ov::op::v0::Constant>(
+ ov::element::i64, ov::Shape{4},
+ std::vector<int64_t>{static_cast<int64_t>(N), static_cast<int64_t>(groups), static_cast<int64_t>(IH),
+ static_cast<int64_t>(IW)});
+ ov::Output<Node> image_reshaped =
+ register_new_node<ov::op::v1::Reshape>(image_input, img_shape_const, false);
+
+ const ov::Shape conv_w_shape = {groups, 1, 1, KH, KW};
+ ov::Output<Node> weight_input;
+ if (auto weight_const = ov::as_type_ptr<ov::op::v0::Constant>(w_trace.get_node_shared_ptr())) {
+ weight_input = register_new_node<ov::op::v0::Constant>(weight_const->get_element_type(), conv_w_shape,
+ weight_const->get_data_ptr());
+ } else {
+ auto shape_const = register_new_node<ov::op::v0::Constant>(
+ ov::element::i64, ov::Shape{5},
+ std::vector<int64_t>{static_cast<int64_t>(groups), 1, 1, static_cast<int64_t>(KH),
+ static_cast<int64_t>(KW)});
+ weight_input = register_new_node<ov::op::v1::Reshape>(w_trace, shape_const, false);
}
- w_trace = n->input_value(0);
- }
- auto weight_const = ov::as_type_ptr<ov::op::v0::Constant>(w_trace.get_node_shared_ptr());
- if (!weight_const) {
- return false;
- }
+ if (weight_input.get_element_type() != image_reshaped.get_element_type()) {
+ weight_input = register_new_node<ov::op::v0::Convert>(weight_input, image_reshaped.get_element_type());
+ }
- // Reshape weight to [OC, IC, 1, KW] (OIHW).
- const auto w_shape = weight_const->get_shape();
- ov::Shape conv_w_shape;
- if (w_shape.size() == 3) {
- conv_w_shape = {w_shape[0], w_shape[1], 1, w_shape[2]};
- } else if (w_shape.size() == 4) {
- conv_w_shape = {w_shape[1], w_shape[2], 1, w_shape[3]};
+ conv_out = register_new_node<ov::op::v1::GroupConvolution>(
+ image_reshaped, weight_input, eip->get_strides(), ov::CoordinateDiff{pb[2], pb[3]},
+ ov::CoordinateDiff{pe[2], pe[3]}, ov::Strides{eip->get_rates()[0], eip->get_rates()[1]},
+ ov::op::PadType::EXPLICIT);
} else {
- return false;
- }
+ if ((ws.size() == 4 && ws[1] == IC && ws[2] == KH && ws[3] == KW) ||
+ (ws.size() == 3 && ws[1] == IC && ws[2] == KW) || (ws.size() == 2 && ws[1] == kernel_spatial_ic)) {
+ OC = ws[0];
+ } else if ((ws.size() == 4 && ws[0] == 1 && ws[2] == IC && ws[3] == KW) ||
+ (ws.size() == 2 && ws[0] == kernel_spatial_ic)) {
+ OC = ws[1];
+ } else if (ws.size() == 3 && ws[0] == 1 && ws[1] == IC && ws[2] == KW) {
+ OC = 1;
+ } else if (ws.size() == 4 && ws[3] == kernel_spatial_ic) {
+ OC = ws[2];
+ } else if (ws.size() == 4 && ws[2] == kernel_spatial_ic) {
+ OC = ws[3];
+ } else if (kernel_spatial_ic && ov::shape_size(ws) % kernel_spatial_ic == 0) {
+ OC = ov::shape_size(ws) / kernel_spatial_ic;
+ } else {
+ return false;
+ }
+
+ const ov::Shape conv_w_shape = {OC, IC, KH, KW};
- auto weight_reshaped = register_new_node<ov::op::v0::Constant>(weight_const->get_element_type(), conv_w_shape,
+ ov::Output<Node> weight_input;
+ if (auto weight_const = ov::as_type_ptr<ov::op::v0::Constant>(w_trace.get_node_shared_ptr())) {
+ weight_input = register_new_node<ov::op::v0::Constant>(weight_const->get_element_type(), conv_w_shape,
weight_const->get_data_ptr());
+ } else {
+ auto shape_const = register_new_node<ov::op::v0::Constant>(
+ ov::element::i64, ov::Shape{4},
+ std::vector<int64_t>{static_cast<int64_t>(OC), static_cast<int64_t>(IC), static_cast<int64_t>(KH),
+ static_cast<int64_t>(KW)});
+ weight_input = register_new_node<ov::op::v1::Reshape>(w_trace, shape_const, false);
+ }
- ov::Output<Node> weight_input = weight_reshaped;
- if (weight_reshaped->get_element_type() != image_input.get_element_type()) {
- weight_input = register_new_node<ov::op::v0::Convert>(weight_reshaped, image_input.get_element_type());
- }
+ if (weight_input.get_element_type() != image_input.get_element_type()) {
+ weight_input = register_new_node<ov::op::v0::Convert>(weight_input, image_input.get_element_type());
+ }
- auto conv = register_new_node<ov::op::v1::Convolution>(
- image_input, weight_input,
- ov::Strides{static_cast<size_t>(eip_strides[0]), static_cast<size_t>(eip_strides[1])},
- ov::CoordinateDiff{pad_h, pad_w}, ov::CoordinateDiff{pad_h, pad_w},
- ov::Strides{static_cast<size_t>(eip_rates[0]), static_cast<size_t>(eip_rates[1])},
- ov::op::PadType::EXPLICIT);
+ conv_out = register_new_node<ov::op::v1::Convolution>(
+ image_input, weight_input, eip->get_strides(), ov::CoordinateDiff{pb[2], pb[3]},
+ ov::CoordinateDiff{pe[2], pe[3]}, ov::Strides{eip->get_rates()[0], eip->get_rates()[1]},
+ ov::op::PadType::EXPLICIT);
+ }
constexpr auto target_type = ov::element::f32;
- ov::Output<Node> conv_out = conv;
if (conv_out.get_element_type() != target_type) {
conv_out = register_new_node<ov::op::v0::Convert>(conv_out, target_type);
}
std::shared_ptr<ov::op::v1::Add> add_node;
ov::Output<Node> bias_input;
- for (const auto & consumer_in : matmul_node->output(0).get_target_inputs()) {
- auto cast = ov::as_type_ptr<ov::op::v0::Convert>(consumer_in.get_node()->shared_from_this());
- if (!cast) {
- continue;
+
+ auto try_fuse_bias = [&](const std::shared_ptr<Node> & n) {
+ auto add = ov::as_type_ptr<ov::op::v1::Add>(n);
+ if (!add) {
+ return false;
}
- for (const auto & add_in : cast->output(0).get_target_inputs()) {
- auto add = ov::as_type_ptr<ov::op::v1::Add>(add_in.get_node()->shared_from_this());
- if (!add) {
- continue;
+ for (size_t i = 0; i < 2; ++i) {
+ if (ov::is_type<ov::op::v0::Constant>(add->get_input_node_shared_ptr(i))) {
+ bias_input = add->input_value(i);
+ add_node = add;
+ return true;
}
- for (size_t i = 0; i < 2; ++i) {
- if (ov::as_type_ptr<ov::op::v0::Constant>(add->input_value(i).get_node_shared_ptr())) {
- bias_input = add->input_value(i);
- add_node = add;
+ }
+ return false;
+ };
+
+ for (const auto & consumer : matmul_node->output(0).get_target_inputs()) {
+ auto n = consumer.get_node()->shared_from_this();
+ if (try_fuse_bias(n)) {
+ break;
+ }
+ if (ov::is_type<ov::op::v0::Convert>(n) || ov::is_type<ov::op::v1::Reshape>(n)) {
+ for (const auto & next : n->output(0).get_target_inputs()) {
+ if (try_fuse_bias(next.get_node()->shared_from_this())) {
break;
}
}
- if (add_node) {
- break;
- }
}
if (add_node) {
break;
}
}
- ov::Output<Node> final_out;
- std::shared_ptr<Node> target_node;
+ ov::Output<Node> final_out = conv_out;
+ std::shared_ptr<Node> target_node = matmul_node;
if (add_node) {
- // Reshape bias [OC, 1] → [1, OC, 1, 1] for NCHW broadcasting.
ov::Output<Node> bias = bias_input;
if (bias.get_element_type() != target_type) {
bias = register_new_node<ov::op::v0::Convert>(bias, target_type);
}
- const auto oc = static_cast<int64_t>(conv_w_shape[0]);
- auto bias_shape = register_new_node<ov::op::v0::Constant>(ov::element::i64, ov::Shape{4},
- std::vector<int64_t>{1, oc, 1, 1});
+ auto bias_shape = register_new_node<ov::op::v0::Constant>(
+ ov::element::i64, ov::Shape{4}, std::vector<int64_t>{1, static_cast<int64_t>(OC), 1, 1});
bias = register_new_node<ov::op::v1::Reshape>(bias, bias_shape, false);
final_out = register_new_node<ov::op::v1::Add>(conv_out, bias);
target_node = add_node;
- } else {
- final_out = conv_out;
- target_node = matmul_node;
}
- // Reshape final output back to the target node's original shape if needed.
+ if (!is_depthwise) {
+ auto perm = register_new_node<ov::op::v0::Constant>(
+ ov::element::i64, ov::Shape{4},
+ weight_is_in0 ? std::vector<int64_t>{1, 0, 2, 3} : std::vector<int64_t>{0, 2, 3, 1});
+ final_out = register_new_node<ov::op::v1::Transpose>(final_out, perm);
+ }
+
auto orig_shape = target_node->get_output_partial_shape(0);
+ if (orig_shape.is_static() && final_out.get_partial_shape().is_static()) {
+ if (ov::shape_size(orig_shape.to_shape()) != ov::shape_size(final_out.get_shape())) {
+ return false;
+ }
+ }
if (orig_shape.is_static() && final_out.get_partial_shape() != orig_shape) {
auto shape_const = register_new_node<ov::op::v0::Constant>(ov::element::i64, ov::Shape{orig_shape.size()},
orig_shape.to_shape());
final_out = register_new_node<ov::op::v1::Reshape>(final_out, shape_const, false);
}
+ auto orig_type = target_node->get_output_element_type(0);
+ if (final_out.get_element_type() != orig_type) {
+ final_out = register_new_node<ov::op::v0::Convert>(final_out, orig_type);
+ }
+
final_out.get_node_shared_ptr()->set_friendly_name(target_node->get_friendly_name());
ov::copy_runtime_info(m.get_matched_nodes(), final_out.get_node_shared_ptr());
ov::replace_node(target_node, final_out.get_node_shared_ptr());
diff --git a/ggml/src/ggml-openvino/openvino/translate_session.cpp b/ggml/src/ggml-openvino/openvino/translate_session.cpp
index e56a4e41d..5a3d11f2d 100644
--- a/ggml/src/ggml-openvino/openvino/translate_session.cpp
+++ b/ggml/src/ggml-openvino/openvino/translate_session.cpp
@@ -5,6 +5,8 @@
#include "ggml-openvino/openvino/node_context.h"
#include "ggml-openvino/openvino/utils.h"
#include "input_model.h"
+#include "pass/fuse_argsort_topk.h"
+#include "pass/fuse_moe_router.h"
#include "pass/fuse_moe_compressed.h"
#include "pass/fuse_to_conv.h"
#include "pass/kv_state_seq_axis.h"
@@ -117,7 +119,7 @@ ov::pass::MakeStateful::ParamResPairs get_kv_param_res_pairs(
return pairs;
}
-void add_sliced_mask_stateful(TensorMap & tensor_map) {
+void add_sliced_mask_stateful(TensorMap & tensor_map, bool imrope) {
auto create_sliced_mask = [&](const std::string & mask_name, const std::string & sliced_name) {
if ((tensor_map.find(mask_name) != tensor_map.end()) &&
(tensor_map.find("token_len_per_seq") != tensor_map.end()) &&
@@ -134,7 +136,12 @@ void add_sliced_mask_stateful(TensorMap & tensor_map) {
auto axes = ov::op::v0::Constant::create(ov::element::i64, {1}, {-1});
auto inp_pos = tensor_map.at("inp_pos").get_node_shared_ptr();
- auto last_inp_pos = std::make_shared<ov::op::v8::Gather>(inp_pos, neg_one, three);
+ // IMROPE's fourth position plane is zero for text; use the token's first plane.
+ ov::Output<ov::Node> last_index = neg_one;
+ if (imrope) {
+ last_index = std::make_shared<ov::op::v1::Add>(token_len_per_seq, neg_one);
+ }
+ auto last_inp_pos = std::make_shared<ov::op::v8::Gather>(inp_pos, last_index, three);
auto last_inp_pos_1d = std::make_shared<ov::op::v1::Reshape>(
last_inp_pos, ov::op::v0::Constant::create(ov::element::i64, {1}, {1}), false);
auto last_inp_pos_cvt = std::make_shared<ov::op::v0::Convert>(last_inp_pos_1d, ov::element::i64);
@@ -242,7 +249,7 @@ void add_rope_sin_cos(TensorMap & tensor_map, GgmlDecoder & ggml_model_decoder)
// Create common patterns
void preprocess(TensorMap & tensor_map, GgmlDecoder & ggml_model_decoder) {
if (ggml_model_decoder.is_stateful()) {
- add_sliced_mask_stateful(tensor_map);
+ add_sliced_mask_stateful(tensor_map, ggml_model_decoder.get_rope_params()[2] == GGML_ROPE_TYPE_IMROPE);
add_position_mask_stateful_swa(tensor_map);
}
// This optimization is error-prone
@@ -300,13 +307,22 @@ std::shared_ptr<Model> TranslateSession::translate_graph(const frontend::InputMo
return ov::OutputVector{};
}
+ const auto & node_output_names = decoder->get_output_names(node_idx);
+ if (operation_type == "GGML_OP_VIEW" && decoder->get_op_case(node_idx) == 2 && node_output_names.size() == 1) {
+ auto direct_output = tensor_map->find(node_output_names[0]);
+ if (direct_output != tensor_map->end()) {
+ // GDN publishes its native attention/state outputs under the two GGML VIEW names.
+ // Keep those mappings instead of rebuilding slices of a packed temporary.
+ return ov::OutputVector{direct_output->second};
+ }
+ }
+
auto it = m_translator_map.find(operation_type);
FRONT_END_OP_CONVERSION_CHECK(it != m_translator_map.end(), "Translation for operation type ", operation_type,
" is not implemented.");
NodeContext node_context(decoder, tensor_map, node_idx, this);
ov::OutputVector converted_outputs = it->second(node_context);
- const auto & node_output_names = decoder->get_output_names(node_idx);
FRONT_END_OP_CONVERSION_CHECK(node_output_names.size() == converted_outputs.size(), "Number of ",
operation_type, " outputs greater than number of converted outputs, which are ",
node_output_names.size(), " and ", converted_outputs.size(), " respectively.");
@@ -468,11 +484,18 @@ std::shared_ptr<Model> TranslateSession::apply_transformations(std::shared_ptr<M
manager.register_pass<ov::pass::MarkDequantization>(
std::vector<ov::element::Type>{ov::element::u8, ov::element::i8, ov::element::u4, ov::element::i4});
manager.register_pass<pass::FuseToConv>();
+ manager.register_pass<pass::FuseArgsortTopK>();
- // MOECompressed has no CPU plugin implementation, so keep the GatherMatmul path
- // everywhere else. Opt-in while the fused path is being brought up.
- if (ggml_openvino_get_device_name() == "GPU" && getenv("GGML_OPENVINO_MOE_OP")) {
+ // MOECompressed has no CPU plugin implementation, so enable it by default only on GPU.
+ // GGML_OPENVINO_MOE_OP=0 keeps the unfused GatherMatmul path for fallback/debugging.
+ if (ggml_openvino_is_gpu() &&
+ ggml_openvino_getenv_int("GGML_OPENVINO_MOE_OP", 1) != 0) {
+ manager.register_pass<pass::FuseMoeRouter>();
manager.register_pass<pass::FuseMoeCompressed>();
+ // Same fusion for models whose gate and up projections share one fused expert
+ // weight (gemma-4). It only matches that shape and only when the experts carry an
+ // integer zero point, so it is a no-op on the separate-gate/up models above.
+ manager.register_pass<pass::FuseMoeCompressedFusedGateUp>();
}
if (ggml_model_decoder->is_stateful()) {
diff --git a/ggml/src/ggml-openvino/openvino/utils.cpp b/ggml/src/ggml-openvino/openvino/utils.cpp
index 98a85e632..40ba938a6 100644
--- a/ggml/src/ggml-openvino/openvino/utils.cpp
+++ b/ggml/src/ggml-openvino/openvino/utils.cpp
@@ -70,11 +70,10 @@ OutputVector rename_outputs_with_suffix(const OutputVector & outputs, const std:
}
namespace {
-ov::Output<ov::Node> rope_yarn_ramp_mix(int n_dims, const float corr_dims[2], float ext_factor) {
- int half_n_dims = n_dims / 2;
- std::vector<float> dim_ids_vec(half_n_dims);
+ov::Output<ov::Node> rope_yarn_ramp_mix(int num_elements, const float corr_dims[2], float ext_factor) {
+ std::vector<float> dim_ids_vec(num_elements);
std::iota(dim_ids_vec.begin(), dim_ids_vec.end(), 0.0f);
- auto dim_ids = ov::op::v0::Constant::create(ov::element::f32, Shape{1, 1, 1, (size_t) half_n_dims}, dim_ids_vec);
+ auto dim_ids = ov::op::v0::Constant::create(ov::element::f32, Shape{1, 1, 1, (size_t) num_elements}, dim_ids_vec);
auto corr_low = ov::op::v0::Constant::create(ov::element::f32, Shape{1, 1, 1, 1}, {corr_dims[0]});
auto corr_high = ov::op::v0::Constant::create(ov::element::f32, Shape{1, 1, 1, 1}, {corr_dims[1]});
auto denom = std::make_shared<ov::op::v1::Maximum>(
@@ -114,8 +113,18 @@ void ggml_rope_yarn_corr_dims(int n_dims,
std::pair<ov::Output<Node>, ov::Output<Node>> make_sin_cos(int32_t * rope_params,
std::shared_ptr<ov::Node> inp_pos,
std::shared_ptr<ov::Node> rope_freqs_weight,
- bool imrope,
- bool stateful) {
+ int mode,
+ bool stateful,
+ int64_t head_dim) {
+ constexpr int TYPE_IMROPE = 2;
+ constexpr int TYPE_VISION = 3;
+ constexpr int TYPE_MROPE = 4;
+
+ const bool is_imrope = (mode == TYPE_IMROPE);
+ const bool is_vision = (mode == TYPE_VISION);
+ const bool is_mrope = (mode == TYPE_MROPE);
+ const bool is_multi = is_imrope || is_vision || is_mrope;
+
if (stateful) {
inp_pos =
std::make_shared<ov::op::v0::Squeeze>(inp_pos, ov::op::v0::Constant::create(ov::element::i64, {1}, {0}));
@@ -123,7 +132,7 @@ std::pair<ov::Output<Node>, ov::Output<Node>> make_sin_cos(int32_t * rope_params
auto pos_perm =
std::make_shared<ov::op::v0::Constant>(ov::element::i64, ov::Shape{3}, std::vector<int64_t>{2, 1, 0});
inp_pos = std::make_shared<ov::op::v1::Transpose>(inp_pos, pos_perm);
- } else if (imrope) {
+ } else if (is_multi) {
inp_pos = std::make_shared<ov::op::v0::Convert>(inp_pos, ov::element::f32);
auto pos_shape = ov::op::v0::Constant::create(ov::element::i64, ov::Shape{5}, {0, 0, 0, 4, -1});
inp_pos = std::make_shared<ov::op::v1::Reshape>(inp_pos, pos_shape, true);
@@ -155,25 +164,102 @@ std::pair<ov::Output<Node>, ov::Output<Node>> make_sin_cos(int32_t * rope_params
const float theta_scale = powf(freq_base, -2.0f / n_dims);
- std::vector<float> factor(n_dims_half);
-
- Output<Node> freq_factors;
-
Output<Node> theta;
float mscale = attn_factor;
- if (imrope) {
- std::vector<int64_t> gather_indices(n_dims_half);
- for (size_t j = 0; j < n_dims_half; j++) {
- gather_indices[j] = j % 3;
- factor[j] = std::pow(theta_scale, j);
+ if (is_multi) {
+ const size_t num_pairs = is_vision ? (n_dims > 0 ? (size_t) n_dims : (size_t) (head_dim / 2)) : n_dims_half;
+ int sections[4];
+ memcpy(sections, rope_params + 11, sizeof(int) * 4);
+ int sect_dims = sections[0] + sections[1] + sections[2] + sections[3];
+ if (sect_dims <= 0) {
+ sect_dims = num_pairs;
+ }
+ int sec_w = sections[0] + sections[1];
+ int sec_e = sec_w + sections[2];
+ std::vector<int64_t> gather_indices(num_pairs);
+ std::vector<float> factor(num_pairs);
+ for (size_t j = 0; j < num_pairs; j++) {
+ int sector = j % sect_dims;
+ int plane = 0;
+ if (is_imrope) {
+ if (sector % 3 == 1 && sector < 3 * sections[1]) {
+ plane = 1;
+ } else if (sector % 3 == 2 && sector < 3 * sections[2]) {
+ plane = 2;
+ } else if (sector % 3 == 0 && sector < 3 * sections[0]) {
+ plane = 0;
+ } else {
+ plane = 3;
+ }
+ factor[j] = std::pow(theta_scale, j);
+ } else if (is_vision) {
+ int p_idx = 0;
+ if (sector < sections[0]) {
+ plane = 0;
+ p_idx = sector;
+ } else if (sector < sections[0] + sections[1]) {
+ plane = 1;
+ p_idx = sector - sections[0];
+ } else if (sector < sec_w + sections[2]) {
+ plane = 2;
+ p_idx = sector - sec_w;
+ } else {
+ plane = 3;
+ p_idx = sector - sec_e;
+ }
+ factor[j] = std::pow(theta_scale, p_idx);
+ } else {
+ if (sector >= sections[0] && sector < sec_w) {
+ plane = 1;
+ } else if (sector >= sec_w && sector < sec_e) {
+ plane = 2;
+ } else if (sector >= sec_e) {
+ plane = 3;
+ } else {
+ plane = 0;
+ }
+ factor[j] = std::pow(theta_scale, j);
+ }
+ gather_indices[j] = plane;
}
auto gather_indices_const =
- std::make_shared<ov::op::v0::Constant>(ov::element::i64, ov::Shape{n_dims_half}, gather_indices);
+ std::make_shared<ov::op::v0::Constant>(ov::element::i64, ov::Shape{num_pairs}, gather_indices);
auto gather_axis = ov::op::v0::Constant::create(ov::element::i32, ov::Shape{}, {4});
inp_pos = std::make_shared<ov::op::v8::Gather>(inp_pos, gather_indices_const, gather_axis);
- auto factor_const = std::make_shared<ov::op::v0::Constant>(ov::element::f32, ov::Shape{n_dims_half}, factor);
- theta = std::make_shared<ov::op::v1::Multiply>(inp_pos, factor_const);
+ Output<Node> factor_node =
+ std::make_shared<ov::op::v0::Constant>(ov::element::f32, ov::Shape{num_pairs}, factor);
+ if (rope_freqs_weight) {
+ Output<Node> rope_factors = std::make_shared<ov::op::v8::Slice>(
+ rope_freqs_weight,
+ ov::op::v0::Constant::create(ov::element::i64, {1}, {0}),
+ ov::op::v0::Constant::create(ov::element::i64, {1}, {(int64_t) num_pairs}),
+ ov::op::v0::Constant::create(ov::element::i64, {1}, {1}),
+ ov::op::v0::Constant::create(ov::element::i64, {1}, {rope_freqs_weight->get_output_partial_shape(0).rank().get_length() - 1}));
+ rope_factors = std::make_shared<ov::op::v1::Reshape>(
+ rope_factors,
+ ov::op::v0::Constant::create(ov::element::i64, {1}, {(int64_t) num_pairs}),
+ false);
+ factor_node = std::make_shared<ov::op::v1::Divide>(factor_node, rope_factors);
+ }
+ auto theta_extrap = std::make_shared<ov::op::v1::Multiply>(inp_pos, factor_node);
+ auto theta_interp = std::make_shared<ov::op::v1::Multiply>(
+ theta_extrap, ov::op::v0::Constant::create(ov::element::f32, {1}, {freq_scale}));
+ if (ext_factor == 0.0f) {
+ theta = theta_interp;
+ } else {
+ float corr_dims[2];
+ ggml_rope_yarn_corr_dims(n_dims, n_ctx_orig, freq_base, beta_fast, beta_slow, corr_dims);
+ auto ramp_mix = rope_yarn_ramp_mix(num_pairs, corr_dims, ext_factor);
+ Output<Node> one = ov::op::v0::Constant::create(ov::element::f32, Shape{1, 1, 1, 1}, {1.0f});
+ auto one_minus_ramp = std::make_shared<ov::op::v1::Subtract>(one, ramp_mix);
+ theta = std::make_shared<ov::op::v1::Add>(
+ std::make_shared<ov::op::v1::Multiply>(theta_interp, one_minus_ramp),
+ std::make_shared<ov::op::v1::Multiply>(theta_extrap, ramp_mix));
+ mscale *= (1.0f + 0.1f * std::log(1.0f / freq_scale));
+ }
} else {
+ std::vector<float> factor(n_dims_half);
+ Output<Node> freq_factors;
float corr_dims[2];
ggml_rope_yarn_corr_dims(n_dims, n_ctx_orig, freq_base, beta_fast, beta_slow, corr_dims);
factor[0] = 1.0f;
@@ -215,7 +301,7 @@ std::pair<ov::Output<Node>, ov::Output<Node>> make_sin_cos(int32_t * rope_params
if (ext_factor == 0.0f) {
theta = theta_interp;
} else {
- auto ramp_mix = rope_yarn_ramp_mix(n_dims, corr_dims, ext_factor);
+ auto ramp_mix = rope_yarn_ramp_mix(n_dims_half, corr_dims, ext_factor);
Output<Node> one;
if (stateful) {
one = ov::op::v0::Constant::create(ov::element::f32, Shape{1, 1, 1}, {1.0f});
@@ -234,7 +320,7 @@ std::pair<ov::Output<Node>, ov::Output<Node>> make_sin_cos(int32_t * rope_params
Output<Node> cos_theta = std::make_shared<ov::op::v0::Cos>(theta);
Output<Node> sin_theta = std::make_shared<ov::op::v0::Sin>(theta);
- if (!imrope) {
+ if (mscale != 1.0f) {
auto mscale_node = ov::op::v0::Constant::create(ov::element::f32, Shape{}, {mscale});
cos_theta = std::make_shared<ov::op::v1::Multiply>(cos_theta, mscale_node);
@@ -305,18 +391,32 @@ ov::Output<ov::Node> process_view_input_new(const NodeContext & context, int inp
// here would re-slice/re-flatten the already-resolved single-plane view against the
// recorded (multi-plane) source strides and emit a constant-target Reshape whose baked
// dims no longer divide the concretized input -> "dimensions do not evenly divide".
+ // A fourth case matters for stateful execution: `expected` comes from ggml metadata and is
+ // always GGML_MAX_DIMS ranks, but stateful drops the leading size-1 batch dim, so the
+ // resolved view is one rank lower. Compare the common trailing dims and require the
+ // leading expected dims we skip to be 1. Without this the rank test below never matches on
+ // the stateful path and every already-resolved MoE expert-plane view is re-sliced.
auto expected_ov_shape = context.get_view_input_ov_shape(input_index, 0);
auto actual_shape = input.get_partial_shape();
if (expected_ov_shape.rank().is_static() && actual_shape.rank().is_static() &&
- expected_ov_shape.rank() == actual_shape.rank()) {
+ expected_ov_shape.rank().get_length() >= actual_shape.rank().get_length()) {
+ const int64_t n_actual = actual_shape.rank().get_length();
+ const int64_t shift = expected_ov_shape.rank().get_length() - n_actual;
bool shapes_match = true;
- for (int64_t i = 0; i < expected_ov_shape.rank().get_length(); ++i) {
- const bool both_dynamic = expected_ov_shape[i].is_dynamic() && actual_shape[i].is_dynamic();
- const bool both_static_equal = expected_ov_shape[i].is_static() && actual_shape[i].is_static() &&
- expected_ov_shape[i] == actual_shape[i];
+ for (int64_t i = 0; i < shift; ++i) {
+ if (!expected_ov_shape[i].is_static() || expected_ov_shape[i].get_length() != 1) {
+ shapes_match = false;
+ break;
+ }
+ }
+ for (int64_t i = 0; i < n_actual && shapes_match; ++i) {
+ const auto & exp = expected_ov_shape[i + shift];
+ const bool both_dynamic = exp.is_dynamic() && actual_shape[i].is_dynamic();
+ const bool both_static_equal =
+ exp.is_static() && actual_shape[i].is_static() && exp == actual_shape[i];
// expected dynamic, actual static: the resolved view already carries the
// concrete size for this fragment; reuse it rather than re-materializing.
- const bool expected_dyn_actual_static = expected_ov_shape[i].is_dynamic() && actual_shape[i].is_static();
+ const bool expected_dyn_actual_static = exp.is_dynamic() && actual_shape[i].is_static();
if (!both_dynamic && !both_static_equal && !expected_dyn_actual_static) {
shapes_match = false;
break;
@@ -399,6 +499,15 @@ ov::Output<ov::Node> process_view_input_new(const NodeContext & context, int inp
const ov::Shape & view_ggml_shape, const ov::PartialShape & view_ov_shape, const std::string & view_name,
size_t view_src_offset, const std::vector<size_t> & view_src_stride, const ov::Shape & view_src_ggml_shape,
const ov::PartialShape & view_src_ov_shape, const std::string & view_src_name) -> ov::Output<ov::Node> {
+ // Stateful execution drops a leading size-1 axis, so `current` can be one rank lower
+ // than the ggml shape metadata (view_stride/view_ggml_shape, always GGML_MAX_DIMS
+ // entries) assumes. Shift any axis index built from that metadata down by the
+ // difference before using it as an OV Slice axis.
+ const auto current_rank = current.get_partial_shape().rank();
+ const int axis_shift = current_rank.is_static() ?
+ static_cast<int>(view_stride.size()) - static_cast<int>(current_rank.get_length()) :
+ 0;
+
auto build_reshape_pattern = [](const ov::PartialShape & target_ov_shape,
const ov::Shape & target_ggml_shape) -> std::vector<int64_t> {
const size_t ndims = target_ggml_shape.size();
@@ -509,7 +618,7 @@ ov::Output<ov::Node> process_view_input_new(const NodeContext & context, int inp
current, ov::op::v0::Constant::create(ov::element::i64, {1}, {begin_val}),
ov::op::v0::Constant::create(ov::element::i64, {1}, {end_val}),
ov::op::v0::Constant::create(ov::element::i64, {1}, {1}),
- ov::op::v0::Constant::create(ov::element::i64, {1}, {slice_dim}));
+ ov::op::v0::Constant::create(ov::element::i64, {1}, {slice_dim - axis_shift}));
if (view_ov_shape.is_static()) {
auto reshaped = std::make_shared<ov::op::v1::Reshape>(
diff --git a/ggml/src/ggml-openvino/openvino/utils.h b/ggml/src/ggml-openvino/openvino/utils.h
index d9858f923..effa1f843 100644
--- a/ggml/src/ggml-openvino/openvino/utils.h
+++ b/ggml/src/ggml-openvino/openvino/utils.h
@@ -4,6 +4,7 @@
#include <memory>
#include <openvino/core/node.hpp>
+#include <openvino/op/convert.hpp>
#include <openvino/op/shape_of.hpp>
#include <openvino/op/slice.hpp>
#include <utility>
@@ -57,8 +58,9 @@ OutputVector rename_outputs_with_suffix(const OutputVector & outputs, const std:
std::pair<ov::Output<Node>, ov::Output<Node>> make_sin_cos(int32_t * rope_params,
std::shared_ptr<ov::Node> inp_pos,
std::shared_ptr<ov::Node> rope_freqs_weight = nullptr,
- bool imrope = false,
- bool stateful = false);
+ int mode = 0,
+ bool stateful = false,
+ int64_t head_dim = 0);
ov::Output<ov::Node> process_view_input(const NodeContext & context, int input_index, int slice_len = 0, int axis = -1);
@@ -69,7 +71,21 @@ template <typename T> OutputVector translate_1to1_match_2_inputs(const NodeConte
num_inputs_check(context, 2, 2);
auto input_0 = process_view_input_new(context, 0);
auto input_1 = process_view_input_new(context, 1);
- auto res = std::make_shared<T>(input_0, input_1);
+
+ auto output_type = context.get_output_type();
+ if (input_0.get_element_type() != input_1.get_element_type()) {
+ if (input_0.get_element_type() != ov::element::f32) {
+ input_0 = std::make_shared<ov::op::v0::Convert>(input_0, ov::element::f32);
+ }
+ if (input_1.get_element_type() != ov::element::f32) {
+ input_1 = std::make_shared<ov::op::v0::Convert>(input_1, ov::element::f32);
+ }
+ }
+
+ ov::Output<ov::Node> res = std::make_shared<T>(input_0, input_1);
+ if (res.get_element_type() != output_type) {
+ res = std::make_shared<ov::op::v0::Convert>(res, output_type);
+ }
return rename_outputs_with_suffix({res}, context.get_name());
}
diff --git a/ggml/src/ggml-openvino/utils.cpp b/ggml/src/ggml-openvino/utils.cpp
index b1ee792fd..67a4123cb 100644
--- a/ggml/src/ggml-openvino/utils.cpp
+++ b/ggml/src/ggml-openvino/utils.cpp
@@ -34,6 +34,7 @@
#include <openvino/runtime/properties.hpp>
#include <openvino/runtime/tensor.hpp>
#include <optional>
+#include <set>
#include <string>
#include <unordered_map>
#include <vector>
@@ -98,27 +99,27 @@ std::optional<ov::Tensor> try_make_kv_sliced_tensor(const std::shared_ptr<GgmlOv
ov::Shape sliced_shape = full_shape;
sliced_shape[2] = static_cast<size_t>(n_kv);
- // Disabling for now as gpu has bug with in-place ScatterUpdate with remote tensors, can re-enable once CVS-186519 is fixed
- // if (ggml_openvino_buffer_is_remote(ggml_tensor)) {
- // auto remote_context = ggml_openvino_get_remote_context();
- // auto gpu_context = remote_context->as<ov::intel_gpu::ocl::ClContext>();
- // return gpu_context.create_tensor(ggml_decoder->get_ov_type(ggml_tensor), sliced_shape, ggml_tensor->data);
- // }
+ if (ggml_openvino_buffer_is_remote(ggml_tensor)) {
+ auto remote_context = ggml_openvino_get_remote_context();
+ auto gpu_context = remote_context->as<ov::intel_gpu::ocl::ClContext>();
+ return gpu_context.create_tensor(ggml_decoder->get_ov_type(ggml_tensor), sliced_shape, ggml_tensor->data);
+ }
return ov::Tensor(GgmlOvDecoder::get_ov_type(ggml_tensor), sliced_shape, ggml_tensor->data);
}
-uint64_t ggml_openvino_model_cache_extra_cfg(const std::string & device, bool stateful) {
+static uint64_t ggml_openvino_model_cache_extra_cfg(bool stateful, bool recurrent) {
const char * manual_gqa_env = ggml_openvino_getenv_str("GGML_OPENVINO_MANUAL_GQA_ATTN");
const bool manual_gqa_enabled = manual_gqa_env != nullptr ?
ggml_openvino_getenv_int("GGML_OPENVINO_MANUAL_GQA_ATTN") > 0 :
- device == "GPU";
+ ggml_openvino_is_gpu();
uint64_t extra_cfg = 1; // Graph-ordinal port names (invalidate older disk-cache blobs).
- extra_cfg = extra_cfg * 131 + (stateful ? 1u : 0u);
+ extra_cfg = extra_cfg * 131 + (stateful ? (recurrent ? 2u : 1u) : 0u);
extra_cfg = extra_cfg * 131 + (ggml_openvino_reduce_compile_mem_enabled() ? 1u : 0u);
extra_cfg = extra_cfg * 131 + (ggml_openvino_getenv_int("GGML_OPENVINO_DISABLE_KV_SLICE") ? 1u : 0u);
extra_cfg = extra_cfg * 131 + (manual_gqa_enabled ? 1u : 0u);
+ extra_cfg = extra_cfg * 131 + (ggml_openvino_getenv_int("GGML_OPENVINO_PROFILING") >= 2 ? 1u : 0u);
return extra_cfg;
}
@@ -134,10 +135,11 @@ std::map<std::string, std::shared_ptr<ov::Node>> get_weight_names(ggml_cgraph *
// miss. Include topology, layouts, op parameters, constant extra inputs and weight
// allocation identities. Never use a sampled weight hash or a graph name alone:
// different models can have identical topology. OV buffer IDs survive address reuse.
-std::string compiled_graph_key(const ggml_cgraph * graph,
- const GgmlOvDecoder & decoder,
- const std::string & device,
- int prefill_chunk_size = 0) {
+static std::string compiled_graph_key(const ggml_cgraph * graph,
+ const GgmlOvDecoder & decoder,
+ const std::string & device,
+ int prefill_chunk_size = 0,
+ bool disk_cache = false) {
std::string key;
auto append = [&key](const auto & value) {
key.append(reinterpret_cast<const char *>(&value), sizeof(value));
@@ -173,7 +175,7 @@ std::string compiled_graph_key(const ggml_cgraph * graph,
const auto * base = tensor->view_src ? tensor->view_src : tensor;
const bool weight = base->buffer && base->buffer->usage == GGML_BACKEND_BUFFER_USAGE_WEIGHTS;
append(weight);
- if (weight) {
+ if (weight && !disk_cache) {
const size_t buffer_id = ggml_backend_openvino_buffer_get_ctx_id(base->buffer);
has_weight_buffer_id |= buffer_id != 0;
append(buffer_id);
@@ -206,7 +208,170 @@ std::string compiled_graph_key(const ggml_cgraph * graph,
}
// Without an allocation generation, pointer reuse could select stale weights.
// Such graphs still get private requests; they simply do not share compilation.
- return has_weight_buffer_id ? key : std::string{};
+ return disk_cache || has_weight_buffer_id ? key : std::string{};
+}
+
+static std::string dynamic_graph_signature(const ggml_cgraph * graph,
+ const GgmlOvDecoder & decoder,
+ const ModelParams & params) {
+ std::string key = "dynamic-1";
+ auto append = [&key](const auto & value) {
+ key.append(reinterpret_cast<const char *>(&value), sizeof(value));
+ };
+ auto append_string = [&](const std::string & value) {
+ append(value.size());
+ key.append(value);
+ };
+ graph_key graph_id(graph, true);
+ append(graph_id.n_nodes);
+ append(graph_id.n_leaves);
+ append_string(graph_id.first_node_name);
+ append_string(graph_id.last_node_name);
+ for (const auto & name : graph_id.input_srcs) {
+ append_string(name);
+ }
+ append(params.n_rs_slots);
+ append(params.has_rs_rollback);
+ append(params.mixed_rope_params);
+ append(params.is_cacheless_attn);
+ for (int layer : params.swa_layers) {
+ append(layer);
+ }
+ for (const auto & [layer, heads] : params.n_heads_kv_per_layer) {
+ append(layer);
+ append(heads);
+ }
+ for (const auto & [name, input] : decoder.get_model_inputs()) {
+ append_string(name);
+ append_string(input.type.get_type_name());
+ }
+ for (const auto & [name, tensor] : decoder.get_model_outputs()) {
+ append_string(name);
+ append(tensor->type);
+ }
+ for (const auto & [name, input] : decoder.get_model_extra_inputs()) {
+ append_string(name);
+ append_string(input.type.get_type_name());
+ append(input.is_parameter);
+ if (!input.is_parameter) {
+ append(input.value);
+ }
+ }
+ return key;
+}
+
+static bool compiled_model_matches_graph(const ov::CompiledModel & model,
+ const std::vector<std::string> & inputs,
+ const std::vector<std::string> & outputs,
+ const GgmlOvDecoder & decoder) {
+ if (model.inputs().size() != inputs.size() || model.outputs().size() != outputs.size()) {
+ return false;
+ }
+ const auto & graph_inputs = decoder.get_model_inputs();
+ const auto & extra_inputs = decoder.get_model_extra_inputs();
+ const auto & graph_outputs = decoder.get_model_outputs();
+ for (size_t i = 0; i < inputs.size(); ++i) {
+ const auto & port = model.input(i);
+ auto graph_it = graph_inputs.find(inputs[i]);
+ if (graph_it != graph_inputs.end()) {
+ if (port.get_element_type() != graph_it->second.type ||
+ !port.get_partial_shape().compatible(graph_it->second.shape)) {
+ return false;
+ }
+ continue;
+ }
+ auto extra_it = extra_inputs.find(inputs[i]);
+ if (extra_it == extra_inputs.end() || port.get_element_type() != extra_it->second.type ||
+ !port.get_partial_shape().compatible(ov::PartialShape(extra_it->second.shape))) {
+ return false;
+ }
+ }
+ for (size_t i = 0; i < outputs.size(); ++i) {
+ auto graph_it = graph_outputs.find(outputs[i]);
+ if (graph_it == graph_outputs.end()) {
+ if (outputs[i].compare(0, 8, "__debug_") != 0) {
+ return false;
+ }
+ continue;
+ }
+ const auto & port = model.output(i);
+ if (port.get_element_type() != GgmlOvDecoder::get_ov_type(graph_it->second) ||
+ !port.get_partial_shape().compatible(ov::PartialShape(GgmlOvDecoder::get_shape(graph_it->second)))) {
+ return false;
+ }
+ }
+ return true;
+}
+
+static std::string ov_profiling_csv_field(const std::string & value) {
+ std::string escaped = "\"";
+ for (char c : value) {
+ escaped += c;
+ if (c == '"') {
+ escaped += c;
+ }
+ }
+ escaped += '"';
+ return escaped;
+}
+
+static void dump_ov_profiling_info(const ov::InferRequest & infer_request) {
+ if (ggml_openvino_getenv_int("GGML_OPENVINO_PROFILING") < 2) {
+ return;
+ }
+
+ static std::atomic<uint64_t> inference_index{0};
+ const uint64_t index = inference_index.fetch_add(1);
+ const std::string path = "openvino_profile_" + std::to_string(index) + ".csv";
+ std::ofstream output(path, std::ios::trunc);
+ if (!output.is_open()) {
+ GGML_LOG_WARN("ggml-openvino: failed to write profiling data to %s\n", path.c_str());
+ return;
+ }
+
+ output << "status,real_time_us,cpu_time_us,start_time_us,node_type,node_name,exec_type\n";
+ int64_t total_real_time = 0;
+ int64_t total_cpu_time = 0;
+ size_t executed_count = 0;
+ for (const auto & info : infer_request.get_profiling_info()) {
+ const char * status = "NOT_RUN";
+ if (info.status == ov::ProfilingInfo::Status::EXECUTED) {
+ status = "EXECUTED";
+ total_real_time += info.real_time.count();
+ total_cpu_time += info.cpu_time.count();
+ executed_count++;
+ } else if (info.status == ov::ProfilingInfo::Status::OPTIMIZED_OUT) {
+ status = "OPTIMIZED_OUT";
+ }
+
+ output << status << ',' << info.real_time.count() << ',' << info.cpu_time.count() << ','
+ << info.start_time.count() << ',' << ov_profiling_csv_field(info.node_type) << ','
+ << ov_profiling_csv_field(info.node_name) << ',' << ov_profiling_csv_field(info.exec_type) << '\n';
+ }
+
+ GGML_LOG_INFO("ggml-openvino: profile %s: %zu executed nodes, %.3f ms device, %.3f ms CPU\n", path.c_str(),
+ executed_count, total_real_time / 1000.0, total_cpu_time / 1000.0);
+}
+
+static void reset_single_slot_recurrent_cache(ggml_cgraph * cgraph,
+ const ModelParams & model_params,
+ const ComputeParams & compute_params) {
+ if (model_params.n_rs_slots != 1 || compute_params.cache_rs_reset_len == 0) {
+ return;
+ }
+ GGML_ASSERT(compute_params.cache_rs_reset_idx == 0 && compute_params.cache_rs_reset_len == 1);
+
+ std::set<ggml_backend_buffer_t> buffers;
+ for (int i = 0; i < cgraph->n_nodes; i++) {
+ auto * node = cgraph->nodes[i];
+ if (node->op == GGML_OP_SCALE && GgmlOvDecoder::is_recurrent_cache(node->view_src) &&
+ node->view_src->ne[1] == 1 && node->view_src->buffer != nullptr) {
+ buffers.insert(node->view_src->buffer);
+ }
+ }
+ for (auto * buffer : buffers) {
+ ggml_backend_buffer_clear(buffer, 0);
+ }
}
ov::Tensor create_ov_output_tensor(const std::shared_ptr<GgmlOvDecoder> & ggml_decoder,
@@ -217,16 +382,7 @@ ov::Tensor create_ov_output_tensor(const std::shared_ptr<GgmlOvDecoder> & ggml_d
return *sliced;
}
- // Disabling for now as gpu has bug with in-place ScatterUpdate with remote tensors, can re-enable once CVS-186519 is fixed
- // if (ggml_tensor->extra != nullptr && !ggml_decoder->is_splited_model()) {
- // auto * extra_base = static_cast<ggml_openvino_extra_base *>(ggml_tensor->extra);
- // if (extra_base->type == ggml_openvino_extra_base::Type::TENSOR) {
- // auto * tensor_extra = static_cast<ggml_openvino_tensor_extra *>(extra_base);
- // return *tensor_extra->tensor;
- // }
- // }
-
- auto output_type = GgmlOvDecoder::get_ov_type(ggml_tensor);
+ auto output_type = ggml_decoder->get_ov_type(ggml_tensor);
ov::Shape output_shape;
void * output_data = ggml_tensor->data;
if (ggml_decoder->is_static()) {
@@ -245,6 +401,21 @@ ov::Tensor create_ov_output_tensor(const std::shared_ptr<GgmlOvDecoder> & ggml_d
output_shape = GgmlOvDecoder::get_shape(ggml_tensor);
}
}
+
+ // The sliced path above covers eligible KV outputs. This also covers recurrent state and full-size KV fallbacks.
+ if (!ggml_openvino_getenv_int("GGML_OPENVINO_DISABLE_REMOTE_OUTPUTS") && !ggml_decoder->is_static() &&
+ !ggml_decoder->is_splited_model() && ggml_openvino_buffer_is_remote(ggml_tensor) &&
+ ggml_tensor->extra != nullptr) {
+ auto * extra_base = static_cast<ggml_openvino_extra_base *>(ggml_tensor->extra);
+ if (extra_base->type == ggml_openvino_extra_base::Type::TENSOR) {
+ auto * tensor_extra = static_cast<ggml_openvino_tensor_extra *>(extra_base);
+ if (tensor_extra->tensor != nullptr && tensor_extra->tensor->get_element_type() == output_type &&
+ tensor_extra->tensor->get_shape() == output_shape) {
+ return *tensor_extra->tensor;
+ }
+ }
+ }
+
ov::Tensor output_tensor(output_type, output_shape, output_data);
return output_tensor;
}
@@ -632,40 +803,63 @@ enum ggml_status ov_graph_compute_dynamic(ggml_cgraph * cgraph, const std::share
auto & core = ov_singleton_core();
const auto & config = ggml_openvino_get_compile_config();
const auto & device = r_ctx->device;
- const auto & stateful = r_ctx->stateful;
static auto is_static = false;
static const bool cache_disabled = ggml_openvino_getenv_int("GGML_OPENVINO_DISABLE_CACHE");
+ auto start_time = ggml_time_us();
+
// is_model_splitted is O(n_nodes^2) plus a create_weight_nodes scan and takes ~20 ms
// on a Llama-1B decode graph. It is called once per graph_compute invocation but the
// graph shape is identical across all decode steps, so memoize by graph_key: compute
// graph_key first (a few hundred us), and if the same key is already in decoder_cache
// we know the graph is not splitted (only not-splitted graphs get inserted there).
+ const int64_t cache_key_start_time = ggml_time_us();
graph_key key(cgraph);
+ const int64_t cache_key_compute_time = ggml_time_us() - cache_key_start_time;
+ int64_t cache_lookup_time = 0;
bool key_seen = false;
if (!cache_disabled) {
+ const int64_t cache_lookup_start_time = ggml_time_us();
std::lock_guard<std::mutex> map_lock(r_ctx->ctx_mutex);
key_seen = r_ctx->decoder_cache.find(key) != r_ctx->decoder_cache.end();
+ cache_lookup_time += ggml_time_us() - cache_lookup_start_time;
}
+ const int64_t model_split_check_start_time = ggml_time_us();
bool model_is_splitted = key_seen ? false : is_model_splitted(cgraph);
+ const int64_t model_split_check_time = ggml_time_us() - model_split_check_start_time;
+ if (ggml_openvino_model_cache_only() && (model_is_splitted || is_naive(cgraph))) {
+ GGML_LOG_ERROR("ggml-openvino: cache-only mode requires a complete model graph\n");
+ return GGML_STATUS_FAILED;
+ }
if (is_naive(cgraph)) {
if (!model_is_splitted) {
return naive_compute(cgraph, core, device, config, *r_ctx->compiled_cache);
}
}
- auto start_time = ggml_time_us();
-
std::shared_ptr<GgmlOvDecoder> ggml_decoder;
std::shared_ptr<ov::InferRequest> infer_request;
ModelParams m_params;
ComputeParams c_params;
+ const int64_t graph_param_compute_start_time = ggml_time_us();
std::tie(m_params, c_params) = GgmlOvDecoder::compute_llm_params(cgraph, is_static);
+ const int64_t graph_param_compute_time = ggml_time_us() - graph_param_compute_start_time;
const bool cache_enabled = !model_is_splitted && !cache_disabled;
+ const bool stateful_recurrent = r_ctx->stateful && m_params.state_size >= 0;
+ const bool stateful_kv_only = r_ctx->stateful && !stateful_recurrent;
+ if (r_ctx->stateful &&
+ (m_params.n_seq > 1 || c_params.n_seq_active > 1 || m_params.n_rs_slots > 1 || m_params.has_rs_rollback)) {
+ GGML_LOG_ERROR("OpenVINO stateful execution requires a single slot without recurrent rollback. Use -np 1.\n");
+ return GGML_STATUS_FAILED;
+ }
+ if (stateful_recurrent && !cache_enabled) {
+ GGML_LOG_ERROR("OpenVINO recurrent stateful execution requires an unsplit graph with model caching enabled.\n");
+ return GGML_STATUS_FAILED;
+ }
bool cache_hit = false;
int64_t decoder_end_time;
@@ -678,6 +872,7 @@ enum ggml_status ov_graph_compute_dynamic(ggml_cgraph * cgraph, const std::share
std::shared_ptr<decoder_runtime_ctx> entry;
ModelParams old_m_params;
+ const int64_t cache_lookup_start_time = ggml_time_us();
if (cache_enabled) {
std::lock_guard<std::mutex> map_lock(r_ctx->ctx_mutex);
auto it = r_ctx->decoder_cache.find(key);
@@ -695,6 +890,7 @@ enum ggml_status ov_graph_compute_dynamic(ggml_cgraph * cgraph, const std::share
entry = std::make_shared<decoder_runtime_ctx>(mutex);
cache_hit = false;
}
+ cache_lookup_time += ggml_time_us() - cache_lookup_start_time;
std::lock_guard<std::mutex> lock(*(entry->mutex));
cache_hit = cache_hit && entry->ptr && r_ctx->infer_request_cache.count(key) != 0;
@@ -714,7 +910,7 @@ enum ggml_status ov_graph_compute_dynamic(ggml_cgraph * cgraph, const std::share
std::map<std::string, std::shared_ptr<ov::Node>> model_weights;
ggml_decoder->set_compute_params(c_params);
ggml_decoder->set_model_params(m_params);
- if (old_m_params.kv_buffer_changed(m_params)) {
+ if (old_m_params.kv_buffer_changed(m_params) || !ggml_decoder->is_bound_to(cgraph)) {
ggml_decoder->update_io(cgraph);
}
ggml_decoder->add_extra_inputs();
@@ -725,7 +921,7 @@ enum ggml_status ov_graph_compute_dynamic(ggml_cgraph * cgraph, const std::share
ov_output_names = r_ctx->ov_output_names_cache.at(key);
}
- if (stateful) {
+ if (stateful_kv_only) {
const auto * inp_pos = get_inp_pos_tensor(cgraph);
int32_t * pos_data = (int32_t *) inp_pos->data;
auto pos_shape = GgmlOvDecoder::get_shape(inp_pos);
@@ -830,67 +1026,85 @@ enum ggml_status ov_graph_compute_dynamic(ggml_cgraph * cgraph, const std::share
auto shared_cache = r_ctx->compiled_cache;
std::unique_lock<std::mutex> compile_lock(shared_cache->mutex);
auto weight_names = get_weight_names(cgraph);
- ggml_decoder = std::make_shared<GgmlOvDecoder>(cgraph, m_params, c_params, weight_names, is_static,
- stateful, model_is_splitted);
- const std::string shared_key = cache_enabled ? compiled_graph_key(cgraph, *ggml_decoder, device) : "";
+ ggml_decoder = std::make_shared<GgmlOvDecoder>(cgraph, m_params, c_params, weight_names,
+ is_static, r_ctx->stateful, model_is_splitted);
+ const std::string model_cache_dir = ggml_openvino_model_cache_dir();
+ const std::string exact_key = cache_enabled ? compiled_graph_key(cgraph, *ggml_decoder, device) : "";
+ uint64_t model_fp = 0;
+ uint64_t exact_fp = 0;
+ std::string blob_path, manifest_path;
+ if (!model_cache_dir.empty() && !model_is_splitted) {
+ const uint64_t extra_cfg =
+ ggml_openvino_model_cache_extra_cfg(r_ctx->stateful, stateful_recurrent);
+ model_fp =
+ ggml_openvino_model_fingerprint(cgraph, device, /*fa=*/true, m_params.rope_params, 16, extra_cfg,
+ dynamic_graph_signature(cgraph, *ggml_decoder, m_params));
+ exact_fp =
+ ggml_openvino_model_fingerprint(cgraph, device, /*fa=*/true, m_params.rope_params, 16, extra_cfg,
+ compiled_graph_key(cgraph, *ggml_decoder, device, 0, true));
+ blob_path = ggml_openvino_model_cache_blob_path(model_cache_dir, model_fp);
+ manifest_path = ggml_openvino_model_cache_manifest_path(model_cache_dir, model_fp);
+ }
+ std::string shared_key =
+ cache_enabled ? (model_fp ? "dynamic:" + std::to_string(model_fp) : exact_key) : "";
ov::CompiledModel shared_model;
bool imported = false;
auto shared_it = shared_cache->graphs.find(shared_key);
- if (!shared_key.empty() && shared_it != shared_cache->graphs.end()) {
+ if (!shared_key.empty() && shared_it != shared_cache->graphs.end() &&
+ (model_fp == 0 || compiled_model_matches_graph(shared_it->second.decode, shared_it->second.input_names,
+ shared_it->second.output_names, *ggml_decoder))) {
shared_model = shared_it->second.decode;
infer_request = std::make_shared<ov::InferRequest>(shared_model.create_infer_request());
ov_input_names = shared_it->second.input_names;
ov_output_names = shared_it->second.output_names;
imported = true;
GGML_LOG_DEBUG("ggml-openvino: shared compiled model HIT (dynamic)\n");
- }
- // Fail fast: a cache-miss recompile feeds weight data to compile_model, but
- // GGML_OPENVINO_RELEASE_WEIGHTS (or GGML_OPENVINO_MEMORY_OPTIMIZE on GPU)
- // may have already dropped the host weight pages
- // (they would read as zeros). That mode requires stable graph shapes.
- if (!imported && ggml_openvino_weight_buffers_released()) {
- GGML_ABORT(
- "ggml-openvino: a new graph needs to be compiled but host weight buffers were already "
- "released via GGML_OPENVINO_RELEASE_WEIGHTS/GGML_OPENVINO_MEMORY_OPTIMIZE. This mode requires "
- "stable graph shapes; disable host weight release for dynamic workloads.");
+ } else if (model_fp && shared_it != shared_cache->graphs.end()) {
+ shared_key = exact_key;
+ shared_it = shared_cache->graphs.find(shared_key);
+ if (shared_it != shared_cache->graphs.end()) {
+ shared_model = shared_it->second.decode;
+ infer_request = std::make_shared<ov::InferRequest>(shared_model.create_infer_request());
+ ov_input_names = shared_it->second.input_names;
+ ov_output_names = shared_it->second.output_names;
+ imported = true;
+ }
+ model_fp = exact_fp;
+ blob_path = ggml_openvino_model_cache_blob_path(model_cache_dir, model_fp);
+ manifest_path = ggml_openvino_model_cache_manifest_path(model_cache_dir, model_fp);
}
if (cache_enabled) {
std::lock_guard<std::mutex> map_lock(r_ctx->ctx_mutex);
r_ctx->infer_request_cache.erase(key);
}
- // Frontend-level compiled-model cache (GGML_OPENVINO_COMPILED_MODEL_CACHE_DIR): if this model
- // was compiled before, import the saved blob and skip requant + convert +
- // compile. Only the dynamic single-model path is cached (split models compile
- // two graphs and are left to the plugin-level ov::cache_dir). The decoder is
- // still needed for I/O mapping, but can be built without weight nodes since
- // the weights are baked into the imported CompiledModel.
- const std::string model_cache_dir = ggml_openvino_model_cache_dir();
- uint64_t model_fp = 0;
- std::string blob_path;
- std::string manifest_path;
- // When the frontend model cache is active it supersedes the plugin-level
- // ov::cache_dir: a blob exported from a model compiled WITH cache_dir cannot
- // be re-imported (import returns an uninitialized model). Strip cache_dir /
- // cache_mode from the config used for the cached compile and the import.
+ // Import standalone blobs with weights before graph conversion and compilation.
+ // Split models use the plugin cache instead. The decoder still maps graph I/O.
ov::AnyMap mc_config = config;
if (!model_cache_dir.empty()) {
mc_config.erase("CACHE_DIR");
- mc_config.erase("CACHE_MODE");
+ mc_config[ov::cache_mode.name()] = ov::CacheMode::OPTIMIZE_SPEED;
}
if (!imported && !model_cache_dir.empty() && !model_is_splitted) {
- const uint64_t extra_cfg = ggml_openvino_model_cache_extra_cfg(device, stateful);
- model_fp =
- ggml_openvino_model_fingerprint(cgraph, device, /*fa=*/true, m_params.rope_params, 16, extra_cfg);
- blob_path = ggml_openvino_model_cache_blob_path(model_cache_dir, model_fp);
- manifest_path = ggml_openvino_model_cache_manifest_path(model_cache_dir, model_fp);
-
- std::ifstream blob_in(blob_path, std::ios::binary);
- bool blob_ok = blob_in.is_open();
- bool manifest_ok =
- blob_ok && ggml_openvino_model_cache_verify_manifest(manifest_path, cgraph, model_fp);
- if (blob_ok && manifest_ok) {
- int64_t import_start = ggml_time_us();
+ const uint64_t preferred_fp = model_fp;
+ bool dynamic_incompatible = preferred_fp == exact_fp;
+ for (int candidate = 0; candidate < 2; ++candidate) {
+ if (imported || (candidate == 1 && preferred_fp == exact_fp)) {
+ break;
+ }
+ const uint64_t candidate_fp = candidate == 0 ? preferred_fp : exact_fp;
+ const std::string candidate_blob =
+ ggml_openvino_model_cache_blob_path(model_cache_dir, candidate_fp);
+ const std::string candidate_manifest =
+ ggml_openvino_model_cache_manifest_path(model_cache_dir, candidate_fp);
+ std::ifstream blob_in(candidate_blob, std::ios::binary);
+ std::vector<std::string> saved_inputs, saved_outputs;
+ if (!blob_in.is_open() ||
+ !ggml_openvino_model_cache_verify_manifest(candidate_manifest, cgraph, candidate_fp,
+ saved_inputs, saved_outputs)) {
+ continue;
+ }
+ const int64_t import_start = ggml_time_us();
try {
ov::CompiledModel cm;
auto remote_context = ggml_openvino_get_remote_context();
@@ -899,50 +1113,78 @@ enum ggml_status ov_graph_compute_dynamic(ggml_cgraph * cgraph, const std::share
} else {
cm = core.import_model(blob_in, device, mc_config);
}
- // Lightweight decoder: names-only weight map (membership is all the
- // decoder needs; weights live in the imported model).
- std::map<std::string, std::shared_ptr<ov::Node>> weight_names;
- for (const auto & n : GgmlOvDecoder::collect_weight_names(cgraph)) {
- weight_names[n] = nullptr;
+ // CPU serialization can rename Results that share a name with a Parameter.
+ if (!saved_inputs.empty() || !saved_outputs.empty()) {
+ ov_input_names = std::move(saved_inputs);
+ ov_output_names = std::move(saved_outputs);
+ } else {
+ for (const auto & p : cm.inputs()) {
+ ov_input_names.push_back(p.get_node()->get_friendly_name());
+ }
+ for (const auto & o : cm.outputs()) {
+ ov_output_names.push_back(o.get_node()->get_friendly_name());
+ }
+ }
+ if (!compiled_model_matches_graph(cm, ov_input_names, ov_output_names, *ggml_decoder)) {
+ dynamic_incompatible |= candidate == 0;
+ ov_input_names.clear();
+ ov_output_names.clear();
+ continue;
}
- ggml_decoder = std::make_shared<GgmlOvDecoder>(cgraph, m_params, c_params, weight_names,
- is_static, stateful, model_is_splitted);
infer_request = std::make_shared<ov::InferRequest>(cm.create_infer_request());
shared_model = cm;
entry->ptr = ggml_decoder;
- // Names must match the decoder's ggml-tensor keys. The non-cached
- // path keys off Parameter/Result *friendly names* (set by the
- // frontend); export_model preserves these, and each compiled-model
- // port's node is exactly that Parameter/Result. Use the port nodes
- // directly (NOT get_runtime_model(), whose graph differs and is
- // unsafe to deref this way).
- for (const auto & p : cm.inputs()) {
- ov_input_names.push_back(p.get_node()->get_friendly_name());
- }
- for (const auto & o : cm.outputs()) {
- ov_output_names.push_back(o.get_node()->get_friendly_name());
- }
imported = true;
+ model_fp = candidate_fp;
+ blob_path = candidate_blob;
+ manifest_path = candidate_manifest;
+ if (candidate_fp == exact_fp) {
+ shared_key = exact_key;
+ shared_it = shared_cache->graphs.find(shared_key);
+ }
if (ggml_openvino_getenv_int("GGML_OPENVINO_PROFILING")) {
GGML_LOG_INFO(" - Model cache import time: %.3f ms \n",
(ggml_time_us() - import_start) / 1000.0);
}
- GGML_LOG_INFO("ggml-openvino: model cache HIT %s\n", blob_path.c_str());
+ GGML_LOG_INFO("ggml-openvino: model cache HIT %s\n", candidate_blob.c_str());
} catch (const std::exception & e) {
- GGML_LOG_WARN("ggml-openvino: model cache import failed (%s), recompiling\n", e.what());
+ GGML_LOG_WARN("ggml-openvino: model cache import failed: %s\n", e.what());
+ dynamic_incompatible |= candidate == 0;
imported = false;
+ ov_input_names.clear();
+ ov_output_names.clear();
+ infer_request.reset();
+ shared_model = {};
}
}
+ if (!imported && dynamic_incompatible) {
+ model_fp = exact_fp;
+ shared_key = exact_key;
+ shared_it = shared_cache->graphs.find(shared_key);
+ blob_path = ggml_openvino_model_cache_blob_path(model_cache_dir, model_fp);
+ manifest_path = ggml_openvino_model_cache_manifest_path(model_cache_dir, model_fp);
+ }
}
std::shared_ptr<ov::Model> model;
+ if (!imported && ggml_openvino_model_cache_only()) {
+ GGML_LOG_ERROR(
+ "ggml-openvino: missing or incompatible compiled model: %s; run once without "
+ "GGML_OPENVINO_COMPILED_MODEL_CACHE_ONLY\n",
+ blob_path.c_str());
+ return GGML_STATUS_FAILED;
+ }
+ if (!imported && ggml_openvino_weight_buffers_released()) {
+ GGML_LOG_ERROR("ggml-openvino: cannot compile a new graph after releasing host weights\n");
+ return GGML_STATUS_FAILED;
+ }
if (imported) {
decoder_end_time = conversion_end_time = compile_end_time = ggml_time_us();
} else {
auto model_weights = GgmlOvDecoder::create_weight_nodes(cgraph);
ggml_decoder = std::make_shared<GgmlOvDecoder>(cgraph, m_params, c_params, model_weights, is_static,
- stateful, model_is_splitted);
+ r_ctx->stateful, model_is_splitted);
decoder_end_time = ggml_time_us();
auto input_model = std::make_shared<ov::frontend::ggml::InputModel>(ggml_decoder);
@@ -969,13 +1211,21 @@ enum ggml_status ov_graph_compute_dynamic(ggml_cgraph * cgraph, const std::share
}
compile_end_time = ggml_time_us();
+ for (const auto & ov_param : model->get_parameters()) {
+ ov_input_names.push_back(ov_param->get_friendly_name());
+ }
+ for (const auto & ov_output : model->get_results()) {
+ ov_output_names.push_back(ov_output->get_friendly_name());
+ }
+
// Export to the frontend model cache for next time. Publish the blob first,
// then the manifest, so a cache hit only sees fully written artifacts.
if (!model_cache_dir.empty() && !model_is_splitted && model_fp != 0) {
try {
- const std::string blob_tmp = blob_path + ".tmp";
- const std::string manifest_tmp = manifest_path + ".tmp";
- if (ggml_openvino_model_cache_write_manifest(manifest_tmp, cgraph, model_fp)) {
+ const std::string blob_tmp = ggml_openvino_model_cache_temp_path(blob_path);
+ const std::string manifest_tmp = ggml_openvino_model_cache_temp_path(manifest_path);
+ if (ggml_openvino_model_cache_write_manifest(manifest_tmp, cgraph, model_fp, ov_input_names,
+ ov_output_names)) {
std::ofstream blob_out(blob_tmp, std::ios::binary | std::ios::trunc);
if (blob_out.is_open()) {
compiled_model.export_model(blob_out);
@@ -1005,12 +1255,6 @@ enum ggml_status ov_graph_compute_dynamic(ggml_cgraph * cgraph, const std::share
shared_model = compiled_model;
entry->ptr = ggml_decoder;
- for (const auto & ov_param : model->get_parameters()) {
- ov_input_names.push_back(ov_param->get_friendly_name());
- }
- for (const auto & ov_output : model->get_results()) {
- ov_output_names.push_back(ov_output->get_friendly_name());
- }
} // end non-imported (compile) path
entry->ptr = ggml_decoder;
@@ -1025,7 +1269,7 @@ enum ggml_status ov_graph_compute_dynamic(ggml_cgraph * cgraph, const std::share
r_ctx->ov_output_names_cache[key] = ov_output_names;
}
- if (stateful && cache_enabled) {
+ if (stateful_kv_only && cache_enabled) {
const auto * inp_pos = get_inp_pos_tensor(cgraph);
auto pos_shape = GgmlOvDecoder::get_shape(inp_pos);
// A freshly compiled model starts with an empty state, so it can only serve a
@@ -1048,6 +1292,22 @@ enum ggml_status ov_graph_compute_dynamic(ggml_cgraph * cgraph, const std::share
}
}
+ if (stateful_recurrent) {
+ const auto * inp_pos = get_inp_pos_tensor(cgraph);
+ const int32_t pos_begin = static_cast<const int32_t *>(inp_pos->data)[0];
+ if (pos_begin == 0) {
+ infer_request->reset_state();
+ } else if (!cache_hit || old_m_params.kv_buffer_changed(m_params) || c_params.cache_rs_reset_len > 0 ||
+ pos_begin < 0 || static_cast<size_t>(pos_begin) != r_ctx->stateful_kv_size) {
+ GGML_LOG_ERROR(
+ "OpenVINO recurrent stateful execution cannot restore or rewind a sequence. Restart at position 0 "
+ "or disable GGML_OPENVINO_STATEFUL_EXECUTION.\n");
+ return GGML_STATUS_FAILED;
+ }
+ } else {
+ reset_single_slot_recurrent_cache(cgraph, m_params, c_params);
+ }
+
for (size_t i = 0; i < ov_input_names.size(); i++) {
const auto & param_name = ov_input_names[i];
auto input_tensor = get_ov_input_tensor(ggml_decoder, param_name);
@@ -1080,6 +1340,12 @@ enum ggml_status ov_graph_compute_dynamic(ggml_cgraph * cgraph, const std::share
ov_raw_infer_start = ggml_time_us();
infer_request->infer();
infer_end_time = ggml_time_us();
+ if (stateful_recurrent) {
+ const auto * inp_pos = get_inp_pos_tensor(cgraph);
+ r_ctx->stateful_kv_size =
+ static_cast<const int32_t *>(inp_pos->data)[0] + get_inp_pos_n_tokens(cgraph, inp_pos);
+ }
+ dump_ov_profiling_info(*infer_request);
if (ggml_openvino_getenv_int("GGML_OPENVINO_DEBUG_OUTPUT") ||
ggml_openvino_getenv_str("GGML_OPENVINO_DEBUG_NODE")) {
@@ -1092,13 +1358,17 @@ enum ggml_status ov_graph_compute_dynamic(ggml_cgraph * cgraph, const std::share
if (ggml_openvino_getenv_int("GGML_OPENVINO_PROFILING")) {
GGML_LOG_INFO("\nGGML OpenVINO Backend: \n");
GGML_LOG_INFO(" - Graph decoder time: %.3f ms \n", (decoder_end_time - start_time) / 1000.0);
+ GGML_LOG_INFO(" - cache key compute time: %.3f ms \n", cache_key_compute_time / 1000.0);
+ GGML_LOG_INFO(" - model split check time: %.3f ms \n", model_split_check_time / 1000.0);
+ GGML_LOG_INFO(" - graph param compute time: %.3f ms \n", graph_param_compute_time / 1000.0);
+ GGML_LOG_INFO(" - cache lookup time: %.3f ms \n", cache_lookup_time / 1000.0);
if (!cache_hit) {
GGML_LOG_INFO(" - Graph conversion time: %.3f ms \n",
(conversion_end_time - decoder_end_time) / 1000.0);
GGML_LOG_INFO(" - Graph compile time: %.3f ms \n", (compile_end_time - conversion_end_time) / 1000.0);
}
GGML_LOG_INFO(" - Graph inference time: %.3f ms \n", (infer_end_time - compile_end_time) / 1000.0);
- GGML_LOG_INFO(" - OV raw infer time: %.3f ms \n", (infer_end_time - ov_raw_infer_start) / 1000.0);
+ GGML_LOG_INFO(" - OV raw infer time: %.3f ms \n", (infer_end_time - ov_raw_infer_start) / 1000.0);
}
}
@@ -1108,7 +1378,7 @@ enum ggml_status ov_graph_compute_dynamic(ggml_cgraph * cgraph, const std::share
// be reading host weights during conversion/compilation. Pin the shared compiled
// models across backend teardown; a later context can create its own request without
// reading the dropped pages. A new, uncached graph still fails fast above.
- if (cache_hit && ggml_openvino_release_weights_enabled(device)) {
+ if (cache_hit && ggml_openvino_release_weights_enabled()) {
std::lock_guard<std::mutex> compile_lock(r_ctx->compiled_cache->mutex);
if (!ggml_openvino_weight_buffers_released()) {
ggml_openvino_release_weight_buffers();
@@ -1218,7 +1488,7 @@ enum ggml_status ov_graph_compute_static(ggml_cgraph * cgraph, const std::shared
ggml_decoder->m_is_prefill = is_prefill;
ggml_decoder->set_model_params(m_params);
ggml_decoder->set_compute_params(c_params);
- if (old_m_params.kv_buffer_changed(m_params)) {
+ if (old_m_params.kv_buffer_changed(m_params) || !ggml_decoder->is_bound_to(cgraph)) {
ggml_decoder->update_io(cgraph);
}
ggml_decoder->add_extra_inputs();
@@ -1356,6 +1626,8 @@ enum ggml_status ov_graph_compute_static(ggml_cgraph * cgraph, const std::shared
}
}
+ reset_single_slot_recurrent_cache(cgraph, m_params, c_params);
+
if (is_prefill) {
auto inp_len = get_inp_pos_n_tokens(cgraph, inp_pos);
for (int chunk_index = 0; chunk_index * prefill_chunk_size < inp_len; chunk_index++) {
diff --git a/ggml/src/ggml-openvino/utils.h b/ggml/src/ggml-openvino/utils.h
index 74c25f0ac..491dbadc2 100644
--- a/ggml/src/ggml-openvino/utils.h
+++ b/ggml/src/ggml-openvino/utils.h
@@ -1,8 +1,9 @@
#include "ggml-decoder.h"
#include "ggml-impl.h"
+#include "ggml-openvino-extra.h"
-#include <algorithm>
#include <cstddef>
+#include <cstdint>
#include <functional>
#include <memory>
#include <mutex>
@@ -10,74 +11,71 @@
#include <openvino/runtime/infer_request.hpp>
#include <string>
#include <unordered_map>
+#include <unordered_set>
#include <utility>
#include <vector>
-// Local execution-cache key. A match still needs the ModelParams compatibility
-// check; this key alone does not identify weights or a compiled model.
+// Cache key for a translated/compiled graph. Node count plus the two end node names identify a
+// graph during inference, where the same few graphs repeat for the whole session. The list of
+// external input names below tells those apart more precisely, but it walks every node and every
+// src slot, so it is built only when GGML_OPENVINO_FULL_GRAPH_KEY is set. Op tests run many small
+// graphs that can share a node count and end names, and need the full key.
struct graph_key {
int n_nodes;
+ int n_leaves;
std::string first_node_name;
std::string last_node_name;
- std::vector<std::string> input_src_names;
+ // Each entry: "<node_idx>:<src_idx>:<op_type>"
+ std::vector<std::string> input_srcs;
- graph_key(const ggml_cgraph * cgraph) : n_nodes(cgraph->n_nodes) {
+ graph_key(const ggml_cgraph * cgraph, bool include_inputs = false) : n_nodes(cgraph->n_nodes), n_leaves(cgraph->n_leafs) {
if (n_nodes > 0) {
first_node_name = cgraph->nodes[0]->name;
last_node_name = cgraph->nodes[n_nodes - 1]->name;
}
- std::unordered_map<const ggml_tensor *, std::string> names;
- auto get_input_key_name = [&names](const ggml_cgraph * graph, const ggml_tensor * tensor) {
- auto it = names.find(tensor);
- if (it == names.end()) {
- it = names.emplace(tensor, GgmlOvDecoder::get_tensor_name(graph, tensor)).first;
- }
- return it->second;
- };
+ static const bool full_key = ggml_openvino_getenv_int("GGML_OPENVINO_FULL_GRAPH_KEY") != 0;
+ if (!full_key && !include_inputs) {
+ return;
+ }
- std::vector<std::string> node_names;
- node_names.reserve(cgraph->n_nodes);
- for (int node_idx = 0; node_idx < cgraph->n_nodes; node_idx++) {
- node_names.emplace_back(cgraph->nodes[node_idx]->name);
+ std::unordered_set<const ggml_tensor *> node_set;
+ node_set.reserve(cgraph->n_nodes);
+ for (int i = 0; i < cgraph->n_nodes; i++) {
+ node_set.insert(cgraph->nodes[i]);
}
for (int node_idx = 0; node_idx < cgraph->n_nodes; node_idx++) {
const ggml_tensor * node = cgraph->nodes[node_idx];
for (int src_idx = 0; src_idx < GGML_MAX_SRC; src_idx++) {
const ggml_tensor * src = node->src[src_idx];
- if (src == nullptr || src->name[0] == '\0') {
- continue;
- }
-
- const std::string src_name = get_input_key_name(cgraph, src);
- if (std::find(node_names.begin(), node_names.end(), src_name) != node_names.end()) {
- continue;
- }
- if (src_name.find("weight") != std::string::npos) {
+ if (src == nullptr || src->buffer == nullptr || node_set.count(src) || src->buffer->usage == GGML_BACKEND_BUFFER_USAGE_WEIGHTS) {
continue;
}
- input_src_names.push_back(std::to_string(node_idx) + ":" + std::to_string(src_idx) + ":" + src_name);
+ input_srcs.push_back(std::to_string(node_idx) + ":" + std::to_string(src_idx) + ":" +
+ GgmlOvDecoder::compute_op_type(node));
}
}
}
bool operator==(const graph_key & other) const {
- return n_nodes == other.n_nodes && first_node_name == other.first_node_name &&
- last_node_name == other.last_node_name && input_src_names == other.input_src_names;
+ return n_nodes == other.n_nodes && n_leaves == other.n_leaves &&
+ first_node_name == other.first_node_name &&
+ last_node_name == other.last_node_name && input_srcs == other.input_srcs;
}
};
struct graph_key_hash {
size_t operator()(const graph_key & key) const {
size_t hash = std::hash<int>{}(key.n_nodes);
+ hash ^= std::hash<int>{}(key.n_leaves) + 0x9e3779b9 + (hash << 6) + (hash >> 2);
if (key.n_nodes > 0) {
hash ^= std::hash<std::string>{}(key.first_node_name) + 0x9e3779b9 + (hash << 6) + (hash >> 2);
hash ^= std::hash<std::string>{}(key.last_node_name) + 0x9e3779b9 + (hash << 6) + (hash >> 2);
}
- for (const auto & input_src_name : key.input_src_names) {
- hash ^= std::hash<std::string>{}(input_src_name) + 0x9e3779b9 + (hash << 6) + (hash >> 2);
+ for (const auto & s : key.input_srcs) {
+ hash ^= std::hash<std::string>{}(s) + 0x9e3779b9 + (hash << 6) + (hash >> 2);
}
return hash;
}