Commit d477de3c for whisper.cpp
commit d477de3ce5aa696a991382df2a2958200ad620a8
Author: Georgi Gerganov <ggerganov@gmail.com>
Date: Mon Oct 5 11:36:25 2026 +0300
llama : fix unexpected graph reallocation in the k-pool models (llama/29958)
* llama : fix unexpected graph reallocation in the k-pool models
Both k-pool models built a graph shape that depends on state the
full-context reserve cannot know:
- qwen4exp branched on inp->cache_safe, which turns false as soon as
llama_memory_seq_cp shares cells (e.g. batched-bench -pps): the QSA
layers swapped scatter+gather for fill+concat and dropped the
new_pool_rep leaf, so the decode graph had 12 fewer nodes than the
reserved one
- glm5-next branched on gather = n_tokens <= 16 && n_kv > n_sel, so the
TG decode built the gather shape (7564 nodes) while the last reserve,
the PP one, had the dense shape (7762 nodes)
Either mismatch forces a decode-time re-reserve that drops the
worst-case sizing and bakes in the current state, so the next state
growth (n_pool, n_kv, n_new) needs more room at an unchanged graph size
and aborts under GGML_SCHED_DEBUG_REALLOC=1. Reproduce with, e.g.:
GGML_SCHED_DEBUG_REALLOC=1 ./bin/llama-batched-bench \
-hf ggml-org/GLM-5.3-Flash-GGUF:Q2_K -npp 2500 -ntg 32 -npl 1,2 \
-c 32768 -pps -kvu
Always scatter+gather the pooled keys, and pick gather from context
constants only: n_ubatch bounds every ubatch, top_k + kpool - 1 bounds
n_sel. Every graph of a context then shares one shape, which the
reserve covers, and the dense path measured faster than the gather path
at 2.5k and 16k context.
Assisted-by: pi:llama.cpp/MiMo-V2.6-Flash-MOPD
* llama : drop the unused k-pool cache_safe graph API
The k-pool graphs no longer branch on cache_safe, so nothing reads
get_kpool_cache_safe() or the conditional new_pool_rep any more: both
models always pass the scatter target, which set_input_kpool now
requires instead of merely preferring.
Also drop the cache_safe copy in kpool_build_sizes(), a sizes-only
helper. The layout and state flag itself stays, it still decides which
pools a layout with shared cells must re-pool.
Assisted-by: pi:llama.cpp/MiMo-V2.6-Flash-MOPD
* tests : add a shared-seq graph reserve regression test
Decode a prompt into seq 0, share its cells with seq 1 via
llama_memory_seq_cp (what llama-batched-bench does for -pps), then keep
decoding both sequences. For the k-pool models sharing clears
cache_safe, which changes the graph topology while the pools keep
growing, so a scheduler that re-reserves with the current state
instead of the worst-case one aborts under GGML_SCHED_DEBUG_REALLOC=1.
The test registration sets that flag, and the test aborts on both
k-pool models before 2220411ec1.
kimi-linear and minimax-01 are skipped: they reserve the final pp graph
with n_seqs = 1 (see [TAG_RESERVE_DIAG_DECAY] in llama-context.cpp), so
every multi-seq graph has a different layout and re-reserves by design.
Assisted-by: pi:llama.cpp/MiMo-V2.6-Flash-MOPD
* cont : add TODOs
* cont : fix comment
* cuda: match the moe weighted reduction on empty ubatches
ggml_cuda_match_moe_weighted_reduction rejected tensors with zero
rows. A ubatch without outputs shrinks the last layer to zero rows
through inp_out_ids, so graph_optimize dropped its alloc dep there and
the scheduler graph lost one node compared to the reserved one. The
scheduler then re-reserved at the size of that ubatch, and the next
ubatch with the same node count but larger tensors aborted under
GGML_SCHED_DEBUG_REALLOC=1.
The compute loop already skips empty nodes before trying any fusion,
so the guard only made the alloc deps depend on the row count.
* tests: build the rollback test only where internal symbols link
The shared-seq case calls llm_arch_from_string, which libllama does
not export through LLAMA_API, so linking test-recurrent-state-rollback
fails on Windows with shared libraries. Its build now sits in the
NOT WIN32 OR NOT BUILD_SHARED_LIBS block, next to test-llama-archs and
the test registration it already lives under.
* tests: skip archs by name in the shared-seq reserve test
The skip of kimi-linear and minimax-01 went through llm_arch_from_string,
which libllama does not export through LLAMA_API, so the test could not
link on Windows with shared libraries. It now compares the
general.architecture string directly, and the test builds on every
platform again.
---------
Co-authored-by: Pascal <admin@serveurperso.com>
diff --git a/ggml/src/ggml-cuda/ggml-cuda.cu b/ggml/src/ggml-cuda/ggml-cuda.cu
index 303d67e4..c7241961 100644
--- a/ggml/src/ggml-cuda/ggml-cuda.cu
+++ b/ggml/src/ggml-cuda/ggml-cuda.cu
@@ -3164,7 +3164,7 @@ static bool ggml_cuda_match_moe_weighted_reduction(
const int n_expert_used = (int) weighted->ne[1];
const int64_t n_tokens = weighted->ne[2] * weighted->ne[3];
- if (n_expert_used < 2 || n_expert_used > MOE_WEIGHTED_REDUCTION_MAX_EXPERTS || n_tokens <= 0) {
+ if (n_expert_used < 2 || n_expert_used > MOE_WEIGHTED_REDUCTION_MAX_EXPERTS) {
return false;
}