Commit dd4c286f3 for llama.cpp
commit dd4c286f3809e73be087f374d9d865ca97c205e4
Author: Yash Raj Pandey <55940078+devYRPauli@users.noreply.github.com>
Date: Fri Oct 2 10:37:11 2026 -0400
ggml-cpu : fix soft_max_back wrong output when dst aliases src1 (#27096)
* ggml-cpu : fix soft_max_back wrong output when dst aliases src1
GGML_OP_SOFT_MAX_BACK is listed in ggml_op_can_inplace, so the graph
allocator may assign dst to alias either src0 (dy) or src1 (y).
The result was built in several steps:
ggml_vec_cpy_f32 (nc, dx, dy);
ggml_vec_acc1_f32 (nc, dx, -dot_y_dy);
ggml_vec_mul_f32 (nc, dx, dx, y);
ggml_vec_scale_f32(nc, dx, scale);
When dst aliases src1, the first step overwrites y and the third step
then reads the overwritten values, so the output is silently wrong.
Aliasing dst with src0 is unaffected. The CUDA kernel completes its
reduction before writing and is already safe.
Replace the sequence with a single fused loop that reads both sources
before writing, which is correct under either aliasing.
Add a regression test that marks dy as a graph output so the allocator
is forced to alias dst with y, asserts that the alias actually
happened, and compares against values computed on the host.
* cont : remove comment
---------
Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>
diff --git a/ggml/src/ggml-cpu/ops.cpp b/ggml/src/ggml-cpu/ops.cpp
index 3e377298f..16a916d20 100644
--- a/ggml/src/ggml-cpu/ops.cpp
+++ b/ggml/src/ggml-cpu/ops.cpp
@@ -6012,11 +6012,10 @@ static void ggml_compute_forward_soft_max_ext_back_f32(
// linear runtime, no additional memory
float dot_y_dy = 0;
- ggml_vec_dot_f32 (nc, &dot_y_dy, 0, y, 0, dy, 0, 1);
- ggml_vec_cpy_f32 (nc, dx, dy);
- ggml_vec_acc1_f32 (nc, dx, -dot_y_dy);
- ggml_vec_mul_f32 (nc, dx, dx, y);
- ggml_vec_scale_f32(nc, dx, scale);
+ ggml_vec_dot_f32(nc, &dot_y_dy, 0, y, 0, dy, 0, 1);
+ for (int i = 0; i < nc; i++) {
+ dx[i] = scale * (dy[i] - dot_y_dy) * y[i];
+ }
#ifndef NDEBUG
for (int i = 0; i < nc; ++i) {