Commit 8e1642198 for llama.cpp

commit 8e1642198dcd4e408f8776222d6ae31b74d01187
Author: Yufeng He <40085740+he-yufeng@users.noreply.github.com>
Date:   Mon Oct 5 17:22:23 2026 +0800

    server: reject partial media truncation (#24076)

    * server: reject partial media truncation

    * server: keep only the keep_first fix

    Drop the mtmd test helper change, which no longer builds since
    clip_image_f32_batch stores its entries by value, and drop the
    vision test: no test fixture reaches a cut between two adjacent
    media chunks with a reused cache (tinygemma3 uses SWA and wraps
    images in text tokens, tinyopenjev and small-test are recurrent),
    so the test passed or failed independently of the fix.

    ---------

    Co-authored-by: Pascal <admin@serveurperso.com>

diff --git a/tools/server/server-common.cpp b/tools/server/server-common.cpp
index be87c45f7..076a85741 100644
--- a/tools/server/server-common.cpp
+++ b/tools/server/server-common.cpp
@@ -667,7 +667,7 @@ void server_tokens::keep_first(size_t n) {
             // note that the case where we keep a full image at the end is allowed:
             //   tokens[n - 1] == LLAMA_TOKEN_NULL && tokens[n] != LLAMA_TOKEN_NULL
             if (tokens[n - 1] == LLAMA_TOKEN_NULL && tokens[n] == LLAMA_TOKEN_NULL) {
-                find_chunk(n - 1); // will throw an error if the token is not begin-of-chunk
+                find_chunk(n); // will throw an error if the cut is not at a chunk boundary
             }
         }
         // remove all image chunks that are not used anymore