mirror of
https://github.com/ggml-org/llama.cpp.git
synced 2026-10-08 22:10:37 +02:00
server: reject partial media truncation (#24076)
* server: reject partial media truncation * server: keep only the keep_first fix Drop the mtmd test helper change, which no longer builds since clip_image_f32_batch stores its entries by value, and drop the vision test: no test fixture reaches a cut between two adjacent media chunks with a reused cache (tinygemma3 uses SWA and wraps images in text tokens, tinyopenjev and small-test are recurrent), so the test passed or failed independently of the fix. --------- Co-authored-by: Pascal <admin@serveurperso.com>
This commit is contained in:
@@ -667,7 +667,7 @@ void server_tokens::keep_first(size_t n) {
|
||||
// note that the case where we keep a full image at the end is allowed:
|
||||
// tokens[n - 1] == LLAMA_TOKEN_NULL && tokens[n] != LLAMA_TOKEN_NULL
|
||||
if (tokens[n - 1] == LLAMA_TOKEN_NULL && tokens[n] == LLAMA_TOKEN_NULL) {
|
||||
find_chunk(n - 1); // will throw an error if the token is not begin-of-chunk
|
||||
find_chunk(n); // will throw an error if the cut is not at a chunk boundary
|
||||
}
|
||||
}
|
||||
// remove all image chunks that are not used anymore
|
||||
|
||||
Reference in New Issue
Block a user