server: fix Windows CI flake in test_completion_unified - #28759
Merged
ServeurpersoCom merged 1 commit intoSep 11, 2026
Merged
ServeurpersoCom merged 1 commit into
ServeurpersoCom merged 1 commit into
Conversation
The expected success table holds when the four requests enter the shared pool together. On a loaded runner they are admitted tens of milliseconds apart, the slot lifetimes overlap differently and the pool overflows while a short request is still resident. The decode failure aborts every slot, so a request the table marks as successful comes back with the context error instead of its generation. Such a request now passes on that error too, while any other status, a different error or a truncated generation still fails the test.
Contributor
Author
|
I sent my agent through every red Server run of the past week to sort the Windows flakes by root cause, here is what it found : |
This was referenced Sep 11, 2026
CISC
approved these changes
Sep 11, 2026
pwilkin
approved these changes
Sep 11, 2026
pl752
pushed a commit
to pl752/llama.cpp
that referenced
this pull request
Sep 15, 2026
…org#28759) The expected success table holds when the four requests enter the shared pool together. On a loaded runner they are admitted tens of milliseconds apart, the slot lifetimes overlap differently and the pool overflows while a short request is still resident. The decode failure aborts every slot, so a request the table marks as successful comes back with the context error instead of its generation. Such a request now passes on that error too, while any other status, a different error or a truncated generation still fails the test.
quimmedes
pushed a commit
to quimmedes/cafe-llama.cpp
that referenced
this pull request
Sep 16, 2026
…org#28759) The expected success table holds when the four requests enter the shared pool together. On a loaded runner they are admitted tens of milliseconds apart, the slot lifetimes overlap differently and the pool overflows while a short request is still resident. The decode failure aborts every slot, so a request the table marks as successful comes back with the context error instead of its generation. Such a request now passes on that error too, while any other status, a different error or a truncated generation still fails the test.
zsogitbe
pushed a commit
to zsogitbe/llama.cpp
that referenced
this pull request
Sep 17, 2026
…org#28759) The expected success table holds when the four requests enter the shared pool together. On a loaded runner they are admitted tens of milliseconds apart, the slot lifetimes overlap differently and the pool overflows while a short request is still resident. The decode failure aborts every slot, so a request the table marks as successful comes back with the context error instead of its generation. Such a request now passes on that error too, while any other status, a different error or a truncated generation still fails the test.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Overview
The four requests only match the expected success table when they enter the shared KV pool together. On the Windows runner they are admitted 30 to 70 ms apart, the short one still holds its cells when the long ones reach their limit, and the overflow aborts every slot at once, so a request the table marks as successful comes back with the context error. It now passes on that error too, and on nothing else.
Additional information
Two parameter sets account for 20 failing jobs over the past week, [90, 90, 40, 90] and [90, 90, 40, 75], same shape in both: admissions at 0.504 / 0.504 / 0.533 / 0.534 holding 90 + 90 + 40 + 39 = 259 cells at the overflow, and 0.501 / 0.501 / 0.538 / 0.569 holding 92 + 92 + 42 + 34 = 260. The table and the predicted_n check are untouched, a truncated generation or any other error still fails. The tolerance can go once the TODO in server-context.cpp lands and an overflow terminates only the largest sequence.
Requirements