Skip to content

server: fix Windows CI flake in test_completion_unified - #28759

Merged
ServeurpersoCom merged 1 commit into
ggml-org:masterfrom
ServeurpersoCom:server-tests-unified-pool-abort
Sep 11, 2026
Merged

ServeurpersoCom merged 1 commit into
ggml-org:masterfrom
ServeurpersoCom:server-tests-unified-pool-abort

Conversation

@ServeurpersoCom

@ServeurpersoCom ServeurpersoCom commented Sep 11, 2026 •

Copy link
Copy Markdown
Contributor

Overview

The four requests only match the expected success table when they enter the shared KV pool together. On the Windows runner they are admitted 30 to 70 ms apart, the short one still holds its cells when the long ones reach their limit, and the overflow aborts every slot at once, so a request the table marks as successful comes back with the context error. It now passes on that error too, and on nothing else.

Additional information

Two parameter sets account for 20 failing jobs over the past week, [90, 90, 40, 90] and [90, 90, 40, 75], same shape in both: admissions at 0.504 / 0.504 / 0.533 / 0.534 holding 90 + 90 + 40 + 39 = 259 cells at the overflow, and 0.501 / 0.501 / 0.538 / 0.569 holding 92 + 92 + 42 + 34 = 260. The table and the predicted_n check are untouched, a truncated generation or any other error still fails. The tolerance can go once the TODO in server-context.cpp lands and an overflow terminates only the largest sequence.

Requirements

The expected success table holds when the four requests enter the shared
pool together. On a loaded runner they are admitted tens of milliseconds
apart, the slot lifetimes overlap differently and the pool overflows
while a short request is still resident. The decode failure aborts every
slot, so a request the table marks as successful comes back with the
context error instead of its generation.

Such a request now passes on that error too, while any other status, a
different error or a truncated generation still fails the test.
@ServeurpersoCom
ServeurpersoCom requested a review from a team as a code owner September 11, 2026 13:18
@ServeurpersoCom

ServeurpersoCom commented Sep 11, 2026 •

Copy link
Copy Markdown
Contributor Author

I sent my agent through every red Server run of the past week to sort the Windows flakes by root cause, here is what it found :

cause                              jobs (7d)  share of router  status

download_finished command lost        42           80%         PR #28747
behind the ANSI reset

DELETE 600s, exit command              5           10%         PR #28755
never sent to the child

DELETE 600s, exit acknowledged         5           10%         PR #28762
but EOF never reached

test_completion_unified               20            -          PR #28759 <- this PR
returning 500

build / infra, misc                   22            -          out of scope

#28747
#28755
#28762

@ServeurpersoCom
ServeurpersoCom merged commit 8172e65 into ggml-org:master Sep 11, 2026
10 checks passed
pl752 pushed a commit to pl752/llama.cpp that referenced this pull request Sep 15, 2026
…org#28759)

The expected success table holds when the four requests enter the shared
pool together. On a loaded runner they are admitted tens of milliseconds
apart, the slot lifetimes overlap differently and the pool overflows
while a short request is still resident. The decode failure aborts every
slot, so a request the table marks as successful comes back with the
context error instead of its generation.

Such a request now passes on that error too, while any other status, a
different error or a truncated generation still fails the test.
quimmedes pushed a commit to quimmedes/cafe-llama.cpp that referenced this pull request Sep 16, 2026
…org#28759)

The expected success table holds when the four requests enter the shared
pool together. On a loaded runner they are admitted tens of milliseconds
apart, the slot lifetimes overlap differently and the pool overflows
while a short request is still resident. The decode failure aborts every
slot, so a request the table marks as successful comes back with the
context error instead of its generation.

Such a request now passes on that error too, while any other status, a
different error or a truncated generation still fails the test.
zsogitbe pushed a commit to zsogitbe/llama.cpp that referenced this pull request Sep 17, 2026
…org#28759)

The expected success table holds when the four requests enter the shared
pool together. On a loaded runner they are admitted tens of milliseconds
apart, the slot lifetimes overlap differently and the pool overflows
while a short request is still resident. The decode failure aborts every
slot, so a request the table marks as successful comes back with the
context error instead of its generation.

Such a request now passes on that error too, while any other status, a
different error or a truncated generation still fails the test.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants