Skip to content

fix(chat): handle out-of-order streamed choice indexes - #3709

Open
HostX0 wants to merge 1 commit into
openai:mainfrom
HostX0:fix/chat-stream-choice-index-order
Open

HostX0 wants to merge 1 commit into
openai:mainfrom
HostX0:fix/chat-stream-choice-index-order

Conversation

@HostX0

@HostX0 HostX0 commented Aug 21, 2026

Copy link
Copy Markdown
  • I understand that this repository is auto-generated and my pull request may not be merged

Changes being requested

Fix Chat Completions stream accumulation when choice chunks arrive out of index order.

ChatCompletionStreamState currently uses choice.index as a Python list offset in the snapshot and event-state collections. That only works when every lower-numbered choice has already appeared. A chunk containing a single choice with index=1, for example, has a one-element choices list, so the initial snapshot path attempts to assign choices[1] and raises IndexError.

This change keeps the API-provided index as the choice identity instead of assuming it is the current list position:

  • key per-choice event state by choice.index;
  • resolve accumulated snapshots by their stored index field;
  • keep newly discovered snapshots ordered by index for the final completion;
  • transform the choices present in the initial chunk by their actual list positions rather than their logical indexes.

The regression test feeds the accumulator a valid 1 -> 0 -> 1 sequence and verifies that the later delta is appended to choice 1, while the final snapshot remains ordered as choices 0 and 1.

Additional context & links

The production change is confined to the handwritten src/openai/lib/streaming/chat/ implementation; no generated API resource or public model changes are required.

@HostX0
HostX0 requested a review from a team as a code owner August 21, 2026 04:43
@HostX0
HostX0 force-pushed the fix/chat-stream-choice-index-order branch from ad1e2a7 to 1440abb Compare August 21, 2026 04:44
@github-actions

Copy link
Copy Markdown
Contributor

Castiron custom code

✅ No new custom-code files detected.

36 mixed files remain; 0 existing customizations changed.

Compared 45341061b382 → 1440abb43768. Generated baselines verified.

36 existing customizations unchanged
  • api.md
  • scripts/castiron/README.md
  • scripts/castiron/custom_code_report.py
  • scripts/castiron/test_custom_code_report.py
  • src/openai/init.py
  • src/openai/_client.py
  • src/openai/resources/audio/transcriptions.py
  • src/openai/resources/audio/translations.py
  • src/openai/resources/beta/beta.py
  • src/openai/resources/beta/responses/responses.py
  • src/openai/resources/beta/threads/runs/runs.py
  • src/openai/resources/beta/threads/threads.py
  • src/openai/resources/chat/completions/completions.py
  • src/openai/resources/embeddings.py
  • src/openai/resources/files.py
  • src/openai/resources/realtime/realtime.py
  • src/openai/resources/responses/responses.py
  • src/openai/resources/uploads/uploads.py
  • src/openai/resources/vector_stores/file_batches.py
  • src/openai/resources/vector_stores/files.py
  • src/openai/resources/videos.py
  • src/openai/resources/webhooks/init.py
  • src/openai/resources/webhooks/webhooks.py
  • src/openai/types/init.py
  • src/openai/types/chat/init.py
  • src/openai/types/chat/chat_completion_message_tool_call.py
  • src/openai/types/fine_tuning/fine_tuning_job_integration.py
  • src/openai/types/responses/init.py
  • src/openai/types/responses/response.py
  • src/openai/types/responses/response_function_web_search.py
  • src/openai/types/responses/response_function_web_search_param.py
  • src/openai/types/responses/tool.py
  • src/openai/types/responses/tool_param.py
  • src/openai/types/webhooks/init.py
  • src/openai/types/websocket_connection_options.py
  • tests/api_resources/test_videos.py

A changed generated baseline means this report cannot reliably identify which handwritten lines changed.

Inspect the custom-code diff

Download the exact patch produced by this run (requires repository access):

gh run download 35490000656 --repo openai/openai-python \
  --name castiron-custom-code-35490000656-1 --dir /tmp/castiron-custom-code-35490000656-1
git apply --stat /tmp/castiron-custom-code-35490000656-1/custom-code.patch
cat /tmp/castiron-custom-code-35490000656-1/custom-code.patch

Or reproduce it from an SDK checkout containing the vendored reporter:

git fetch --no-tags origin 45341061b382473ee8a8e8c796ef8819089abbb6 1440abb43768e5bc5073069667251afe2e1f2d3e
python3 scripts/castiron/custom_code_report.py report \
  --base 45341061b382473ee8a8e8c796ef8819089abbb6 \
  --head 1440abb43768e5bc5073069667251afe2e1f2d3e --fetch --require-head-hash --public \
  --out /tmp/castiron-custom-code-1440abb43768
cat /tmp/castiron-custom-code-1440abb43768/custom-code.patch

This is the current full custom patch for mixed files, not an attribution of only the handwritten lines changed by this PR.

Full report and patch

@jnohclee-rgb

Copy link
Copy Markdown

AI-assisted offline accumulator comparison at 1440abb43768e5bc5073069667251afe2e1f2d3e against parent 45341061b382473ee8a8e8c796ef8819089abbb6, on Python 3.14.7 / Pydantic 2.13.5. I extended the sequence coverage beyond the existing single-choice 1 → 0 → 1 regression, using fictional typed chunks and the real ChatCompletionStreamState.handle_chunk().

Five sequences per revision:

  • initial chunk contains choices [1,0], followed by deltas and reverse-order finishes: same correct sorted final snapshots and independent content.done events on both revisions;
  • sparse text choices: first index4, then index2, then a chunk updating both, then finishes: parent IndexError; head final [(2,"Bb"),(4,"Dd")], with exactly one content.done event for each accumulated value;
  • multi-choice refusal deltas [1,0] with reverse-order finishes: correct independent refusal snapshots and done events on both;
  • sparse refusal choices: index3, then index1, both updated and finished: parent IndexError; head final [(1,"R1x"),(3,"R3y")], with exactly one refusal.done event for each value;
  • ordered text control [0,1]: complete event payloads and final snapshots unchanged.

The sparse cases also exercise the new per-index event-state dictionary: finishing one choice does not suppress the other choice's done event.

Minimal sparse-text input sequence (all chunks have id/model "fictional", created=0 and object "chat.completion.chunk"):

choices_by_chunk = [
    [{"index": 4, "delta": {"content": "D"}, "finish_reason": None}],
    [{"index": 2, "delta": {"content": "B"}, "finish_reason": None}],
    [{"index": 4, "delta": {"content": "d"}, "finish_reason": None},
     {"index": 2, "delta": {"content": "b"}, "finish_reason": None}],
    [{"index": 2, "delta": {}, "finish_reason": "stop"},
     {"index": 4, "delta": {}, "finish_reason": "stop"}],
]

Each list was passed through ChatCompletionChunk.construct(...) then list(state.handle_chunk(chunk)). These are intentionally synthetic sparse-index cases, not a claim that the production API emits them. No network or live model calls occurred. Coverage is accumulator-level, not sync/async HTTP integration, tool-call indexing, or the full suite.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants