Skip to content

fix(streaming): restore HTTP/1.1 connection reuse after [DONE] - #4041

Draft
marcuswood-oai wants to merge 2 commits into
mainfrom
fix/streaming-connection-reuse-4040
Draft

marcuswood-oai wants to merge 2 commits into
mainfrom
fix/streaming-connection-reuse-4040

Conversation

@marcuswood-oai

@marcuswood-oai marcuswood-oai commented Oct 8, 2026 •

Copy link
Copy Markdown
Contributor

Changes being requested

Fully consumed HTTP/1.1 Chat Completions streams currently close at [DONE] before the transport reads the HTTP body ending, so each request opens a new connection. Resume the existing byte iterator after [DONE] to let the transport return completed connections to the pool.

The drain runs only for HTTP/1.1 after the completion marker and discards bytes without parsing trailing SSE data. Transport errors during that cleanup are suppressed; cancellation and unexpected errors still propagate, and the existing finally closes the response. Early exits and errors before [DONE] do not drain. HTTP/2 keeps its existing behavior.

This deliberately uses the request's existing read timeout. There is no independent cleanup deadline: a delayed ending adds completion latency, and disabling read timeouts can permit an indefinite wait. Tokens are still yielded as they arrive.

Validation

Three focused tests cover sync/async HTTP/1.1 connection reuse and failed-cleanup recovery, HTTP/2 and discarded trailing bytes, and cancellation during the drain. They use the existing HTTPX2/legacy HTTPX test matrix.

  • Streaming suite with both transports: 215 passed, 4 skipped.
  • Pydantic v1 streaming/chat suite: 78 passed, 31 skipped.
  • Ruff, repository Pyright, mypy, and Castiron custom-code budget passed.

Synthetic TCP/TLS benchmark

40 requests per condition, same SDK version and transport dependencies before/after. Medians below use an immediate HTTP ending; these are synthetic measurements, not production latency estimates.

Case Before After
Sync, fresh client per request 4.19 ms 4.35 ms
Async, fresh client per request 6.44 ms 6.52 ms
Sync, reused client 4.47 ms 1.71 ms
Async, reused client 6.00 ms 2.43 ms
HTTP/1.1 connections for 40 requests on one client 40 1

Deliberately delaying the HTTP ending by 5 ms or 30 ms adds approximately that wait after the final token. HTTP/2 continues to reuse one connection and complete promptly without draining.

Additional context & links

Fixes #4040. Related to #3440.

@github-actions

github-actions Bot commented Oct 8, 2026 •

Copy link
Copy Markdown
Contributor

Castiron custom code

Evaluated main: 9301e319ea33ef28fba380f39a289dedc14652c1.

✅ No new custom-code files detected.

46 mixed files remain; 0 existing customizations changed.

Compared 9301e319ea33 → 83833ea92b0d. Generated baselines verified.

46 existing customizations unchanged
  • api.md
  • src/openai/init.py
  • src/openai/_client.py
  • src/openai/resources/audio/transcriptions.py
  • src/openai/resources/audio/translations.py
  • src/openai/resources/beta/agents/environments/files.py
  • src/openai/resources/beta/agents/sessions/artifacts.py
  • src/openai/resources/beta/agents/sessions/sessions.py
  • src/openai/resources/beta/beta.py
  • src/openai/resources/beta/responses/responses.py
  • src/openai/resources/beta/threads/runs/runs.py
  • src/openai/resources/beta/threads/threads.py
  • src/openai/resources/chat/completions/completions.py
  • src/openai/resources/embeddings.py
  • src/openai/resources/files.py
  • src/openai/resources/live/forks.py
  • src/openai/resources/live/live.py
  • src/openai/resources/live/sideband.py
  • src/openai/resources/realtime/api.md
  • src/openai/resources/realtime/realtime.py
  • src/openai/resources/responses/responses.py
  • src/openai/resources/uploads/uploads.py
  • src/openai/resources/vector_stores/file_batches.py
  • src/openai/resources/vector_stores/files.py
  • src/openai/resources/videos.py
  • src/openai/resources/webhooks/init.py
  • src/openai/resources/webhooks/webhooks.py
  • src/openai/types/beta/agent_session_message.py
  • src/openai/types/chat/init.py
  • src/openai/types/chat/chat_completion_message_tool_call.py
  • src/openai/types/fine_tuning/fine_tuning_job_integration.py
  • src/openai/types/realtime/conversation_item_input_audio_transcription_delta_event.py
  • src/openai/types/realtime/realtime_error_event.py
  • src/openai/types/responses/init.py
  • src/openai/types/responses/response.py
  • src/openai/types/responses/response_function_web_search.py
  • src/openai/types/responses/response_function_web_search_param.py
  • src/openai/types/responses/responses_client_event.py
  • src/openai/types/responses/responses_client_event_param.py
  • src/openai/types/responses/tool.py

6 more in the full report.

A changed generated baseline means this report cannot reliably identify which handwritten lines changed.

Inspect the custom-code diff

Download the exact patch produced by this run (requires repository access):

gh run download 37822801555 --repo openai/openai-python \
  --name castiron-custom-code-37822801555-1 --dir /tmp/castiron-custom-code-37822801555-1
git apply --stat /tmp/castiron-custom-code-37822801555-1/custom-code.patch
cat /tmp/castiron-custom-code-37822801555-1/custom-code.patch

Or reproduce it from an SDK checkout containing the vendored reporter:

git fetch --no-tags origin 9301e319ea33ef28fba380f39a289dedc14652c1 83833ea92b0da18e51e32f1d39650aeff691430f
python3 scripts/castiron/custom_code_report.py report \
  --base 9301e319ea33ef28fba380f39a289dedc14652c1 \
  --head 83833ea92b0da18e51e32f1d39650aeff691430f --fetch --require-head-hash --public \
  --out /tmp/castiron-custom-code-83833ea92b0d
cat /tmp/castiron-custom-code-83833ea92b0d/custom-code.patch

This is the current full custom patch for mixed files, not an attribution of only the handwritten lines changed by this PR.

Full report and patch

Comment thread tests/test_streaming.py Fixed
Comment thread tests/test_streaming.py
await asyncio.wait_for(draining.wait(), timeout=5)
task.cancel()
with pytest.raises(asyncio.CancelledError):
await task

@markstuart-oai markstuart-oai left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed 83833ea92b0da18e51e32f1d39650aeff691430f. No actionable findings.

The HTTP/1.1 cleanup resumes the active byte iterator after [DONE] without parsing trailing SSE data. Transport failures are suppressed only during cleanup; cancellation still propagates and the response closes in finally. HTTP/2 and early-close paths remain unchanged.

The documented timeout tradeoff remains: this uses the request's per-read timeout, with no total cleanup deadline. A delayed ending adds completion latency; continuous trailing bytes or disabled read timeouts can keep iteration open indefinitely.

Source-only review; I did not run tests or the benchmark locally. Hosted Python 3.10, Python 3.14, HTTPX2, build, lint, and CodeQL checks passed for this head.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Streaming responses never return their connection to the pool (closed before the chunked terminator is read)

2 participants