Skip to content

Responses create spends CPU proportional to the body size transforming a body it returns unchanged #4021

Description

@zdurm

Confirm this is an issue with the Python library and not an underlying OpenAI API

  • This is an issue with the Python library

Describe the bug

AsyncResponses.create runs async_maybe_transform over the whole request body before every request. With a long input and many tools, that transform costs tens of milliseconds of CPU per request, and the cost grows linearly with the body size. For a body of plain dicts (no aliased keys, no format fields, no file inputs) the transform returns a body equal to its input, so the CPU time does not change the request.

The transform runs on the event loop. In an asyncio app that shares the loop with latency-sensitive work, such as streaming audio, each request blocks that work for the whole transform. A chat app sends the full history on every turn, so the block grows as the conversation gets longer. Responses.create (sync) runs the same transform.

Median of 11 runs, from the script under Code snippets:

input items tools transform (ms) output == input
30 10 4.6 True
90 40 14.5 True
180 86 29.7 True
360 86 54.0 True

cProfile of one request with 180 input items and 86 tools: 21,068 calls to _utils/_transform.py:_async_transform_recursive, 7,857 calls to _async_transform_typeddict, and about 74,000 calls to typing.get_origin. The time goes to the walk over every nested dict and list, not to any field the transform rewrites.

A possible direction: decide once per type, and cache, whether its annotations contain any PropertyInfo (alias, format, discriminator). When a type contains none, return the data without walking it.

To Reproduce

  1. Build a Responses request body of plain dicts with a long input (user messages, function_call and function_call_output items) and 86 function tools.
  2. Time async_maybe_transform(body, ResponseCreateParamsStreaming), the call AsyncResponses.create makes before it sends the request.
  3. Compare the transformed body with the input.

The script under Code snippets does this with synthetic data and makes no network call. Its output:

input_items  tools  transform_ms  output_equals_input
         30     10           4.6  True
         90     40          14.5  True
        180     86          29.7  True
        360     86          54.0  True

Code snippets

import asyncio
import time

from openai._utils import async_maybe_transform
from openai.types.responses import response_create_params


def build_body(turns: int, n_tools: int) -> dict:
    items = []
    for i in range(turns):
        items.append({"role": "user", "content": [{"type": "input_text", "text": "lorem ipsum " * 100}]})
        items.append({"type": "function_call", "call_id": f"call_{i}", "name": "lookup", "arguments": '{"q": "x"}'})
        items.append({"type": "function_call_output", "call_id": f"call_{i}", "output": "result " * 200})
    tools = [
        {
            "type": "function",
            "name": f"tool_{i}",
            "description": "does a thing " * 10,
            "strict": False,
            "parameters": {
                "type": "object",
                "properties": {f"arg_{j}": {"type": "string", "description": "an argument"} for j in range(8)},
                "required": ["arg_0"],
            },
        }
        for i in range(n_tools)
    ]
    return {"model": "gpt-4.1", "input": items, "tools": tools, "stream": True}


async def median_ms(body: dict) -> tuple[float, bool]:
    # The same call AsyncResponses.create makes before sending the request.
    for _ in range(3):
        out = await async_maybe_transform(body, response_create_params.ResponseCreateParamsStreaming)
    runs = []
    for _ in range(11):
        start = time.perf_counter()
        out = await async_maybe_transform(body, response_create_params.ResponseCreateParamsStreaming)
        runs.append((time.perf_counter() - start) * 1000)
    return sorted(runs)[5], out == body


async def main() -> None:
    print("input_items  tools  transform_ms  output_equals_input")
    for turns, n_tools in [(10, 10), (30, 40), (60, 86), (120, 86)]:
        body = build_body(turns, n_tools)
        ms, same = await median_ms(body)
        print(f"{len(body['input']):>11}  {n_tools:>5}  {ms:>12.1f}  {same}")


asyncio.run(main())

OS

macOS 26.5 (arm64)

Python version

Python v3.14.3

Library version

openai v3.24.0

Activity

  1. added
    bugSomething isn't working
    on Oct 2, 2026
  2. added a commit that references this issue on Oct 5, 2026
    07dedeb
  3. jnohclee-rgb commented on Oct 5, 2026

    @jnohclee-rgb

    I independently measured the transformation helper on main becc1d20eed83c1b8d85e15dc131a372d9dc7813 (SDK 3.24.0, Python 3.14.7, macOS). All payloads were fictional; no client, API request or model was used. After two warm-ups, seven runs per shape gave median wall times of 4.55 ms (30 input items/10 tools), 14.23 ms (90/40), 28.65 ms (180/86), and 51.18 ms (360/86). Every transformed body compared equal to its input. These are host-specific synthetic timings, not production latency.

    The 180/86 profile recorded 95,579 get_origin calls, 7,857 _async_transform_typeddict calls and 21,068 _async_transform_recursive calls. An Annotated alias/date control still correctly renamed the field and serialized datetime, so bypassing transformation wholesale would be unsafe. A narrower optimization should retain those behaviors. The exact offline diagnostic is below.

    To run the diagnostic, save it under a scripts/ directory and create a sibling results/ directory. Use the dependency versions above and the pinned SDK source on PYTHONPATH.

    from pathlib import Path
    import asyncio,cProfile,json,statistics,time,importlib.metadata
    from datetime import datetime,timezone
    from typing_extensions import TypedDict,Annotated
    from openai._utils import async_maybe_transform,PropertyInfo
    from openai.types.responses import response_create_params
    
    def body(turns,tool_count):
        items=[]
        for i in range(turns):
            items.extend([{'role':'user','content':[{'type':'input_text','text':'fictional '*100}]},{'type':'function_call','call_id':f'fictional_call_{i}','name':'lookup','arguments':'{"q":"fictional"}'},{'type':'function_call_output','call_id':f'fictional_call_{i}','output':'fictional result '*200}])
        tools=[{'type':'function','name':f'fictional_tool_{i}','description':'fictional '*10,'strict':False,'parameters':{'type':'object','properties':{f'arg_{j}':{'type':'string','description':'fictional'} for j in range(8)},'required':['arg_0']}} for i in range(tool_count)]
        return {'model':'fictional-unused-model','input':items,'tools':tools,'stream':True}
    
    class TransformControl(TypedDict):
        field:Annotated[str,PropertyInfo(alias='api_field')]
        timestamp:Annotated[datetime,PropertyInfo(format='iso8601')]
    
    async def main():
        rows=[]
        for turns,tools in ((10,10),(30,40),(60,86),(120,86)):
            value=body(turns,tools)
            for _ in range(2):await async_maybe_transform(value,response_create_params.ResponseCreateParamsStreaming)
            durations=[];cpu=[]
            for _ in range(7):
                t=time.perf_counter();c=time.process_time();out=await async_maybe_transform(value,response_create_params.ResponseCreateParamsStreaming)
                durations.append((time.perf_counter()-t)*1000);cpu.append((time.process_time()-c)*1000)
                assert out==value
            rows.append({'input_items':len(value['input']),'tools':tools,'median_wall_ms':statistics.median(durations),'median_process_cpu_ms':statistics.median(cpu),'output_equals_input':True})
        profile=cProfile.Profile();profile.enable();await async_maybe_transform(body(60,86),response_create_params.ResponseCreateParamsStreaming);profile.disable()
        counts={}
        for entry in profile.getstats():
            if hasattr(entry.code,'co_name') and entry.code.co_name in {'_async_transform_recursive','_async_transform_typeddict','get_origin'}:counts[entry.code.co_name]=counts.get(entry.code.co_name,0)+entry.callcount
        dt=datetime(2026,1,1,tzinfo=timezone.utc)
        control=await async_maybe_transform({'field':'fictional','timestamp':dt},TransformControl)
        assert control=={'api_field':'fictional','timestamp':dt.isoformat()}
        result={'sdk_version':importlib.metadata.version('openai'),'rows':rows,'profile_180_items_86_tools':counts,'alias_and_datetime_control':control,'live_api_calls':0,'scope':'Offline synthetic request-body transformation on Mini. No client, request or model execution.'}
        (Path(__file__).resolve().parents[1]/'results/transform-cost.json').write_text(json.dumps(result,indent=2)+'\n')
        print(json.dumps(result))
    asyncio.run(main())

    Diagnostic assistance: prepared with Codex and independently executed against the pinned source; all inputs are fictional.

  4. MuhammadUsman-Khan commented on Oct 11, 2026

    @MuhammadUsman-Khan

    Summary

    Implemented the optimization suggested in the issue on my fork — branch perf/4021-transform-cache.

    Approach

    • Precompute a per-type dispatch plan (cached via lru_cache) so the hot loop performs no typing introspection.
    • All-TypedDict unions folded into a single merged-key pass, collapsing the per-value union fan-out.
    • Metadata-free types take a fast path; anything carrying PropertyInfo (aliases, iso8601/base64 formats) runs the original code unchanged.

    Benchmarks

    The issue's repro script, medians of 11 runs:

    input items tools before (ms) after (ms)
    30 10 18.0 2.5
    90 40 52.6 7.8
    180 86 107.5 15.9
    360 86 190.2 27.3

    Output is identical in every case (output == input).

    Correctness

    • All 64 tests in tests/test_transform.py pass.
    • Differential-fuzzed against pristine main: 874 cases (aliases, iso8601/base64 formats, NOT_GIVEN, pydantic models, nested unions, unannotated keys, 400 randomized type/data pairs, sync + async) — zero behavioral differences.
    • ruff check and ruff format clean.

    Note

    I'd open this as a PR but the repo limits PR creation to collaborators — happy for a maintainer to open one from the branch, or I can submit it if the restriction is lifted.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't working

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions