Repository navigation
Responses create spends CPU proportional to the body size transforming a body it returns unchanged #4021
Description
Activity
- added a commit that references this issue
on Oct 5, 2026 I independently measured the transformation helper on main
becc1d20eed83c1b8d85e15dc131a372d9dc7813(SDK 3.24.0, Python 3.14.7, macOS). All payloads were fictional; no client, API request or model was used. After two warm-ups, seven runs per shape gave median wall times of 4.55 ms (30 input items/10 tools), 14.23 ms (90/40), 28.65 ms (180/86), and 51.18 ms (360/86). Every transformed body compared equal to its input. These are host-specific synthetic timings, not production latency.The 180/86 profile recorded 95,579 get_origin calls, 7,857 _async_transform_typeddict calls and 21,068 _async_transform_recursive calls. An Annotated alias/date control still correctly renamed the field and serialized datetime, so bypassing transformation wholesale would be unsafe. A narrower optimization should retain those behaviors. The exact offline diagnostic is below.
To run the diagnostic, save it under a
scripts/directory and create a siblingresults/directory. Use the dependency versions above and the pinned SDK source onPYTHONPATH.from pathlib import Path import asyncio,cProfile,json,statistics,time,importlib.metadata from datetime import datetime,timezone from typing_extensions import TypedDict,Annotated from openai._utils import async_maybe_transform,PropertyInfo from openai.types.responses import response_create_params def body(turns,tool_count): items=[] for i in range(turns): items.extend([{'role':'user','content':[{'type':'input_text','text':'fictional '*100}]},{'type':'function_call','call_id':f'fictional_call_{i}','name':'lookup','arguments':'{"q":"fictional"}'},{'type':'function_call_output','call_id':f'fictional_call_{i}','output':'fictional result '*200}]) tools=[{'type':'function','name':f'fictional_tool_{i}','description':'fictional '*10,'strict':False,'parameters':{'type':'object','properties':{f'arg_{j}':{'type':'string','description':'fictional'} for j in range(8)},'required':['arg_0']}} for i in range(tool_count)] return {'model':'fictional-unused-model','input':items,'tools':tools,'stream':True} class TransformControl(TypedDict): field:Annotated[str,PropertyInfo(alias='api_field')] timestamp:Annotated[datetime,PropertyInfo(format='iso8601')] async def main(): rows=[] for turns,tools in ((10,10),(30,40),(60,86),(120,86)): value=body(turns,tools) for _ in range(2):await async_maybe_transform(value,response_create_params.ResponseCreateParamsStreaming) durations=[];cpu=[] for _ in range(7): t=time.perf_counter();c=time.process_time();out=await async_maybe_transform(value,response_create_params.ResponseCreateParamsStreaming) durations.append((time.perf_counter()-t)*1000);cpu.append((time.process_time()-c)*1000) assert out==value rows.append({'input_items':len(value['input']),'tools':tools,'median_wall_ms':statistics.median(durations),'median_process_cpu_ms':statistics.median(cpu),'output_equals_input':True}) profile=cProfile.Profile();profile.enable();await async_maybe_transform(body(60,86),response_create_params.ResponseCreateParamsStreaming);profile.disable() counts={} for entry in profile.getstats(): if hasattr(entry.code,'co_name') and entry.code.co_name in {'_async_transform_recursive','_async_transform_typeddict','get_origin'}:counts[entry.code.co_name]=counts.get(entry.code.co_name,0)+entry.callcount dt=datetime(2026,1,1,tzinfo=timezone.utc) control=await async_maybe_transform({'field':'fictional','timestamp':dt},TransformControl) assert control=={'api_field':'fictional','timestamp':dt.isoformat()} result={'sdk_version':importlib.metadata.version('openai'),'rows':rows,'profile_180_items_86_tools':counts,'alias_and_datetime_control':control,'live_api_calls':0,'scope':'Offline synthetic request-body transformation on Mini. No client, request or model execution.'} (Path(__file__).resolve().parents[1]/'results/transform-cost.json').write_text(json.dumps(result,indent=2)+'\n') print(json.dumps(result)) asyncio.run(main())
Diagnostic assistance: prepared with Codex and independently executed against the pinned source; all inputs are fictional.
Summary
Implemented the optimization suggested in the issue on my fork — branch
perf/4021-transform-cache.Approach
- Precompute a per-type dispatch plan (cached via
lru_cache) so the hot loop performs notypingintrospection. - All-
TypedDictunions folded into a single merged-key pass, collapsing the per-value union fan-out. - Metadata-free types take a fast path; anything carrying
PropertyInfo(aliases,iso8601/base64formats) runs the original code unchanged.
Benchmarks
The issue's repro script, medians of 11 runs:
input items tools before (ms) after (ms) 30 10 18.0 2.5 90 40 52.6 7.8 180 86 107.5 15.9 360 86 190.2 27.3 Output is identical in every case (
output == input).Correctness
- All 64 tests in
tests/test_transform.pypass. - Differential-fuzzed against pristine
main: 874 cases (aliases,iso8601/base64formats,NOT_GIVEN, pydantic models, nested unions, unannotated keys, 400 randomized type/data pairs, sync + async) — zero behavioral differences. ruff checkandruff formatclean.
Note
I'd open this as a PR but the repo limits PR creation to collaborators — happy for a maintainer to open one from the branch, or I can submit it if the restriction is lifted.
- Precompute a per-type dispatch plan (cached via
Confirm this is an issue with the Python library and not an underlying OpenAI API
Describe the bug
AsyncResponses.createrunsasync_maybe_transformover the whole request body before every request. With a longinputand manytools, that transform costs tens of milliseconds of CPU per request, and the cost grows linearly with the body size. For a body of plain dicts (no aliased keys, noformatfields, no file inputs) the transform returns a body equal to its input, so the CPU time does not change the request.The transform runs on the event loop. In an asyncio app that shares the loop with latency-sensitive work, such as streaming audio, each request blocks that work for the whole transform. A chat app sends the full history on every turn, so the block grows as the conversation gets longer.
Responses.create(sync) runs the same transform.Median of 11 runs, from the script under Code snippets:
cProfileof one request with 180 input items and 86 tools: 21,068 calls to_utils/_transform.py:_async_transform_recursive, 7,857 calls to_async_transform_typeddict, and about 74,000 calls totyping.get_origin. The time goes to the walk over every nested dict and list, not to any field the transform rewrites.A possible direction: decide once per type, and cache, whether its annotations contain any
PropertyInfo(alias, format, discriminator). When a type contains none, return the data without walking it.To Reproduce
input(user messages,function_callandfunction_call_outputitems) and 86 functiontools.async_maybe_transform(body, ResponseCreateParamsStreaming), the callAsyncResponses.createmakes before it sends the request.The script under Code snippets does this with synthetic data and makes no network call. Its output:
Code snippets
OS
macOS 26.5 (arm64)
Python version
Python v3.14.3
Library version
openai v3.24.0