Repository navigation
Conversation
Signed-off-by: 1fanwang <1fannnw@gmail.com>
Signed-off-by: 1fanwang <1fannnw@gmail.com>
|
Independent offline verification (AI-assisted): I extended the verifier check to actual SDK Compared main Runnable separately against each source tree (pass its import ast,importlib,json,pathlib,sys
from typing import Annotated,Any,Sequence,Literal
from openai.types import Model
path=pathlib.Path(sys.argv[1]);tree=ast.parse(path.read_text());names={'evaluate_forwardref','assert_matches_model','assert_matches_type','_assert_list_type'}
# Execute unchanged verifier functions and their existing imports; omit rich display/server/test infrastructure.
body=[n for n in tree.body if isinstance(n,(ast.Import,ast.ImportFrom)) and not (isinstance(n,ast.Import) and any(a.name=='rich' for a in n.names)) or isinstance(n,ast.Assign) and any(isinstance(t,ast.Name) and t.id=='BaseModelT' for t in n.targets) or isinstance(n,ast.FunctionDef) and n.name in names]
ns={};exec(compile(ast.Module(body=body,type_ignores=[]),str(path),'exec'),ns)
valid=Model.model_construct(id='fictional',created=0,object='model',owned_by='fictional');invalid=Model.model_construct(id='fictional',created='wrong',object='model',owned_by='fictional')
shapes=[('list',list[Model],lambda m:[m]),('sequence_list',Sequence[list[Model]],lambda m:([m],)),('dict_list',dict[str,list[Model]],lambda m:{'first':[m]}),('list_dict_sequence',list[dict[str,Sequence[Model]]],lambda m:[{'first':(m,)}]),('nullable_list',list[Model|None],lambda m:[None,m]),('annotated_list',Annotated[list[Model],'metadata'],lambda m:[m])]
cases=[(name+'_'+('valid' if good else 'invalid'),annotation,wrap(valid if good else invalid),good) for name,annotation,wrap in shapes for good in [False,True]]
cases += [('direct_any',Any,{'unknown':object()},True),('list_any',list[Any],[1,'fictional',None],True),('sequence_any',Sequence[Any],(1,'fictional',None),True),('literal_valid',list[Literal['fixture']],['fixture'],True),('literal_invalid',list[Literal['fixture']],['wrong'],False)]
rows=[]
for name,annotation,value,expected in cases:
try:ns['assert_matches_type'](annotation,value,path=['response']);accepted=True;trace_paths=[]
except AssertionError as exc:
accepted=False;trace_paths=[];tb=exc.__traceback__
while tb:
frame_path=tb.tb_frame.f_locals.get('path')
if isinstance(frame_path,list):trace_paths.append(list(frame_path))
tb=tb.tb_next
rows.append({'case':name,'accepted':accepted,'expected_accepted':expected,'diagnostic_paths':trace_paths,'pass':accepted==expected})
print(json.dumps({'source_verifier':str(path),'cases':rows,'passed':sum(r['pass'] for r in rows),'failed':sum(not r['pass'] for r in rows),'scope':'Unchanged AST-selected test verifier functions; actual SDK Model constructed without coercion, nested model/union/Annotated/Any/Literal boundaries; no full test suite, server or SDK parser behavior claim.'})) |
Changes being requested
API tests can pass when response fields inside lists or sequences have the wrong types. Their verifier calls a static typing helper that does not validate values when the tests run.
Check each member with the existing recursive verifier and retain its position in the diagnostic path. Unconstrained annotations still accept any value. This changes test verification, not the SDK's response-parsing behavior.
Additional context & links
The added regression starts a real local HTTP server and calls the sync and async clients. In loose-validation mode, an invalid model-creation timestamp reaches the verifier as a string. The verifier now rejects it; the valid integer response still passes.
Testing Done
The baseline is 69a2c1d.
Set up the project's locked environment with
uv sync --locked. Create a separate checkout at that baseline, copy the added regression module into it, and share the environment:git worktree add --detach ../openai-type-assertions-before 69a2c1db6feacf32be6693809e7cab1c3b49cad7 cp tests/test_utils/test_assert_matches_type.py ../openai-type-assertions-before/tests/test_utils/ ln -s "$PWD/.venv" ../openai-type-assertions-before/.venvRun the same command from the baseline and PR checkouts:
Raw output before the fix:
Both invalid-response cases fail with:
Raw output after the fix:
The baseline exits 1; the patched checkout exits 0. The same HTTP checks also pass with Pydantic v1.