Skip to content

feat(providers): translate chat completions onto Responses-only upstreams (chatgpt/Codex backend) #1084

Description

@weselben

Part of #723 (dialect independence across providers). This issue specifies the first slice: the chat-to-Responses adapter and its wiring into the chatgpt provider.

Problem

GoModel translates Responses API requests onto chat completions. See ResponsesViaChat in internal/providers/responses_adapter.go:364. The reverse direction does not exist. A client that calls /v1/chat/completions cannot use an upstream that serves only the Responses API.

This case already exists in the repo. The chatgpt provider targets the ChatGPT Codex backend at https://chatgpt.com/backend-api/codex. That backend serves only the Responses API. Its chat surface returns an error today (internal/providers/chatgpt/chatgpt.go:142). Chat-completions clients cannot use these models.

Current state

  • The adapter works in one direction only: Responses request to chat upstream (ResponsesViaChat, StreamResponsesViaChat).
  • The machinery for the reverse direction already exists:
    • Shape converters: responses_converter.go, responses_input.go, responses_output.go.
    • History reconstruction from stored snapshots: previous_response.go.
    • ID minting: responses_converter.go:47 mints resp_<uuid> for translated streams.
  • The chatgpt provider is the first consumer. It is Responses-native today. Its chat surface is unsupported.

Proposal

Add a ChatViaResponses helper that mirrors ResponsesViaChat. Wire it into ChatCompletion and StreamChatCompletion of the chatgpt provider.

  • Convert messages to Responses input items. This reverses the responses_input.go mapping.
  • Convert Responses output items to chat choices: text, tool calls, finish reasons. Map reasoning to reasoning_content, per repo convention.
  • Convert Responses SSE events to chat completion chunks. This reverses OpenAIResponsesStreamConverter. The Codex backend is streaming-only. A non-streaming chat request must collect the stream before it replies.
  • Map usage fields: input_tokens to prompt_tokens, output_tokens to completion_tokens. Map input_tokens_details.cached_tokens to prompt_tokens_details.cached_tokens.
  • Respect the strict parameter allowlist of the backend. Translate fields that keep their meaning. Reject fields that do not, such as logprobs and logit_bias. Do not drop fields silently. Postel's law applies in both directions.

IDs and conversation state

Minted IDs make the response shape correct. They are not resource handles.

  • Mint chatcmpl-<uuid> for the client-facing reply. This keeps the OpenAI chat shape. Chat clients do not send this ID back. Chat completions is stateless: each request carries the full messages array.
  • Keep the upstream resp_ ID internal. Use it for diagnostics and logs.
  • Server-side chaining cannot exist on the Codex backend. The backend pins store: false. It rejects store: true and previous_response_id (internal/providers/chatgpt/request.go). An ID without a stored response behind it is a label, not a handle.
  • Consequence: no client gets previous_response_id chaining through Codex. Full history travels in the request body.
  • Chat-path snapshots are out of scope here. If added later, they are gateway-side history reconstruction. They need their own lookup and ownership design.

Out of scope

  • Non-streaming translated Responses replies keep the upstream chatcmpl- ID. Translated streams mint resp_ IDs. Minting resp_ for non-stream replies too would make the translated surface consistent. This is a separate change.

This issue was drafted with AI assistance.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

No labels
No labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions