Skip to content

fix(stream): output correct SSE format for OpenAI vs Anthropic endpoints - #1

Open
keggin-CHN wants to merge 1 commit into
yukmakoto:masterfrom
keggin-CHN:fix/openai-stream-sse-format
Open

keggin-CHN wants to merge 1 commit into
yukmakoto:masterfrom
keggin-CHN:fix/openai-stream-sse-format

Conversation

@keggin-CHN

@keggin-CHN keggin-CHN commented May 10, 2026 •

Copy link
Copy Markdown

Bug Description

When a client calls POST /v1/chat/completions (OpenAI-compatible endpoint) with streaming enabled, the server incorrectly returns Anthropic-format SSE events instead of OpenAI-format chunks.

Clients receive events like:

{"type":"message_start","message":{...}}
{"type":"content_block_delta","index":0,"delta":{"type":"text_delta","text":"..."}}
{"type":"message_stop"}

But OpenAI-compatible clients expect:

data: {"id":"chatcmpl-xxx","object":"chat.completion.chunk","choices":[{"index":0,"delta":{"content":"..."},"finish_reason":null}]}
data: [DONE]

This caused errors in OpenAI-compatible clients (e.g. Cherry Studio):

  • Type validation failed: missing choices[]
  • JSON parsing failed: Unexpected non-whitespace character after JSON (extra })

Root Cause

convertAndSendSSE and emitTextDelta always emitted Anthropic SSE format regardless of the is_anthropic flag passed to handleStreamProxy. The flag was only used to build the upstream request payload, not to format the downstream response.

Additionally, a }} inside a writeAll call (which is NOT a format string) emitted two literal } characters, producing invalid JSON in every streaming chunk.

Fix

  • Thread the is_anthropic / model / chat_id context through convertAndSendSSE and emitTextDelta
  • POST /v1/chat/completions (is_anthropic=false): emit standard OpenAI SSE chunks (choices[].delta.content) and terminate with data: [DONE]
  • POST /v1/messages (is_anthropic=true): keep existing Anthropic SSE format unchanged
  • Skip Anthropic-only events (content_block_start, content_block_stop, ping) when serving OpenAI clients
  • Fix extra } in writeAll that produced malformed JSON

Testing

Verified with Cherry Studio calling POST /v1/chat/completions with model gpt-5.5 on Windows — streaming now works correctly end-to-end without any JSON parse errors.

Summary by CodeRabbit

  • Improvements
    • Enhanced streaming response support to handle both Anthropic and OpenAI-compatible event formats within the same pathway.
    • Improved event formatting for streaming responses, ensuring proper response structure and end-of-stream signaling based on provider type.
    • Better compatibility with OpenAI-style event streams while maintaining support for Anthropic responses.

Review Change Stack

When POST /v1/chat/completions is called, the streaming response was
incorrectly outputting Anthropic-format SSE events (message_start,
content_block_delta, etc.) instead of OpenAI-format chunks.

This caused clients expecting OpenAI SSE format to fail with JSON
parsing errors like:
  Type validation failed: missing choices[]
  JSON parsing failed: Unexpected non-whitespace character (extra })

Changes:
- Pass is_anthropic flag through to convertAndSendSSE and emitTextDelta
- /v1/chat/completions: emit {choices:[{delta:{content:...}}]} chunks + [DONE]
- /v1/messages: keep existing Anthropic SSE format unchanged
- Skip irrelevant events (content_block_start, ping) for OpenAI clients
- Fix extra closing brace in writeAll (}} was not escaped, produced invalid JSON)
@coderabbitai

coderabbitai Bot commented May 10, 2026 •

Copy link
Copy Markdown
📝 Walkthrough

Walkthrough

src/stream.zig extends the SSE streaming proxy to emit either Anthropic or OpenAI-compatible event formats from a single codepath. Handshake, content conversion, and termination logic now branch by format, while all internal delta emission passes format context to ensure consistency across multiple upstream JSON shapes.

Changes

Streaming Format Duality

Layer / File(s) Summary
Chat ID Generation
src/stream.zig
Upstream model and stable OpenAI-style chat_id (timestamp-based) are extracted early for use in downstream chunk/event IDs.
Routing Signatures
src/stream.zig
convertAndSendSSE and emitTextDelta are extended to accept is_anthropic, model, and chat_id parameters, enabling format-aware event generation.
Opening Handshake
src/stream.zig
SSE header emission branches by format: Anthropic sends message_start with model; OpenAI sends initial assistant role chat.completion.chunk containing chat_id, model, and empty delta.
Content Routing
src/stream.zig
For OpenAI-format clients, Anthropic-only content block events (content_block_start, content_block_stop, ping) are filtered out, and content_block_delta is converted to OpenAI choices[0].delta.content chunks; Anthropic-format clients preserve original passthrough.
Delta Emission
src/stream.zig
All internal calls to emitTextDelta now pass format context, ensuring deltas from multiple upstream sources (xAI, Gemini, etc.) produce consistent Anthropic or OpenAI chunk formats.
Closing Handshake & Integration
src/stream.zig
End-of-stream logic branches by format: Anthropic emits message_delta and message_stop; OpenAI emits final chat.completion.chunk with finish_reason: "stop" followed by [DONE] marker. Main stream loop invokes convertAndSendSSE with full context.

🎯 4 (Complex) | ⏱️ ~45 minutes

🐰 Streams now sing in dual tongues,
Anthropic whispers, OpenAI speaks,
One pipe, two formats, no forks required—
Chat IDs sprout like carrots from timestamps,
And the proxy hops merrily between their dialects! 🐇✨

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title clearly and specifically describes the main fix: outputting the correct SSE format based on whether the endpoint is OpenAI or Anthropic compatible. It directly summarizes the primary change in the changeset.
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests

Tip

💬 Introducing Slack Agent: The best way for teams to turn conversations into code.

Slack Agent is built on CodeRabbit's deep understanding of your code, so your team can collaborate across the entire SDLC without losing context.

  • Generate code and open pull requests
  • Plan features and break down work
  • Investigate incidents and troubleshoot customer tickets together
  • Automate recurring tasks and respond to alerts with triggers
  • Summarize progress and report instantly

Built for teams:

  • Shared memory across your entire org—no repeating context
  • Per-thread sandboxes to safely plan and execute work
  • Governance built-in—scoped access, auditability, and budget controls

One agent for your entire SDLC. Right inside Slack.

👉 Get started


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🧹 Nitpick comments (1)
src/stream.zig (1)

305-319: ⚡ Quick win

Consider routing OpenAI text_delta through emitTextDelta to avoid duplication.

Lines 313-316 emit the exact same OpenAI chunk shape as emitTextDelta (Lines 408-410). Funneling this branch through the helper keeps the OpenAI payload format defined in a single place and avoids divergence later (e.g., when adding created, fingerprints, or chunked tool_calls).

♻️ Proposed consolidation
             if (std.mem.eql(u8, event_type, "content_block_delta")) {
                 const delta = obj.object.get("delta") orelse return;
                 if (delta != .object) return;
                 if (!is_anthropic) {
-                    // OpenAI format: extract text and emit choices[].delta.content
                     const delta_type = switch (delta.object.get("type") orelse return) { .string => |s| s, else => return };
                     if (std.mem.eql(u8, delta_type, "text_delta")) {
                         const text = switch (delta.object.get("text") orelse return) { .string => |s| s, else => return };
-                        var buf: std.io.Writer.Allocating = .init(allocator);
-                        defer buf.deinit();
-                        const w = &buf.writer;
-                        try w.print("data: {{\"id\":\"{s}\",\"object\":\"chat.completion.chunk\",\"model\":\"{s}\",\"choices\":[{{\"index\":0,\"delta\":{{\"content\":", .{ chat_id, model });
-                        try std.json.Stringify.encodeJsonString(text, .{}, w);
-                        try w.writeAll("},\"finish_reason\":null}]}\n\n");
-                        try socket.send(client_stream, buf.written());
+                        try emitTextDelta(client_stream, text, block_index, is_anthropic, model, chat_id, allocator);
                     }
                     return;
                 }
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@src/stream.zig` around lines 305 - 319, The OpenAI `text_delta` branch
duplicates the chunk-emission logic; replace the inline writer/JSON assembly in
the if (std.mem.eql(u8, delta_type, "text_delta")) block with a call to the
existing helper emitTextDelta so the OpenAI payload format is centralized.
Locate the block guarded by is_anthropic == false and delta_type ==
"text_delta", extract the text string as currently done, then call emitTextDelta
passing the same client_stream, chat_id, model, and text (and any other
contextual items emitTextDelta requires) instead of constructing and sending the
buffer locally; remove the duplicated writer/print/json encode/send code to
avoid divergence.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Nitpick comments:
In `@src/stream.zig`:
- Around line 305-319: The OpenAI `text_delta` branch duplicates the
chunk-emission logic; replace the inline writer/JSON assembly in the if
(std.mem.eql(u8, delta_type, "text_delta")) block with a call to the existing
helper emitTextDelta so the OpenAI payload format is centralized. Locate the
block guarded by is_anthropic == false and delta_type == "text_delta", extract
the text string as currently done, then call emitTextDelta passing the same
client_stream, chat_id, model, and text (and any other contextual items
emitTextDelta requires) instead of constructing and sending the buffer locally;
remove the duplicated writer/print/json encode/send code to avoid divergence.

ℹ️ Review info
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro

Run ID: 594006cf-d228-4f8b-b69a-a7000856107d

📥 Commits

Reviewing files that changed from the base of the PR and between f65fc92 and 461506b.

📒 Files selected for processing (1)
  • src/stream.zig

@Jackenychen

Copy link
Copy Markdown

感谢大佬,cherrystudio能用了

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants