fix(stream): output correct SSE format for OpenAI vs Anthropic endpoints - #1
keggin-CHN wants to merge 1 commit into
Conversation
When POST /v1/chat/completions is called, the streaming response was
incorrectly outputting Anthropic-format SSE events (message_start,
content_block_delta, etc.) instead of OpenAI-format chunks.
This caused clients expecting OpenAI SSE format to fail with JSON
parsing errors like:
Type validation failed: missing choices[]
JSON parsing failed: Unexpected non-whitespace character (extra })
Changes:
- Pass is_anthropic flag through to convertAndSendSSE and emitTextDelta
- /v1/chat/completions: emit {choices:[{delta:{content:...}}]} chunks + [DONE]
- /v1/messages: keep existing Anthropic SSE format unchanged
- Skip irrelevant events (content_block_start, ping) for OpenAI clients
- Fix extra closing brace in writeAll (}} was not escaped, produced invalid JSON)
📝 WalkthroughWalkthrough
ChangesStreaming Format Duality
🎯 4 (Complex) | ⏱️ ~45 minutes
🚥 Pre-merge checks | ✅ 5✅ Passed checks (5 passed)
✏️ Tip: You can configure your own custom pre-merge checks in the settings. ✨ Finishing Touches🧪 Generate unit tests (beta)
Tip 💬 Introducing Slack Agent: The best way for teams to turn conversations into code.Slack Agent is built on CodeRabbit's deep understanding of your code, so your team can collaborate across the entire SDLC without losing context.
Built for teams:
One agent for your entire SDLC. Right inside Slack. Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
There was a problem hiding this comment.
🧹 Nitpick comments (1)
src/stream.zig (1)
305-319: ⚡ Quick winConsider routing OpenAI
text_deltathroughemitTextDeltato avoid duplication.Lines 313-316 emit the exact same OpenAI chunk shape as
emitTextDelta(Lines 408-410). Funneling this branch through the helper keeps the OpenAI payload format defined in a single place and avoids divergence later (e.g., when addingcreated, fingerprints, or chunked tool_calls).♻️ Proposed consolidation
if (std.mem.eql(u8, event_type, "content_block_delta")) { const delta = obj.object.get("delta") orelse return; if (delta != .object) return; if (!is_anthropic) { - // OpenAI format: extract text and emit choices[].delta.content const delta_type = switch (delta.object.get("type") orelse return) { .string => |s| s, else => return }; if (std.mem.eql(u8, delta_type, "text_delta")) { const text = switch (delta.object.get("text") orelse return) { .string => |s| s, else => return }; - var buf: std.io.Writer.Allocating = .init(allocator); - defer buf.deinit(); - const w = &buf.writer; - try w.print("data: {{\"id\":\"{s}\",\"object\":\"chat.completion.chunk\",\"model\":\"{s}\",\"choices\":[{{\"index\":0,\"delta\":{{\"content\":", .{ chat_id, model }); - try std.json.Stringify.encodeJsonString(text, .{}, w); - try w.writeAll("},\"finish_reason\":null}]}\n\n"); - try socket.send(client_stream, buf.written()); + try emitTextDelta(client_stream, text, block_index, is_anthropic, model, chat_id, allocator); } return; }🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@src/stream.zig` around lines 305 - 319, The OpenAI `text_delta` branch duplicates the chunk-emission logic; replace the inline writer/JSON assembly in the if (std.mem.eql(u8, delta_type, "text_delta")) block with a call to the existing helper emitTextDelta so the OpenAI payload format is centralized. Locate the block guarded by is_anthropic == false and delta_type == "text_delta", extract the text string as currently done, then call emitTextDelta passing the same client_stream, chat_id, model, and text (and any other contextual items emitTextDelta requires) instead of constructing and sending the buffer locally; remove the duplicated writer/print/json encode/send code to avoid divergence.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Nitpick comments:
In `@src/stream.zig`:
- Around line 305-319: The OpenAI `text_delta` branch duplicates the
chunk-emission logic; replace the inline writer/JSON assembly in the if
(std.mem.eql(u8, delta_type, "text_delta")) block with a call to the existing
helper emitTextDelta so the OpenAI payload format is centralized. Locate the
block guarded by is_anthropic == false and delta_type == "text_delta", extract
the text string as currently done, then call emitTextDelta passing the same
client_stream, chat_id, model, and text (and any other contextual items
emitTextDelta requires) instead of constructing and sending the buffer locally;
remove the duplicated writer/print/json encode/send code to avoid divergence.
|
感谢大佬,cherrystudio能用了 |
Bug Description
When a client calls
POST /v1/chat/completions(OpenAI-compatible endpoint) with streaming enabled, the server incorrectly returns Anthropic-format SSE events instead of OpenAI-format chunks.Clients receive events like:
But OpenAI-compatible clients expect:
This caused errors in OpenAI-compatible clients (e.g. Cherry Studio):
Type validation failed: missing choices[]JSON parsing failed: Unexpected non-whitespace character after JSON(extra})Root Cause
convertAndSendSSEandemitTextDeltaalways emitted Anthropic SSE format regardless of theis_anthropicflag passed tohandleStreamProxy. The flag was only used to build the upstream request payload, not to format the downstream response.Additionally, a
}}inside awriteAllcall (which is NOT a format string) emitted two literal}characters, producing invalid JSON in every streaming chunk.Fix
is_anthropic/model/chat_idcontext throughconvertAndSendSSEandemitTextDeltaPOST /v1/chat/completions(is_anthropic=false): emit standard OpenAI SSE chunks (choices[].delta.content) and terminate withdata: [DONE]POST /v1/messages(is_anthropic=true): keep existing Anthropic SSE format unchangedcontent_block_start,content_block_stop,ping) when serving OpenAI clients}inwriteAllthat produced malformed JSONTesting
Verified with Cherry Studio calling
POST /v1/chat/completionswith modelgpt-5.5on Windows — streaming now works correctly end-to-end without any JSON parse errors.Summary by CodeRabbit