Repository navigation
fix(failover): fail over on relayed upstream failures wearing 4xx statuses - #607
Conversation
…tuses
Aggregator-style providers can relay a transient failure of their own
upstream as a client error: OpenCode Zen returns 400 "Upstream request
failed" when its upstream breaks, so requests with configured failover
targets still died with the raw 400.
ShouldAttemptFailover now treats a message that blames an upstream
("upstream" plus failure phrasing) as failover-eligible regardless of
the status code, mirroring the existing model-availability heuristic.
The three inline fragment-scan loops are extracted into named fragment
lists checked via one containsAny helper.
Closes #605
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: Organization UI Review profile: ASSERTIVE Plan: Pro Plus Run ID: 📒 Files selected for processing (3)
📝 WalkthroughWalkthroughFailover detection now centralizes message matching and recognizes selected upstream transient failure messages relayed with 4xx statuses. Tests cover matching and non-matching upstream errors, and documentation describes the new trigger. ChangesFailover detection
Estimated code review effort: 2 (Simple) | ~10 minutes Possibly related PRs
Poem
🚥 Pre-merge checks | ✅ 5✅ Passed checks (5 passed)
✨ Finishing Touches📝 Generate docstrings
🧪 Generate unit tests (beta)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
|
Preview deployment for your docs. Learn more about Mintlify Previews.
💡 Tip: Enable Workflows to automatically generate PRs for you. |
|
Codecov Report✅ All modified and coverable lines are covered by tests. 📢 Thoughts on this report? Let us know! |
Confidence Score: 5/5The PR appears safe to merge with no actionable defects identified. The new predicate reaches the intended relayed upstream failures, retains the established failover behavior, and does not alter the existing model-availability or status-based checks.
What T-Rex did
Reviews (1): Last reviewed commit: "fix(failover): fail over on relayed upst..." | Re-trigger Greptile |
|
Preview deployment for your docs. Learn more about Mintlify Previews.
💡 Tip: Enable Workflows to automatically generate PRs for you. |
Summary
Closes #605.
OpenCode Zen (and other aggregator-style providers) can relay a transient failure of their upstream as a 4xx client error — e.g.
400 "Error from provider (Console Go): Upstream request failed". Failover previously ran only on5xx,429, and model-availability errors, so a request whose primary model had configured failover targets still returned the raw 400.ShouldAttemptFailovernow also fires when the error message blames an upstream: it must contain"upstream"plus failure phrasing (failed/error/unavailable/timed out/timeout), regardless of the status code the provider chose. Messages that merely mention an upstream without failure phrasing (e.g."parameter tools is not supported by the upstream provider") still do not fail over, so genuine validation errors don't sweep targets.No new configuration: this follows the same message-heuristic pattern as the existing model-availability trigger, keeping the default behavior right without a knob.
Changes
internal/gateway/failover.go: new upstream-failure trigger; the three inline fragment-scan loops are extracted into named fragment lists (modelUnavailableFragments,upstreamFailureFragments,retiredModel404Fragments) checked via onecontainsAnyhelper.internal/gateway/failover_test.go: table cases for the exact reported OpenCode message, upstream timeout/unavailable, and a no-failover guard for upstream capability errors.docs/features/failover.mdx: "When It Runs" documents the new trigger.User-visible impact
🤖 Generated with Claude Code
Summary by CodeRabbit