Skip to content

feat(eu): route AI through OpenRouter's EU endpoint and turn chat back on - #1052

Merged
Makisuo merged 1 commit into
mainfrom
feat/eu-openrouter-in-region
Sep 24, 2026
Merged

Makisuo merged 1 commit into
mainfrom
feat/eu-openrouter-in-region

Conversation

@Makisuo

@Makisuo Makisuo commented Sep 24, 2026 •

Copy link
Copy Markdown
Collaborator

The EU instance shipped with AI off (#997) because model calls carry customer spans and logs to model providers. OpenRouter's in-region routing fixes that: requests to eu.openrouter.ai are decrypted and served only by providers inside the EU, and a model with no EU provider returns a 404 instead of hopping to the US.

Changes

  • apps/ai/src/platform/Llm.ts: with MAPLE_REGION=eu (already on the AI worker via selfObservabilityEnv), every OpenRouter call goes to https://eu.openrouter.ai/api/v1: chat, reviews, embeddings and decisions. US stays on the global endpoint.
  • EU default model: openai/gpt-6-luna for chat, triage and reviews. The EU catalogue (66 models on 2026-09-25) serves none of the US defaults (glm-5.3-flash, deepseek-v4.1-flash, Jev). MAPLE_TRIAGE_MODEL_OPENROUTER / MAPLE_REVIEW_MODEL_OPENROUTER still override it. Its context limits are added to the table.
  • Decision model: Jev has no EU provider, so the EU instance has no decision model unless MAPLE_DECISION_MODEL is set. The triage route returns "no decision model is served in this region" without calling out. The gate already treats no verdict as "investigate".
  • Web: reverts fix(web): disable AI chat on the EU instance #997's gating (aiChatEnabled): the header button, command palette action, C shortcut, widget fix action and /chat are back on EU.
  • docs/eu-region-plan.md: Phase 4 rewritten for the new setup.

Before deploying prod-eu

  • The account behind prod-eu's OPENROUTER_API_KEY must be on OpenRouter's Business or Enterprise plan. Otherwise every EU model call fails.
  • MAPLE_LLM_PROVIDER must stay unset on prod-eu: Workers AI has no region pin.
  • Unverified: whether openai/text-embedding-3-small is served in the EU. The public embeddings listing ignores region=eu. If it isn't, the review feedback filter's embeddings fail there.

Tests

Llm.test.ts covers the EU URL and default models, the US instance staying on the global endpoint, and the decision model being unset in the EU. apps/web typecheck is clean.


View with [code]smith Autofix with [code]smith
Need help on this PR? Tag @codesmith-bot with what you need. Autofix is disabled.

Summary by CodeRabbit

  • New Features

    • AI chat, the “Ask Maple AI” command, and dashboard widget-fix actions are now available in the EU instance.
    • EU AI requests are routed through an EU-based service and use models available in that region.
  • Bug Fixes

    • When no decision model is available in the EU, incident triage now reports that limitation instead of attempting classification. The investigation flow treats missing verdicts as requiring investigation.

…k on

The EU instance shipped with AI off (#997) because model calls would carry
customer spans and logs to US providers. OpenRouter's in-region routing
closes that: requests to eu.openrouter.ai are decrypted and served only by
providers inside the EU, and a model with no EU provider is a 404 rather
than a hop to the US.

With MAPLE_REGION=eu, every OpenRouter call (chat, reviews, embeddings,
decisions) goes to https://eu.openrouter.ai/api/v1. The EU catalogue serves
none of the US defaults, so the EU instance defaults to openai/gpt-6-luna
for chat, triage and reviews. Jev has no EU provider, so the EU has no
decision model unless MAPLE_DECISION_MODEL is set; the triage route says so
without calling out, and the gate reads no verdict as "investigate".

Reverts the web gating from #997 so every chat entry point is back on EU.
@maple-review-bot

maple-review-bot Bot commented Sep 24, 2026 •

Copy link
Copy Markdown

Maple review

Confidence Quality Open findings Commit
3/5 · needs attention 100/100 · excellent none 76d64e9

Why 3/5: I verified the EU endpoint/model plumbing in apps/ai/src/platform/Llm.ts and its env wiring, but the truncated web hunks (dashboard-layout.tsx, chat.tsx, region.ts) and the EU behaviour of the non-chat consumers of DecisionModel were not fully read.

Warning

This review ended early; what follows is what it established.

The EU instance now routes every OpenRouter call to https://eu.openrouter.ai/api/v1 when MAPLE_REGION=eu and switches its chat, triage and review defaults to openai/gpt-6-luna, which has an EU provider; Jev's decision model has none, so resolveDecisionModel returns undefined there and the triage route fails closed. The web change undoes #997's instance guard, re-enabling the global chat sheet, the ⌘K action, widget "fix with AI" and the header button everywhere. As written it looks safe: the region flag is derived into the AI worker's env, the model ids and limits are registered, and the new EU path is covered by tests.

  • layerLlm now builds its OpenRouterClient against openRouterApiUrl(env), picking the EU in-region endpoint from MAPLE_REGION (Llm.ts:505, 72-75)
  • resolveTriageModel and resolveReviewModel fall back to EU_DEFAULT_OPENROUTER_MODEL / EU_DEFAULT_REVIEW_MODEL in the EU, and MODEL_LIMITS gains openai/gpt-6-luna (Llm.ts:208-210, 408-409)
  • resolveDecisionModel returns undefined in the EU unless MAPLE_DECISION_MODEL is set, so triage.http.ts answers IncidentTriageModelError instead of calling a model with no EU provider
  • Web chat is re-enabled on all instances: aiChatEnabled is deleted from lib/region.ts and its guards are removed from the chat sheet, command palette, widget actions and dashboard layout

What was checked

  • MAPLE_REGION actually reaches the AI worker's LlmEnv: it is derived in selfObservabilityEnv (packages/infra/src/env.ts:222) and that env is merged into the worker's bindings (apps/ai/src/worker.ts:104)
  • The EU branch of resolveTriageModel is the only caller of the new default and keeps MAPLE_TRIAGE_MODEL_OPENROUTER as the override (apps/ai/src/platform/Llm.ts:408)
  • Removing aiChatEnabled from lib/region.ts leaves no dangling references — a repo-wide grep for the symbol returns nothing at the head SHA
  • The EU routing and model-id assertions are exercised by the new tests (https://eu.openrouter.ai/api/v1/chat/completions, openai/gpt-6-luna)
Observability coverage: 1 of 1 changes observable
Change Kind Observable Evidence
OpenRouter chat/completions call routed by region outbound HTTP call yes Existing client instrumentation: instrumentedModel wraps every model layer with LanguageModel spans and provider attribution (apps/ai/src/platform/Llm.ts:343-352), so the EU URL swap is visible on the same spans as before

Updated on every push. Resolve a thread or reply "won't fix" to dismiss a finding, or mention @maple to ask about one. Confidence is the reviewer's judgement of merge risk, capped at 2 by a critical finding and 3 by a warning. Quality: 100, minus 25 per critical finding, 10 per warning and 2 per note still open.

@coderabbitai

coderabbitai Bot commented Sep 24, 2026

Copy link
Copy Markdown

Review in Change Stack →

Navigate logical layers of code changes, visualize relationships, and explore their blast radius.

Note

Currently processing new changes in this PR. This may take a few minutes, please wait...

⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Advanced

Run ID: a136537b-88ea-437b-b56f-5274c2ec438b

📥 Commits

Reviewing files that changed from the base of the PR and between 1cbb111 and 76d64e9.

📒 Files selected for processing (10)
  • apps/ai/src/platform/Llm.test.ts
  • apps/ai/src/platform/Llm.ts
  • apps/ai/src/routes/internal/triage.http.ts
  • apps/web/src/components/chat/global-chat-sheet.tsx
  • apps/web/src/components/command-palette/command-palette.tsx
  • apps/web/src/components/dashboard-builder/widgets/widget-actions-context.tsx
  • apps/web/src/components/layout/dashboard-layout.tsx
  • apps/web/src/lib/region.ts
  • apps/web/src/routes/chat.tsx
  • docs/eu-region-plan.md
 ______________________________________________________
< Review complete: I laughed, I cried, I filed issues. >
 ------------------------------------------------------
  \
   \   \
        \ /\
        ( )
      .( o ).
✨ Finishing Touches
📝 Generate docstrings
  • Commit to this branch
  • Create a new PR
🧪 Generate unit tests (beta)
  • Commit to this branch
  • Create a new PR

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@Makisuo
Makisuo merged commit b9c5b73 into main Sep 24, 2026
41 of 42 checks passed
@Makisuo
Makisuo deleted the feat/eu-openrouter-in-region branch September 24, 2026 22:26
Makisuo added a commit that referenced this pull request Sep 24, 2026
… reviews short (#1053)

* fix(pr-review): stop small reviews running out of budget before reading the diff

A 9-file review (#1052) was cut off after 18 model calls: the 800k token
budget counts every re-sent prompt, cache reads included, so at ~45k a
prompt any review got about 18 calls. The close-out then showed the agent
each earlier tool result cut to 4,000 characters, so the web diffs it had
read arrived truncated, and the close-out prompt capped confidence at 3.

- PR_REVIEW_BUDGET: 500 tool calls, 20M tokens; the 10 minute wall clock
  stays the runaway guard.
- reviewCallBudget: min(500, max(40, 10 * files + 20)), was capped at 60.
- review_files children: 100 calls, 4M tokens (was 16 and 250k).
- Close-out keeps up to 60k characters per tool result, above the 50k
  tool output cap, so it sees what the pass saw.

* fix(ai): make every agent budget a runaway guard instead of a pace

The engine's token count includes every re-sent prompt, cache reads
too, so it grows with the square of a run's length and says little
about spend. Sized from p95 traffic it cut normal runs short: a review
at 18 calls, a chat turn at 23. Every budget now clears maxToolCalls
calls at a full live context, so the call cap or the wall clock binds
first, and a test holds all four agents to that.

- Investigation: 200 calls, 25.6M tokens (was 100 and 1.2M). A run cost
  about $0.04 at 79% cache reads.
- PR reply: 200 calls, 25.6M tokens (was 40 and 600k).
- PR review: 64M tokens, 500 calls.
- The review no longer states a call budget in pr_changed_files, and the
  prompt's "Spending your calls" section, which told it to read narrow
  line ranges and never a whole file, becomes "Reading the change": read
  every reviewed diff, batch, and read beyond the diff whenever a
  decision depends on it.
- The close-out caps confidence at 3 only when a reviewed diff went
  unread, not on every early end.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant