Skip to content

fix(agent-runtime): send adaptive thinking to effort-only Claude models - #949

Merged
vastsa merged 2 commits into
vastsa:mainfrom
onerentop:fix/anthropic-adaptive-thinking-compat
Sep 23, 2026
Merged

vastsa merged 2 commits into
vastsa:mainfrom
onerentop:fix/anthropic-adaptive-thinking-compat

Conversation

@onerentop

Copy link
Copy Markdown
Contributor

问题

使用 Claude Opus 5.5(claude-opus-5-5)并开启思考时,请求被拒绝:

400 {"error":{"message":"claude-opus-5-5 requires adaptive thinking; omit thinking or use thinking.type=adaptive and output_config.effort","type":"invalid_request_error"},"type":"error"}

根因

pi-ai 的 Anthropic 适配器只有在 model.compat.forceAdaptiveThinking === true 时才发送 thinking: {type:"adaptive"} + output_config.effort,否则回落到 thinking: {type:"enabled", budget_tokens}。PI-Desktop 的模型配置只来自 models.dev(ADR 0134),modelConfigFromModelsDev 不产生任何 compat,因此所有 Anthropic Messages 行在开启思考时都走 budget 路径。Opus 4.7+ / Opus 5.x / Fable 等只支持 adaptive 的模型会直接 400。

改动

buildProviderModel(packages/agent-runtime/src/provider-binding.ts):wire API 为 anthropic-messages 时,从 models.dev 自己发布的 reasoning_options 推导标志——有 effort 选项且没有 budget_tokens 选项 ⇒ compat.forceAdaptiveThinking: true。

  • 不读取 pi-ai 内置模型目录,models.dev 仍是唯一元数据来源(符合 ADR 0134,无需新例外)。
  • 仍发布 budget_tokens 的模型(Opus 4.5/4.6、Sonnet 4.x、Haiku 4.5)请求形状不变。
  • 显式的目录 compat 记录被保留(合并而非覆盖)。
  • xhigh/max:models.dev 的 effort 值已经经 thinkingLevelMapFromModelsDev 映射到 thinkingLevelMap(例如 xhigh: "xhigh"、max: "max"),开启该标志后 pi-ai 的 mapThinkingLevelToEffort 会直接使用,不会回落为 high。
  • 由于 models.dev 的别名匹配已在目录层完成(reasoning_options 随匹配到的记录一起下发),-thinking / -agent / 命名空间等变体无需额外匹配逻辑。

与 #908 的关系:#908 从 pi-ai 内置目录采纳 adaptive 元数据,review 中指出与 ADR 0134 冲突,并建议"将 adaptive 映射纳入 models.dev-owned 配置"。本 PR 就是这一方向的实现,范围更小(+21 行实现)。

Spec:docs/spec/03-runtime/11-provider-model-system.md 及 zh-CN 镜像补充了该规则。

测试

provider-binding.test.ts 新增 3 条:

  1. effort-only 的 claude-opus-5-5 → 请求体 thinking.type === "adaptive"、无 budget_tokens、output_config: {effort:"medium"}(去掉修复后该用例失败,实际为 "enabled");
  2. 仅发布 budget_tokens 的模型 → 仍是 thinking.type === "enabled" + budget_tokens;
  3. 显式 compat.forceAdaptiveThinking 被保留。

验证

  • vitest run src/provider-binding.test.ts:27/27 通过
  • agent-runtime 全量:1005/1006;native-pi-session 的 fork 用例与 hosted-search-compaction(找不到 pi-coding-agent dist)在未改动基线上同样失败,Windows 本机环境既有问题
  • tsc --noEmit(agent-runtime)、biome lint、check-architecture、docs/check-locales(80 对)、docs/check-docs:通过
  • check-pr-base-main:通过

真实环境验证:

Task candidate: v0.15.6 (d2c9b6b1e) + 本修复,重新打包 sidecar.js
Base main:      ea5890b94(本分支基线);v0.15.6 源码重打包与已安装 sidecar.js 逐字节一致
E2E suites:     NOT RUN — 自动化 E2E 需要付费 Anthropic API;以手动真实请求替代
Result:         已安装的 PI-Desktop(Windows)替换 sidecar 后,Opus 5.5 + 思考开启正常回复;修复前同配置稳定 400
Environment:    Windows 11,PI-Desktop 0.15.6 安装包,Anthropic provider(用户已配置)

剩余风险:未对第三方 Anthropic 兼容网关做真实请求;若网关对 effort-only 模型仍只接受 budget 思考,需要该网关在 models.dev 中发布 budget_tokens。

🤖 Generated with Claude Code

Opus 5.5 (and other Claude models that only publish an effort ladder,
such as Opus 4.7+, Opus 5 and Fable) reject budget thinking with
"requires adaptive thinking; omit thinking or use thinking.type=adaptive".

pi-ai only emits adaptive thinking when compat.forceAdaptiveThinking is
set. PI-Desktop builds model configs from models.dev, which carries no
pi-ai compat record, so every Anthropic Messages request with thinking
enabled fell back to thinking.type=enabled with budget_tokens.

Derive the flag in buildProviderModel from the models.dev reasoning
options: an effort option without a budget_tokens option means the
model only accepts adaptive thinking. Models that still publish
budget_tokens keep budget thinking, and explicit catalog compat wins.
Keep the zh-CN provider model system spec aligned with the English
source so the locale pair check stays green for the Anthropic adaptive
thinking rule.
@vastsa
vastsa merged commit cb8a87d into vastsa:main Sep 23, 2026
vastsa pushed a commit that referenced this pull request Sep 25, 2026
Since 41d15d5 a custom Anthropic Messages gateway (vendorKey custom,
unknown base URL) no longer resolves a models.dev record for a Claude
id that several publishers list, because the lookup refuses to pick one
of them without a provider identity. The row then falls back to the
generic model shape with no reasoning options, so the adaptive-thinking
derivation from #949 cannot fire and Opus 5.5 turns fail again with
"requires adaptive thinking" (400).

Which thinking shape a Claude id accepts is a property of the model,
stated by Anthropic's own record, not of the deployment. When the
lookup misses on an Anthropic Messages row, take only reasoning_options
and the derived thinkingLevelMap from Anthropic's record for exactly
the same id. Limits and modalities stay generic, so the ambiguity rule
from 41d15d5 still holds; aliases, renamed ids, other wire APIs and
non-Claude ids served over the Anthropic protocol are unchanged.

catalogModelConfigFor centralizes the repeated "catalog record or
generic shape" choice for the session and subagent launch paths.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants