Skip to content

fix(providers): stop static snapshots from constraining model discovery #1584

Description

@Astro-Han

Problem

PR #1572 made provider model discovery remote-first, but checked-in fallback snapshots still constrain model availability in two ways.

Live discovery is bounded by the snapshot

Seven providers use protocol discovery with filter: 'fallback-models':

  • xai-oauth
  • zenmux
  • nvidia
  • deepinfra
  • groq
  • openrouter
  • alibaba

For these providers, filterDiscoveredModels intersects the live response with fallbackModels. Refresh can observe that a known model disappeared or is unavailable to the current account, but it can never surface a new model ID. A successful provider response is therefore less authoritative than the bundled recovery snapshot.

Static-only providers have no escape hatch

A provider with modelDiscovery.kind: 'fallback' correctly hides remote refresh when its saved inference credential has no callable model-list endpoint. However, the settings UI also offers no supported way to add an exact model ID manually. Users must wait for a Maka release whenever the provider adds a model that is not yet in the bundled snapshot.

Volcengine Ark direct API and Volcengine Ark Agent Plan currently fall into this category. Agent Plan was checked with a valid dedicated key: both /api/plan/v3/models and /api/plan/v1/models pass gateway authentication and then return 404. A permanently failing refresh button would not restore the missing capability.

The combined result is a capability regression: users can have valid access to a new model while Maka prevents them from selecting it.

Refs #1571 and #1572.

Desired outcome

Make fallback inventory, live discovery, capability filtering, and explicit user choices independent concerns.

  • Treat fallbackModels as bootstrap, offline recovery, recommendation, and display metadata—not as an allowlist over a successful live catalog.
  • Audit every filter: 'fallback-models' provider and replace snapshot membership checks with provider-backed language/tool compatibility filtering where the response exposes enough information.
  • Preserve a returned model with explicit unknown capabilities when the provider cannot describe them, rather than dropping it solely because its ID is new to Maka.
  • Continue excluding models that are explicitly known to be incompatible with Maka's chat runtime.
  • Give verified static-only providers a manual exact-model-ID path. Persist those entries as explicit user choices, not as provider-fetched facts.
  • Preserve last-known-good catalogs on discovery failure and reject empty or malformed responses.
  • Keep model IDs exact through discovery, persistence, selection, and inference.
  • Add contracts proving that:
    • a successful live response can introduce a compatible ID absent from the bundled snapshot;
    • an explicit user model survives restart and catalog refresh;
    • provenance still distinguishes fallback, provider-fetched, and user-entered models.

Non-goals

  • Do not expose refresh actions that are known to call nonexistent endpoints.
  • Do not remove filtering in a way that presents embedding, image-generation, audio, or otherwise incompatible models as chat choices.
  • Do not add Volcengine control-plane AK/SK or account-login credentials as part of this issue.
  • Do not treat a manually entered model as proof that the provider advertised or validated it.

Alternatives considered

  • Updating bundled model arrays more frequently only shortens the staleness window.
  • Blindly trusting every /models entry can expose unsupported modalities.
  • Showing a refresh button for a verified 404 endpoint creates noise without restoring model selection.
  • Persisting user-entered IDs as fetched models corrupts provenance and makes reconciliation unsafe.
中文对照

问题

PR #1572 已经把大多数 provider 改成远端优先发现,但静态模型快照仍在两个地方限制实际能力。

拉取成功也发现不了新模型

以下七个 provider 使用 filter: 'fallback-models':

  • xai-oauth
  • zenmux
  • nvidia
  • deepinfra
  • groq
  • openrouter
  • alibaba

它们会请求真实的模型列表,但 filterDiscoveredModels 随后只保留 fallbackModels 中已有的 ID。刷新可以确认老模型是否下线、当前账号是否仍有权限,却无法发现任何 Maka 尚未收录的新模型。

这使静态恢复快照反过来成为远端目录的上限:即使 provider 已经明确返回一个新模型,Maka 仍会把它丢掉。

无拉取接口时,用户也无法手工补充

对于 modelDiscovery.kind: 'fallback' 的 provider,如果推理 Key 确实没有可调用的模型列表接口,隐藏刷新按钮是合理的;但当前设置页也没有手工录入精确模型 ID 的入口。上游新增模型后,用户只能等待 Maka 发布新的内置快照。

目前火山方舟直连和 Agent Plan 属于这一类。Agent Plan 已用有效专属 Key 验证:/api/plan/v3/models 与 /api/plan/v1/models 都通过网关鉴权后返回 404。为这种端点保留一个必然失败的刷新按钮,并不能补回缺失的模型选择能力。

最终表现是:账号和 Key 明明已经可以调用新模型,Maka 却因为本地静态目录尚未更新而不允许用户选择它。这属于能力退化。

关联 #1571、#1572。

目标

把内置快照、远端发现、能力过滤和用户显式选择拆成互不替代的事实来源。

  • fallbackModels 只负责首次配置、离线恢复、推荐顺序和展示元数据,不再充当成功远端目录的 allowlist。
  • 逐一审计七个 filter: 'fallback-models' provider;如果接口提供模型类型或工具能力信息,就按 provider 事实过滤,而不是按静态 ID 过滤。
  • 当接口只返回模型 ID、无法说明能力时,保留该模型并明确标记能力未知;不能仅因 Maka 尚未收录这个 ID 就静默丢弃。
  • 对已经明确属于非对话、非 Maka 运行协议的模型继续过滤,避免把向量、图片、音频等模型放进聊天选择器。
  • 为确认没有模型列表接口的 provider 提供手工输入精确模型 ID 的能力;这类记录必须标记为用户选择,不能伪装成 provider 拉取结果。
  • 拉取失败时继续保留最后一次成功目录;空响应或畸形响应不得覆盖可用缓存。
  • 模型 ID 在拉取、持久化、选择和推理过程中保持原样。
  • 增加合约测试,确保:
    • 远端成功返回的新模型即使不在内置快照中,也能进入目录;
    • 用户手工添加的模型在重启和后续刷新后仍然存在;
    • fallback、provider 拉取和用户录入三种来源不会混淆。

不在本任务范围内

  • 不为已经确认返回 404 的端点展示刷新按钮。
  • 不通过取消全部过滤来放入向量、图片、音频或其他不兼容模型。
  • 不在本任务中引入火山方舟控制面 AK/SK 或账号登录凭据。
  • 用户手工输入模型 ID 不代表 provider 已公开、验证或承诺支持该模型。

已考虑的替代方案

  • 更频繁地更新内置列表只能缩短过期窗口,不能消除结构性限制。
  • 无条件接受 /models 的全部结果可能把不兼容模态暴露给聊天运行时。
  • 对已确认 404 的接口显示刷新按钮,只会产生重复错误。
  • 把用户录入的模型保存为“远端拉取”会破坏来源信息,并使后续目录合并不可靠。

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

bugSomething isn't working

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions