Skip to content

fix(onnx): stop writing a trailing newline into cache refs (#1803) - #1805

Open
fengyue-xve wants to merge 1 commit into
TencentCloud:developfrom
fengyue-xve:fix/onnx-hf-cache-ref
Open

fengyue-xve wants to merge 1 commit into
TencentCloud:developfrom
fengyue-xve:fix/onnx-hf-cache-ref

Conversation

@fengyue-xve

Copy link
Copy Markdown

Summary

修复 #1803:本地 ONNX 嵌入模型(如 jinaai/jina-embeddings-v2-base-zh)下载成功后,文档建立索引仍报「Could not load model jinaai/jina-embeddings-v2-base-zh from any source.」

根因:COS 下载器 hf_cache_snapshot_dir() 写 refs/main 时带了尾部换行("cos-mirror\n")。huggingface_hub 解析 ref 时原样读取文件内容(不做 strip()),于是本地查找的快照目录变成 snapshots/"cos-mirror\n" —— 该目录并不存在,本地加载直接失败;之后 fastembed 回退到联网下载(NAS / 内网环境访问不了 huggingface.co),重试耗尽后抛出该错误。

对照实验(本机 huggingface_hub 2.1.1,同一缓存目录、同一 commit 结构):

refs/main 内容 snapshot_download(local_files_only=True)
"cos-mirror\n" 抛 LocalEntryNotFoundError(完整复现报错链路)
"cos-mirror" 正常返回 snapshots/cos-mirror

修复(两点,缺一不可):

  1. 写入端:hf_cache_snapshot_dir() 去掉换行 —— 新下载的缓存不再写坏 ref。
  2. 加载端自愈:新增 repair_hf_cache_refs(),在 _build_text_embedding() 构建 fastembed 实例前,把旧版本写坏的 ref 原地修正 —— 已受影响的用户(包括 issue 作者)升级后无需重新下载即可恢复。

为什么必须加自愈:坏缓存里 ONNX 权重文件是存在的,is_model_downloaded() 判定为「已下载」,用户在界面重新点下载会直接跳过 —— 只修写入端触达不到存量缓存。

Fixes #1803

Target branch

  • Base is develop (feature / fix — default)
  • Base is main (release/* or hotfix/* only)

Type of change

  • Bug fix
  • New feature
  • Breaking change
  • Documentation
  • Refactor / chore
  • Release / hotfix

Test plan

新增 / 更新的测试(tests/unit/agents/):

  • test_cos_download_writes_hf_cache:断言收紧为精确匹配 refs/main == "cos-mirror"(原断言里的 .strip() 恰好掩盖了本 bug);
  • test_repair_hf_cache_refs_strips_trailing_newline、test_repair_hf_cache_refs_leaves_healthy_cache_untouched:自愈函数红绿验证;
  • test_build_text_embedding_repairs_stale_cache_refs:加载网关在把缓存交给 fastembed 之前完成修复(断言写在 fake TextEmbedding.__init__ 里,固定调用顺序)。

双向验证:以上用例在修复前代码上全部失败(红),修复后全部通过(绿)。

本地门禁(Windows):

  • ruff check src tests / ruff format --check src tests / mypy src/octop:通过;

  • pytest tests/unit/agents(含本次改动,串行):687 passed, 5 skipped;

  • 全量 pytest -n auto -m "not live":本机 Windows 存在既有环境问题 —— xdist 多进程并行时随机有 worker 进程硬崩、INTERNALERROR 中断整次运行(崩溃时正在执行的测试每次不同;在未改动的干净基线上同样复现,且发生在与本 PR 毫不相干的 test_connectors.py 上),崩溃还会吞掉失败明细,因此以 --junitxml 逐个取出失败用例;

  • 对本机并行运行中出现的全部失败用例逐条串行复跑(PYTHONUTF8=1,同时消除本机 GBK 默认编码造成的 fixture 读取伪失败):除 1 例依赖外网可达性的用例(test_probe_tencent_mirror_unreachable,本机网络能连通该镜像,断言前提不成立)外全部通过,且干净基线上结果一致。

  • make all passes locally(上述并行崩溃与网络用例均为本机既有环境问题,已在干净基线上对照复现;本 PR 涉及的 tests/unit/agents 全量绿灯)

  • Added/updated tests

Checklist

  • Updated CHANGELOG.md (if user-facing)
  • README / docs updated (if needed)

huggingface_hub resolves refs/main verbatim, so the newline written by
hf_cache_snapshot_dir made local loads look for a snapshot directory
literally named "cos-mirror\n". That directory never exists, so loads
failed and fastembed fell back to the network, ending in "Could not
load model ... from any source." on machines that cannot reach
huggingface.co.

Write the ref without the newline and repair refs left by older
versions when loading, so affected caches recover without a re-download.
@fengyue-xve
fengyue-xve force-pushed the fix/onnx-hf-cache-ref branch from 312f398 to 6a6ca8c Compare October 8, 2026 13:05
@fengyue-xve

Copy link
Copy Markdown
Author

Both jobs are red on the unrelated flaky test_cold_target_install_budget_preserves_regular_timeout (float wobble in the 900s budget assertion); fix proposed in #1826.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant