Skip to content

[fix] Align DeepSeek-V4.1 prompt and tool-call encoding - #10236

Merged
tastelikefeet merged 5 commits into
modelscope:mainfrom
taking-lying-flat:fix/deepseek-v41-agent-protocol
Sep 25, 2026
Merged

tastelikefeet merged 5 commits into
modelscope:mainfrom
taking-lying-flat:fix/deepseek-v41-agent-protocol

Conversation

@taking-lying-flat

@taking-lying-flat taking-lying-flat commented Sep 23, 2026 •

Copy link
Copy Markdown
Contributor

PR type

  • Bug Fix
  • New Feature
  • Document Updates
  • More Models or Datasets Support

PR information

DeepSeek-V4.1 currently selects the V4 agent, so training formats tool calls with V4 tags and inference fails to parse official V4.1 calls. Its inherited reasoning settings also emit V4 effort prompts instead of the V4.1 numeric budget.

Align the DeepSeek templates with the official V4.1 encoder:

  • Register a V4.1 agent using the spaced calls, invoke, and parameter tags consistently for tool instructions, history formatting and parsing. V4 keeps its existing tags and output.
  • Handle reasoning budgets 1–100 and low/high/max aliases (50/75/100), with a default of 75 in thinking mode. Explicitly disabling thinking remains effective when an effort is supplied.
  • Preserve V4.1 system delimiters and assistant channel markers, including historical tool turns. Fold mid-conversation system queries locally while retaining training loss masks.

The final diff contains three production files: the DeepSeek agent, its registry entry, and the DeepSeek model template. Shared template bases, input conversion, dataset processing, inference engine and response dataclasses are identical to the base revision. Native provider-field conversion, general mixed-role/content-block compatibility and separate tool-namespace response fields are outside this PR's scope.

Experiment results

Using real tokenizer/config snapshots and official encoder revision dba1be0a40aa45a94ad051997016db3960a90277:

  • 147 passed: V4/V4.1 single and parallel tool-call round-trips, string/JSON argument types, official parser comparisons, history/training serialization, image preprocessing, reasoning settings, system prompts and loss masks. Reasoning-history fixtures use Swift's canonical <think>...</think> representation.
  • 12 passed: the existing repository V4 Flash and V4.1 tests, with model IDs redirected to local tokenizer/config snapshots and assertions unchanged.
  • 34 passed, 48 subtests passed: existing shared tool/schema and streaming-client tests, plus local HTTP and SSE tool-call round-trips through InferClient. Model output is simulated for transport tests.
  • 323 supported comparisons exactly match upstream, covering 34 other agent templates and 22 model templates. The 80 additional attempted combinations already fail on the baseline; there are no new success/error transitions.
  • flake8, isort, yapf and git diff --check passed. The diff was checked to contain only the three files above.

Tests and audit artifacts are kept outside the production-only PR. Live model generation and GPU training were not run.

@tastelikefeet

Copy link
Copy Markdown
Collaborator

Hi, thanks for pointing out the problem! Can you help to add some tests to keep it right? for example, agent template/reasoning efforts and history thinking/multi-turn chats/tool call parsing

Comment thread swift/template/templates/deepseek.py Outdated
messages = []
for is_assistant, group in groupby(inputs.messages, key=lambda m: m['role'] == 'assistant'):
queries = list(group)
if is_assistant or not any(m['role'] == 'system' for m in queries):

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Why not just find the first assistant? groupby maybe too heavy.

@tastelikefeet
tastelikefeet merged commit 9d56934 into modelscope:main Sep 25, 2026
3 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants