Yiduo Jia · Muzhi Zhu · Jinchuan Shi · Hao Zhong · Yuling Xi · Ke Liu · Hao Chen†
🎓 Zhejiang University, State Key Lab of CAD & CG
† Corresponding author
Abstract · Method · Results · Ablations & Analysis · Case Study · Citation
Ultra-long video temporal grounding requires balancing long-range evidence search with fine-grained event understanding under a limited visual budget, yet existing agentic methods still rely largely on predefined policies and tool capabilities. Motivated by this, we propose a novel policy–tool coevolution framework that jointly evolves high-level policies and executable media tools from the agentic reasoning trajectories of a VLM, forming a reusable skill without updating model parameters.
During evolution, an external skill updater distills transferable task experience in long-video temporal grounding, accordingly refining the orchestration of long-range image-based and fine-grained video-based observations. Alongside these policy updates, the updater employs its coding capabilities to upgrade existing tools or create new ones, adapting the tools to long-video evidence acquisition.
Equipped with the evolved skill, the VLM autonomously orchestrates tools under the guidance of the evolved policy, coordinating image and video observations for agentic inference without relying on a separate, stronger planning model.
Extensive experiments spanning five benchmarks and three VLMs show that policy–tool coevolution consistently improves temporal grounding accuracy in ultra-long videos while reducing visual token cost at inference, and that the evolved skill yields substantial performance gains on general long-video QA without additional task-specific evolution, demonstrating the effectiveness and generalizability of our framework for long-video understanding.
Policy–Tool Coevolution. We jointly evolve high-level policies and executable media tools from a VLM's execution trajectories, distilling ultra-long video grounding experience into a reusable external skill without updating model parameters.
- 🎞️ Coordinated Image–Video Observation. We evolve the orchestration of complementary image-based and video-based observations, enabling the VLM to acquire evidence autonomously without manually predefined coordination strategies or a stronger external planner.
- 📈 Accuracy Gains and Cost Reduction. Experiments spanning five benchmarks and three VLMs show that policy–tool coevolution consistently improves temporal grounding accuracy in ultra-long videos while reducing visual token cost at inference. The evolved skill also yields substantial performance gains on general long-video QA without additional task-specific evolution.
A policy–tool coevolution framework that distills task experience into a reusable external skill for ultra-long video temporal grounding.
A frozen VLM executes ultra-long video temporal grounding tasks with an external skill comprising high-level policies and executable media tools. Across evolution rounds, an external skill updater distills transferable experience from execution trajectories and task feedback, refining policies for task planning and observation orchestration while synthesizing code to modify existing media tools or create new ones.
Image-based observations can provide compact coverage of extended temporal ranges and facilitate comparisons across distant candidate regions, whereas video-based observations preserve local temporal continuity for reasoning about motion, event order, and temporal boundaries.
At inference, the VLM autonomously orchestrates media tools under the guidance of the evolved policy, using accumulated evidence to decide on subsequent observations without relying on a separate, stronger planning model.
Coevolution improves grounding accuracy while simultaneously reducing visual token cost. The results below compare the base and evolved skills on Qwen3.5-27B, with relative gains and cost reductions measured against the base skill.
| Benchmark | Metric | Base Skill | Evolved Skill | Relative Gain | Cost Reduction |
|---|---|---|---|---|---|
| VUE-LVTR | IoU AUC | 0.4107 | 0.5137 | ↑ 25.1% | ↓ 30.0% |
| ExtremeWhenBench | mIoU | 0.1507 | 0.2636 | ↑ 74.9% | ↓ 11.4% |
| CoMET-Bench | mIoU | 0.1355 | 0.1635 | ↑ 20.7% | ↓ 18.9% |
Consistent gains are also observed when skills are evolved separately on different VLMs. Directly applying the evolved skill to general long-video QA also improves performance without additional task-specific evolution.
Generalization across VLMs on temporal grounding. Bold: results with the evolved skill. Evolution Gain reports absolute changes, with ↑ indicating relative improvements.
| VLM | Method | VUE-LVTR | ExtremeWhenBench | CoMET-Bench | ||||
|---|---|---|---|---|---|---|---|---|
| Precision AUC ↑ | Recall AUC ↑ | IoU AUC ↑ | mIoU ↑ | Recall@0.5 ↑ | mIoU ↑ | Rejection-F1 ↑ | ||
| Qwen3.5-9B | + Base Skill | 0.3240 | 0.3425 | 0.2405 | 0.0626 | 0.0576 | 0.0672 | 67.51 |
| + Evolved Skill | 0.4287 | 0.4707 | 0.3244 | 0.1247 | 0.1170 | 0.0852 | 68.80 | |
| Evolution Gain | +0.1047 ↑ 32.3% | +0.1282 ↑ 37.4% | +0.0839 ↑ 34.9% | +0.0621 ↑ 99.2% | +0.0594 ↑ 103.1% | +0.0180 ↑ 26.8% | +1.29 ↑ 1.9% | |
| Qwen3.5-27B | + Base Skill | 0.4706 | 0.4789 | 0.4107 | 0.1507 | 0.1478 | 0.1355 | 66.13 |
| + Evolved Skill | 0.5918 | 0.5773 | 0.5137 | 0.2636 | 0.2785 | 0.1635 | 77.06 | |
| Evolution Gain | +0.1212 ↑ 25.8% | +0.0984 ↑ 20.5% | +0.1030 ↑ 25.1% | +0.1129 ↑ 74.9% | +0.1307 ↑ 88.4% | +0.0280 ↑ 20.7% | +10.93 ↑ 16.5% | |
| Qwen3.6-27B | + Base Skill | 0.5098 | 0.5465 | 0.4506 | 0.1691 | 0.1694 | 0.1312 | 68.53 |
| + Evolved Skill | 0.5519 | 0.5727 | 0.4985 | 0.2109 | 0.2129 | 0.1477 | 72.37 | |
| Evolution Gain | +0.0421 ↑ 8.3% | +0.0262 ↑ 4.8% | +0.0479 ↑ 10.6% | +0.0418 ↑ 24.7% | +0.0435 ↑ 25.7% | +0.0165 ↑ 12.6% | +3.84 ↑ 5.6% | |
Direct transfer to long-video QA. Overall accuracy is reported in percent. Evolution Gain reports improvements in percentage points, with ↑ indicating relative improvements.
| VLM | Method | LVBench | LSDBench |
|---|---|---|---|
| Overall Acc. (%) ↑ | Overall Acc. (%) ↑ | ||
| Qwen3.5-9B | + Base Skill | 40.09 | 49.08 |
| + Evolved Skill | 45.84 | 53.68 | |
| Evolution Gain | +5.75 ↑ 14.3% | +4.60 ↑ 9.4% | |
| Qwen3.5-27B | + Base Skill | 45.45 | 61.27 |
| + Evolved Skill | 54.10 | 69.25 | |
| Evolution Gain | +8.65 ↑ 19.0% | +7.98 ↑ 13.0% | |
| Qwen3.6-27B | + Base Skill | 48.55 | 61.12 |
| + Evolved Skill | 54.36 | 65.64 | |
| Evolution Gain | +5.81 ↑ 12.0% | +4.52 ↑ 7.4% |
Joint evolution achieves the best accuracy–cost combination. Policy-only and tool-only variants both improve grounding accuracy over the base skill, confirming policies and tools as effective evolution targets.
On the VUE-LVTR held-out set with Qwen3.5-27B, joint evolution achieves the following relative improvements over the policy-only and tool-only variants:
| Comparison | IoU AUC Improvement | Visual Token Reduction |
|---|---|---|
| vs. Policy-only | ↑ 5.0% | ↓ 40.6% |
| vs. Tool-only | ↑ 16.7% | ↓ 25.5% |
Both single-modality variants improve grounding performance through evolution. The evolved image+video skill outperforms both single-modality skills across all grounding metrics while using fewer visual tokens, underscoring the benefits of strategically coordinating their complementary strengths.
Evolved skills on the VUE-LVTR held-out set with Qwen3.5-27B:
| Observation Mode | IoU AUC ↑ | Visual Tokens (k/query) ↓ |
|---|---|---|
| Image only | 0.4374 | 194.60 |
| Video only | 0.4772 | 221.98 |
| Image + video | 0.5137 | 141.89 |
As evolution proceeds, the skill achieves higher grounding accuracy with lower visual token cost more consistently. Candidate omissions, boundary errors, and dynamic ambiguities exposed in trajectories drive coordinated adjustments to observation capabilities and their orchestration.
We analyze the execution trajectories of the base and evolved skills on two ExtremeWhenBench queries to illustrate how policy–tool coevolution changes evidence acquisition for action and scene localization in ultra-long videos.
| Candidate Localization and Motion Verification | Long-Range Search and Boundary Refinement | |
|---|---|---|
| Video duration | 50.6 minutes | 100.1 minutes |
| Temporal grounding IoU | 0.00 → 1.00 | 0.00 → 1.00 |
| Visual token cost | 198.68k → 96.81k | 237.60k → 128.82k |
| Cost reduction | ↓ 51.27% | ↓ 45.78% |
| Recorded trajectories | View the trajectories → | View the trajectories → |
The first case illustrates how better candidate localization allows video observation to focus on verifying query-specific motion, reducing repeated inspection of an incorrect region. The second highlights the importance of discovering a brief target during global search before investing in local boundary refinement.
If you find CoEvoWhen useful for your research, please cite: