Skip to content

Latest commit

 

History

1 Commit

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 

Repository files navigation

CoEvoWhen: Policy-Tool Coevolution for Ultra-Long Video Temporal Grounding

Yiduo Jia · Muzhi Zhu · Jinchuan Shi · Hao Zhong · Yuling Xi · Ke Liu · Hao Chen†

🎓 Zhejiang University, State Key Lab of CAD & CG
† Corresponding author

Paper: Arxiv Link Project: Page


Abstract · Method · Results · Ablations & Analysis · Case Study · Citation

📖 Abstract

Ultra-long video temporal grounding requires balancing long-range evidence search with fine-grained event understanding under a limited visual budget, yet existing agentic methods still rely largely on predefined policies and tool capabilities. Motivated by this, we propose a novel policy–tool coevolution framework that jointly evolves high-level policies and executable media tools from the agentic reasoning trajectories of a VLM, forming a reusable skill without updating model parameters.

During evolution, an external skill updater distills transferable task experience in long-video temporal grounding, accordingly refining the orchestration of long-range image-based and fine-grained video-based observations. Alongside these policy updates, the updater employs its coding capabilities to upgrade existing tools or create new ones, adapting the tools to long-video evidence acquisition.

Equipped with the evolved skill, the VLM autonomously orchestrates tools under the guidance of the evolved policy, coordinating image and video observations for agentic inference without relying on a separate, stronger planning model.

Extensive experiments spanning five benchmarks and three VLMs show that policy–tool coevolution consistently improves temporal grounding accuracy in ultra-long videos while reducing visual token cost at inference, and that the evolved skill yields substantial performance gains on general long-video QA without additional task-specific evolution, demonstrating the effectiveness and generalizability of our framework for long-video understanding.

✨ Highlights

  • Policy–Tool Coevolution. We jointly evolve high-level policies and executable media tools from a VLM's execution trajectories, distilling ultra-long video grounding experience into a reusable external skill without updating model parameters.
  • 🎞️ Coordinated Image–Video Observation. We evolve the orchestration of complementary image-based and video-based observations, enabling the VLM to acquire evidence autonomously without manually predefined coordination strategies or a stronger external planner.
  • 📈 Accuracy Gains and Cost Reduction. Experiments spanning five benchmarks and three VLMs show that policy–tool coevolution consistently improves temporal grounding accuracy in ultra-long videos while reducing visual token cost at inference. The evolved skill also yields substantial performance gains on general long-video QA without additional task-specific evolution.

🚀 CoEvoWhen Framework

A policy–tool coevolution framework that distills task experience into a reusable external skill for ultra-long video temporal grounding.

Policy–Tool Coevolution

A frozen VLM executes ultra-long video temporal grounding tasks with an external skill comprising high-level policies and executable media tools. Across evolution rounds, an external skill updater distills transferable experience from execution trajectories and task feedback, refining policies for task planning and observation orchestration while synthesizing code to modify existing media tools or create new ones.

Policy-Guided Agentic Inference

Image-based observations can provide compact coverage of extended temporal ranges and facilitate comparisons across distant candidate regions, whereas video-based observations preserve local temporal continuity for reasoning about motion, event order, and temporal boundaries.

At inference, the VLM autonomously orchestrates media tools under the guidance of the evolved policy, using accumulated evidence to decide on subsequent observations without relying on a separate, stronger planning model.

📊 Quantitative Results

Ultra-Long Video Temporal Grounding

Coevolution improves grounding accuracy while simultaneously reducing visual token cost. The results below compare the base and evolved skills on Qwen3.5-27B, with relative gains and cost reductions measured against the base skill.

Benchmark Metric Base Skill Evolved Skill Relative Gain Cost Reduction
VUE-LVTR IoU AUC 0.4107 0.5137 ↑ 25.1% ↓ 30.0%
ExtremeWhenBench mIoU 0.1507 0.2636 ↑ 74.9% ↓ 11.4%
CoMET-Bench mIoU 0.1355 0.1635 ↑ 20.7% ↓ 18.9%

Cross-VLM Generalization and Cross-Task Transfer

Consistent gains are also observed when skills are evolved separately on different VLMs. Directly applying the evolved skill to general long-video QA also improves performance without additional task-specific evolution.

Generalization across VLMs on temporal grounding. Bold: results with the evolved skill. Evolution Gain reports absolute changes, with ↑ indicating relative improvements.

VLMMethodVUE-LVTRExtremeWhenBenchCoMET-Bench
Precision
AUC ↑
Recall
AUC ↑
IoU
AUC ↑
mIoU ↑Recall@0.5 ↑mIoU ↑Rejection-F1 ↑
Qwen3.5-9B+ Base Skill0.32400.34250.24050.06260.05760.067267.51
+ Evolved Skill0.42870.47070.32440.12470.11700.085268.80
Evolution Gain+0.1047
↑ 32.3%
+0.1282
↑ 37.4%
+0.0839
↑ 34.9%
+0.0621
↑ 99.2%
+0.0594
↑ 103.1%
+0.0180
↑ 26.8%
+1.29
↑ 1.9%
Qwen3.5-27B+ Base Skill0.47060.47890.41070.15070.14780.135566.13
+ Evolved Skill0.59180.57730.51370.26360.27850.163577.06
Evolution Gain+0.1212
↑ 25.8%
+0.0984
↑ 20.5%
+0.1030
↑ 25.1%
+0.1129
↑ 74.9%
+0.1307
↑ 88.4%
+0.0280
↑ 20.7%
+10.93
↑ 16.5%
Qwen3.6-27B+ Base Skill0.50980.54650.45060.16910.16940.131268.53
+ Evolved Skill0.55190.57270.49850.21090.21290.147772.37
Evolution Gain+0.0421
↑ 8.3%
+0.0262
↑ 4.8%
+0.0479
↑ 10.6%
+0.0418
↑ 24.7%
+0.0435
↑ 25.7%
+0.0165
↑ 12.6%
+3.84
↑ 5.6%

Direct transfer to long-video QA. Overall accuracy is reported in percent. Evolution Gain reports improvements in percentage points, with ↑ indicating relative improvements.

VLMMethodLVBenchLSDBench
Overall Acc. (%) ↑Overall Acc. (%) ↑
Qwen3.5-9B+ Base Skill40.0949.08
+ Evolved Skill45.8453.68
Evolution Gain+5.75
↑ 14.3%
+4.60
↑ 9.4%
Qwen3.5-27B+ Base Skill45.4561.27
+ Evolved Skill54.1069.25
Evolution Gain+8.65
↑ 19.0%
+7.98
↑ 13.0%
Qwen3.6-27B+ Base Skill48.5561.12
+ Evolved Skill54.3665.64
Evolution Gain+5.81
↑ 12.0%
+4.52
↑ 7.4%

💡 Ablations & Analysis

Ablation on Policy–Tool Coevolution

Joint evolution achieves the best accuracy–cost combination. Policy-only and tool-only variants both improve grounding accuracy over the base skill, confirming policies and tools as effective evolution targets.

On the VUE-LVTR held-out set with Qwen3.5-27B, joint evolution achieves the following relative improvements over the policy-only and tool-only variants:

Comparison IoU AUC Improvement Visual Token Reduction
vs. Policy-only ↑ 5.0% ↓ 40.6%
vs. Tool-only ↑ 16.7% ↓ 25.5%

Ablation on Image–Video Coordination

Both single-modality variants improve grounding performance through evolution. The evolved image+video skill outperforms both single-modality skills across all grounding metrics while using fewer visual tokens, underscoring the benefits of strategically coordinating their complementary strengths.

Evolved skills on the VUE-LVTR held-out set with Qwen3.5-27B:

Observation Mode IoU AUC ↑ Visual Tokens (k/query) ↓
Image only 0.4374 194.60
Video only 0.4772 221.98
Image + video 0.5137 141.89

Analysis of Skill Evolution

Online performance–cost dynamics during skill evolution.   Execution feedback drives coordinated updates to high-level policies and executable media tools.

As evolution proceeds, the skill achieves higher grounding accuracy with lower visual token cost more consistently. Candidate omissions, boundary errors, and dynamic ambiguities exposed in trajectories drive coordinated adjustments to observation capabilities and their orchestration.

🎬 Case Study

We analyze the execution trajectories of the base and evolved skills on two ExtremeWhenBench queries to illustrate how policy–tool coevolution changes evidence acquisition for action and scene localization in ultra-long videos.

Candidate Localization and Motion Verification Long-Range Search and Boundary Refinement
Scooping fries with a metal skimmer A couple walking along a tree-lined path in golden light
Video duration 50.6 minutes 100.1 minutes
Temporal grounding IoU 0.00 → 1.00 0.00 → 1.00
Visual token cost 198.68k → 96.81k 237.60k → 128.82k
Cost reduction ↓ 51.27% ↓ 45.78%
Recorded trajectories View the trajectories → View the trajectories →

The first case illustrates how better candidate localization allows video observation to focus on verifying query-specific motion, reducing repeated inspection of an incorrect region. The second highlights the importance of discovering a brief target during global search before investing in local boundary refinement.

📜 Citation

If you find CoEvoWhen useful for your research, please cite:

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages