Skip to content

Repository files navigation

VBVR-Pro: A Scalable and Verifiable Suite for Native Visual Reasoning

Overview

Native visual reasoning, i.e., reasoning through visual generation, has recently emerged as a promising direction for studying visual intelligence beyond language. Yet progress remains bottlenecked by the lack of scalable training tasks, reliable feedback, and controlled comparisons across generative substrates. In this work, we introduce VBVR-Pro, a closed-loop testbed that makes native visual reasoning through generation trainable, verifiable, optimizable, and experimentally controllable. 1) Task scaling. VBVR-Pro turns visual reasoning into a controlled task space of 300 procedurally generated tasks. Models trained on VBVR-Pro show strong transfer beyond the proposed suite across six held-out visual reasoning benchmarks such as RISE-Video, MME-CoF-Pro, and BabyVision. Further analysis validates that these gains reflect visual reasoning rather than instruction-pattern fitting. 2) Verifiable rewards. VBVR-Pro provides verifiable reward scorers for task-grounded evaluation. Through a systematic study of leading MLLMs as judges, we identify recurring failure modes of the prevalent VLM-as-a-judge paradigm. In contrast, the proposed scorers are grounded on verifiable task-specific rules, achieve fine-grained alignment with human judgments. Importantly, they serve as reliable reward signals for large-scale multi-task reinforcement learning and demonstrate stronger post-RL performance across visual reasoning tasks. 3) Mechanism study. VBVR-Pro enables controlled modality studies across more than 30 image, video, and interleaved generators. Our analysis shows that video generation remains strongest for tasks requiring persistent spatiotemporal state tracking, while interleaved generation provides a compute-efficient alternative by externalizing intermediate visual states. Critically, ablations and probing confirm the presence of vision-native trajectories, that are a more crucial substrate than explicit linguistic chains of thought for visual reasoning. We release all data, models, scorers, and code to facilitate future research.

The models are presented in the paper VBVR-Pro: A Scalable and Verifiable Suite for Native Visual Reasoning.

Models Zoo

Model Base Architecture Other Remarks
Image Generation Models
VBVR-Pro-BAGELBAGEL-7B-MoTComplete model
VBVR-Pro-FLUX2-devFLUX.2-devComplete model, Diffusers format
VBVR-Pro-FLUX2-dev-diffsynthFLUX.2-devLoRA model, DiffSynth format
VBVR-Pro-Qwen-Image-EditQwen-Image-Edit-2511Complete model, Diffusers format
VBVR-Pro-Qwen-Image-Edit-diffsynthQwen-Image-Edit-2511LoRA model, DiffSynth format
Interleaved Image Generation Models
VBVR-Pro-ThinkMorphThinkMorph-7BComplete model
VBVR-Pro-SenseNova-U1SenseNova-U1-8B-MoTComplete model
Video Generation Models
VBVR-Pro-LTX2.3LTX-Video-2.3Complete model, Diffusers format
VBVR-Pro-LTX2.3-diffsynthLTX-Video-2.3LoRA model, DiffSynth format
VBVR-Pro-Wan2.1-I2V-14BWan2.1-I2V-14B-720PComplete model, Diffusers format
VBVR-Pro-Wan2.1-I2V-14B-diffsynthWan2.1-I2V-14B-720PLoRA model, DiffSynth format
VBVR-Pro-Wan2.2-I2V-A14BWan2.2-I2V-A14BComplete model, Diffusers format
VBVR-Pro-Wan2.2-I2V-A14B-diffsynthWan2.2-I2V-A14BLoRA model, DiffSynth format
VBVR-Pro-Wan2.2-TI2V-5BWan2.2-TI2V-5BComplete model, Diffusers format
VBVR-Pro-Wan2.2-TI2V-5B-diffsynthWan2.2-TI2V-5BLoRA model, DiffSynth format
VBVR-Pro-Wan2.2-TI2V-5B-RLVRWan2.2-TI2V-5BComplete model, RL with verifiable rewards
VBVR-Pro-Wan2.2-TI2V-5B-RLVLM-Qwen3.6-27B-RewardWan2.2-TI2V-5BComplete model, RL with Qwen3.6-27B VLM rewards

VBVR-Pro Benchmark Results

Models Overall In-Domain by Category Out-of-Domain by Category
Avg.Abst.Know.Perc.Spat.Trans. Avg.Abst.Know.Perc.Spat.Trans.
Image Generation Models
Proprietary Models
Qwen-Image-2.00.3130.2480.2690.1960.2250.1700.1320.3780.3410.2350.3910.3840.080
Seedream-5.0-Pro0.5570.4850.5180.3120.5090.4010.2170.6290.5070.4550.6610.5590.202
Open-source Models
BAGEL-7B-MoT0.0890.0660.0390.0850.0670.0460.0270.1110.2010.0310.0730.0280.121
FLUX.2-dev0.1570.1080.0880.1090.0720.1000.0660.2060.1970.1650.1840.2410.077
Qwen-Image-Edit0.1340.1080.0920.0820.1000.1090.0560.1590.1760.0630.1410.1820.082
Strong Baselines
VBVR-Pro-BAGEL0.1720.1680.1990.1050.1100.2130.0550.1760.2540.1040.1480.0150.145
VBVR-Pro-FLUX.20.4070.4840.4830.3230.3670.4490.3360.3300.3610.2720.2550.4540.128
VBVR-Pro-Qwen-Image0.3220.3320.2980.2170.1930.4310.2220.3110.3410.2390.2330.4130.181
Interleaved Image Generation Models
Proprietary Models
GPT-Image-20.5070.4280.4560.3180.4280.2060.3000.5870.3980.4130.6330.4800.303
Nano Banana Pro0.5640.4800.5180.4220.5120.2850.1740.6480.5530.4990.6570.5850.220
Open-source Models
ThinkMorph-7B0.1540.1130.1000.0820.1010.1480.0310.1950.1760.1660.1630.2530.103
VBVR-SenseNova-U10.4080.4690.3560.3130.3730.3860.4770.3470.2910.3170.2750.4800.238
SenseNova-U1-8B-MoT0.5650.5330.5010.3950.5440.3550.3490.5970.4480.4950.5330.7170.401
Strong Baselines
VBVR-Pro-ThinkMorph0.3730.4020.4030.3440.2380.4540.1840.3440.3670.2240.2380.5350.257
VBVR-Pro-SenseNova-U10.6380.8110.6480.6950.6210.7700.5410.4640.4800.3280.3440.5580.408
Video Generation Models
Proprietary Models
Veo 3.10.3090.3120.2750.2990.2520.2670.1570.3050.3050.2330.2520.3120.219
Kling V30.3920.3560.2130.3260.3200.3550.2290.4270.2940.5640.3750.2420.412
SeedDance 2.00.4990.4510.3380.3610.3530.4680.3080.5470.3690.5110.4780.5380.532
Open-source Models
HunyuanVideo-I2V0.0540.0540.0230.0640.0150.0840.0320.0530.0880.0140.0280.0620.055
CogVideoX1.5-5B-I2V0.0850.1000.0610.1180.0690.0920.0600.0700.1250.0380.0510.0400.024
Wan2.1-I2V-14B0.1000.1050.0520.1250.0910.1020.0520.0950.1120.0730.0710.1230.044
Wan2.2-TI2V-5B0.0940.0660.0290.0730.0500.0830.0310.1220.1560.0520.1060.0630.099
Wan2.2-I2V-14B-720P0.1820.1570.0820.1310.1100.1610.1560.2070.2240.1390.1400.1950.273
LTX2.3-I2AV0.1120.1060.0620.1090.0700.1330.0550.1190.1610.1350.0860.0910.050
VBVR-Wan2.20.5170.5480.2370.4990.3340.5660.5910.4860.3100.3430.3450.7320.684
Strong Baselines
VBVR-Pro-LTX2.30.4250.5270.4090.5100.3460.4600.3900.3240.3810.1080.2010.4770.386
VBVR-Pro-Wan2.1-I2V-14B0.5620.7300.6170.5800.4520.6760.6230.3950.4100.3050.2300.6170.439
VBVR-Pro-Wan2.2-TI2V-5B0.4700.6410.5280.5560.3730.5650.5570.3000.3330.1270.1610.5050.409
VBVR-Pro-Wan2.2-I2V-14B0.6700.8080.6320.6850.5560.7510.6360.5320.4790.4180.3500.6790.690

Installation

We recommend using uv to manage the unified Python 3.10 environment.

uv installation guide: https://docs.astral.sh/uv/getting-started/installation/#installing-uv

git clone https://github.com/Video-Reason/VBVR-Pro.git
cd VBVR-Pro/
uv sync --extra cu124 # or one of [cu118|cu121|cu124|cu126|cu128|cu129]
source .venv/bin/activate

Select the extra for the CUDA wheel build you want; the extras are mutually exclusive. For faster attention on BAGEL and ThinkMorph, install the optional FlashAttention extension:

uv sync --extra cu124 --extra flash-attn

Download Models

Merged models (self-contained, recommended)

The merged (complete) models do not require separate base model downloads:

# Image generation
hf download Video-Reason/VBVR-Pro-FLUX2-dev --local-dir models/VBVR-Pro-FLUX2-dev
hf download Video-Reason/VBVR-Pro-Qwen-Image-Edit --local-dir models/VBVR-Pro-Qwen-Image-Edit
hf download Video-Reason/VBVR-Pro-BAGEL --local-dir models/VBVR-Pro-BAGEL

# Interleaved image generation
hf download Video-Reason/VBVR-Pro-ThinkMorph --local-dir models/VBVR-Pro-ThinkMorph
hf download Video-Reason/VBVR-Pro-SenseNova-U1 --local-dir models/VBVR-Pro-SenseNova-U1

# Video generation
hf download Video-Reason/VBVR-Pro-LTX2.3 --local-dir models/VBVR-Pro-LTX2.3
hf download Video-Reason/VBVR-Pro-Wan2.1-I2V-14B --local-dir models/VBVR-Pro-Wan2.1-I2V-14B
hf download Video-Reason/VBVR-Pro-Wan2.2-I2V-A14B --local-dir models/VBVR-Pro-Wan2.2-I2V-A14B
hf download Video-Reason/VBVR-Pro-Wan2.2-TI2V-5B --local-dir models/VBVR-Pro-Wan2.2-TI2V-5B
hf download Video-Reason/VBVR-Pro-Wan2.2-TI2V-5B-RLVR --local-dir models/VBVR-Pro-Wan2.2-TI2V-5B-RLVR
hf download Video-Reason/VBVR-Pro-Wan2.2-TI2V-5B-RLVLM-Qwen3.6-27B-Reward --local-dir models/VBVR-Pro-Wan2.2-TI2V-5B-RLVLM-Qwen3.6-27B-Reward

DiffSynth LoRA models (require base model downloads)

For -diffsynth releases, download the corresponding base models first:

export MODELS_DIR="${PWD}/models"
mkdir -p "${MODELS_DIR}"

hf download black-forest-labs/FLUX.2-dev \
  --local-dir "${MODELS_DIR}/black-forest-labs/FLUX.2-dev"
hf download Qwen/Qwen-Image-Edit-2511 \
  --local-dir "${MODELS_DIR}/Qwen/Qwen-Image-Edit-2511"
hf download DiffSynth-Studio/LTX-2.3-Repackage \
  --local-dir "${MODELS_DIR}/DiffSynth-Studio/LTX-2.3-Repackage"
hf download google/gemma-3-12b-it-qat-q4_0-unquantized \
  --local-dir "${MODELS_DIR}/google/gemma-3-12b-it-qat-q4_0-unquantized"
hf download Wan-AI/Wan2.1-I2V-14B-720P \
  --local-dir "${MODELS_DIR}/Wan-AI/Wan2.1-I2V-14B-720P"
hf download Wan-AI/Wan2.2-I2V-A14B \
  --local-dir "${MODELS_DIR}/Wan-AI/Wan2.2-I2V-A14B"
hf download Wan-AI/Wan2.2-TI2V-5B \
  --local-dir "${MODELS_DIR}/Wan-AI/Wan2.2-TI2V-5B"

Inference

All models are invoked through the unified example.py CLI:

# Image editing (FLUX.2 merged)
python example.py \
  --model_path models/VBVR-Pro-FLUX2-dev \
  --image_paths first_frame.png \
  --prompt "Move the red block to the left of the blue block." \
  --output flux_edit.png

# Video generation (Wan2.2 merged)
python example.py \
  --model_path models/VBVR-Pro-Wan2.2-I2V-A14B \
  --image_paths first_frame.png \
  --prompt "The subject walks toward the doorway." \
  --num_frames 81 --width 832 --height 480 \
  --output wan.mp4

# Audio-video generation (LTX-2.3 merged)
python example.py \
  --model_path models/VBVR-Pro-LTX2.3 \
  --image_paths first_frame.png \
  --prompt "The machine starts and makes a quiet mechanical hum." \
  --num_frames 49 --fps 24 \
  --output ltx.mp4

# Interleaved generation with reasoning (ThinkMorph)
python example.py \
  --model_path models/VBVR-Pro-ThinkMorph \
  --image_paths first_frame.png \
  --prompt "Solve the task step by step." \
  --think --max_rounds 10 --num_images 2 \
  --output thinkmorph_outputs

# DiffSynth LoRA (downloads base on first use)
python example.py \
  --model_path models/VBVR-Pro-Qwen-Image-Edit-diffsynth \
  --image_paths first_frame.png \
  --prompt "Place the cup on the empty shelf." \
  --output qwen_lora.png

Use python example.py --list-models for the full list of supported model types.

Training

For reproducible data preparation and distributed training, use the guide for the intended stage:

Development

Install the repository-wide lint environment and run the same checks as CI:

uv sync --frozen --only-group dev --no-install-project --inexact
.venv/bin/ruff check --output-format=github .
.venv/bin/ruff format --check .

The checks cover Python code in the root project and rl_training/. External repositories under vbvr_pro_models/ are excluded.

Download Evaluation Data

hf download Video-Reason/VBVR-Pro-Bench --repo-type dataset --local-dir data/VBVR-Pro-Bench

After generating outputs, evaluate using VBVR-Pro-Bench and submit results to the leaderboard.

License

VBVR-Pro source code, scripts, configuration files, task-specific scoring software, and the rl_training package are licensed under the Apache License 2.0. VBVR-Pro data and benchmark materials are separately licensed under CC BY-NC 4.0. Model weights and third-party materials remain subject to their applicable model-card and upstream terms. See LICENSE.md for details.

Citation

@misc{xu2026vbvrproscalableverifiablesuite,
      title={VBVR-Pro: A Scalable and Verifiable Suite for Native Visual Reasoning},
      author={Junxiang Xu and Ruisi Wang and Fanyi Pu and Maijunxian Wang and Ran Ji and Tongxi Zhou and Chenyang Gu and Jing Zuo and Hongcan Xiao and Yimeng Geng and Wanqi Yin and Wei Chen and Oscar Qian and Zhengan Yan and Ziqi Huang and Haiwen Diao and Liang Pan and Bo Li and Xiangyu Fan and Dezhi Luo and Fengyuan Yu and Zehong Zhao and Qingying Gao and Tinghui Zhu and Yilan Zhang and Jingqi Tong and Pinyuan Feng and Zhengze Jiang and Letian Wang and Ziyu Guo and Renrui Zhang and Jieneng Chen and Sonia Joseph and Constantin Venhoff and Saman Motamed and Mengyue Yang and Chandra Sripada and Alan Yuille and Philip Torr and Lvmin Zhang and Vikash Kumar and Daniel Khashabi and Nikolaus Kriegeskorte and Raphaël Millière and Vincent C. Müller and Anyi Rao and Quan Wang and Ziwei Liu and Dahua Lin and Lei Yang and Hokin Deng and Zhongang Cai},
      year={2026},
      eprint={2608.26105},
      archivePrefix={arXiv},
      primaryClass={cs.CV},
      url={https://arxiv.org/abs/2608.26105},
}

Acknowledgements

This project includes code modified from DiffSynth-Studio, BAGEL, and ThinkMorph. We gratefully acknowledge the authors and contributors for their open research.

About

VBVR-Pro: A Scalable and Verifiable Suite for Native Visual Reasoning

Resources

Stars

22 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages