Skip to content
ljx1002Public

About

TAPFormer is a model that fuses images and events for high-frame-rate tracking any point (pixel) .

Resources

Stars

47 stars

Watchers

2 watching

Forks

Latest commit

 

History

24 Commits

Folders and files

Repository files navigation

TAPFormer: Robust Arbitrary Point Tracking via Transient Asynchronous Fusion of Frames and Events (CVPR 2026)

arXiv arXiv Project Page Dataset Hugging Face Page

Jiaxiong Liu1, Zhen Tan1, Jinpu Zhang1, Yi Zhou2, Hui Shen1, Xieyuanli Chen1, Dewen Hu1

1National University of Defense Technology    2Hunan University

TAPFormer teaser omage

Key Features

  • The first real-world TAP benchmark covering challenging conditions with synchronized frame–event data.
  • A novel asynchronous fusion paradigm that explicitly models temporal continuity between frames and events via Transient Asynchronous Fusion (TAF).
  • State-of-the-art performance across multiple datasets in TAP and feature point tracking tasks.

Example Predictions

Example 1 Example 2 Example 3
Example 4 Example 5 Example 6
Example 7 Example 8 Example 9

Installation

Requirements

  • Python 3.7+
  • CUDA-capable GPU (recommended) or CPU
  • PyTorch 1.8+

Setup

  1. Clone the repository:
git clone <repository-url>
cd tapformer
  1. Set up the environment:
conda create --name TAPFormer python=3.10
conda activate TAPFormer
  1. Install dependencies:
pip install -r requirements.txt
  1. (Optional) Install flow-vis for optical flow visualization:
pip install git+https://github.com/tomrunia/OpticalFlow_Visualization.git

Note: If you encounter issues with flow-vis, it's only needed for optical flow visualization mode and can be skipped if you don't use that feature.

Quick Start

1. Test Sequences and Pretrained Weights

We build the first benchmark for multimodal arbitrary point tracking, including a synthetic frame–event training set and manually annotated real-world test sequences, providing a comprehensive platform for future research. We also evaluate our model on the feature point tracking benchmarks EDS and EC. The updated EDS datasets ground truth annotations can be downloaded here.

Furthermore, we also provide the network weights trained on the FE-FastKub dataset

2. Prepare Your Data

To generate input event representations, run the following script file to generate event representations for the corresponding dataset:

data_pretation\real\InivTAP\genarate_event_represent_InivTAP.py
data_pretation\real\DrivTAP\genarate_event_represent_DrivTAP.py
data_pretation\real\genrate_EFrame_for_EDS_EC.py

Ensure your dataset is organized in the following structure:

dataset_dir/
├── eds_subseq/
│   └── sequence_name/
│       ├── events/
│       ├── images_corrected/
│       └── sequence_name.gt.txt
├── ec_subseq/
│   └── sequence_name/
│       ├── events/
│       ├── images_corrected/
│       └── track.gt.txt
├── InivTAP/
│   └── sequence_name/
│       ├── events/
│       ├── images_corrected/
│       └── annotations.npy
└── DrivTAP/
    └── sequence_name/
        ├── events/
        ├── images_corrected/
        └── annotations.npy

3. Configure Your Settings

Edit the YAML configuration files in the config/ directory:

  • config/config_eds_ec.yaml - For EDS and EC datasets
  • config/config_InivTAP_DrivTAP.yaml - For InivTAP and DrivTAP dataset

Key configuration options:

  • dataset_dir: Path to your dataset directory
  • ckpt_root: Path to model checkpoint
  • eval_datasets_*: List of sequences to evaluate
  • visualization.enable: Enable/disable visualization
  • output.save_results: Save evaluation results
  • output.save_trajectory: Save trajectory files

4. Run Evaluation

EDS and EC Datasets

python test_EDS_EC.py --config config/config_eds_ec.yaml

TAPFormer Dataset

python test_InivTAP_DrivTAP.py --config config/config_InivTAP_DrivTAP.yaml

Output

When enabled, the evaluation script generates:

  • Visualization videos: Tracked points overlaid on input frames
  • Trajectory files: Predicted trajectories in text format
  • Result files: Evaluation metrics (mean error, age, etc.)

Output files are saved in:

  • EDS: output/eval_eds_subseq/{sequence_name}/
  • EC: output/eval_ec_subseq/{sequence_name}/
  • InivTAP and DrivTAP: output/eval_InivTAP_DrivTAP_subseq/{sequence_name}/

Citation

If you use this code in your research, please cite:

@article{liu2026tapformer,
  title={TAPFormer: Robust Arbitrary Point Tracking via Transient Asynchronous Fusion of Frames and Events},
  author={Liu, Jiaxiong and Tan, Zhen and Zhang, Jinpu and Zhou, Yi and Shen, Hui and Chen, Xieyuanli and Hu, Dewen},
  journal={arXiv preprint arXiv:2603.04989},
  year={2026}
}
@inproceedings{liu2025tracking,
  title={Tracking any point with frame-event fusion network at high frame rate},
  author={Liu, Jiaxiong and Wang, Bo and Tan, Zhen and Zhang, Jinpu and Shen, Hui and Hu, Dewen},
  booktitle={2025 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)},
  pages={18834--18840},
  year={2025},
  organization={IEEE}
}

Acknowledgments

We gratefully appreciate the following repositories and thank the authors for their excellent work:

License

See the LICENSE file for details about the license under which this code is made available.

About

TAPFormer is a model that fuses images and events for high-frame-rate tracking any point (pixel) .

Resources

Stars

47 stars

Watchers

2 watching

Forks

Releases

Packages

Contributors

Languages