Beyond Static Frames: Temporal Aggregate-and-Restore Vision Transformer for Human Pose Estimation

News

2026-04-09: 🌟TAR-ViTPose has been selected as a CVPR 2026 Highlight.
2026-02-21: 🎉This repo is the official implementation for Beyond Static Frames: Temporal Aggregate-and-Restore Vision Transformer for Human Pose Estimation[arXiv]. The paper has been accepted to CVPR 2026.

Introduction

Vision Transformers (ViTs) have recently achieved stateof-the-art performance in 2D human pose estimation due to their strong global modeling capability. However, existing ViT-based pose estimators are designed for static images and process each frame independently, thereby ignoring the temporal coherence that exists in video sequences. This limitation often results in unstable predictions, especially in challenging scenes involving motion blur, occlusion, or defocus. In this paper, we propose TAR-ViTPose, a novel Temporal Aggregate-and-Restore Vision Transformer tailored for video-based 2D human pose estimation. TAR-ViTPose enhances static ViT representations by aggregating temporal cues across frames in a plug-and-play manner, leading to more robust and accurate pose estimation. To effectively aggregate joint-specific features that are temporally aligned across frames, we introduce a joint-centric temporal aggregation (JTA) that assigns each joint a learnable query token to selectively attend to its corresponding regions from neighboring frames. Furthermore, we develop a global restoring attention (GRA) to restore the aggregated temporal features back into the token sequence of the current frame, enriching its pose representation while fully preserving global context for precise keypoint localization. Extensive experiments demonstrate that TAR-ViTPose substantially improves upon the single-frame baseline ViTPose, achieving a +2.3 mAP gain on the PoseTrack2017 benchmark. Moreover, our approach outperforms existing stateof-the-art video-based methods, while also achieving a noticeably higher runtime frame rate in real-world applications.

Weights Download

We provide the model weights trained by the method in this paper, which can be downloaded here. https://drive.google.com/drive/folders/1zmQaK56IHOJIf1AQ8ZWZbq0UtBRmsPuI?usp=drive_link

Visualizations

Here are some qualitative results from both the PoseTrack dataset and real-world scenarios:

Video Demo

Environment

The code is developed and tested under the following environment:

Python 3.11.6
PyTorch 2.0.1
CUDA 11.8

conda create -n tarvitpose pyhton=3.11.6
conda activate tarvitpose
conda install pytorch==2.0.1 torchvision==0.15.2 torchaudio==2.0.2 pytorch-cuda=11.8 -c pytorch -c nvidia
cd tarvitpose
conda env update --file environment.yml

Usage

To download some auxiliary materials, please refer to DCPose.

Follow the MMPose instruction to install the mmpose.

Training

cd tools
python train.py --config ../configs/posetrack17/tarvitpose.yaml

Evaluation

cd tools
python val.py --config ../configs/posetrack17/tarvitpose.yaml --weights_path /path/to/weights.pt

Video Inference

cd tools
python inference.py -c ../configs/posetrack17/tarvitpose.yaml -w /path/to/weights.pt -i input_video.mp4

Citations

If you find our paper useful in your research, please consider citing:

@InProceedings{fang_2026_tarvitpose,
    author    = {Fang, Hongwei and Cai, Jiahang and Wang, Xun and Yang, Wenwu},
    title     = {Beyond Static Frames: Temporal Aggregate-and-Restore Vision Transformer for Human Pose Estimation},
    booktitle = {Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)},
    year      = {2026},
}

Acknowledgment

Our codes are mainly based on Poseidon and MMPose. Many thanks to the authors!

Name		Name	Last commit message	Last commit date
Latest commit History 12 Commits
configs/posetrack17		configs/posetrack17
core		core
datasets		datasets
demo		demo
engine		engine
models		models
posetimation		posetimation
tools		tools
utils		utils
.gitignore		.gitignore
LICENSE		LICENSE
README.md		README.md
environment.yml		environment.yml

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Repository files navigation

Beyond Static Frames: Temporal Aggregate-and-Restore Vision Transformer for Human Pose Estimation

News

Introduction

Weights Download

Visualizations

Video Demo

Environment

Usage

Training

Evaluation

Video Inference

Citations

Acknowledgment

About

Uh oh!

Releases

Packages

Uh oh!

Contributors

Uh oh!

Languages

Folders and files

Latest commit

History

Repository files navigation

Beyond Static Frames: Temporal Aggregate-and-Restore Vision Transformer for Human Pose Estimation

News

Introduction

Weights Download

Visualizations

Video Demo

Environment

Usage

Training

Evaluation

Video Inference

Citations

Acknowledgment

About

Resources

License

Uh oh!

Stars

Watchers

Forks

Releases

Packages 0

Uh oh!

Contributors

Uh oh!

Languages

Packages