TensorX
返回文献探索

Paper · arXiv 2501.05453

An Empirical Study of Autoregressive Pre-training from Videos

Jathushan Rajasegaran, Ilija Radosavovic, Rahul Ravishankar, Yossi Gandelsman, Christoph Feichtenhofer, Jitendra Malik

39 upvotesJanuary 9, 2025arXiv 预印本
AI 摘要

Autoregressive pre-training on a massive dataset of videos and images using transformer models yields competitive performance across various tasks and exhibits scaling behavior similar to language models.

autoregressive pre-trainingtransformer modelsvisual tokensimage recognitionvideo classificationobject trackingroboticsscaling curveslanguage models

Abstract

We empirically study autoregressive pre-training from videos. To perform our study, we construct a series of autoregressive video models, called Toto. We treat videos as sequences of visual tokens and train transformer models to autoregressively predict future tokens. Our models are pre-trained on a diverse dataset of videos and images comprising over 1 trillion visual tokens. We explore different architectural, training, and inference design choices. We evaluate the learned visual representations on a range of downstream tasks including image recognition, video classification, object tracking, and robotics. Our results demonstrate that, despite minimal inductive biases, autoregressive pre-training leads to competitive performance across all benchmarks. Finally, we find that scaling our video models results in similar scaling curves to those seen in language models, albeit with a different rate. More details at https://brjathu.github.io/toto/

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
An Empirical Study of Autoregressive Pre-training from Videos | TensorX