TensorX
返回文献探索

Paper · arXiv 2504.05298

One-Minute Video Generation with Test-Time Training

Karan Dalal, Daniel Koceja, Gashon Hussein, Jiarui Xu, Yue Zhao, Youjin Song, Shihao Han, Ka Chun Cheung, Jan Kautz, Carlos Guestrin, Tatsunori Hashimoto, Sanmi Koyejo, Yejin Choi, Yu Sun, Xiaolong Wang

110 upvotesApril 7, 2025arXiv 预印本
AI 摘要

Test-Time Training (TTT) layers enable pre-trained Transformers to generate coherent one-minute videos from text storyboards, outperforming alternatives like Mamba~2 and Gated DeltaNet.

self-attention layersMamba layersTest-Time Training (TTT) layerspre-trained Transformerone-minute videostext storyboardsTom and Jerry cartoonsElo pointssliding-window attention layers

Abstract

Transformers today still struggle to generate one-minute videos because self-attention layers are inefficient for long context. Alternatives such as Mamba layers struggle with complex multi-scene stories because their hidden states are less expressive. We experiment with Test-Time Training (TTT) layers, whose hidden states themselves can be neural networks, therefore more expressive. Adding TTT layers into a pre-trained Transformer enables it to generate one-minute videos from text storyboards. For proof of concept, we curate a dataset based on Tom and Jerry cartoons. Compared to baselines such as Mamba~2, Gated DeltaNet, and sliding-window attention layers, TTT layers generate much more coherent videos that tell complex stories, leading by 34 Elo points in a human evaluation of 100 videos per method. Although promising, results still contain artifacts, likely due to the limited capability of the pre-trained 5B model. The efficiency of our implementation can also be improved. We have only experimented with one-minute videos due to resource constraints, but the approach can be extended to longer videos and more complex stories. Sample videos, code and annotations are available at: https://test-time-training.github.io/video-dit

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
One-Minute Video Generation with Test-Time Training | TensorX