TensorX
返回文献探索

Paper · arXiv 2504.12626

Packing Input Frame Context in Next-Frame Prediction Models for Video Generation

Lvmin Zhang, Maneesh Agrawala

52 upvotesApril 17, 2025arXiv 预印本
AI 摘要

FramePack, a neural network for video generation, compresses frames to manage transformer context length and enhances video diffusion models with increased batch sizes and improved frame prediction.

neural networkFramePacknext-frame predictiontransformercontext lengthvideo diffusioncomputation bottleneckimage diffusionbatch sizeanti-drifting samplinginverted temporal orderexposure biasdiffusion schedulersflow shift

Abstract

We present a neural network structure, FramePack, to train next-frame (or next-frame-section) prediction models for video generation. The FramePack compresses input frames to make the transformer context length a fixed number regardless of the video length. As a result, we are able to process a large number of frames using video diffusion with computation bottleneck similar to image diffusion. This also makes the training video batch sizes significantly higher (batch sizes become comparable to image diffusion training). We also propose an anti-drifting sampling method that generates frames in inverted temporal order with early-established endpoints to avoid exposure bias (error accumulation over iterations). Finally, we show that existing video diffusion models can be finetuned with FramePack, and their visual quality may be improved because the next-frame prediction supports more balanced diffusion schedulers with less extreme flow shift timesteps.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Packing Input Frame Context in Next-Frame Prediction Models for Video Generation | TensorX