TensorX
返回文献探索

Paper · arXiv 2511.15700

First Frame Is the Place to Go for Video Content Customization

Jingxi Chen, Zongxia Li, Zhichao Liu, Guangyao Shi, Xiyang Wu, Fuxiao Liu, Cornelia Fermuller, Brandon Y. Feng, Yiannis Aloimonos

54 upvotesNovember 19, 2025arXiv 预印本
AI 摘要

Video generation models use the first frame as a conceptual memory buffer, enabling robust customization with minimal training examples.

conceptual memory buffervideo generation modelsreference-based video customization

Abstract

What role does the first frame play in video generation models? Traditionally, it's viewed as the spatial-temporal starting point of a video, merely a seed for subsequent animation. In this work, we reveal a fundamentally different perspective: video models implicitly treat the first frame as a conceptual memory buffer that stores visual entities for later reuse during generation. Leveraging this insight, we show that it's possible to achieve robust and generalized video content customization in diverse scenarios, using only 20-50 training examples without architectural changes or large-scale finetuning. This unveils a powerful, overlooked capability of video generation models for reference-based video customization.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
First Frame Is the Place to Go for Video Content Customization | TensorX