TensorX
返回文献探索

Paper · arXiv 2502.05173

VideoRoPE: What Makes for Good Video Rotary Position Embedding?

Xilin Wei, Xiaoran Liu, Yuhang Zang, Xiaoyi Dong, Pan Zhang, Yuhang Cao, Jian Tong, Haodong Duan, Qipeng Guo, Jiaqi Wang, Xipeng Qiu, Dahua Lin

64 upvotesFebruary 7, 2025arXiv 预印本
AI 摘要

VideoRoPE, a 3D position embedding that preserves spatio-temporal relationships, improves performance on video tasks by addressing limitations in previous RoPE variants.

Rotary Position EmbeddingRoPElong-context capabilitiesV-NIAH-Ddiagonal layoutlow-frequency temporal allocationadjustable temporal spacinglong video retrievalvideo understandingvideo hallucination

Abstract

While Rotary Position Embedding (RoPE) and its variants are widely adopted for their long-context capabilities, the extension of the 1D RoPE to video, with its complex spatio-temporal structure, remains an open challenge. This work first introduces a comprehensive analysis that identifies four key characteristics essential for the effective adaptation of RoPE to video, which have not been fully considered in prior work. As part of our analysis, we introduce a challenging V-NIAH-D (Visual Needle-In-A-Haystack with Distractors) task, which adds periodic distractors into V-NIAH. The V-NIAH-D task demonstrates that previous RoPE variants, lacking appropriate temporal dimension allocation, are easily misled by distractors. Based on our analysis, we introduce VideoRoPE, with a 3D structure designed to preserve spatio-temporal relationships. VideoRoPE features low-frequency temporal allocation to mitigate periodic oscillations, a diagonal layout to maintain spatial symmetry, and adjustable temporal spacing to decouple temporal and spatial indexing. VideoRoPE consistently surpasses previous RoPE variants, across diverse downstream tasks such as long video retrieval, video understanding, and video hallucination. Our code will be available at https://github.com/Wiselnn570/VideoRoPE{https://github.com/Wiselnn570/VideoRoPE}.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
VideoRoPE: What Makes for Good Video Rotary Position Embedding? | TensorX