TensorX
返回文献探索

Paper · arXiv 2307.16449

MovieChat: From Dense Token to Sparse Memory for Long Video Understanding

Enxin Song, Wenhao Chai, Guanhong Wang, Yucheng Zhang, Haoyang Zhou, Feiyang Wu, Xun Guo, Tian Ye, Yan Lu, Jenq-Neng Hwang, Gaoang Wang

17 upvotesJuly 31, 2023arXiv 预印本
AI 摘要

A video understanding system using a dual-memory mechanism based on the Atkinson-Shiffrin model, leveraging Transformers, achieves leading performance on long videos.

video foundation modelslarge language modelslong videoscomputation complexitymemory costlong-term temporal connectionAtkinson-Shiffrin memory modelshort-term memorylong-term memoryTransformersMovieChat

Abstract

Recently, integrating video foundation models and large language models to build a video understanding system overcoming the limitations of specific pre-defined vision tasks. Yet, existing systems can only handle videos with very few frames. For long videos, the computation complexity, memory cost, and long-term temporal connection are the remaining challenges. Inspired by Atkinson-Shiffrin memory model, we develop an memory mechanism including a rapidly updated short-term memory and a compact thus sustained long-term memory. We employ tokens in Transformers as the carriers of memory. MovieChat achieves state-of-the-art performace in long video understanding.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号