TensorX
返回文献探索

Paper · arXiv 2305.06355

VideoChat: Chat-Centric Video Understanding

KunChang Li, Yinan He, Yi Wang, Yizhuo Li, Wenhai Wang, Ping Luo, Yali Wang, Limin Wang, Yu Qiao

3 upvotesMay 10, 2023arXiv 预印本
AI 摘要

VideoChat combines video foundation models and large language models with a learnable neural interface for tasks such as spatiotemporal reasoning and causal relationship inference, supported by a new video-centric instruction dataset.

video foundation modelslarge language modelsneural interfacespatiotemporal reasoningevent localizationcausal relationship inferenceinstructional tuningvideo-centric instruction dataset

Abstract

In this study, we initiate an exploration into video understanding by introducing VideoChat, an end-to-end chat-centric video understanding system. It integrates video foundation models and large language models via a learnable neural interface, excelling in spatiotemporal reasoning, event localization, and causal relationship inference. To instructively tune this system, we propose a video-centric instruction dataset, composed of thousands of videos matched with detailed descriptions and conversations. This dataset emphasizes spatiotemporal reasoning and causal relationships, providing a valuable asset for training chat-centric video understanding systems. Preliminary qualitative experiments reveal our system's potential across a broad spectrum of video applications and set the standard for future research. Access our code and data at https://github.com/OpenGVLab/Ask-Anything

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
VideoChat: Chat-Centric Video Understanding | TensorX