TensorX
返回文献探索

Paper · arXiv 2603.14935

Video-CoE: Reinforcing Video Event Prediction via Chain of Events

Qile Su, Jing Tang, Rui Chen, Lei Sun, Xiangxiang Chu

90 upvotesMarch 16, 2026arXiv 预印本
AI 摘要

A new Chain of Events paradigm is introduced for video event prediction that improves temporal modeling and logical reasoning in multimodal language models through structured event chains and enhanced training protocols.

video event predictionmultimodal language modelstemporal modelinglogical reasoningevent chainstraining protocols

Abstract

Despite advances in the application of MLLMs for various video tasks, video event prediction (VEP) remains relatively underexplored. VEP requires the model to perform fine-grained temporal modeling of videos and establish logical relationships between videos and future events, which current MLLMs still struggle with. In this work, we first present a comprehensive evaluation of current leading MLLMs on the VEP task, revealing the reasons behind their inaccurate predictions, including lack of logical reasoning ability for future events prediction and insufficient utilization of visual information. To address these challenges, we propose Chain of Events (CoE) paradigm, which constructs temporal event chains to implicitly enforce MLLM focusing on the visual content and the logical connections between videos and future events, incentivizing model's reasoning capability with multiple training protocols. Experimental results on public benchmarks demonstrate that our method outperforms both leading open-source and commercial MLLMs, establishing a new state-of-the-art on the VEP task. Codes and models will be released soon.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号