TensorX
返回文献探索

Paper · arXiv 2410.02713

Video Instruction Tuning With Synthetic Data

Yuanhan Zhang, Jinming Wu, Wei Li, Bo Li, Zejun Ma, Ziwei Liu, Chunyuan Li

41 upvotesOctober 3, 2024arXiv 预印本
AI 摘要

A synthetic dataset, LLaVA-Video-178K, and a new video large multimodal model, LLaVA-Video, achieve strong performance in video instruction-following tasks.

video large multimodal modelssynthetic datasetLLaVA-Video-178Kdetailed captioningopen-ended question-answeringmultiple-choice QAvisual instruction tuningvideo benchmarks

Abstract

The development of video large multimodal models (LMMs) has been hindered by the difficulty of curating large amounts of high-quality raw data from the web. To address this, we propose an alternative approach by creating a high-quality synthetic dataset specifically for video instruction-following, namely LLaVA-Video-178K. This dataset includes key tasks such as detailed captioning, open-ended question-answering (QA), and multiple-choice QA. By training on this dataset, in combination with existing visual instruction tuning data, we introduce LLaVA-Video, a new video LMM. Our experiments demonstrate that LLaVA-Video achieves strong performance across various video benchmarks, highlighting the effectiveness of our dataset. We plan to release the dataset, its generation pipeline, and the model checkpoints.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Video Instruction Tuning With Synthetic Data | TensorX