TensorX
返回文献探索

Paper · arXiv 2607.28509

RefCaptioner: Multi-Reference Image-Grounded Video Captioning

Tengfei Liu, Yang Shi, Yuran Wang, Xiaohan Zhang, Yuqing Wen, Yuqi Tang, Qixun Wang, Zhuoran Zhang, Xuanyu Zhu, Weihong Lin, Xinlei Yu, Yujie Wei, Xinwei Long, Fengxiang Wang, Xinlong Chen, Yue Ding, Jialu Chen, Haotian Wang, Yuanxing Zhang

30 upvotesJuly 30, 2026arXiv 预印本
AI 摘要

RefCaptioner is a two-stage framework for multi-reference image-grounded video captioning that improves phrase-level grounding and factual consistency, supported by a new benchmark and training corpus.

multi-reference image-grounded video captioningRefCaptionermixed-data SFTHierarchical Coverage-Discounted GRPOphrase-level bindingdistractor rejectioncross-reference consistencyMRVBench

Abstract

Existing video captioning models generate natural descriptions of video content but cannot explicitly ground local visual elements to multiple reference images. We introduce multi-reference image-grounded video captioning, a new task requiring factual video descriptions with phrase-level reference grounding, and propose RefCaptioner, a two-stage post-training framework for this task. RefCaptioner combines mixed-data SFT with Hierarchical Coverage-Discounted GRPO to jointly improve reference selection, phrase-level binding, distractor rejection, and cross-reference consistency while preserving general video-captioning ability. To support training, we construct a corpus containing 20,000 videos and 171,354 reference images. We further introduce MRVBench, a benchmark for evaluating caption factuality and multi-reference grounding on both real-world and AI-generated videos. Experiments show that RefCaptioner achieves the best overall performance among the open-source models while remaining competitive on standard video captioning benchmarks. Human evaluation further confirms that its captions are preferred by annotators and enable more source-faithful video reconstruction with both open-source and proprietary video generators.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
RefCaptioner: Multi-Reference Image-Grounded Video Captioning | TensorX