TensorX
返回文献探索

Paper · arXiv 2606.13141

Rethinking RAG in Long Videos: What to Retrieve and How to Use It?

Yuho Lee, Jisu Shin, Nicole Hee-Yeon Kim, Jihwan Bang, Juntae Lee, Kyuwoong Hwang, Fatih Porikli, Hwanjun Song

36 upvotesJune 11, 2026arXiv 预印本
AI 摘要

VideoRAG systems are extended to handle long egocentric videos with multi-modal retrieval across temporal granularities, addressing limitations in existing benchmarks and methods through a new benchmark and chunk-adaptive reranking approach.

retrieval-augmented generationVideoRAGegocentric videomulti-modal retrievaltemporal granularitiesV-RAGBenchchunk-adaptive rerankinginterleaved evidence form

Abstract

Retrieval-augmented generation is moving beyond text into long, egocentric video, where systems must select query-relevant chunks across multiple modalities and temporal granularities. Yet progress in VideoRAG is limited by two gaps: existing benchmarks allow queries to be answered without the video, obscuring retrieval errors, and prior methods apply a single modality-granularity configuration per query, ignoring chunk-level variability. We address both by introducing V-RAGBench, a benchmark of langlequery, evidence chunk, answerrangle triplets that enables faithful, decoupled evaluation of retrieval and generation, and CARVE, a simple method that runs parallel retrievers across configurations and employs chunk-adaptive reranking to identify the winning configuration for each chunk. Each chunk then enters the generator under its winning configuration selected during retrieval, yielding an interleaved evidence form where the chunk-level decision propagates across both stages. CARVE outperforms eight recent VideoRAG baselines, with the chunks supplied to the generator interleaving multiple configurations rather than sharing a single one, a behavior unattainable by query-level methods.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Rethinking RAG in Long Videos: What to Retrieve and How to Use It? | TensorX