Vript: A Video Is Worth Thousands of Words
Dongjie Yang, Suyuan Huang, Chengqiang Lu +5 authors
Vript introduces a high-quality video-text dataset with detailed captions including camera operations, enhancing video captioning and generation, and proposes Vriptor, a top-performing model with Vript-Hard, a new benchmark for video understanding challenges.