VidLA: Video-Language Alignment at Scale
Mamshad Nayeem Rizve, Fan Fei, Jayakrishnan Unnikrishnan +5 authors
VidLA, a simple yet effective video-language alignment approach, uses temporally hierarchical data tokens and a two-tower architecture to enhance performance with large, semantically aligned datasets and surpass state-of-the-art methods.