Demystifing Video Reasoning
Ruisi Wang, Zhongang Cai, Fanyi Pu +11 authors
Diffusion-based video models demonstrate reasoning capabilities through denoising steps rather than frame sequences, exhibiting behaviors like working memory, self-correction, and perception-before-action within specialized transformer layers.