TensorX
返回文献探索

Paper · arXiv 2307.08506

Does Visual Pretraining Help End-to-End Reasoning?

Chen Sun, Calvin Luo, Xingyi Zhou, Anurag Arnab, Cordelia Schmid

7 upvotesJuly 17, 2023arXiv 预印本
AI 摘要

End-to-end visual reasoning is achieved using a self-supervised transformer-based framework that compresses video frames into tokens and reconstructs them, outperforming supervised pretraining methods.

end-to-end learningvisual reasoningvisual pretrainingcompositional generalizationself-supervised frameworktransformer networkvideo framestemporal contextreconstruction losscompact representationtemporal dynamicsobject permanenceCATERACREsupervised pretrainingimage classificationexplicit object detection

Abstract

We aim to investigate whether end-to-end learning of visual reasoning can be achieved with general-purpose neural networks, with the help of visual pretraining. A positive result would refute the common belief that explicit visual abstraction (e.g. object detection) is essential for compositional generalization on visual reasoning, and confirm the feasibility of a neural network "generalist" to solve visual recognition and reasoning tasks. We propose a simple and general self-supervised framework which "compresses" each video frame into a small set of tokens with a transformer network, and reconstructs the remaining frames based on the compressed temporal context. To minimize the reconstruction loss, the network must learn a compact representation for each image, as well as capture temporal dynamics and object permanence from temporal context. We perform evaluation on two visual reasoning benchmarks, CATER and ACRE. We observe that pretraining is essential to achieve compositional generalization for end-to-end visual reasoning. Our proposed framework outperforms traditional supervised pretraining, including image classification and explicit object detection, by large margins.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Does Visual Pretraining Help End-to-End Reasoning? | TensorX