TensorX
返回文献探索

Paper · arXiv 2411.19458

Multiview Equivariance Improves 3D Correspondence Understanding with Minimal Feature Finetuning

Yang You, Yixin Li, Congyue Deng, Yue Wang, Leonidas Guibas

6 upvotesNovember 29, 2024arXiv 预印本
AI 摘要

Evaluating and enhancing 3D equivariance in ViT-based models improves performance in tasks like pose estimation, tracking, and semantic transfer through simple fine-tuning with 3D correspondences.

ViT family3D equivariant featuressemantic embeddingspose estimationtrackingsemantic transferfinetuning strategy3D correspondences

Abstract

Vision foundation models, particularly the ViT family, have revolutionized image understanding by providing rich semantic features. However, despite their success in 2D comprehension, their abilities on grasping 3D spatial relationships are still unclear. In this work, we evaluate and enhance the 3D awareness of ViT-based models. We begin by systematically assessing their ability to learn 3D equivariant features, specifically examining the consistency of semantic embeddings across different viewpoints. Our findings indicate that improved 3D equivariance leads to better performance on various downstream tasks, including pose estimation, tracking, and semantic transfer. Building on this insight, we propose a simple yet effective finetuning strategy based on 3D correspondences, which significantly enhances the 3D correspondence understanding of existing vision models. Remarkably, even finetuning on a single object for just one iteration results in substantial performance gains. All code and resources will be made publicly available to support further advancements in 3D-aware vision models. Our code is available at https://github.com/qq456cvb/3DCorrEnhance.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号