TensorX
返回文献探索

Paper · arXiv 2305.09664

Understanding 3D Object Interaction from a Single Image

Shengyi Qian, David F. Fouhey

2 upvotesMay 16, 2023arXiv 预印本
AI 摘要

A transformer-based model predicts 3D object location, properties, and affordance, improving machine interaction with 3D scenes and objects.

transformer-based model3D locationphysical propertiesaffordanceegocentric videosindoor imagesrobotics data

Abstract

Humans can easily understand a single image as depicting multiple potential objects permitting interaction. We use this skill to plan our interactions with the world and accelerate understanding new objects without engaging in interaction. In this paper, we would like to endow machines with the similar ability, so that intelligent agents can better explore the 3D scene or manipulate objects. Our approach is a transformer-based model that predicts the 3D location, physical properties and affordance of objects. To power this model, we collect a dataset with Internet videos, egocentric videos and indoor images to train and validate our approach. Our model yields strong performance on our data, and generalizes well to robotics data.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Understanding 3D Object Interaction from a Single Image | TensorX