TensorX
返回文献探索

Paper · arXiv 2412.14015

Prompting Depth Anything for 4K Resolution Accurate Metric Depth Estimation

Haotong Lin, Sida Peng, Jingxiao Chen, Songyou Peng, Jiaming Sun, Minghuan Liu, Hujun Bao, Jiashi Feng, Xiaowei Zhou, Bingyi Kang

12 upvotesDecember 18, 2024arXiv 预印本
AI 摘要

Prompt Depth Anything integrates LiDAR data into depth foundation models for high-resolution depth estimation, achieving state-of-the-art results on ARKitScenes and ScanNet++ datasets and improving downstream applications.

promptingdepth foundation modelsPrompt Depth AnythingLiDARdepth decodersynthetic data LiDAR simulationpseudo GT depth generationARKitScenesScanNet++3D reconstructiongeneralized robotic grasping

Abstract

Prompts play a critical role in unleashing the power of language and vision foundation models for specific tasks. For the first time, we introduce prompting into depth foundation models, creating a new paradigm for metric depth estimation termed Prompt Depth Anything. Specifically, we use a low-cost LiDAR as the prompt to guide the Depth Anything model for accurate metric depth output, achieving up to 4K resolution. Our approach centers on a concise prompt fusion design that integrates the LiDAR at multiple scales within the depth decoder. To address training challenges posed by limited datasets containing both LiDAR depth and precise GT depth, we propose a scalable data pipeline that includes synthetic data LiDAR simulation and real data pseudo GT depth generation. Our approach sets new state-of-the-arts on the ARKitScenes and ScanNet++ datasets and benefits downstream applications, including 3D reconstruction and generalized robotic grasping.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号