TensorX
返回文献探索

Paper · arXiv 2312.10763

M3DBench: Let's Instruct Large Models with Multi-modal 3D Prompts

Mingsheng Li, Xin Chen, Chi Zhang, Sijin Chen, Hongyuan Zhu, Fukun Yin, Gang Yu, Tao Chen

18 upvotesDecember 17, 2023arXiv 预印本
AI 摘要

M3DBench is a comprehensive 3D instruction-following dataset enabling a broad range of multimodal 3D tasks and serves as a new benchmark for assessing model performance in understanding 3D prompts.

Large Language ModelsMultimodal Language ModelsM3DBenchinstruction-following datasetmultimodal instructions3D tasksregion levelscene levelinstruction-response pairsbenchmark

Abstract

Recently, 3D understanding has become popular to facilitate autonomous agents to perform further decisionmaking. However, existing 3D datasets and methods are often limited to specific tasks. On the other hand, recent progress in Large Language Models (LLMs) and Multimodal Language Models (MLMs) have demonstrated exceptional general language and imagery tasking performance. Therefore, it is interesting to unlock MLM's potential to be 3D generalist for wider tasks. However, current MLMs' research has been less focused on 3D tasks due to a lack of large-scale 3D instruction-following datasets. In this work, we introduce a comprehensive 3D instructionfollowing dataset called M3DBench, which possesses the following characteristics: 1) It supports general multimodal instructions interleaved with text, images, 3D objects, and other visual prompts. 2) It unifies diverse 3D tasks at both region and scene levels, covering a variety of fundamental abilities in real-world 3D environments. 3) It is a large-scale 3D instruction-following dataset with over 320k instruction-response pairs. Furthermore, we establish a new benchmark for assessing the performance of large models in understanding multi-modal 3D prompts. Extensive experiments demonstrate the effectiveness of our dataset and baseline, supporting general 3D-centric tasks, which can inspire future research.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
M3DBench: Let's Instruct Large Models with Multi-modal 3D Prompts | TensorX