TensorX
返回文献探索

Paper · arXiv 2311.03226

LDM3D-VR: Latent Diffusion Model for 3D VR

Gabriela Ben Melech Stan, Diana Wofk, Estelle Aflalo, Shao-Yen Tseng, Zhipeng Cai, Michael Paulitsch, Vasudev Lal

9 upvotesNovember 6, 2023arXiv 预印本
AI 摘要

LDM3D-VR, comprising LDM3D-pano and LDM3D-SR, generates high-resolution RGBD from textual prompts and low-resolution inputs, demonstrating state-of-the-art performance in virtual reality applications.

latent diffusion modelsRGBDvirtual realityLDM3D-VRLDM3D-panoLDM3D-SRpretrained modelsdatasetsdepth mapscaptions

Abstract

Latent diffusion models have proven to be state-of-the-art in the creation and manipulation of visual outputs. However, as far as we know, the generation of depth maps jointly with RGB is still limited. We introduce LDM3D-VR, a suite of diffusion models targeting virtual reality development that includes LDM3D-pano and LDM3D-SR. These models enable the generation of panoramic RGBD based on textual prompts and the upscaling of low-resolution inputs to high-resolution RGBD, respectively. Our models are fine-tuned from existing pretrained models on datasets containing panoramic/high-resolution RGB images, depth maps and captions. Both models are evaluated in comparison to existing related methods.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
LDM3D-VR: Latent Diffusion Model for 3D VR | TensorX