TensorX
返回文献探索

Paper · arXiv 2310.19784

CustomNet: Zero-shot Object Customization with Variable-Viewpoints in Text-to-Image Diffusion Models

Ziyang Yuan, Mingdeng Cao, Xintao Wang, Zhongang Qi, Chun Yuan, Ying Shan

10 upvotesOctober 30, 2023arXiv 预印本
AI 摘要

CustomNet enhances text-to-image generation by integrating 3D novel view synthesis to customize object viewpoints, locations, and backgrounds, ensuring identity preservation and diverse outputs without test-time optimization.

text-to-image generation3D novel view synthesisobject customizationviewpointlocationbackgroundidentity preservationdataset construction pipelinezero-shot customization

Abstract

Incorporating a customized object into image generation presents an attractive feature in text-to-image generation. However, existing optimization-based and encoder-based methods are hindered by drawbacks such as time-consuming optimization, insufficient identity preservation, and a prevalent copy-pasting effect. To overcome these limitations, we introduce CustomNet, a novel object customization approach that explicitly incorporates 3D novel view synthesis capabilities into the object customization process. This integration facilitates the adjustment of spatial position relationships and viewpoints, yielding diverse outputs while effectively preserving object identity. Moreover, we introduce delicate designs to enable location control and flexible background control through textual descriptions or specific user-defined images, overcoming the limitations of existing 3D novel view synthesis methods. We further leverage a dataset construction pipeline that can better handle real-world objects and complex backgrounds. Equipped with these designs, our method facilitates zero-shot object customization without test-time optimization, offering simultaneous control over the viewpoints, location, and background. As a result, our CustomNet ensures enhanced identity preservation and generates diverse, harmonious outputs.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
CustomNet: Zero-shot Object Customization with Variable-Viewpoints in Text-to-Image Diffusion Models | TensorX