TensorX
返回文献探索

Paper · arXiv 2502.17157

DICEPTION: A Generalist Diffusion Model for Visual Perceptual Tasks

Canyu Zhao, Mingyu Liu, Huanyi Zheng, Muzhi Zhu, Zhiyue Zhao, Hao Chen, Tong He, Chunhua Shen

52 upvotesFebruary 24, 2025arXiv 预印本
AI 摘要

DICEPTION, a text-to-image diffusion model, achieves state-of-the-art performance in multiple perception tasks using minimal data and computational resources by leveraging color encoding and conditional image generation.

text-to-image diffusion modelsperception taskscolor encodingconditional image generation

Abstract

Our primary goal here is to create a good, generalist perception model that can tackle multiple tasks, within limits on computational resources and training data. To achieve this, we resort to text-to-image diffusion models pre-trained on billions of images. Our exhaustive evaluation metrics demonstrate that DICEPTION effectively tackles multiple perception tasks, achieving performance on par with state-of-the-art models. We achieve results on par with SAM-vit-h using only 0.06% of their data (e.g., 600K vs. 1B pixel-level annotated images). Inspired by Wang et al., DICEPTION formulates the outputs of various perception tasks using color encoding; and we show that the strategy of assigning random colors to different instances is highly effective in both entity segmentation and semantic segmentation. Unifying various perception tasks as conditional image generation enables us to fully leverage pre-trained text-to-image models. Thus, DICEPTION can be efficiently trained at a cost of orders of magnitude lower, compared to conventional models that were trained from scratch. When adapting our model to other tasks, it only requires fine-tuning on as few as 50 images and 1% of its parameters. DICEPTION provides valuable insights and a more promising solution for visual generalist models.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
DICEPTION: A Generalist Diffusion Model for Visual Perceptual Tasks | TensorX