TensorX
返回文献探索

Paper · arXiv 2603.01068

LLaDA-o: An Effective and Length-Adaptive Omni Diffusion Model

Zebin You, Xiaolu Zhang, Jun Zhou, Chongxuan Li, Ji-Rong Wen

22 upvotesMarch 1, 2026arXiv 预印本
AI 摘要

LLaDA-o is an omni diffusion model that uses a Mixture of Diffusion framework to jointly handle text understanding and visual generation through a shared attention backbone, achieving state-of-the-art performance in multimodal tasks.

Mixture of Diffusionomni diffusion modeldiscrete masked diffusioncontinuous diffusionattention backbonelength adaptation strategymultimodal understandingmultimodal generationDPG-Bench

Abstract

We present LLaDA-o, an effective and length-adaptive omni diffusion model for multimodal understanding and generation. LLaDA-o is built on a Mixture of Diffusion (MoD) framework that decouples discrete masked diffusion for text understanding and continuous diffusion for visual generation, while coupling them through a shared, simple, and efficient attention backbone that reduces redundant computation for fixed conditions. Building on MoD, we further introduce a data-centric length adaptation strategy that enables flexible-length decoding in multimodal settings without architectural changes. Extensive experiments show that LLaDA-o achieves state-of-the-art performance among omni-diffusion models on multimodal understanding and generation benchmarks, and reaches 87.04 on DPG-Bench for text-to-image generation, supporting the effectiveness of unified omni diffusion modeling. Code is available at https://github.com/ML-GSAI/LLaDA-o.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
LLaDA-o: An Effective and Length-Adaptive Omni Diffusion Model | TensorX