TensorX
返回文献探索

Paper · arXiv 2607.29679

Scaling Properties of Text Conditioning in Visual Generation

Zilong Chen, Chaorui Deng, Kunchang Li, Hongyi Yuan, Haoqi Fan

41 upvotesJuly 31, 2026arXiv 预印本
AI 摘要

Structured language in prompts drives diffusion loss scaling, enabling improved visual generation through annotated prompts and a trained prompter.

diffusion lossstructured languageGPGEDdiffusabilitypromptabilitysupervised fine-tuningcold-startverifier-gated on-policy distillationcompositional reasoning

Abstract

We study empirical scaling properties for text conditioning in visual generation. Such properties have rarely been measured because diffusion loss does not scale with the number of tokens in natural-language prompts. Surprisingly, we find that the converged diffusion loss scales with the amount of structured language in the prompt. To quantify structured language, we adapt two complementary measures: a white-box likelihood metric (GPG) and a black-box attribute metric (ED). Across controlled training runs, the converged diffusion loss decreases approximately linearly with GPG and follows a power law with ED. Guided by these scaling properties, we improve diffusability by constructing structured prompts with semantic and geometric annotations derived from images, and improve promptability by training a prompter through supervised fine-tuning, cold-start, and verifier-gated on-policy distillation. The resulting system outperforms all evaluated open-weight models on nearly every compositional, reasoning, and world-knowledge benchmark, while matching or surpassing the strongest closed-weight models on most evaluations.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Scaling Properties of Text Conditioning in Visual Generation | TensorX