TensorX
返回文献探索

Paper · arXiv 2503.21749

LeX-Art: Rethinking Text Generation via Scalable High-Quality Data Synthesis

Shitian Zhao, Qilong Wu, Xinyue Li, Bo Zhang, Ming Li, Qi Qin, Dongyang Liu, Kaipeng Zhang, Hongsheng Li, Yu Qiao, Peng Gao, Bin Fu, Zhen Li

26 upvotesMarch 27, 2025arXiv 预印本
AI 摘要

A suite called LeX-Art for high-quality text-image synthesis includes data-centric pipeline, prompt enrichment, and text-to-image models, achieving state-of-the-art performance with a new benchmark and metric.

Deepseek-R1LeX-10KLeX-EnhancerLeX-FLUXLeX-LuminaLeX-BenchPairwise Normalized Edit Distance (PNED)CreateBench

Abstract

We introduce LeX-Art, a comprehensive suite for high-quality text-image synthesis that systematically bridges the gap between prompt expressiveness and text rendering fidelity. Our approach follows a data-centric paradigm, constructing a high-quality data synthesis pipeline based on Deepseek-R1 to curate LeX-10K, a dataset of 10K high-resolution, aesthetically refined 1024times1024 images. Beyond dataset construction, we develop LeX-Enhancer, a robust prompt enrichment model, and train two text-to-image models, LeX-FLUX and LeX-Lumina, achieving state-of-the-art text rendering performance. To systematically evaluate visual text generation, we introduce LeX-Bench, a benchmark that assesses fidelity, aesthetics, and alignment, complemented by Pairwise Normalized Edit Distance (PNED), a novel metric for robust text accuracy evaluation. Experiments demonstrate significant improvements, with LeX-Lumina achieving a 79.81% PNED gain on CreateBench, and LeX-FLUX outperforming baselines in color (+3.18%), positional (+4.45%), and font accuracy (+3.81%). Our codes, models, datasets, and demo are publicly available.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
LeX-Art: Rethinking Text Generation via Scalable High-Quality Data Synthesis | TensorX