TensorX
返回文献探索

Paper · arXiv 2403.09622

Glyph-ByT5: A Customized Text Encoder for Accurate Visual Text Rendering

Zeyu Liu, Weicong Liang, Zhanhao Liang, Chong Luo, Ji Li, Gao Huang, Yuhui Yuan

17 upvotesMarch 14, 2024arXiv 预印本
AI 摘要

Customizing the ByT5 text encoder with Glyph-ByT5 and integrating it with SDXL significantly improves text rendering accuracy and paragraph rendering in design and real images.

text encoderGlyph-ByT5ByT5fine-tuningpaired glyph-text datasetSDXLGlyph-SDXLtext rendering accuracyautomated multi-line layoutsscene text rendering

Abstract

Visual text rendering poses a fundamental challenge for contemporary text-to-image generation models, with the core problem lying in text encoder deficiencies. To achieve accurate text rendering, we identify two crucial requirements for text encoders: character awareness and alignment with glyphs. Our solution involves crafting a series of customized text encoder, Glyph-ByT5, by fine-tuning the character-aware ByT5 encoder using a meticulously curated paired glyph-text dataset. We present an effective method for integrating Glyph-ByT5 with SDXL, resulting in the creation of the Glyph-SDXL model for design image generation. This significantly enhances text rendering accuracy, improving it from less than 20% to nearly 90% on our design image benchmark. Noteworthy is Glyph-SDXL's newfound ability for text paragraph rendering, achieving high spelling accuracy for tens to hundreds of characters with automated multi-line layouts. Finally, through fine-tuning Glyph-SDXL with a small set of high-quality, photorealistic images featuring visual text, we showcase a substantial improvement in scene text rendering capabilities in open-domain real images. These compelling outcomes aim to encourage further exploration in designing customized text encoders for diverse and challenging tasks.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号