TensorX
返回文献探索

Paper · arXiv 2411.17176

ChatGen: Automatic Text-to-Image Generation From FreeStyle Chatting

Chengyou Jia, Changliang Xia, Zhuohang Dang, Weijia Wu, Hangwei Qian, Minnan Luo

24 upvotesNovember 26, 2024arXiv 预印本
AI 摘要

ChatGen-Evo, a multi-stage evolution strategy, automates text-to-image generation by improving accuracy and image quality through systematic evaluation on ChatGenBench.

text-to-imageAutomatic T2IChatGenBenchChatGen-Evomulti-stage evolution strategy

Abstract

Despite the significant advancements in text-to-image (T2I) generative models, users often face a trial-and-error challenge in practical scenarios. This challenge arises from the complexity and uncertainty of tedious steps such as crafting suitable prompts, selecting appropriate models, and configuring specific arguments, making users resort to labor-intensive attempts for desired images. This paper proposes Automatic T2I generation, which aims to automate these tedious steps, allowing users to simply describe their needs in a freestyle chatting way. To systematically study this problem, we first introduce ChatGenBench, a novel benchmark designed for Automatic T2I. It features high-quality paired data with diverse freestyle inputs, enabling comprehensive evaluation of automatic T2I models across all steps. Additionally, recognizing Automatic T2I as a complex multi-step reasoning task, we propose ChatGen-Evo, a multi-stage evolution strategy that progressively equips models with essential automation skills. Through extensive evaluation across step-wise accuracy and image quality, ChatGen-Evo significantly enhances performance over various baselines. Our evaluation also uncovers valuable insights for advancing automatic T2I. All our data, code, and models will be available in https://chengyou-jia.github.io/ChatGen-Home

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
ChatGen: Automatic Text-to-Image Generation From FreeStyle Chatting | TensorX