TensorX
返回文献探索

Paper · arXiv 2410.01731

ComfyGen: Prompt-Adaptive Workflows for Text-to-Image Generation

Rinon Gal, Adi Haviv, Yuval Alaluf, Amit H. Bermano, Daniel Cohen-Or, Gal Chechik

16 upvotesOctober 2, 2024arXiv 预印本
AI 摘要

Automatic generation of text-to-image workflows based on user prompts enhances image quality compared to monolithic models or generic workflows.

prompt-adaptive workflow generationLLM-based approachestuning-based methodtraining-free methodprompt-dependent flow prediction

Abstract

The practical use of text-to-image generation has evolved from simple, monolithic models to complex workflows that combine multiple specialized components. While workflow-based approaches can lead to improved image quality, crafting effective workflows requires significant expertise, owing to the large number of available components, their complex inter-dependence, and their dependence on the generation prompt. Here, we introduce the novel task of prompt-adaptive workflow generation, where the goal is to automatically tailor a workflow to each user prompt. We propose two LLM-based approaches to tackle this task: a tuning-based method that learns from user-preference data, and a training-free method that uses the LLM to select existing flows. Both approaches lead to improved image quality when compared to monolithic models or generic, prompt-independent workflows. Our work shows that prompt-dependent flow prediction offers a new pathway to improving text-to-image generation quality, complementing existing research directions in the field.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
ComfyGen: Prompt-Adaptive Workflows for Text-to-Image Generation | TensorX