TensorX
返回文献探索

Paper · arXiv 2307.06350

T2I-CompBench: A Comprehensive Benchmark for Open-world Compositional Text-to-image Generation

Kaiyi Huang, Kaiyue Sun, Enze Xie, Zhenguo Li, Xihui Liu

7 upvotesJuly 12, 2023arXiv 预印本
AI 摘要

T2I-CompBench is a benchmark for assessing compositional text-to-image generation, introducing evaluation metrics and GORS, a reward-driven fine-tuning method.

text-to-image modelsGenerative mOdel fine-tuning with Reward-driven Sample selectionGORScompositional text-to-image generationattribute bindingobject relationshipscomplex compositionscolor bindingshape bindingtexture bindingspatial relationshipsnon-spatial relationships

Abstract

Despite the stunning ability to generate high-quality images by recent text-to-image models, current approaches often struggle to effectively compose objects with different attributes and relationships into a complex and coherent scene. We propose T2I-CompBench, a comprehensive benchmark for open-world compositional text-to-image generation, consisting of 6,000 compositional text prompts from 3 categories (attribute binding, object relationships, and complex compositions) and 6 sub-categories (color binding, shape binding, texture binding, spatial relationships, non-spatial relationships, and complex compositions). We further propose several evaluation metrics specifically designed to evaluate compositional text-to-image generation. We introduce a new approach, Generative mOdel fine-tuning with Reward-driven Sample selection (GORS), to boost the compositional text-to-image generation abilities of pretrained text-to-image models. Extensive experiments and evaluations are conducted to benchmark previous methods on T2I-CompBench, and to validate the effectiveness of our proposed evaluation metrics and GORS approach. Project page is available at https://karine-h.github.io/T2I-CompBench/.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号