TensorX
返回文献探索

Paper · arXiv 2502.21318

How far can we go with ImageNet for Text-to-Image generation?

L. Degeorge, A. Ghosh, N. Dufour, D. Picard, V. Kalogeiton

26 upvotesFebruary 28, 2025arXiv 预印本
AI 摘要

Strategic data augmentation of small, well-curated datasets can achieve competitive performance in text-to-image generation with significantly fewer parameters and training images compared to large-scale models.

text-to-image generationImageNetdata augmentationSD-XLGenEvalDPGBench

Abstract

Recent text-to-image (T2I) generation models have achieved remarkable results by training on billion-scale datasets, following a `bigger is better' paradigm that prioritizes data quantity over quality. We challenge this established paradigm by demonstrating that strategic data augmentation of small, well-curated datasets can match or outperform models trained on massive web-scraped collections. Using only ImageNet enhanced with well-designed text and image augmentations, we achieve a +2 overall score over SD-XL on GenEval and +5 on DPGBench while using just 1/10th the parameters and 1/1000th the training images. Our results suggest that strategic data augmentation, rather than massive datasets, could offer a more sustainable path forward for T2I generation.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
How far can we go with ImageNet for Text-to-Image generation? | TensorX