TensorX
返回文献探索

Paper · arXiv 2406.08478

What If We Recaption Billions of Web Images with LLaMA-3?

Xianhang Li, Haoqin Tu, Mude Hui, Zeyu Wang, Bingchen Zhao, Junfei Xiao, Sucheng Ren, Jieru Mei, Qing Liu, Huangjie Zheng, Yuyin Zhou, Cihang Xie

43 upvotesJune 12, 2024arXiv 预印本
AI 摘要

Enhancing image-text pairs through recaptioning with LLaMA-3 improves model performance in vision-language tasks, including zero-shot retrieval and text-to-image generation.

LLaMA-3LLaVA-1.5DataComp-1BRecap-DataComp-1BCLIPcross-modal retrievaltext-to-image Diffusion Transformers

Abstract

Web-crawled image-text pairs are inherently noisy. Prior studies demonstrate that semantically aligning and enriching textual descriptions of these pairs can significantly enhance model training across various vision-language tasks, particularly text-to-image generation. However, large-scale investigations in this area remain predominantly closed-source. Our paper aims to bridge this community effort, leveraging the powerful and open-sourced LLaMA-3, a GPT-4 level LLM. Our recaptioning pipeline is simple: first, we fine-tune a LLaMA-3-8B powered LLaVA-1.5 and then employ it to recaption 1.3 billion images from the DataComp-1B dataset. Our empirical results confirm that this enhanced dataset, Recap-DataComp-1B, offers substantial benefits in training advanced vision-language models. For discriminative models like CLIP, we observe enhanced zero-shot performance in cross-modal retrieval tasks. For generative models like text-to-image Diffusion Transformers, the generated images exhibit a significant improvement in alignment with users' text instructions, especially in following complex queries. Our project page is https://www.haqtu.me/Recap-Datacomp-1B/

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
What If We Recaption Billions of Web Images with LLaMA-3? | TensorX