TensorX
返回文献探索

Paper · arXiv 2306.03514

Recognize Anything: A Strong Image Tagging Model

Youcai Zhang, Xinyu Huang, Jinyu Ma, Zhaoyang Li, Zhaochuan Luo, Yanchun Xie, Yuzhuo Qin, Tong Luo, Yaqian Li, Shilong Liu, Yandong Guo, Lei Zhang

12 upvotesJune 6, 2023arXiv 预印本
AI 摘要

The Recognize Anything Model (RAM) is a powerful image tagging model trained using large-scale image-text pairs, achieving superior zero-shot performance compared to supervised models and Google API.

foundation modelimage tagginglarge-scale image-text pairsannotation-free image tagstext semantic parsingcaption and tagging tasksdata enginezero-shot performanceCLIPBLIP

Abstract

We present the Recognize Anything Model (RAM): a strong foundation model for image tagging. RAM can recognize any common category with high accuracy. RAM introduces a new paradigm for image tagging, leveraging large-scale image-text pairs for training instead of manual annotations. The development of RAM comprises four key steps. Firstly, annotation-free image tags are obtained at scale through automatic text semantic parsing. Subsequently, a preliminary model is trained for automatic annotation by unifying the caption and tagging tasks, supervised by the original texts and parsed tags, respectively. Thirdly, a data engine is employed to generate additional annotations and clean incorrect ones. Lastly, the model is retrained with the processed data and fine-tuned using a smaller but higher-quality dataset. We evaluate the tagging capabilities of RAM on numerous benchmarks and observe impressive zero-shot performance, significantly outperforming CLIP and BLIP. Remarkably, RAM even surpasses the fully supervised manners and exhibits competitive performance with the Google API. We are releasing the RAM at https://recognize-anything.github.io/ to foster the advancements of large models in computer vision.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Recognize Anything: A Strong Image Tagging Model | TensorX