TensorX
返回文献探索

Paper · arXiv 2402.05937

InstaGen: Enhancing Object Detection by Training on Synthetic Dataset

Chengjian Feng, Yujie Zhong, Zequn Jie, Weidi Xie, Lin Ma

13 upvotesFebruary 8, 2024arXiv 预印本
AI 摘要

A diffusion model enhanced with an instance-level grounding head serves as a synthetic data generator to improve object detection performance, especially in open-vocabulary and data-sparse scenarios.

diffusion modelsinstance-level groundingtext embeddingregional visual featureself-trainingInstaGenobject detectionsynthetic datasetopen-vocabularydata-sparse

Abstract

In this paper, we introduce a novel paradigm to enhance the ability of object detector, e.g., expanding categories or improving detection performance, by training on synthetic dataset generated from diffusion models. Specifically, we integrate an instance-level grounding head into a pre-trained, generative diffusion model, to augment it with the ability of localising arbitrary instances in the generated images. The grounding head is trained to align the text embedding of category names with the regional visual feature of the diffusion model, using supervision from an off-the-shelf object detector, and a novel self-training scheme on (novel) categories not covered by the detector. This enhanced version of diffusion model, termed as InstaGen, can serve as a data synthesizer for object detection. We conduct thorough experiments to show that, object detector can be enhanced while training on the synthetic dataset from InstaGen, demonstrating superior performance over existing state-of-the-art methods in open-vocabulary (+4.5 AP) and data-sparse (+1.2 to 5.2 AP) scenarios.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号