TensorX
返回文献探索

Paper · arXiv 2404.07973

Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models

Haotian Zhang, Haoxuan You, Philipp Dufter, Bowen Zhang, Chen Chen, Hong-You Chen, Tsu-Jui Fu, William Yang Wang, Shih-Fu Chang, Zhe Gan, Yinfei Yang

33 upvotesApril 11, 2024arXiv 预印本
AI 摘要

Ferret-v2 improves on Ferret by incorporating flexible high-resolution grounding, multi-granularity visual encoding with DINOv2, and a three-stage training paradigm, leading to enhanced visual understanding and performance.

regional understandingLarge Language Model (LLM)visual encoderDINOv2multi-granularity visual encodinghigh-resolution groundingthree-stage training paradigmimage-caption alignmenthigh-resolution dense alignment

Abstract

While Ferret seamlessly integrates regional understanding into the Large Language Model (LLM) to facilitate its referring and grounding capability, it poses certain limitations: constrained by the pre-trained fixed visual encoder and failed to perform well on broader tasks. In this work, we unveil Ferret-v2, a significant upgrade to Ferret, with three key designs. (1) Any resolution grounding and referring: A flexible approach that effortlessly handles higher image resolution, improving the model's ability to process and understand images in greater detail. (2) Multi-granularity visual encoding: By integrating the additional DINOv2 encoder, the model learns better and diverse underlying contexts for global and fine-grained visual information. (3) A three-stage training paradigm: Besides image-caption alignment, an additional stage is proposed for high-resolution dense alignment before the final instruction tuning. Experiments show that Ferret-v2 provides substantial improvements over Ferret and other state-of-the-art methods, thanks to its high-resolution scaling and fine-grained visual processing.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models | TensorX