TensorX
返回文献探索

Paper · arXiv 2402.07865

Prismatic VLMs: Investigating the Design Space of Visually-Conditioned Language Models

Siddharth Karamcheti, Suraj Nair, Ashwin Balakrishna, Percy Liang, Thomas Kollar, Dorsa Sadigh

14 upvotesFebruary 12, 2024arXiv 预印本
AI 摘要

The paper provides standardized evaluations and investigations into visually-conditioned language models (VLMs), exploring factors affecting their performance and offering resources like a unified evaluation framework and optimized training code.

visually-conditioned language modelsVLMsvisual dialoguescene understandingrobotic task planningLLaVaInstructBLIPPaLI-3image preprocessingarchitectureoptimizationvisual question answeringobject localizationhallucinationpretrained visual representationsinstruct-tuned language modelsunified frameworkVLM trainingcheckpointsstate-of-the-art

Abstract

Visually-conditioned language models (VLMs) have seen growing adoption in applications such as visual dialogue, scene understanding, and robotic task planning; adoption that has fueled a wealth of new models such as LLaVa, InstructBLIP, and PaLI-3. Despite the volume of new releases, key design decisions around image preprocessing, architecture, and optimization are under-explored, making it challenging to understand what factors account for model performance - a challenge further complicated by the lack of objective, consistent evaluations. To address these gaps, we first compile a suite of standardized evaluations spanning visual question answering, object localization from language, and targeted challenge sets that probe properties such as hallucination; evaluations that provide calibrated, fine-grained insight into a VLM's capabilities. Second, we rigorously investigate VLMs along key design axes, including pretrained visual representations and quantifying the tradeoffs of using base vs. instruct-tuned language models, amongst others. We couple our analysis with three resource contributions: (1) a unified framework for evaluating VLMs, (2) optimized, flexible code for VLM training, and (3) checkpoints for all models, including a family of VLMs at the 7-13B scale that strictly outperform InstructBLIP and LLaVa v1.5, the state-of-the-art in open-source VLMs.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Prismatic VLMs: Investigating the Design Space of Visually-Conditioned Language Models | TensorX