TensorX
返回文献探索

Paper · arXiv 2403.12596

Chart-based Reasoning: Transferring Capabilities from LLMs to VLMs

Victor Carbune, Hassan Mansoor, Fangyu Liu, Rahul Aralikatte, Gilles Baechler, Jindong Chen, Abhanshu Sharma

12 upvotesMarch 19, 2024arXiv 预印本
AI 摘要

A method transfers reasoning capabilities from large-language models to vision-language models, achieving state-of-the-art performance on ChartQA and superior performance on PlotQA and FigureQA by enhancing chart representation and training with a larger, synthesized dataset and multitask loss.

vision-language modelslarge-language modelsChartQAPaLI3-5BplotQAFigureQAchart-to-table translationreasoning tracesmultitask lossChartPaLI-5BPaLIX-55BOCR systemprogram-of-thought promptGemini UltraGPT-4V

Abstract

Vision-language models (VLMs) are achieving increasingly strong performance on multimodal tasks. However, reasoning capabilities remain limited particularly for smaller VLMs, while those of large-language models (LLMs) have seen numerous improvements. We propose a technique to transfer capabilities from LLMs to VLMs. On the recently introduced ChartQA, our method obtains state-of-the-art performance when applied on the PaLI3-5B VLM by chen2023pali3, while also enabling much better performance on PlotQA and FigureQA. We first improve the chart representation by continuing the pre-training stage using an improved version of the chart-to-table translation task by liu2023deplot. We then propose constructing a 20x larger dataset than the original training set. To improve general reasoning capabilities and improve numerical operations, we synthesize reasoning traces using the table representation of charts. Lastly, our model is fine-tuned using the multitask loss introduced by hsieh2023distilling. Our variant ChartPaLI-5B outperforms even 10x larger models such as PaLIX-55B without using an upstream OCR system, while keeping inference time constant compared to the PaLI3-5B baseline. When rationales are further refined with a simple program-of-thought prompt chen2023program, our model outperforms the recently introduced Gemini Ultra and GPT-4V.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号