TensorX
返回文献探索

Paper · arXiv 2408.07246

Seeing and Understanding: Bridging Vision with Chemical Knowledge Via ChemVLM

Junxian Li, Di Zhang, Xunzhi Wang, Zeying Hao, Jingdi Lei, Qian Tan, Cai Zhou, Wei Liu, Weiyun Wang, Zhe Chen, Wenhai Wang, Wei Li, Shufei Zhang, Mao Su, Wanli Ouyang, Yuqiang Li, Dongzhan Zhou

22 upvotesAugust 14, 2024arXiv 预印本
AI 摘要

ChemVLM, a multimodal language model combiningtext and image understanding in chemistry, achieves state-of-the-art performance using VIT-MLP-LLM, ChemLLM-20B, and InternVIT-6B architectures.

VIT-MLP-LLMChemLLM-20BInternVIT-6Bmultimodal language modelchemical text knowledgeimage encoderbilingual multimodal question-answering dataset

Abstract

In this technical report, we propose ChemVLM, the first open-source multimodal large language model dedicated to the fields of chemistry, designed to address the incompatibility between chemical image understanding and text analysis. Built upon the VIT-MLP-LLM architecture, we leverage ChemLLM-20B as the foundational large model, endowing our model with robust capabilities in understanding and utilizing chemical text knowledge. Additionally, we employ InternVIT-6B as a powerful image encoder. We have curated high-quality data from the chemical domain, including molecules, reaction formulas, and chemistry examination data, and compiled these into a bilingual multimodal question-answering dataset. We test the performance of our model on multiple open-source benchmarks and three custom evaluation sets. Experimental results demonstrate that our model achieves excellent performance, securing state-of-the-art results in five out of six involved tasks. Our model can be found at https://huggingface.co/AI4Chem/ChemVLM-26B.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Seeing and Understanding: Bridging Vision with Chemical Knowledge Via ChemVLM | TensorX