TensorX
返回文献探索

Paper · arXiv 2502.13449

Mol-LLaMA: Towards General Understanding of Molecules in Large Molecular Language Model

Dongki Kim, Wonbin Lee, Sung Ju Hwang

44 upvotesFebruary 19, 2025arXiv 预印本
AI 摘要

Mol-LLaMA, a multi-modal instruction-tuned large molecular language model, integrates various molecular encoders to understand general features and provide detailed responses for molecular analysis tasks.

large molecular language modelmulti-modal instruction tuningmolecular structuresmolecular encodersmolecular representations

Abstract

Understanding molecules is key to understanding organisms and driving advances in drug discovery, requiring interdisciplinary knowledge across chemistry and biology. Although large molecular language models have achieved notable success in interpreting molecular structures, their instruction datasets are limited to the specific knowledge from task-oriented datasets and do not fully cover the fundamental characteristics of molecules, hindering their abilities as general-purpose molecular assistants. To address this issue, we propose Mol-LLaMA, a large molecular language model that grasps the general knowledge centered on molecules via multi-modal instruction tuning. To this end, we design key data types that encompass the fundamental features of molecules, incorporating essential knowledge from molecular structures. In addition, to improve understanding of molecular features, we introduce a module that integrates complementary information from different molecular encoders, leveraging the distinct advantages of different molecular representations. Our experimental results demonstrate that Mol-LLaMA is capable of comprehending the general features of molecules and generating relevant responses to users' queries with detailed explanations, implying its potential as a general-purpose assistant for molecular analysis.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Mol-LLaMA: Towards General Understanding of Molecules in Large Molecular Language Model | TensorX