TensorX
返回文献探索

Paper · arXiv 2410.13782

DPLM-2: A Multimodal Diffusion Protein Language Model

Xinyou Wang, Zaixiang Zheng, Fei Ye, Dongyu Xue, Shujian Huang, Quanquan Gu

21 upvotesOctober 17, 2024arXiv 预印本
AI 摘要

DPLM-2 is a multimodal protein foundation model that simultaneously generates amino acid sequences and their 3D structures using a discrete diffusion approach, improving performance in tasks like folding and scaffold prediction.

discrete diffusion protein language modelDPLMmultimodal protein foundation model3D coordinatesdiscrete tokenslookup-free quantization-based tokenizerjoint distributionsequence-based protein language modelswarm-up strategyevolutionary datastructural inductive biasesconditional generation tasksfoldinginverse foldingscaffoldingmultimodal motif inputsstructure-aware representations

Abstract

Proteins are essential macromolecules defined by their amino acid sequences, which determine their three-dimensional structures and, consequently, their functions in all living organisms. Therefore, generative protein modeling necessitates a multimodal approach to simultaneously model, understand, and generate both sequences and structures. However, existing methods typically use separate models for each modality, limiting their ability to capture the intricate relationships between sequence and structure. This results in suboptimal performance in tasks that requires joint understanding and generation of both modalities. In this paper, we introduce DPLM-2, a multimodal protein foundation model that extends discrete diffusion protein language model (DPLM) to accommodate both sequences and structures. To enable structural learning with the language model, 3D coordinates are converted to discrete tokens using a lookup-free quantization-based tokenizer. By training on both experimental and high-quality synthetic structures, DPLM-2 learns the joint distribution of sequence and structure, as well as their marginals and conditionals. We also implement an efficient warm-up strategy to exploit the connection between large-scale evolutionary data and structural inductive biases from pre-trained sequence-based protein language models. Empirical evaluation shows that DPLM-2 can simultaneously generate highly compatible amino acid sequences and their corresponding 3D structures eliminating the need for a two-stage generation approach. Moreover, DPLM-2 demonstrates competitive performance in various conditional generation tasks, including folding, inverse folding, and scaffolding with multimodal motif inputs, as well as providing structure-aware representations for predictive tasks.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号