TensorX
返回文献探索

Paper · arXiv 2504.16427

Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark

Hanlei Zhang, Zhuohang Li, Yeshuang Zhu, Hua Xu, Peiwu Wang, Haige Zhu, Jie Zhou, Jinchao Zhang

18 upvotesApril 23, 2025arXiv 预印本
AI 摘要

MMLA benchmark assesses multimodal large language models' understanding of human language semantics across various core dimensions, highlighting limitations and offering a foundation for future research.

multimodal language modelsbenchmarkmultimodal semanticszero-shot inferencesupervised fine-tuninginstruction tuningsentimentspeaking stylecommunication behavior

Abstract

Multimodal language analysis is a rapidly evolving field that leverages multiple modalities to enhance the understanding of high-level semantics underlying human conversational utterances. Despite its significance, little research has investigated the capability of multimodal large language models (MLLMs) to comprehend cognitive-level semantics. In this paper, we introduce MMLA, a comprehensive benchmark specifically designed to address this gap. MMLA comprises over 61K multimodal utterances drawn from both staged and real-world scenarios, covering six core dimensions of multimodal semantics: intent, emotion, dialogue act, sentiment, speaking style, and communication behavior. We evaluate eight mainstream branches of LLMs and MLLMs using three methods: zero-shot inference, supervised fine-tuning, and instruction tuning. Extensive experiments reveal that even fine-tuned models achieve only about 60%~70% accuracy, underscoring the limitations of current MLLMs in understanding complex human language. We believe that MMLA will serve as a solid foundation for exploring the potential of large language models in multimodal language analysis and provide valuable resources to advance this field. The datasets and code are open-sourced at https://github.com/thuiar/MMLA.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark | TensorX