TensorX
返回文献探索

Paper · arXiv 2506.23009

MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models

Jian Chen, Wenye Ma, Penghang Liu, Wei Wang, Tengwei Song, Ming Li, Chenguang Wang, Ruiyi Zhang, Changyou Chen

11 upvotesJune 28, 2025arXiv 预印本
AI 摘要

MusiXQA, a new dataset for evaluating MLLMs on music sheet understanding, reveals limitations and enables the development of Phi-3-MusiX, an improved MLLM for this task.

Multimodal Large Language ModelsMLLMsMusiXQAMusiXTeXvisual QA tasksPhi-3-MusiXGPT-based methods

Abstract

Multimodal Large Language Models (MLLMs) have achieved remarkable visual reasoning abilities in natural images, text-rich documents, and graphic designs. However, their ability to interpret music sheets remains underexplored. To bridge this gap, we introduce MusiXQA, the first comprehensive dataset for evaluating and advancing MLLMs in music sheet understanding. MusiXQA features high-quality synthetic music sheets generated via MusiXTeX, with structured annotations covering note pitch and duration, chords, clefs, key/time signatures, and text, enabling diverse visual QA tasks. Through extensive evaluations, we reveal significant limitations of current state-of-the-art MLLMs in this domain. Beyond benchmarking, we developed Phi-3-MusiX, an MLLM fine-tuned on our dataset, achieving significant performance gains over GPT-based methods. The proposed dataset and model establish a foundation for future advances in MLLMs for music sheet understanding. Code, data, and model will be released upon acceptance.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models | TensorX