TensorX
返回文献探索

Paper · arXiv 2607.08168

MuScriptor: An Open Model for Multi-Instrument Music Transcription

Simon Rouard, Michael Krause, Axel Roebel, Carl-Johann Simon-Gabriel, Alexandre Défossez

22 upvotesJuly 9, 2026arXiv 预印本
AI 摘要

A multi-instrument transcription model combines synthetic pre-training, real-audio fine-tuning, reinforcement learning, and instrument conditioning to transcribe diverse real-world music.

automatic music transcriptionsynthetic data pre-trainingfine-tuningreinforcement learninginstrument presence conditioningmulti-instrument transcription

Abstract

Existing methods for automatic music transcription are often limited to single-instrument recordings or fail on complex, real music mixes. Although previous work utilizes synthetic training data, the resulting models generalize poorly, leading to largely unusable transcription output in realistic, multi-instrument settings. In this work, we analyze the effectiveness of synthetic data for pre-training while combining it with fine-tuning on real music audio and post-training using reinforcement learning. We further introduce conditioning on instrument presence to customize transcriptions. Finally, we release MuScriptor, an open-weight multi-instrument music transcription model that works on real-world music recordings from across a diverse range of musical genres.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号