TensorX
返回文献探索

Paper · arXiv 2405.09062

Naturalistic Music Decoding from EEG Data via Latent Diffusion Models

Emilian Postolache, Natalia Polouliakh, Hiroaki Kitano, Akima Connelly, Emanuele Rodolà, Taketo Akama

10 upvotesMay 15, 2024arXiv 预印本
AI 摘要

Latent diffusion models are employed to reconstruct naturalistic music from EEG recordings using an end-to-end training approach and evaluating with neural embedding-based metrics.

latent diffusion modelsgenerative modelsmusic reconstructionEEG recordingsNMED-T datasetneural embedding-based metricssong classificationneural decodingbrain-computer interfaces

Abstract

In this article, we explore the potential of using latent diffusion models, a family of powerful generative models, for the task of reconstructing naturalistic music from electroencephalogram (EEG) recordings. Unlike simpler music with limited timbres, such as MIDI-generated tunes or monophonic pieces, the focus here is on intricate music featuring a diverse array of instruments, voices, and effects, rich in harmonics and timbre. This study represents an initial foray into achieving general music reconstruction of high-quality using non-invasive EEG data, employing an end-to-end training approach directly on raw data without the need for manual pre-processing and channel selection. We train our models on the public NMED-T dataset and perform quantitative evaluation proposing neural embedding-based metrics. We additionally perform song classification based on the generated tracks. Our work contributes to the ongoing research in neural decoding and brain-computer interfaces, offering insights into the feasibility of using EEG data for complex auditory information reconstruction.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Naturalistic Music Decoding from EEG Data via Latent Diffusion Models | TensorX