TensorX
返回文献探索

Paper · arXiv 2511.03601

Step-Audio-EditX Technical Report

Chao Yan, Boyong Wu, Peng Yang, Pengfei Tan, Guoqiang Hu, Yuxin Zhang, Xiangyu, Zhang, Fei Tian, Xuerui Yang, Xiangyu Zhang, Daxin Jiang, Gang Yu

30 upvotesNovember 5, 2025arXiv 预印本
AI 摘要

Step-Audio-EditX, an open-source LLM-based audio model, excels in expressive and iterative audio editing and zero-shot TTS using large-margin synthetic data.

LLM-based audio modelexpressive audio editingiterative audio editingemotion editingspeaking styleparalinguisticszero-shot text-to-speechlarge-margin synthetic datalarge-margin learningrepresentation-level disentanglement

Abstract

We present Step-Audio-EditX, the first open-source LLM-based audio model excelling at expressive and iterative audio editing encompassing emotion, speaking style, and paralinguistics alongside robust zero-shot text-to-speech (TTS) capabilities.Our core innovation lies in leveraging only large-margin synthetic data, which circumvents the need for embedding-based priors or auxiliary modules. This large-margin learning approach enables both iterative control and high expressivity across voices, and represents a fundamental pivot from the conventional focus on representation-level disentanglement. Evaluation results demonstrate that Step-Audio-EditX surpasses both MiniMax-2.6-hd and Doubao-Seed-TTS-2.0 in emotion editing and other fine-grained control tasks.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号