TensorX
返回文献探索

Paper · arXiv 2510.04162

Drax: Speech Recognition with Discrete Flow Matching

Aviv Navon, Aviv Shamsian, Neta Glazer, Yael Segal-Feldman, Gill Hetz, Joseph Keshet, Ethan Fetaya

28 upvotesOctober 5, 2025arXiv 预印本
AI 摘要

Drax, a discrete flow matching framework for ASR, achieves state-of-the-art recognition accuracy with improved efficiency by constructing an audio-conditioned probability path.

diffusion modelsflow-based modelsnon-autoregressive modelsdiscrete flow matchingaudio-conditioned probability pathgeneralization gapdivergencescumulative velocity errorsASRrecognition accuracyaccuracy-efficiency trade-offs

Abstract

Diffusion and flow-based non-autoregressive (NAR) models have shown strong promise in large language modeling, however, their potential for automatic speech recognition (ASR) remains largely unexplored. We propose Drax, a discrete flow matching framework for ASR that enables efficient parallel decoding. To better align training with inference, we construct an audio-conditioned probability path that guides the model through trajectories resembling likely intermediate inference errors, rather than direct random noise to target transitions. Our theoretical analysis links the generalization gap to divergences between training and inference occupancies, controlled by cumulative velocity errors, thereby motivating our design choice. Empirical evaluation demonstrates that our approach attains recognition accuracy on par with state-of-the-art speech models while offering improved accuracy-efficiency trade-offs, highlighting discrete flow matching as a promising direction for advancing NAR ASR.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Drax: Speech Recognition with Discrete Flow Matching | TensorX