TensorX
返回文献探索

Paper · arXiv 2508.11598

Representing Speech Through Autoregressive Prediction of Cochlear Tokens

Greta Tuckute, Klemen Kotar, Evelina Fedorenko, Daniel L. K. Yamins

18 upvotesAugust 15, 2025arXiv 预印本
AI 摘要

AuriStream, a biologically inspired two-stage model, encodes speech using cochlear tokens and an autoregressive sequence model, achieving state-of-the-art performance on speech tasks and generating interpretable audio continuations.

cochlear tokensautoregressive sequence modelhuman auditory processing hierarchyhuman cochlealexical semanticsspectrogram space

Abstract

We introduce AuriStream, a biologically inspired model for encoding speech via a two-stage framework inspired by the human auditory processing hierarchy. The first stage transforms raw audio into a time-frequency representation based on the human cochlea, from which we extract discrete cochlear tokens. The second stage applies an autoregressive sequence model over the cochlear tokens. AuriStream learns meaningful phoneme and word representations, and state-of-the-art lexical semantics. AuriStream shows competitive performance on diverse downstream SUPERB speech tasks. Complementing AuriStream's strong representational capabilities, it generates continuations of audio which can be visualized in a spectrogram space and decoded back into audio, providing insights into the model's predictions. In summary, we present a two-stage framework for speech representation learning to advance the development of more human-like models that efficiently handle a range of speech-based tasks.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Representing Speech Through Autoregressive Prediction of Cochlear Tokens | TensorX