Step-Audio 2 Technical Report
Boyong Wu, Chao Yan, Chen Hu +106 authors
Step-Audio~2, an end-to-end multi-modal large language model, integrates latent audio encoding and reinforcement learning to achieve state-of-the-art performance in ASR, audio understanding, and speech conversation, incorporating discrete audio token generation and retrieval-augmented generation.