Qwen2.5-Omni Technical Report
Jin Xu, Zhifang Guo, Jinzheng He +11 authors
Qwen2.5-Omni is a multimodal model that processes text, images, audio, and video in a streaming fashion and generates text and speech using a dual-track architecture, achieving state-of-the-art performance on multimodal benchmarks.