TensorX
返回文献探索

Paper · arXiv 2607.10387

GigaChat Audio: Time-aware Large Audio Language Model

Aleksandr Kutsakov, Mariia Sadovina, Georgii Gospodinov, Alexandr Maximenko, Oleg Kutuzov, Pavel Bogomolov, Fyodor Minkin

37 upvotesJuly 11, 2026arXiv 预印本
AI 摘要

A time-aware audio LLM answers questions with explicit timestamps across long recordings by interleaving periodic time markers with continuous audio tokens, achieving strong temporal grounding and supporting time-anchored descriptions.

audio-conditioned LLMstemporal groundingtime-aware audio LLMtime markersaudio tokenssynthetic supervisiontime-anchored fragment descriptionstime representationtokenizationduration-mixture design

Abstract

Temporal grounding in long recordings remains challenging for audio-conditioned LLMs. We present a time-aware audio LLM that answers questions with explicit timestamps over up to 120 minutes of input. Our approach interleaves periodic time markers with continuous audio tokens using large-scale synthetic supervision from a cascaded pipeline. Our model achieves strong temporal-grounding accuracy on short and long benchmarks and supports time-anchored fragment descriptions and summaries. Extensive ablations examine how time representation, marker frequency, tokenization, and duration-mixture design affect accuracy and computational cost. We release model weights and datasets to support further research on time-aware audio understanding, available at https://huggingface.co/ai-sage/GigaChat3.1-Audio-10B-A1.8B.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
GigaChat Audio: Time-aware Large Audio Language Model | TensorX