TensorX
返回文献探索

Paper · arXiv 2608.03812

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models

Wanshun Su, Yang Shi, Feihu Liu, Ziwen Yu, Yan Min, Zhuoran Zhang, Qixun Wang, Haotian Wang, Shixuan Liu, Yuanxing Zhang, Peng Wu, Chengfu Huo, Liang Ding

28 upvotesAugust 4, 2026arXiv 预印本
AI 摘要

OmniPack is a training-free framework that combines structural pre-LLM token compression with query-guided semantic refinement inside the LLM to reduce computational overhead in omni-modal models while preserving performance.

omni-modal large language modelstoken compressionpre-LLM compressioninner-LLM compressionstructural redundancymodality-specific importanceglobal coveragesimilarity-aware mergingaudio-visual collaborationtextual guidanceFLOPs

Abstract

Omni-modal large language models (Omni-LLMs) have achieved remarkable performance on audio-visual understanding tasks, but processing long and highly redundant visual and audio token sequences incurs substantial computational overhead, demanding aggressive token compression for efficient deployment. Existing methods often degrade at low token budgets: pre-LLM compression may discard structurally important and globally distributed evidence, whereas inner-LLM compression often underexploits query-conditioned audio-visual collaboration. To address these limitations, we propose OmniPack, a training-free framework that coordinates structural compression before the LLM with task-relevant semantic refinement within the LLM. Before the LLM, OmniPack removes structural redundancy through modality-specific importance, global coverage, and similarity-aware merging. After sufficient multimodal interaction, it further consolidates diverse, task-relevant representations through textual guidance and audio-visual collaboration. Extensive experiments on five benchmarks with three Omni-LLM backbones demonstrate that OmniPack consistently achieves the best performance-efficiency trade-off across diverse retention ratios, outperforming all existing methods. Notably, on Qwen2.5-Omni-7B, OmniPack preserves 98.0% of the original performance while reducing FLOPs to 16.7%, and still retains 92.9% of the original performance with only 6.8% of the original FLOPs.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models | TensorX