TensorX
返回文献探索

Paper · arXiv 2508.02317

VeOmni: Scaling Any Modality Model Training with Model-Centric Distributed Recipe Zoo

Qianli Ma, Yaowei Zheng, Zhelun Shi, Zhongkai Zhao, Bin Jia, Ziyue Huang, Zhiqi Lin, Youjie Li, Jiacheng Yang, Yanghua Peng, Zhi Zhang, Xin Liu

23 upvotesAugust 4, 2025arXiv 预印本
AI 摘要

A modular training framework accelerates the development of omni-modal LLMs through efficient 3D parallelism and flexible configuration.

large language modelsomni-modal understandingomni-modal generationheterogeneous model architecturesparallel logicmodel-centric distributed recipescommunicationcomputation3D parallelismflexible configuration interfacemixture-of-expertstokens/sec/GPU throughputcontext lengths

Abstract

Recent advances in large language models (LLMs) have driven impressive progress in omni-modal understanding and generation. However, training omni-modal LLMs remains a significant challenge due to the heterogeneous model architectures required to process diverse modalities, necessitating sophisticated system design for efficient large-scale training. Existing frameworks typically entangle model definition with parallel logic, incurring limited scalability and substantial engineering overhead for end-to-end omni-modal training. % We present \veomni, a modular and efficient training framework to accelerate the development of omni-modal LLMs. \veomni introduces model-centric distributed recipes that decouples communication from computation, enabling efficient 3D parallelism on omni-modal LLMs. \veomni also features a flexible configuration interface supporting seamless integration of new modalities with minimal code change. % Using \veomni, a omni-modal mixture-of-experts (MoE) model with 30B parameters can be trained with over 2,800 tokens/sec/GPU throughput and scale to 160K context lengths via 3D parallelism on 128 GPUs, showcasing its superior efficiency and scalability for training large omni-modal LLMs.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
VeOmni: Scaling Any Modality Model Training with Model-Centric Distributed Recipe Zoo | TensorX