TensorX
返回文献探索

Paper · arXiv 2504.07951

Scaling Laws for Native Multimodal Models Scaling Laws for Native Multimodal Models

Mustafa Shukor, Enrico Fini, Victor Guilherme Turrisi da Costa, Matthieu Cord, Joshua Susskind, Alaaeldin El-Nouby

31 upvotesApril 10, 2025arXiv 预印本
AI 摘要

Early-fusion architectures, which do not rely on image encoders, outperform late-fusion models in native multimodal architectures, especially when enhanced with Mixture of Experts (MoEs).

native multimodal modelsearly-fusionlate-fusionMixture of Experts

Abstract

Building general-purpose models that can effectively perceive the world through multimodal signals has been a long-standing goal. Current approaches involve integrating separately pre-trained components, such as connecting vision encoders to LLMs and continuing multimodal training. While such approaches exhibit remarkable sample efficiency, it remains an open question whether such late-fusion architectures are inherently superior. In this work, we revisit the architectural design of native multimodal models (NMMs)--those trained from the ground up on all modalities--and conduct an extensive scaling laws study, spanning 457 trained models with different architectures and training mixtures. Our investigation reveals no inherent advantage to late-fusion architectures over early-fusion ones, which do not rely on image encoders. On the contrary, early-fusion exhibits stronger performance at lower parameter counts, is more efficient to train, and is easier to deploy. Motivated by the strong performance of the early-fusion architectures, we show that incorporating Mixture of Experts (MoEs) allows for models that learn modality-specific weights, significantly enhancing performance.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号