TensorX
返回文献探索

Paper · arXiv 2305.06324

Alternating Gradient Descent and Mixture-of-Experts for Integrated Multimodal Perception

Hassan Akbari, Dan Kondratyuk, Yin Cui, Rachel Hornung, Huisheng Wang, Hartwig Adam

1 upvotesMay 10, 2023arXiv 预印本
AI 摘要

Integrated Multimodal Perception (IMP) combines Alternating Gradient Descent and Mixture-of-Experts in a single Transformer encoder to improve multimodal understanding and achieve state-of-the-art zero-shot video classification.

Integrated Multimodal PerceptionIMPTransformer encoderAlternating Gradient DescentAGDMixture-of-ExpertsMoEmultimodal understandingimage classificationvideo classificationimage-text retrievalvideo-text retrievalzero-shot video classificationKinetics-400Kinetics-600Kinetics-700

Abstract

We present Integrated Multimodal Perception (IMP), a simple and scalable multimodal multi-task training and modeling approach. IMP integrates multimodal inputs including image, video, text, and audio into a single Transformer encoder with minimal modality-specific components. IMP makes use of a novel design that combines Alternating Gradient Descent (AGD) and Mixture-of-Experts (MoE) for efficient model \& task scaling. We conduct extensive empirical studies about IMP and reveal the following key insights: 1) performing gradient descent updates by alternating on diverse heterogeneous modalities, loss functions, and tasks, while also varying input resolutions, efficiently improves multimodal understanding. 2) model sparsification with MoE on a single modality-agnostic encoder substantially improves the performance, outperforming dense models that use modality-specific encoders or additional fusion layers and greatly mitigating the conflicts between modalities. IMP achieves competitive performance on a wide range of downstream tasks including image classification, video classification, image-text, and video-text retrieval. Most notably, we train a sparse IMP-MoE-L focusing on video tasks that achieves new state-of-the-art in zero-shot video classification. Our model achieves 77.0% on Kinetics-400, 76.8% on Kinetics-600, and 76.8% on Kinetics-700 zero-shot classification accuracy, improving the previous state-of-the-art by +5%, +6.7%, and +5.8%, respectively, while using only 15% of their total training computational cost.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Alternating Gradient Descent and Mixture-of-Experts for Integrated Multimodal Perception | TensorX