TensorX
返回文献探索

Paper · arXiv 2504.20996

X-Fusion: Introducing New Modality to Frozen Large Language Models

Sicheng Mo, Thao Nguyen, Xun Huang, Siddharth Srinivasan Iyer, Yijun Li, Yuchen Liu, Abhishek Tandon, Eli Shechtman, Krishna Kumar Singh, Yong Jae Lee, Bolei Zhou, Yuheng Li

13 upvotesApril 29, 2025arXiv 预印本
AI 摘要

X-Fusion enhances pretrained LLMs for multimodal tasks by integrating vision-specific information through a dual-tower design and freezes the LLM's parameters for consistency in understanding and generation.

Large Language Models (LLMs)multimodal tasksdual-tower designmodality-specific weightsvision-specific informationimage-to-texttext-to-imageunderstanding-focused dataimage data noisefeature alignment

Abstract

We propose X-Fusion, a framework that extends pretrained Large Language Models (LLMs) for multimodal tasks while preserving their language capabilities. X-Fusion employs a dual-tower design with modality-specific weights, keeping the LLM's parameters frozen while integrating vision-specific information for both understanding and generation. Our experiments demonstrate that X-Fusion consistently outperforms alternative architectures on both image-to-text and text-to-image tasks. We find that incorporating understanding-focused data improves generation quality, reducing image data noise enhances overall performance, and feature alignment accelerates convergence for smaller models but has minimal impact on larger ones. Our findings provide valuable insights into building efficient unified multimodal models.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号