MM1.5: Methods, Analysis & Insights from Multimodal LLM Fine-tuning
Haotian Zhang, Mingfei Gao, Zhe Gan +20 authors
MM1.5 enhances multimodal language models with a data-centric approach, supporting diverse tasks like text-rich image understanding and video comprehension through careful data curation and training strategies.