InternVL 2.5, an advanced multimodal large language model, showcases competitive performance across various benchmarks, including multimodal reasoning and understanding, and is the first open-source model to surpass 70% on the MMMU benchmark using Chain-of-Thought reasoning.
multimodal large language modelvision encoderslanguage modelsdataset sizesHF ↗arXiv ↗
A 14-billion parameter language model surpasses its teacher model in STEM-focused QA through strategic use of synthetic data, improved data quality, and enhanced training techniques.
Shibo Hao, Sainbayar Sukhbaatar, DiJia Su +4 authors
A new reasoning paradigm, Coconut, utilizes continuous representations in latent space to enhance LLM performance on complex reasoning tasks, particularly those requiring backtracking.
ProcessBench evaluates models' ability to identify errors in mathematical reasoning steps, showing that existing process reward models struggle with difficult problems and underperform compared to critic models and a fine-tuned PRM.
STIV, a text-image-conditioned video generation method integrating Diffusion Transformer and classifier-free guidance, achieves state-of-the-art performance in text-to-video, text-image-to-video, and image-to-video tasks.
The paper provides a framework for understanding and evaluating different types of memory in reinforcement learning agents, using cognitive science-inspired definitions and a standardized experimental methodology.
A benchmark called Geoperception is introduced to evaluate MLLMs' geometric description accuracy, leading to the development of Euclid, a model optimized for low-level geometric perception using synthetic data and a data curriculum.
53multimodal large language modelslow-level visual perceptionHF ↗arXiv ↗
LG AI Research, Soyoung An, Kyunghoon Bae +30 authors
EXAONE 3.5 language models, available in three configurations, demonstrate exceptional instruction following, long-context comprehension, and competitive performance across various benchmarks.
51instruction-tuned language modelslong-context comprehensionHF ↗arXiv ↗
LatentLM integrates continuous and discrete data using causal Transformers, VAEs, and next-token diffusion, achieving superior performance in multimodal tasks like image generation, large language model integration, and text-to-speech synthesis.
Zhisheng Zhong, Chengyao Wang, Yuqi Liu +12 authors
As Multi-modal Large Language Models (MLLMs) evolve, expanding beyond
single-domain capabilities is essential to meet the demands for more versatile
and efficient AI. However, previous omni-models have insufficiently explored
speech, neglecting its integration with multi-modality. We introduce Lyra, an
efficient MLLM that enhances multimodal abilities, including advanced
long-speech comprehension, sound understanding, cross-modality efficiency, and
seamless speech interaction. To achieve efficiency and speech-centric
capabilities, Lyra employs three strategies: (1) leveraging existing
open-source large models and a proposed multi-modality LoRA to reduce training
costs and data requirements; (2) using a latent multi-modality regularizer and
extractor to strengthen the relationship between speech and other modalities,
thereby enhancing model performance; and (3) constructing a high-quality,
extensive dataset that includes 1.5M multi-modal (language, vision, audio) data
samples and 12K long speech samples, enabling Lyra to handle complex long
speech inputs and achieve more robust omni-cognition. Compared to other
omni-methods, Lyra achieves state-of-the-art performance on various
vision-language, vision-speech, and speech-language benchmarks, while also
using fewer computational resources and less training data.
48Multi-modal Large Language Models (MLLMs)omni-modelsHF ↗arXiv ↗
DiffSensei, a diffusion-based framework with a multimodal large language model, enhances manga generation by precisely controlling multi-character appearances and interactions using masked cross-attention and text cues.
48diffusion-basedmultimodal large language modelHF ↗arXiv ↗
A human-curated benchmark (CodeArena) and a large synthetic instruction corpus (SynCode-Instruct) are introduced to evaluate code LLMs based on human preference alignment, revealing performance differences between open-source and proprietary models.
48code large language modelscode generationHF ↗arXiv ↗
LiFT fine-tunes text-to-video generative models using human feedback to improve video alignment with text descriptions, demonstrating superior performance over a larger model.
A large-scale multimodal instruction-tuning dataset with intermediate rationales improves reasoning capabilities in large language models, achieving state-of-the-art performance across benchmarks.
46multimodal large language modelsinstruction-tuning datasetHF ↗arXiv ↗
A plug-and-play module enhances text-to-video models for multi-camera video generation, incorporating 6 DoF camera poses and ensuring dynamic consistency across viewpoints using a hybrid training scheme and multi-view synchronization.
SwiftEdit performs instant text-guided image editing with high efficiency using a one-step inversion framework and mask-guided editing with attention rescaling.
POINTS1.5, an enhanced vision-language model, incorporates a dynamic high-resolution vision encoder, bilingual support, and rigorous dataset filtering methods to achieve superior performance.
APOLLO, an approximated gradient scaling optimizer for LLMs, reduces memory usage without compromising performance, enhancing throughput and scalability.
Leffa, a flow field learning method integrated into diffusion models, enhances controllable person image generation by addressing fine-grained detail distortions.
SnapGen, a compact and fast text-to-image diffusion model, generates high-resolution images on mobile devices using design optimizations, knowledge distillation, and adversarial guidance.
AgentTrek synthesizes high-quality GUI agent trajectories using web tutorials, improving their performance cost-effectively compared to human annotation.
Kasra Arabi, Benjamin Feuer, R. Teal Witter +2 authors
The proposed two-stage watermarking framework using Fourier patterns and diffusion model's initial noise enhances robustness to forgery and removal in AI-generated images.
ACDiT, an Autoregressive blockwise Conditional Diffusion Transformer, bridges visual generation and autoregressive modeling, demonstrating effectiveness in image and video generation and potential for visual understanding tasks.
UniReal treats image generation and editing tasks as discontinuous video generation to capture visual variations and consistency, learning from large-scale video data.
30unifying approachdiscontinuous video generationHF ↗arXiv ↗
Maya is an open-source multimodal multilingual model that addresses the limitations of Vision-Language Models in handling low-resource languages and cultural nuances by introducing a multilingual pretraining dataset and analyzing/removing toxicity.
Thibault Le Sellier De Chezelles, Maxime Gasse, Alexandre Drouin +17 authors
BrowserGym provides a standardized environment for evaluating web agents using LLMs, facilitating consistent benchmarking and improving agent development and analysis.
24Large Language Modelsgym-like environmentHF ↗arXiv ↗
Nicolas Dufour, David Picard, Vicky Kalogeiton +1 authors
A generative geolocation approach using diffusion and Riemannian flow matching achieves state-of-the-art performance on visual geolocation tasks and introduces probabilistic geolocation with new metrics.
A generative pre-training approach using motion token sequences from video data enhances robot learning, enabling transfer to real manipulation tasks with improved robustness and efficiency.
22Large Language Modelsautoregressive pre-trainingHF ↗arXiv ↗