PIXART-{\delta} integrates Latent Consistency Model (LCM) and ControlNet into PIXART-{\alpha} for accelerated, high-quality text-to-image synthesis with efficient training and inference.
49Latent Consistency Model (LCM)ControlNetHF ↗arXiv ↗
Prompt-aligned personalization improves text alignment and enhances personalized image creation with complex prompts by using score distillation sampling.
A survey examining the design, performance, and future directions of MultiModal Large Language Models (MM-LLMs) to enhance reasoning and MM task capabilities.
49MultiModal Large Language ModelsMM-LLMsHF ↗arXiv ↗
VidEgoThink, a comprehensive benchmark for evaluating egocentric video understanding capabilities, highlights the need for advancements in Multi-modal Large Language Models for effective application in Embodied AI.
49Multi-modal Large Language ModelsMLLMsHF ↗arXiv ↗
A novel method for generating mathematical code with reasoning steps enhances mathematical reasoning abilities in large language models using a comprehensive dataset named MathCode-Pile.
AndroidLab provides a systematic framework for training and evaluating Android agents, supporting both large language models and multimodal models, and improves their task success rates.
Parth Sarthi, Salman Abdullah, Aditi Tuli +3 authors
The RAPTOR model enhances retrieval-augmented language models by recursively embedding, clustering, and summarizing text chunks, leading to better performance on question-answering tasks involving complex reasoning.
48retrieval-augmented language modelsRecursive embeddingHF ↗arXiv ↗
DiffSensei, a diffusion-based framework with a multimodal large language model, enhances manga generation by precisely controlling multi-character appearances and interactions using masked cross-attention and text cues.
48diffusion-basedmultimodal large language modelHF ↗arXiv ↗
A theoretical framework called Internal Consistency and a Self-Feedback framework with Self-Evaluation and Self-Update modules are presented to address reasoning and hallucination in large language models by assessing coherence across model layers.
48large language modelsself-consistencyHF ↗arXiv ↗
Alex Havrilla, Yuqing Du, Sharath Chandra Raparthy +6 authors
Multiple RLHF algorithms, including Expert Iteration, PPO, and Return-Conditioned RL, similarly improve LLM reasoning with similar sample complexity, with Expert Iteration slightly outperforming others.
48Reinforcement Learning from Human FeedbackRLHFHF ↗arXiv ↗
A human-curated benchmark (CodeArena) and a large synthetic instruction corpus (SynCode-Instruct) are introduced to evaluate code LLMs based on human preference alignment, revealing performance differences between open-source and proprietary models.
48code large language modelscode generationHF ↗arXiv ↗
Zhisheng Zhong, Chengyao Wang, Yuqi Liu +12 authors
As Multi-modal Large Language Models (MLLMs) evolve, expanding beyond
single-domain capabilities is essential to meet the demands for more versatile
and efficient AI. However, previous omni-models have insufficiently explored
speech, neglecting its integration with multi-modality. We introduce Lyra, an
efficient MLLM that enhances multimodal abilities, including advanced
long-speech comprehension, sound understanding, cross-modality efficiency, and
seamless speech interaction. To achieve efficiency and speech-centric
capabilities, Lyra employs three strategies: (1) leveraging existing
open-source large models and a proposed multi-modality LoRA to reduce training
costs and data requirements; (2) using a latent multi-modality regularizer and
extractor to strengthen the relationship between speech and other modalities,
thereby enhancing model performance; and (3) constructing a high-quality,
extensive dataset that includes 1.5M multi-modal (language, vision, audio) data
samples and 12K long speech samples, enabling Lyra to handle complex long
speech inputs and achieve more robust omni-cognition. Compared to other
omni-methods, Lyra achieves state-of-the-art performance on various
vision-language, vision-speech, and speech-language benchmarks, while also
using fewer computational resources and less training data.
48Multi-modal Large Language Models (MLLMs)omni-modelsHF ↗arXiv ↗
Enrico Fini, Mustafa Shukor, Xiujun Li +13 authors
AIMV2, a multimodal vision encoder paired with an autoregressive decoder, achieves superior performance in vision benchmarks and multimodal image understanding tasks compared to contrastive models.
Johannes Schmude, Sujit Roy, Will Trojak +26 authors
Prithvi WxC, a 2.3 billion parameter foundation model using an encoder-decoder architecture with transformer concepts, addresses weather forecasting, downscaling, and extreme events estimation with a mixed masked reconstruction and forecasting objective.
Rogerio Bonatti, Dan Zhao, Francesco Bonacci +8 authors
Windows Agent Arena provides a general, reproducible environment for evaluating multi-modal agent performance on Windows OS tasks, with scalability and parallelizability features.
LLM-generated research ideas are perceived as more novel than those from human experts but are deemed slightly less feasible, based on blind evaluations by NLP researchers.
A new pipeline and dataset enable long-context LLMs to generate responses with precise sentence-level citations, improving their trustworthiness and outperforming existing models.
An on-device Planner-Action framework, utilizing Phi-3 Mini for planning and Octopus for execution, employs model fine-tuning and multi-LoRA training to enhance efficiency and multi-domain handling on resource-constrained devices.
Ahmet Üstün, Viraat Aryabumi, Zheng-Xin Yong +14 authors
Aya, a multilingual generative language model supporting over 50% lower-resourced languages, excels in both generative and discriminative tasks across 99 languages, and provides extensive evaluations and open-source resources.
The Flexible Vision Transformer adapts to varied image resolutions and aspect ratios through dynamic tokenization and extrapolation techniques, outperforming traditional methods.
Jupinder Parmar, Shrimai Prabhumoye, Joseph Jennings +24 authors
Nemotron-4 15B, a large multilingual language model, excels in English, multilingual, and coding tasks, demonstrating superior performance in multilingual capabilities compared to larger and specialized models.
48large multilingual language modeldownstream evaluation areasHF ↗arXiv ↗
Fu-Yun Wang, Zhaoyang Huang, Alexander William Bergman +9 authors
The Phased Consistency Model (PCM) addresses limitations in Latent Consistency Models (LCM) and outperforms them in high-resolution, text-conditioned image and few-step text-to-video generation.
MAP-Neo, a fully open-sourced bilingual LLM with 7B parameters, achieves performance comparable to existing state-of-the-art LLMs, enhancing transparency and fostering innovation in the open research community.
VIF-RAG-QA and FollowRAG Benchmark enhance Instruction-Following alignment in RAG systems through automated instruction synthesis, verification, and evaluation.
48Retrieval-Augmented Generation (RAG)Large Language Models (LLMs)HF ↗arXiv ↗
F5-TTS, a fully non-autoregressive text-to-speech system, improves E2 TTS by leveraging ConvNeXt and Sway Sampling for better performance and efficiency.
Hadas Orgad, Michael Toker, Zorik Gekhman +4 authors
LLMs internally encode detailed truthfulness information about their outputs, which can be used to detect and predict errors but shows variability across datasets and does not always align with the final output.
StructRAG enhances LLMs by identifying optimal structure types, reconstructing documents into these structures, and inferring answers for knowledge-intensive reasoning tasks, achieving state-of-the-art performance.
JudgeBench is a benchmark for evaluating LLM-based judges using objective correctness for challenging tasks like knowledge, reasoning, math, and coding.
A semi-supervised fine-tuning framework named SemiEvol enhances LLM adaptation using both labeled and unlabeled data, showing improved performance through bi-level knowledge propagation and collaborative learning.
47supervised fine-tuninglarge language modelsHF ↗arXiv ↗