A benchmark (SEED-Bench-R1) evaluates reinforcement learning versus supervised fine-tuning for multimodal large language models in video understanding, highlighting RL's data efficiency and superior performance but identifying limitations in logical coherence and visual cue processing.
38Chain of Thought (COT)Large Language Models (LLMs)HF ↗arXiv ↗
Alexander Martin, Reno Kriz, William Gantt Walden +5 authors
WikiVideo proposes Collaborative Article Generation (CAG) to enhance high-level event summarization from videos by integrating r1-style reasoning and VideoLLM for better inferences compared to state-of-the-art VideoLLMs.
Giulio Starace, Oliver Jaffe, Dane Sherburn +10 authors
PaperBench evaluates AI agents' ability to replicate state-of-the-art AI research by decomposing replication tasks into graded sub-tasks, using both LLM-based and human judges to assess performance.
ReaRec enhances sequential recommendation by using inference-time multi-step reasoning, improving model performance on long-tail items and user preference evolution.
CodeARC introduces an interactive evaluation framework to assess the ability of large language model agents in inductive program synthesis using differential testing and iterative refinement.
35inductive program synthesisprogramming by exampleHF ↗arXiv ↗
Visual self-supervised learning matches language-supervised visual pretraining performance on VQA and vision benchmarks when both are trained on the same dataset and scaled appropriately.
A transparent reinforcement learning framework and standardized evaluation are introduced for vision-language models, showing RL's superiority in generalization over supervised fine-tuning.
Team Cohere, Aakanksha, Arash Ahmadian +223 authors
Command A, a multilingual large language model, uses decentralized training with self-refinement and model merging to achieve efficient and top-performing Retrieval Augmented Generation for enterprise use.
GeometryCrafter uses a point map VAE and video diffusion model to estimate high-fidelity, temporally coherent depth maps from open-world videos, enhancing 3D reconstruction and camera parameter estimation.
OThink-MR1, an advanced MLLM using GRPO-D, enhances reinforcement learning performance and demonstrates superior cross-task generalization compared to supervised fine-tuning.
29Multimodal Large Language Models (MLLMs)supervised fine-tuning (SFT)HF ↗arXiv ↗
A novel Shifted Thinking Window method trains LLMs on code-related reasoning trajectories to reduce excessive thinking tokens while maintaining performance and efficient test-time scaling.
Agent S2, a compositional framework using Mixture-of-Grounding and Proactive Hierarchical Planning, achieves state-of-the-art performance in computer use automation across various benchmarks and operating systems.
The paper presents RIG, an end-to-end agent policy that integrates reasoning and imagination, significantly improving sample efficiency and generalization in complex environments through joint reasoning and action outcome prediction.
A visualization tool named Landscape of Thoughts helps inspect and analyze reasoning paths of large language models on multi-choice datasets, identifying model strengths, correct answers, and reasoning inconsistencies.
26large language modelschain-of-thoughtHF ↗arXiv ↗
Tommie Kerssies, Niccolò Cavagnero, Alexander Hermans +5 authors
The Encoder-only Mask Transformer (EoMT) achieves state-of-the-art image segmentation accuracy by learning task-specific components through large-scale pre-training, while maintaining superior prediction speed compared to models with additional architectural complexity.
Summarization refinement faces challenges when extending to multi-dimension.
In this paper, we introduce ReFeed, a powerful summarization refinement
pipeline that enhances multiple dimensions through reflective reasoning on
feedback. To achieve this, we release SumFeed-CoT, a large-scale Long-CoT-based
dataset optimized for training a lightweight model with reflective reasoning.
Our experiments reveal how the number of dimensions, feedback exposure, and
reasoning policy influence refinement performance, highlighting reflective
reasoning and simultaneously addressing multiple feedback is crucial to
mitigate trade-off between dimensions. Furthermore, ReFeed is robust to noisy
feedback and feedback order. Lastly, our finding emphasizes that creating data
with a proper goal and guideline constitutes a fundamental pillar of effective
reasoning. The dataset and model will be released.
Existing large language models exhibit significant recitation behavior when conditions are subtly altered, suggesting they may not possess human-like intelligence.
Reinforcement learning with verifiable rewards extended to diverse domains using model-based soft scoring outperforms existing LLMs in free-form answer settings with reliable reward signals.
24reinforcement learning with verifiable rewardslarge language modelsHF ↗arXiv ↗