MinerU-Diffusion is a diffusion-based framework that replaces autoregressive decoding with parallel diffusion denoising for document OCR, improving robustness and decoding speed.
Intern-S1-Pro is a one-trillion-parameter scientific multimodal foundation model that enhances general and scientific capabilities through advanced agent functionalities and specialized task mastery across multiple scientific disciplines.
134multimodal foundation modelreinforcement learningHF ↗arXiv ↗
Omni-WorldBench addresses the lack of comprehensive evaluation for interactive 4D world models by introducing a benchmark that assesses temporal dynamics and causal interaction effects across diverse scenarios.
daVinci-MagiHuman is an open-source audio-video generative model that synchronizes text, video, and audio through a single-stream Transformer architecture, achieving high-quality human-centric content generation with efficient inference capabilities.
125audio-video generative foundation modelsingle-stream TransformerHF ↗arXiv ↗
A diffusion framework called PixelSmile is introduced that disentangles facial expression semantics through symmetric joint training and contrastive learning to enable precise, controllable, and fine-grained expression editing with robust identity preservation.
HopChain is a scalable framework that generates multi-hop vision-language reasoning data to enhance VLMs' long-chain reasoning capabilities across diverse benchmarks.
Astrolabe is an efficient online reinforcement learning framework for distilled autoregressive video models that improves generation quality through forward-process RL formulation and streaming training with multi-reward objectives.
109autoregressive video modelsreinforcement learningHF ↗arXiv ↗
OpenResearcher presents a reproducible pipeline for training deep research agents using offline search environments and synthesized trajectories, achieving improved accuracy on benchmark tasks.
102deep research agentslong-horizon trajectoriesHF ↗arXiv ↗
Xiangru Jian, Shravan Nayak, Kevin Qinghong Lin +5 authors
CUA-Suite introduces a large-scale ecosystem of expert video demonstrations and annotations for computer-use agents, providing continuous screen recordings and detailed reasoning annotations to advance desktop automation capabilities.
99computer-use agentsexpert video demonstrationsHF ↗arXiv ↗
WildWorld is a large-scale dataset for action-conditioned world modeling that provides explicit state annotations from a photorealistic game, enabling better understanding of latent-state dynamics and long-horizon consistency.
93dynamical systems theoryreinforcement learningHF ↗arXiv ↗
AwaRes is a spatial-on-demand framework for vision-language models that dynamically retrieves high-resolution image segments based on query needs, using tool-calling and multi-turn reinforcement learning with composite rewards.
A 560-billion-parameter Mixture-of-Experts model advances formal reasoning in Lean4 through tool-integrated reasoning with a hybrid framework and hierarchical policy optimization for stable training on long-horizon tasks.
Diffusion Transformers can be enhanced through a parameter-efficient calibration approach that improves generative quality while reducing inference steps.
Alexander H. Liu, Alexis Tacnet, Andy Ehrenberg +184 authors
Voxtral TTS is a multilingual text-to-speech model that generates natural speech from short reference audio using a hybrid architecture combining semantic token generation and flow-matching for acoustic tokens.
SpecEyes accelerates agentic multimodal large language models by using a lightweight speculative planner with cognitive gating and heterogeneous parallel processing to reduce latency and improve throughput.
62agentic multimodal large language modelsiterative visual tool invocationHF ↗arXiv ↗
Self-distillation in large language models can degrade mathematical reasoning performance by suppressing uncertainty expression, particularly affecting out-of-distribution tasks.
59self-distillationlarge language modelsHF ↗arXiv ↗
A large-scale dataset and open-source model are developed to improve image restoration performance and close the gap with closed-source alternatives, with a dedicated benchmark for real-world degradation evaluation.
Ling Yue, Kushal Raj Bhandari, Ching-Yun Ko +6 authors
LLM-based systems use executable workflows that interleave various computational components, with recent approaches organized by workflow structure determination timing and optimization dimensions.
Memory Sparse Attention (MSA) enables large language models to process extremely long contexts with linear complexity and high efficiency through innovations like sparse attention and document-wise RoPE.
53long-term memorylarge language modelsHF ↗arXiv ↗
Optical flow models trained on high-quality data often degrade severely when confronted with real-world corruptions such as blur, noise, and compression artifacts. To overcome this limitation, we formulate Degradation-Aware Optical Flow, a new task targeting accurate dense correspondence estimation from real-world corrupted videos. Our key insight is that the intermediate representations of image restoration diffusion models are inherently corruption-aware but lack temporal awareness. To address this limitation, we lift the model to attend across adjacent frames via full spatio-temporal attention, and empirically demonstrate that the resulting features exhibit zero-shot correspondence capabilities. Based on this finding, we present DA-Flow, a hybrid architecture that fuses these diffusion features with convolutional features within an iterative refinement framework. DA-Flow substantially outperforms existing optical flow methods under severe degradation across multiple benchmarks.
Jenny Zhang, Bingchen Zhao, Wannan Yang +5 authors
Hyperagents represent a self-referential framework that integrates task and meta-agents into a single editable program, enabling metacognitive self-modification and open-ended improvement across diverse computational domains.
51self-improving AI systemsDarwin Gödel MachineHF ↗arXiv ↗
TerraScope is a unified vision-language model that enables pixel-grounded geospatial reasoning through modality-flexible and multi-temporal capabilities, evaluated on a new benchmark with detailed visual reasoning outputs.
VideoDetective framework improves long video understanding by integrating query-to-segment relevance and inter-segment affinity through visual-temporal graphs and hypothesis verification loops.
50multimodal large language modelsvideo understandingHF ↗arXiv ↗
Geometric Latent Diffusion (GLD) framework utilizes geometric foundation models' feature space as latent space for novel view synthesis, achieving superior 2D and 3D performance while reducing training time significantly.
Byungwoo Jeon, Dongyoung Kim, Huiwon Jang +2 authors
SpatialBoost enhances vision encoders' 3D spatial awareness by integrating linguistic 3D spatial knowledge through a multi-turn Chain-of-Thought reasoning process with Large Language Models.
A two-stage self-evolving mobile GUI agent named UI-Voyager is proposed, featuring rejection fine-tuning and group relative self-distillation to improve efficiency and performance in GUI automation tasks.
EVA is an efficient reinforcement learning framework for video understanding that enables adaptive reasoning through iterative planning and attention mechanisms, outperforming existing methods on multiple video benchmarks.
44multimodal large language modelsreinforcement learningHF ↗arXiv ↗
Chuanrui Zhang, Minghan Qin, Yuang Wang +3 authors
A unified multimodal large language model framework called SIMART is proposed for generating articulated 3D assets with reduced tokenization overhead and improved simulation readiness.
Thomas De Min, Subhankar Roy, Stéphane Lathuilière +2 authors
MLLMs demonstrate limited proactive behavior in requesting user interventions for challenging tasks, with performance hindered by conversational context and in-context learning biases, though reinforcement learning fine-tuning shows potential for learning such behaviors.
Woosung Koh, Jeyoung Jeon, Youngjin Song +4 authors
Multi-task supervised fine-tuning with heterogeneous learning dynamics benefits from an iterative overfitting-aware search algorithm that improves performance across diverse datasets and compute budgets.