TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

606 篇论文 · 按点赞排序

361

Self-rewarding correction for mathematical reasoning

Wei Xiong, Hanning Zhang, Chenlu Ye +3 authors

Self-rewarding reasoning large language models independently generate and correct their outputs during inference using a two-stage algorithmic framework, enhancing performance without external feedback.

82self-rewarding reasoninglarge language modelsHF ↗arXiv ↗
362

RM-R1: Reward Modeling as Reasoning

Xiusi Chen, Gaotang Li, Ziqi Wang +9 authors

Reasoning Reward Models (ReasRMs) enhance reward modeling for large language models by integrating reasoning tasks, improving interpretability and performance.

81reward modelingreinforcement learning from human feedback (RLHF)HF ↗arXiv ↗
363

MiMo-VL Technical Report

Xiaomi LLM-Core Team, Zihao Yue, Zhenru Lin +71 authors

MiMo-VL-7B-SFT and MiMo-VL-7B-RL provide state-of-the-art general visual understanding and multimodal reasoning through four-stage pre-training and Mixed On-policy Reinforcement Learning, outperforming models with up to 78B parameters.

81vision-language modelsmultimodal reasoningHF ↗arXiv ↗
365

VerlTool: Towards Holistic Agentic Reinforcement Learning with Tool Use

Dongfu Jiang, Yi Lu, Zhuofeng Li +9 authors

VerlTool is a unified and modular framework for Agentic Reinforcement Learning with Tool use, addressing inefficiencies in existing approaches and providing competitive performance across multiple domains.

81Reinforcement Learning with Verifiable RewardsAgentic Reinforcement Learning with Tool useHF ↗arXiv ↗
367

FineVision: Open Data Is All You Need

Luis Wiedmann, Orr Zohar, Amir Mahla +6 authors

FineVision, a large-scale and curated dataset, enhances vision-language models through rigorous data collection, de-duplication, and human oversight, leading to improved performance.

81vision-language modelsFineVisionHF ↗arXiv ↗
368

Should We Still Pretrain Encoders with Masked Language Modeling?

Hippolyte Gisserot-Boukhlef, Nicolas Boizard, Manuel Faysse +5 authors

A biphasic training strategy combining Causal Language Modeling and Masked Language Modeling yields optimal text representation performance, especially when initialized with pretrained CLM models.

81Masked Language ModelingCausal Language ModelingHF ↗arXiv ↗
370

EuroBERT: Scaling Multilingual Encoders for European Languages

Nicolas Boizard, Hippolyte Gisserot-Boukhlef, Duarte M. Alves +16 authors

EuroBERT, a family of multilingual encoders covering European and global languages, outperforms existing models across various tasks and supports long sequences, surpassing traditional bidirectional encoders.

81bidirectional encoder modelsgenerative decoder-only modelsHF ↗arXiv ↗
371

Thyme: Think Beyond Images

Yi-Fan Zhang, Xingyu Lu, Shukang Yin +17 authors

Thyme, a novel paradigm, enables MLLMs to autonomously perform image manipulations and computations, enhancing performance in perception and reasoning tasks through a two-stage training strategy and GRPO-ATS algorithm.

81MLLMsthink with imagesHF ↗arXiv ↗
372

UniVideo: Unified Understanding, Generation, and Editing for Videos

Cong Wei, Quande Liu, Zixuan Ye +5 authors

UniVideo, a dual-stream framework combining a Multimodal Large Language Model and a Multimodal DiT, extends unified modeling to video generation and editing, achieving state-of-the-art performance and supporting task composition and generalization.

81Multimodal Large Language ModelMultimodal DiTHF ↗arXiv ↗
373

Distilling LLM Agent into Small Models with Retrieval and Code Tools

Minki Kang, Jongwon Jeong, Seanie Lee +2 authors

Agent Distillation transfers reasoning and task-solving capabilities from large language models to smaller models using enhanced prompts and self-consistent actions, matching performance of larger models on various reasoning tasks.

81Large language modelssmall language modelsHF ↗arXiv ↗
374

IterResearch: Rethinking Long-Horizon Agents via Markovian State Reconstruction

Guoxin Chen, Zile Qiao, Xuanzhong Chen +13 authors

IterResearch, an iterative deep-research paradigm, improves long-horizon reasoning by reformulating it as a Markov Decision Process with strategic workspace reconstruction and Efficiency-Aware Policy Optimization, achieving better performance and interaction scaling compared to existing agents.

80deep-research agentsdynamic reasoningHF ↗arXiv ↗
378

MIRIX: Multi-Agent Memory System for LLM-Based Agents

Yu Wang, Xi Chen

MIRIX, a modular multi-agent memory system, enhances language models' memory capabilities by integrating diverse memory types and a dynamic framework, achieving superior performance in multimodal and long-form conversation benchmarks.

80modularmulti-agent memory systemHF ↗arXiv ↗
380

Video-R1: Reinforcing Video Reasoning in MLLMs

Kaituo Feng, Kaixiong Gong, Bohao Li +5 authors

Video-R1, leveraging rule-based reinforcement learning and temporal information, enhances video reasoning in multimodal large language models using a combination of video and image data.

79rule-based reinforcement learningRLHF ↗arXiv ↗
384

Scaling Law for Quantization-Aware Training

Mengzhao Chen, Chaoyi Zhang, Jing Liu +8 authors

A unified scaling law for quantization-aware training (QAT) identifies key factors affecting quantization error, leading to improvements through mixed-precision quantization.

79quantization-aware trainingQATHF ↗arXiv ↗
388

OmniGen2: Exploration to Advanced Multimodal Generation

Chenyuan Wu, Pengfei Zheng, Ruiran Yan +19 authors

OmniGen2, a versatile generative model, introduces dual decoding pathways for text and images, preserves original text generation, and achieves competitive results with a new subject-driven benchmark.

79decoding pathwaysunshared parametersHF ↗arXiv ↗
389

Multimodal Evaluation of Russian-language Architectures

Artem Chervyakov, Ulyana Isaeva, Anton Emelyanov +15 authors

Mera Multi is an open multimodal evaluation framework for Russian-spoken architectures, addressing the lack of such benchmarks with 18 newly constructed tasks and a methodology to prevent benchmark leakage.

79multimodal large language modelsMera MultiHF ↗arXiv ↗
13 / 21

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号