TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

本月最热323

Group Sequence Policy Optimization

Chujie Zheng, Shixuan Liu, Mingze Li +9 authors

Group Sequence Policy Optimization (GSPO) is a reinforcement learning algorithm that improves training efficiency and performance of large language models by using sequence-level importance ratios and operations.

Group Sequence Policy OptimizationGSPOreinforcement learninglarge language modelsHF ↗arXiv ↗

50 篇论文 · 按点赞排序

04

MemOS: A Memory OS for AI System

Zhiyu Li, Shichao Song, Chenyang Xi +36 authors

MemOS, a memory operating system for Large Language Models, addresses memory management challenges by unifying plaintext, activation-based, and parameter-level memories, enabling efficient storage, retrieval, and continual learning.

170Large Language ModelsArtificial General IntelligenceHF ↗arXiv ↗
05

Agentic Reinforced Policy Optimization

Guanting Dong, Hangyu Mao, Kai Ma +11 authors

Agentic Reinforced Policy Optimization (ARPO) enhances multi-turn reasoning in large language models by balancing long-horizon capabilities and tool interactions, using entropy-based adaptive rollouts and advantage attribution.

162reinforcement learningverifiable rewardsHF ↗arXiv ↗
06

Scaling RL to Long Videos

Yukang Chen, Wei Huang, Baifeng Shi +11 authors

A framework scales vision-language models for long video reasoning using reinforcement learning, achieving strong performance on benchmarks and demonstrating consistent gains with increased video frames.

160vision-language modelsreinforcement learningHF ↗arXiv ↗
09

GUI-G^2: Gaussian Reward Modeling for GUI Grounding

Fei Tang, Zhangxuan Gu, Zhengxi Lu +9 authors

A new reward framework, GUI-G$^2$, models GUI elements as continuous Gaussian distributions to improve autonomous interaction through dense gradient signals, outperforming existing methods in spatial reasoning tasks.

135reinforcement learningbinary rewardsHF ↗arXiv ↗
10

Kwai Keye-VL Technical Report

Kwai Keye Team, Biao Yang, Bin Wen +57 authors

Kwai Keye-VL, an 8-billion-parameter multimodal model, excels in short-video understanding and general vision-language tasks through a comprehensive pre-training and post-training process, including a five-mode data mixture and reinforcement learning.

133Multimodal Large Language ModelsKwai Keye-VLHF ↗arXiv ↗
12

nablaNABLA: Neighborhood Adaptive Block-Level Attention

Dmitrii Mikhailov, Aleksey Letunovskiy, Maria Kovaleva +6 authors

NABLA, a Neighborhood Adaptive Block-Level Attention mechanism, enhances video diffusion transformers by reducing computational overhead without significantly impacting generative quality or visual fidelity.

126transformer-based architecturesvideo generationHF ↗arXiv ↗
14

T-LoRA: Single Image Diffusion Model Customization Without Overfitting

Vera Soboleva, Aibek Alanov, Andrey Kuznetsov +1 authors

T-LoRA, a timestep-dependent low-rank adaptation framework, enhances diffusion model personalization with a dynamic fine-tuning strategy and orthogonal initialization, improving concept fidelity and text alignment in data-limited settings.

121diffusion model fine-tuningoverfittingHF ↗arXiv ↗
15

SingLoRA: Low Rank Adaptation Using a Single Matrix

David Bensaïd, Noam Rotstein, Roy Velich +2 authors

SingLoRA, a reformulated low-rank adaptation method, enhances parameter-efficient fine-tuning by learning a single low-rank matrix and its transpose, ensuring stable optimization and reducing parameter count.

116Low-Rank AdaptationLoRAHF ↗arXiv ↗
16

Test-Time Scaling with Reflective Generative Model

Zixiao Wang, Yuxin Wang, Xiaorui Wang +8 authors

MetaStone-S1, a reflective generative model using a self-supervised process reward model, achieves high performance with reduced parameters and supports test time scaling.

108reflective generative modelself-supervised process reward modelHF ↗arXiv ↗
17

4KAgent: Agentic Any Image to 4K Super-Resolution

Yushen Zuo, Qi Zheng, Mingyang Wu +10 authors

4KAgent, a unified agentic super-resolution system, enhances low-resolution images to 4K using a profiling module, perception agent, and restoration agent, achieving state-of-the-art performance across various imaging domains.

107agentic super-resolutionProfilingHF ↗arXiv ↗
19

A Survey on Latent Reasoning

Rui-Jie Zhu, Tianhao Peng, Tianhao Cheng +30 authors

Latent reasoning in Large Language Models (LLMs) performs multi-step inference in continuous hidden states, enhancing reasoning capabilities without token-level supervision, and includes methodologies like activation-based recurrence and infinite-depth reasoning via masked diffusion models.

95chain-of-thought (CoT)latent reasoningHF ↗arXiv ↗
20

Yume: An Interactive World Generation Model

Xiaofeng Mao, Shaoheng Lin, Zhen Li +7 authors

A framework for generating and exploring interactive, high-fidelity video worlds from images using a Masked Video Diffusion Transformer, advanced sampling techniques, and model acceleration.

92camera motion quantizationMasked Video Diffusion TransformerHF ↗arXiv ↗
25

The Invisible Leash: Why RLVR May Not Escape Its Origin

Fang Wu, Weihao Xuan, Ximing Lu +2 authors

Reinforcement Learning with Verifiable Rewards (RLVR) enhances precision but may limit exploration and discovery of new solutions, suggesting potential limits to its effectiveness in expanding reasoning capabilities.

85Reinforcement Learning with Verifiable RewardsRLVRHF ↗arXiv ↗
27

Should We Still Pretrain Encoders with Masked Language Modeling?

Hippolyte Gisserot-Boukhlef, Nicolas Boizard, Manuel Faysse +5 authors

A biphasic training strategy combining Causal Language Modeling and Masked Language Modeling yields optimal text representation performance, especially when initialized with pretrained CLM models.

81Masked Language ModelingCausal Language ModelingHF ↗arXiv ↗
29

MIRIX: Multi-Agent Memory System for LLM-Based Agents

Yu Wang, Xi Chen

MIRIX, a modular multi-agent memory system, enhances language models' memory capabilities by integrating diverse memory types and a dynamic framework, achieving superior performance in multimodal and long-form conversation benchmarks.

80modularmulti-agent memory systemHF ↗arXiv ↗
1 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号