TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

686 篇论文 · 按点赞排序

244

Harder Is Better: Boosting Mathematical Reasoning via Difficulty-Aware GRPO and Multi-Aspect Question Reformulation

Yanqi Dai, Yuxiang Ji, Xiao Zhang +3 authors

MathForge enhances mathematical reasoning in large models through a dual framework combining difficulty-aware policy optimization and multi-aspect question reformulation to address limitations in existing reinforcement learning methods.

119Reinforcement Learning with Verifiable RewardsGroup Relative Policy OptimizationHF ↗arXiv ↗
245

Geometric Action Model for Robot Policy Learning

Jisang Han, Seonghu Jeon, Jaewoo Jung +7 authors

A geometric action model leverages pretrained geometric foundation models to enable language-conditioned manipulation policies with improved accuracy, robustness, and efficiency in 3D physical environments.

119vision-language-action modelsvideo world-action modelsHF ↗arXiv ↗
246

Stealing Reasoning Traces from Proprietary LLM APIs

Alexander Panfilov, David Schmotz, Ilia Shumailov +5 authors

Encrypted reasoning traces shared across sessions and models can be intercepted and injected into weaker models to extract proprietary reasoning, private data, hidden hazards, and hidden prompts.

119chain-of-thoughtencrypted reasoning tracesHF ↗arXiv ↗
247

PixelSmile: Toward Fine-Grained Facial Expression Editing

Jiabin Hua, Hengyuan Xu, Aojie Li +4 authors

A diffusion framework called PixelSmile is introduced that disentangles facial expression semantics through symmetric joint training and contrastive learning to enable precise, controllable, and fine-grained expression editing with robust identity preservation.

118diffusion frameworkfacial expression editingHF ↗arXiv ↗
248

LatentPress: Context Compression Beyond Text and Vision

Zhengze Zhou, Hejian Sang

LatentPress compresses conversational and document context into continuous memory tokens read directly by a frozen decoder, achieving high compression with faster inference and improved accuracy over text or OCR methods.

118continuous memory tokensfrozen decoderHF ↗arXiv ↗
255

Qwen-Image-2.0 Technical Report

Bing Zhao, Chenfei Wu, Deqing Li +72 authors

Qwen-Image-2.0 is an advanced image generation model that combines high-fidelity synthesis with precise editing capabilities through a unified framework using Qwen3-VL as condition encoder and Multimodal Diffusion Transformer for joint modeling.

116multimodal diffusion transformercondition encoderHF ↗arXiv ↗
260

Self-Distilled Agentic Reinforcement Learning

Zhengxi Lu, Zhiyuan Yao, Zhuowen Han +8 authors

SDAR enhances reinforcement learning for multi-turn agent training by integrating self-distillation through a sigmoid gate that selectively strengthens positive token-level guidance while mitigating negative teacher rejections.

116Reinforcement learningon-policy self-distillationHF ↗arXiv ↗
263

Controlled Self-Evolution for Algorithmic Code Optimization

Tu Hu, Ronghao Chen, Shuo Zhang +9 authors

Controlled Self-Evolution method improves code generation through diversified initialization, feedback-guided genetic evolution, and hierarchical memory to enhance exploration efficiency and solution quality.

115self-evolution methodsgenerate-verify-refine cyclesHF ↗arXiv ↗
265

DOPD: Dual On-policy Distillation

Xinlei Yu, Gen Li, Qingyi Si +13 authors

DOPD addresses privilege illusion in on-policy distillation by dynamically routing token-level supervision between teacher and student policies based on advantage gaps and probabilities, improving capability transfer in large and vision-language models.

114on-policy distillationtoken-level signalsHF ↗arXiv ↗
267

Golden Goose: A Simple Trick to Synthesize Unlimited RLVR Tasks from Unverifiable Internet Text

Ximing Lu, David Acuna, Jaehun Jung +12 authors

Golden Goose synthesizes unlimited RLVR tasks from unverifiable internet text by creating multiple-choice question-answering versions of fill-in-the-middle tasks, enabling large-scale training and achieving state-of-the-art results in cybersecurity and other domains.

113Reinforcement Learning with Verifiable RewardsLarge Language ModelsHF ↗arXiv ↗
270

Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe

Yaxuan Li, Yuxin Zuo, Bingxiang He +8 authors

On-policy distillation dynamics in large language models depend on compatible thinking patterns between teacher and student models, with successful distillation characterized by alignment on high-probability tokens and requiring teachers to provide novel capabilities beyond student training data.

113on-policy distillationlarge language modelsHF ↗arXiv ↗
9 / 23

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号