TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Jun 1 – Jun 7, 2026

50 篇论文 · 按点赞排序

32

Task-Focused Memorization for Multimodal Agents

Tao Zou, Yichen He, Tian Qiu +2 authors

A reinforcement-learning-based framework called TaskMem is introduced to dynamically determine what information to store in long-term memory for multimodal agents, improving performance on streaming video benchmarks.

41multimodal agentslong-term memoryHF ↗arXiv ↗
34

Qwen-Image-Flash: Beyond Objective Design

Tianhe Wu, Kun Yan, Zikai Zhou +21 authors

Few-step distillation for visual generative models benefits from systematic investigation of training recipes beyond just distillation objectives, leading to improved student performance through optimized data composition, teacher guidance, and task mixture.

39few-step distillationvisual generative modelsHF ↗arXiv ↗
37

dMoE: dLLMs with Learnable Block Experts

Sicheng Feng, Zigeng Chen, Gongfan Fang +2 authors

Diffusion large language models combined with mixture-of-experts architectures face a mismatch between block parallel decoding and token-level expert selection, which dMoE addresses by aggregating token-level distributions into block-level routing to reduce activated experts and improve efficiency.

38Diffusion Large Language Modelsautoregressive modelsHF ↗arXiv ↗
39

Draft-OPD: On-Policy Distillation for Speculative Draft Models

Haodi Lei, Yafy Li, Haoran Zhang +8 authors

Speculative decoding uses a lightweight draft model to accelerate large language model inference, but supervised fine-tuning plateaus due to offline-to-inference mismatch, which is addressed through on-policy distillation with target-assisted rollouts and error replay.

37speculative decodingdraft modelHF ↗arXiv ↗
40

NITP: Next Implicit Token Prediction for LLM Pre-training

Xiangdong Zhang, Debing Zhang, Shaofeng Zhang +3 authors

Next Implicit Token Prediction enhances language model training by adding dense continuous supervision in representation space, improving generalization and performance across model sizes with minimal computational overhead.

37next-token predictionlanguage modelsHF ↗arXiv ↗
44

Where to Look: Can Foundation Models Reach a Target Viewpoint Through Active Exploration?

Liyang Li, Muzhi Zhu, Zhiyue Zhao +5 authors

Target Viewpoint Reproduction task challenges foundation models to actively adjust 3D viewpoints to match target images, revealing limitations in visual history processing and embodied movement mapping, with a unified post-training framework improving success rates through various training methods.

32Target Viewpoint ReproductionTVRBenchHF ↗arXiv ↗
46

Streaming Communication in Multi-Agent Reasoning

Zhen Yang, Xiaogang Xu, Wen Wang +3 authors

StreamMA enables efficient multi-agent reasoning by streaming intermediate results and leveraging reliable early steps to improve both latency and effectiveness across various reasoning tasks.

30multi-agent reasoning systemsgenerate-then-transfer paradigmHF ↗arXiv ↗
50

Self-Distilled Policy Gradient

Yifeng Liu, Shiyuan Zhang, Yifan Zhang +1 authors

A self-distilled policy-gradient framework combines on-policy self-distillation with verifier advantages and KL regularization to improve reinforcement learning stability and performance.

28self-distillationpolicy-gradientHF ↗arXiv ↗
2 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号