TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Aug 24 – Aug 30, 2026

50 篇论文 · 按点赞排序

36

OmniAssistBench: Assistant-style Interaction Benchmark for Omni-LLMs

Xianyun Sun, Chaoyou Fu, Zhengye Zhang +6 authors

OmniAssistBench evaluates real-time interactive video assistants by reverse-engineering multi-turn interaction videos, revealing that current omni-modal models struggle with visual prompts, context retention, and timely responses.

31omni-modal large language modelsOmniAssistBenchHF ↗arXiv ↗
41

RISE: Adaptive Imagination for World Action Models

Hongbo Lu, Liang Yao, Chenghao He +5 authors

RISE adaptively decides when to continue or stop imagination rollouts for planning by weighing expected benefit against cost, supported by a counterfactual driving dataset with expert annotations.

26World Action Modelsadaptive imaginationHF ↗arXiv ↗
42

ReWorld: An Interactive World Model with Long-Horizon Memory

Zhifei Chen, Luozhou Wang, Guibao Shen +8 authors

ReWorld separates short-horizon control and long-horizon memory during training, then bounds both at inference via mixed attention windows, a pose-indexed landmark bank, and distribution-matching LoRA distillation to enable real-time interactive world modeling with strong action fidelity and long-range recall.

24mixed per-head attention windowsglobal headsHF ↗arXiv ↗
46

AutoResearch: Insight In, Hallucination Out

Yiming Ren, Xiang Liu, Qumeng Sun +4 authors

AutoResearch is a two-stage autonomous system that grounds research ideas through integrated generation and evidence-based execution to improve experimental reliability and measurable outcomes.

22cross-modal retrievalsystems optimizationHF ↗arXiv ↗
48

Best Practice Critic Optimization

Penghui Qi, Xiangxin Zhou, Wee Sun Lee

BPCO stabilizes critic-based reinforcement learning for language models by combining bounded value predictions, Monte Carlo targets, and adaptive advantage estimation, matching group-based methods with single-response sampling.

18GRPODPPOHF ↗arXiv ↗
50

On-policy Distillation with Verifiable Reward

Wenze Lin, Jiale Zhao, Xitai Jiang +5 authors

OPDVR integrates on-policy distillation with verifiable rewards via a ReLU-gated implicit reward reformulation, improving reasoning performance without extra hyperparameters.

17Reinforcement Learning with Verifiable Rewardson-policy distillationHF ↗arXiv ↗
2 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号