TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Jul 20 – Jul 26, 2026

50 篇论文 · 按点赞排序

32

GigaChat Audio: Time-aware Large Audio Language Model

Aleksandr Kutsakov, Mariia Sadovina, Georgii Gospodinov +4 authors

A time-aware audio LLM answers questions with explicit timestamps across long recordings by interleaving periodic time markers with continuous audio tokens, achieving strong temporal grounding and supporting time-anchored descriptions.

37audio-conditioned LLMstemporal groundingHF ↗arXiv ↗
33

NVIDIA-labs OO Agents: Native Python Object-Oriented Agents

Paul Furgale, Severin Klingler, James Nolan +12 authors

NOOA treats AI agents as Python objects whose methods and fields define actions, state, and prompts, enabling deterministic testing and LLM-driven runtime completion within a unified programming model.

36agent-as-a-Python-objectLLM-driven agent loopHF ↗arXiv ↗
34

Self Gradient Forcing: Native Long Video Extrapolation

Junhao Zhuang, Shiyi Zhang, Yuxuan Bian +11 authors

Self Gradient Forcing restores missing supervision for historical memory writing in autoregressive video diffusion by using a two-pass strategy that trains context key-value representations via future-frame losses without full rollout backpropagation.

36Self Forcingautoregressive video diffusionHF ↗arXiv ↗
36

Qwen-Music Technical Report

Jin Xu, Kangdi Wang, Ruibin Yuan +24 authors

Qwen-Music is a music generation system that uses tokenized semantic representations, autoregressive language modeling with melody planning, and generative rendering to produce high-fidelity songs from text or reference audio.

34autoregressive modelingchain-of-thoughtHF ↗arXiv ↗
39

RecGPT-V3 Technical Report

Bowen Zheng, Chao Yi, Dian Chen +21 authors

RecGPT-V3 improves large-scale recommendation by using persistent user memory, hybrid text and semantic-ID reasoning, and compressed latent reasoning to boost engagement and cut serving costs.

30large language modelsrecommender systemsHF ↗arXiv ↗
40

Understanding Reasoning from Pretraining to Post-Training

Jingyan Shen, Ang Li, Salman Rahman +4 authors

Using chess and math as controlled testbeds, the study shows that pretraining loss predicts reinforcement learning gains and that RL amplifies or surfaces correct reasoning differently by difficulty, linking pretraining scale to post-training reasoning performance.

30reinforcement learninglarge language modelsHF ↗arXiv ↗
41

Group Entropy-Controlled Policy Optimization

Guangran Cheng, Chengqi Lyu, Songyang Gao +2 authors

GEPO extends GRPO with group-level entropy-conditioned advantage shaping to balance exploration and exploitation across heterogeneous tasks during LLM reinforcement learning.

29entropy controlreinforcement learningHF ↗arXiv ↗
43

Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction

Tencent WorkBuddy Bench Team, Siqi Cai, Shaopeng Chen +35 authors

We introduce Tencent WorkBuddy Bench, a multi-domain evaluation suite for coding agents; this report documents its construction methodology, scoring protocol, and a cross-model leaderboard. At its core is a unified evaluation framework for constructing and running distribution-informed coding-agent tasks across four work domains - Code, Web, Office, and Security. Rather than adapting public issue text, every task is reverse-engineered from a real commit, pull request, or business scenario and rewritten as a short, colloquial, role-played request, so that a task's prompt is not recoverable by web-searching the underlying issue, pull request, or commit thread. Because the dataset is released openly - task directories, environment images, evaluation harness, tests, and reference solutions - contamination resistance rests on this construction together with dataset versioning rather than on secrecy. The four subsets - repository-level engineering, front-end development, office and business workflows, and red-/blue-team security - probe complementary facets of real work, each with its own verification style. All are packaged in a uniform task-directory format and run, under a uniform and reproducible protocol, on two agent harnesses (CodeBuddy Code and Claude Code); the full open release makes the benchmark reproducible end to end and directly auditable, since any third party can re-run each task and inspect its content. Because each subset uses a different scoring instrument, scores are not comparable across subsets and the suite reports no suite-wide average. We report a cross-model leaderboard across several model families.

26HF ↗arXiv ↗
44

Recursive Harness Self-Improvement

Hyunin Lee, Jinglue Xu, Jeffrey Seely +3 authors

Recursive Harness Self-Improvement optimizes user-built agent harnesses via iterative prompt-level refinement to enhance execution traces and agent performance with minimal computation.

26Recursive Harness Self-Improvementharness-in-the-loop learningHF ↗arXiv ↗
46

LLMs Get Lost in Evolving User Intent

Jihoon Tack, Philippe Laban, Jennifer Neville

Large language models show significant performance drops when user intent evolves across multi-turn conversations, revealing a critical gap in tracking dynamic goals.

25LLMsmulti-turn conversationHF ↗arXiv ↗
48

An Exam for Active Observers

Jiarui Zhang, Muzi Tao, Shangshang Wang +3 authors

ActiveVision reveals that current multimodal large language models lack robust active visual observation, performing far below humans on tasks requiring repeated perception.

25multimodal large language modelsactive observationHF ↗arXiv ↗
49

SciForma: Structure-Faithful Generation of Scientific Diagrams

Yuxuan Luo, Peng Zhang, Xinjie Zhang +3 authors

SciForma improves scientific diagram generation by decomposing structural quality into component, arrow, and text axes, using multi-dimensional preference optimization and iterative editing to achieve high structural fidelity.

24Supervised fine-tuningMulti-Dimensional Conjunctive Preference OptimizationHF ↗arXiv ↗
50

Color Pass-Through via Camera-Display Coupling

Ruikang Li, Molin Li, Jiarui Wu +3 authors

A learned end-to-end framework couples smartphone cameras and displays to reduce color reproduction errors across the full capture-to-display pipeline.

21end-to-end learned frameworkend-to-end optimizationHF ↗arXiv ↗
2 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号