TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Aug 11 – Aug 17, 2025

50 篇论文 · 按点赞排序

31

OpenCUA: Open Foundations for Computer-Use Agents

Xinyuan Wang, Bowen Wang, Dunjie Lu +36 authors

OpenCUA is an open-source framework for vision-language models as computer-use agents, featuring an annotation infrastructure, a large-scale dataset, and a scalable pipeline that achieves state-of-the-art performance.

33vision-language modelscomputer-use agentsHF ↗arXiv ↗
36

Reinforcement Learning in Vision: A Survey

Weijia Wu, Chen Gao, Joya Chen +6 authors

This survey synthesizes recent advancements in visual reinforcement learning, covering policy optimization strategies, thematic pillars, and evaluation protocols, while highlighting open challenges.

30reinforcement learningvisual intelligenceHF ↗arXiv ↗
40

Train Long, Think Short: Curriculum Learning for Efficient Reasoning

Hasan Abed Al Kader Hammoud, Kumail Alhamoud, Abed Hammoud +3 authors

A curriculum learning strategy using Group Relative Policy Optimization (GRPO) enhances the reasoning abilities of large language models by progressively tightening token budgets, improving accuracy and token efficiency.

27Group Relative Policy Optimization (GRPO)curriculum learningHF ↗arXiv ↗
42

StableAvatar: Infinite-Length Audio-Driven Avatar Video Generation

Shuyuan Tu, Yueming Pan, Yinming Huang +6 authors

StableAvatar, an end-to-end video diffusion transformer, synthesizes infinite-length high-quality audio-driven avatar videos with natural synchronization and identity consistency using a Time-step-aware Audio Adapter and Audio Native Guidance Mechanism.

26diffusion modelsaudio-driven avatar video generationHF ↗arXiv ↗
46

Spectrum Projection Score: Aligning Retrieved Summaries with Reader Models in Retrieval-Augmented Generation

Zhanghao Hu, Qinglin Zhu, Siya Qi +3 authors

Large Language Models (LLMs) have shown improved generation performance through retrieval-augmented generation (RAG) following the retriever-reader paradigm, which supplements model inputs with externally retrieved knowledge. However, prior work often evaluates RAG holistically, assessing the retriever and reader jointly, making it difficult to isolate the true contribution of retrieval, particularly given the prompt sensitivity of LLMs used as readers. We introduce Spectrum Projection Score (SPS), a lightweight, supervision-free metric that allows the reader to gauge the semantic alignment of a retrieved summary with its hidden representation by comparing the area formed by generated tokens from the summary, and the principal directions of subspace in the reader and to measure the relevance. Building on SPS we present xCompress, an inference time controller framework that dynamically samples, ranks, and compresses retrieval summary candidates. Extensive experiments on five QA benchmarks with four open source LLMs show that SPS not only enhances performance across a range of tasks but also provides a principled perspective on the interaction between retrieval and generation.

19HF ↗arXiv ↗
2 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号