TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

606 篇论文 · 按点赞排序

513

UI-TARS: Pioneering Automated GUI Interaction with Native Agents

Yujia Qin, Yining Ye, Junjie Fang +32 authors

UI-TARS, a native GUI agent model using screenshots as input, outperforms commercial models in various benchmarks through enhanced perception, unified action modeling, system-2 reasoning, and iterative training with reflective online traces.

64native GUI agent modelcontext-aware understandingHF ↗arXiv ↗
515

ReTool: Reinforcement Learning for Strategic Tool Use in LLMs

Jiazhan Feng, Shijue Huang, Xingwei Qu +6 authors

ReTool, a tool-integrated learning framework, enhances reasoning models with real-time code execution and reinforcement learning, significantly improving performance in structured problem-solving tasks like mathematical reasoning.

63reasoning modelsreinforcement learningHF ↗arXiv ↗
516

One RL to See Them All: Visual Triple Unified Reinforcement Learning

Yan Ma, Linge Du, Xuyang Shen +7 authors

A unified reinforcement learning system, V-Triune, combines visual reasoning and perception tasks in vision-language models through a single training pipeline, achieving significant improvements across various tasks.

63visual triple unified reinforcement learningsample-level data formattingHF ↗arXiv ↗
517

S*: Test Time Scaling for Code Generation

Dacheng Li, Shiyi Cao, Chengkun Cao +6 authors

A hybrid test-time scaling framework improves code generation coverage and accuracy across various models and domains.

63hybrid test-time scaling frameworkparallel scalingHF ↗arXiv ↗
519

DR Tulu: Reinforcement Learning with Evolving Rubrics for Deep Research

Rulin Shao, Akari Asai, Shannon Zejiang Shen +18 authors

Reinforcement Learning with Evolving Rubrics (RLER) enables training of deep research models for long-form tasks, outperforming existing models and proprietary systems while being more cost-effective.

63Reinforcement Learning with Verifiable Rewards (RLVR)Reinforcement Learning with Evolving Rubrics (RLER)HF ↗arXiv ↗
524

Ovis-U1 Technical Report

Guo-Hua Wang, Shanshan Zhao, Xinjie Zhang +9 authors

Ovis-U1, a 3-billion-parameter unified model, integrates multimodal understanding, text-to-image generation, and image editing using a diffusion-based visual decoder and bidirectional token refiner, achieving state-of-the-art performance across various benchmarks.

63diffusion-based visual decoderbidirectional token refinerHF ↗arXiv ↗
525

4DNeX: Feed-Forward 4D Generative Modeling Made Easy

Zhaoxi Chen, Tianqi Liu, Long Zhuo +6 authors

4DNeX generates high-quality dynamic 3D scene representations from a single image using a fine-tuned pretrained video diffusion model, outperforming existing methods in efficiency and generalizability.

62feed-forward framework4D scene representationsHF ↗arXiv ↗
530

LIMO: Less is More for Reasoning

Yixin Ye, Zhen Huang, Yang Xiao +3 authors

LIMO, a new model, achieves high mathematical reasoning performance using minimal training data, challenging the notion that extensive datasets are necessary for complex reasoning.

62LIMOLIMO HypothesisHF ↗arXiv ↗
531

Process Reinforcement through Implicit Rewards

Ganqu Cui, Lifan Yuan, Zefan Wang +20 authors

PRIME leverages implicit process rewards to improve the reinforcement learning of large language models, achieving better performance with less data compared to traditional methods.

62dense process rewardssparse outcome-level rewardsHF ↗arXiv ↗
532

Enhancing Human-Like Responses in Large Language Models

Ethem Yağız Çalık, Talha Rüzgar Akkuş

Advancements in enhancing natural language understanding, conversational coherence, and emotional intelligence in large language models improve user interactions and expand AI applications, while future research will address ethical implications and biases.

62large language modelsfine-tuningHF ↗arXiv ↗
539

Towards Best Practices for Open Datasets for LLM Training

Stefan Baack, Stella Biderman, Kasia Odrozek +36 authors

Many AI companies are training their large language models (LLMs) on data without the permission of the copyright owners. The permissibility of doing so varies by jurisdiction: in countries like the EU and Japan, this is allowed under certain restrictions, while in the United States, the legal landscape is more ambiguous. Regardless of the legal status, concerns from creative producers have led to several high-profile copyright lawsuits, and the threat of litigation is commonly cited as a reason for the recent trend towards minimizing the information shared about training datasets by both corporate and public interest actors. This trend in limiting data information causes harm by hindering transparency, accountability, and innovation in the broader ecosystem by denying researchers, auditors, and impacted individuals access to the information needed to understand AI models. While this could be mitigated by training language models on open access and public domain data, at the time of writing, there are no such models (trained at a meaningful scale) due to the substantial technical and sociological challenges in assembling the necessary corpus. These challenges include incomplete and unreliable metadata, the cost and complexity of digitizing physical records, and the diverse set of legal and technical skills required to ensure relevance and responsibility in a quickly changing landscape. Building towards a future where AI systems can be trained on openly licensed data that is responsibly curated and governed requires collaboration across legal, technical, and policy domains, along with investments in metadata standards, digitization, and fostering a culture of openness.

61HF ↗arXiv ↗
18 / 21

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号