TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Nov 24 – Nov 30, 2025
本周最热173

General Agentic Memory Via Deep Research

B. Y. Yan, Chaofan Li, Hongjin Qian +2 authors

GAM, a novel framework that employs JIT compilation principles, improves memory efficiency and task completion by leveraging a lightweight memorizer and researcher in conjunction with reinforcement learning.

general agentic memoryGAMjust-in time compilationJIT compilationHF ↗arXiv ↗

50 篇论文 · 按点赞排序

02

SAM 3: Segment Anything with Concepts

Nicolas Carion, Laura Gustafson, Yuan-Ting Hu +35 authors

Segment Anything Model 3 achieves state-of-the-art performance in promptable concept segmentation and tracking by leveraging a unified model architecture with decoupled recognition and localization.

138Segment Anything ModelPromptable Concept SegmentationHF ↗arXiv ↗
03

Latent Collaboration in Multi-Agent Systems

Jiaru Zou, Xiyuan Yang, Ruizhong Qiu +10 authors

LatentMAS enables efficient, lossless collaboration among LLM agents in latent space, improving performance and reducing computational costs compared to text-based methods.

131multi-agent systemslarge language modelsHF ↗arXiv ↗
08

Unveiling Intrinsic Dimension of Texts: from Academic Abstract to Creative Story

Vladislav Pedashenko, Laida Kushnareva, Yana Khassan Nibal +5 authors

The study explores intrinsic dimension in large language models through cross-encoder analysis, linguistic features, and sparse autoencoders, revealing its independence from entropy, genre-specific stratification, and causal features related to text type.

91intrinsic dimensioncross-encoder analysisHF ↗arXiv ↗
09

Multimodal Evaluation of Russian-language Architectures

Artem Chervyakov, Ulyana Isaeva, Anton Emelyanov +15 authors

Mera Multi is an open multimodal evaluation framework for Russian-spoken architectures, addressing the lack of such benchmarks with 18 newly constructed tasks and a methodology to prevent benchmark leakage.

79multimodal large language modelsMera MultiHF ↗arXiv ↗
11

DR Tulu: Reinforcement Learning with Evolving Rubrics for Deep Research

Rulin Shao, Akari Asai, Shannon Zejiang Shen +18 authors

Reinforcement Learning with Evolving Rubrics (RLER) enables training of deep research models for long-form tasks, outperforming existing models and proprietary systems while being more cost-effective.

64Reinforcement Learning with Verifiable Rewards (RLVR)Reinforcement Learning with Evolving Rubrics (RLER)HF ↗arXiv ↗
12

MedSAM3: Delving into Segment Anything with Medical Concepts

Anglin Liu, Rundong Xue, Xu R. Cao +5 authors

MedSAM-3, a text-promptable medical segmentation model fine-tuned on SAM 3 architecture, achieves superior performance across various medical imaging modalities using semantic conceptual labels and multimodal large language models.

56Segment Anything Model (SAM)text promptableHF ↗arXiv ↗
18

Soft Adaptive Policy Optimization

Chang Gao, Chujie Zheng, Xiong-Hui Chen +7 authors

Soft Adaptive Policy Optimization (SAPO) enhances the stability and performance of reinforcement learning in large language models by adaptively attenuating off-policy updates with a smooth, temperature-controlled gate, leading to improved training stability and performance.

44reinforcement learninglarge language modelsHF ↗arXiv ↗
20

UltraFlux: Data-Model Co-Design for High-quality Native 4K Text-to-Image Generation across Diverse Aspect Ratios

Tian Ye, Song Fei, Lei Zhu

UltraFlux, a Flux-based DiT trained on a 4K dataset, addresses failures in diffusion transformers at 4K resolution through enhanced positional encoding, improved VAE compression, gradient rebalancing, and aesthetic curriculum learning, achieving superior performance compared to existing models.

38diffusion transformerstext-to-image generationHF ↗arXiv ↗
27

GigaWorld-0: World Models as Data Engine to Empower Embodied AI

GigaWorld Team, Angen Ye, Boyuan Wang +22 authors

GigaWorld-0 is a unified world model framework that integrates video generation and 3D modeling to produce high-quality, diverse, and physically plausible VLA data, enabling strong real-world performance in embodied AI without real-world training.

30GigaWorld-0GigaWorld-0-VideoHF ↗arXiv ↗
30

HunyuanVideo 1.5 Technical Report

Bing Wu, Chang Zou, Changlin Li +78 authors

HunyuanVideo 1.5 is a lightweight video generation model with state-of-the-art visual quality and motion coherence, using a DiT architecture with SSTA and an efficient video super-resolution network.

29DiT architectureselective and sliding tile attentionHF ↗arXiv ↗
1 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号