TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Jul 7 – Jul 13, 2025
本周最热170

MemOS: A Memory OS for AI System

Zhiyu Li, Shichao Song, Chenyang Xi +36 authors

MemOS, a memory operating system for Large Language Models, addresses memory management challenges by unifying plaintext, activation-based, and parameter-level memories, enabling efficient storage, retrieval, and continual learning.

Large Language ModelsArtificial General Intelligencememory managementlong-context reasoningHF ↗arXiv ↗

50 篇论文 · 按点赞排序

02

Scaling RL to Long Videos

Yukang Chen, Wei Huang, Baifeng Shi +11 authors

A framework scales vision-language models for long video reasoning using reinforcement learning, achieving strong performance on benchmarks and demonstrating consistent gains with increased video frames.

160vision-language modelsreinforcement learningHF ↗arXiv ↗
03

T-LoRA: Single Image Diffusion Model Customization Without Overfitting

Vera Soboleva, Aibek Alanov, Andrey Kuznetsov +1 authors

T-LoRA, a timestep-dependent low-rank adaptation framework, enhances diffusion model personalization with a dynamic fine-tuning strategy and orthogonal initialization, improving concept fidelity and text alignment in data-limited settings.

121diffusion model fine-tuningoverfittingHF ↗arXiv ↗
04

SingLoRA: Low Rank Adaptation Using a Single Matrix

David Bensaïd, Noam Rotstein, Roy Velich +2 authors

SingLoRA, a reformulated low-rank adaptation method, enhances parameter-efficient fine-tuning by learning a single low-rank matrix and its transpose, ensuring stable optimization and reducing parameter count.

116Low-Rank AdaptationLoRAHF ↗arXiv ↗
05

4KAgent: Agentic Any Image to 4K Super-Resolution

Yushen Zuo, Qi Zheng, Mingyang Wu +10 authors

4KAgent, a unified agentic super-resolution system, enhances low-resolution images to 4K using a profiling module, perception agent, and restoration agent, achieving state-of-the-art performance across various imaging domains.

107agentic super-resolutionProfilingHF ↗arXiv ↗
06

A Survey on Latent Reasoning

Rui-Jie Zhu, Tianhao Peng, Tianhao Cheng +30 authors

Latent reasoning in Large Language Models (LLMs) performs multi-step inference in continuous hidden states, enhancing reasoning capabilities without token-level supervision, and includes methodologies like activation-based recurrence and infinite-depth reasoning via masked diffusion models.

95chain-of-thought (CoT)latent reasoningHF ↗arXiv ↗
07

Should We Still Pretrain Encoders with Masked Language Modeling?

Hippolyte Gisserot-Boukhlef, Nicolas Boizard, Manuel Faysse +5 authors

A biphasic training strategy combining Causal Language Modeling and Masked Language Modeling yields optimal text representation performance, especially when initialized with pretrained CLM models.

81Masked Language ModelingCausal Language ModelingHF ↗arXiv ↗
08

MIRIX: Multi-Agent Memory System for LLM-Based Agents

Yu Wang, Xi Chen

MIRIX, a modular multi-agent memory system, enhances language models' memory capabilities by integrating diverse memory types and a dynamic framework, achieving superior performance in multimodal and long-form conversation benchmarks.

80modularmulti-agent memory systemHF ↗arXiv ↗
09

Skywork-R1V3 Technical Report

Wei Shen, Jiangbo Pei, Yi Peng +7 authors

Skywork-R1V3, an open-source vision-language model, enhances visual reasoning through a post-training reinforcement learning framework, achieving state-of-the-art performance on multimodal reasoning tasks.

73vision-language modelvisual reasoningHF ↗arXiv ↗
13

How to Train Your LLM Web Agent: A Statistical Diagnosis

Dheeraj Vattikonda, Santhoshi Ravichandran, Emiliano Penaloza +13 authors

A study on compute allocation for post-training LLM-based web agents finds that combining supervised fine-tuning with on-policy reinforcement learning improves performance and reduces computational costs compared to using either method alone.

52LLM-based web agentssupervised fine-tuningHF ↗arXiv ↗
15

Perception-Aware Policy Optimization for Multimodal Reasoning

Zhenhailong Wang, Xuehang Guo, Sofia Stoica +8 authors

PAPO, an extension of GRPO, enhances multimodal reasoning by integrating perception-aware supervision, leading to improved performance and reduced perception errors in tasks with high vision dependency.

48Reinforcement Learning with Verifiable RewardsLarge Language ModelsHF ↗arXiv ↗
27

PyVision: Agentic Vision with Dynamic Tooling

Shitian Zhao, Haoquan Zhang, Shaoheng Lin +4 authors

PyVision enables MLLMs to autonomously generate, execute, and refine Python-based tools for visual reasoning, achieving significant performance improvements across benchmarks.

33MLLMsvisual reasoningHF ↗arXiv ↗
28

StreamDiT: Real-Time Streaming Text-to-Video Generation

Akio Kodaira, Tingbo Hou, Ji Hou +2 authors

StreamDiT, a streaming video generation model using transformer-based diffusion with flow matching and adaLN DiT, achieves real-time performance at 16 FPS with 4B parameters and multistep distillation.

33transformer-based diffusion modelsflow matchingHF ↗arXiv ↗
1 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号