TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

December 2025
本月最热307

From Code Foundation Models to Agents and Applications: A Practical Guide to Code Intelligence

Jian Yang, Xianglong Liu, Weifeng Lv +68 authors

A comprehensive guide to code LLMs, covering their lifecycle from data curation to deployment, including techniques, trade-offs, and research-practice gaps.

Transformer-based architecturesHumanEvalprompting paradigmscode pre-trainingHF ↗arXiv ↗

50 篇论文 · 按点赞排序

07

Kling-Omni Technical Report

Kling Team, Jialu Chen, Yuanzheng Ci +65 authors

Kling-Omni is an end-to-end generative framework that unifies video generation, editing, and reasoning from multimodal inputs to produce high-fidelity video content.

173generative frameworkmultimodal visual language inputsHF ↗arXiv ↗
08

Qwen3-VL Technical Report

Shuai Bai, Yuxuan Cai, Ruizhe Chen +61 authors

Qwen3-VL, a vision-language model, excels in text and multimodal understanding through advanced architectures and larger contexts, achieving superior performance across benchmarks.

164vision-language modelinterleaved contextsHF ↗arXiv ↗
09

Memory in the Age of AI Agents

Yuyang Hu, Shichun Liu, Yanwei Yue +44 authors

This survey provides an updated overview of agent memory research, distinguishing its forms, functions, and dynamics, and highlights emerging research directions.

160agent memoryLLM memoryHF ↗arXiv ↗
11

Step-GUI Technical Report

Haolong Yan, Jia Wang, Xin Huang +94 authors

A self-evolving training pipeline with the Calibrated Step Reward System and GUI-MCP protocol improve GUI automation efficiency, accuracy, and privacy in real-world scenarios.

134multimodal large language modelsGUI automationHF ↗arXiv ↗
16

MMGR: Multi-Modal Generative Reasoning

Zefan Cai, Haoyi Qiu, Tianyi Ma +9 authors

MMGR evaluates video and image models across reasoning abilities in multiple domains, revealing performance gaps and highlighting limitations in perceptual data and causal correctness.

121Frechet Video Distance (FVD)MMGRHF ↗arXiv ↗
20

Adaptation of Agentic AI

Pengcheng Jiang, Jiacheng Lin, Zhiyi Shi +31 authors

This paper presents a framework for agent and tool adaptation in agentic AI systems, clarifying design strategies and identifying open challenges for improving AI capabilities.

111agentic AI systemsfoundation modelsHF ↗arXiv ↗
21

Stabilizing Reinforcement Learning with LLMs: Formulation and Practices

Chujie Zheng, Kai Dang, Bowen Yu +7 authors

The paper provides a theoretical foundation for optimizing sequence-level rewards in reinforcement learning using token-level objectives, highlighting the importance of techniques like importance sampling correction, clipping, and Routing Replay for stabilizing training, especially with large language models.

109reinforcement learninglarge language modelsHF ↗arXiv ↗
22

Towards Scalable Pre-training of Visual Tokenizers for Generation

Jingfeng Yao, Yuda Song, Yucong Zhou +1 authors

A unified visual tokenizer pre-training framework (VTP) improves generative performance by optimizing image-text contrastive, self-supervised, and reconstruction losses, leading to better scaling properties and higher zero-shot accuracy and faster convergence.

108latent spacevisual tokenizersHF ↗arXiv ↗
27

SemanticGen: Video Generation in Semantic Space

Jianhong Bai, Xiaoshi Wu, Xintao Wang +9 authors

SemanticGen addresses slow convergence and computational costs in video generation by using a two-stage diffusion model approach that first generates semantic features and then VAE latents, leading to faster convergence and high-quality results.

95VAE spaceVAE decoderHF ↗arXiv ↗
30

Next-Embedding Prediction Makes Strong Vision Learners

Sihan Xu, Ziqiao Ma, Wenhao Chai +5 authors

Generative pretraining using next embedding prediction outperforms traditional self-supervised methods in visual learning tasks, achieving high accuracy on ImageNet and effective transfer to semantic segmentation.

91generative pretrainingpredictive tasksHF ↗arXiv ↗
1 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号