TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Dec 15 – Dec 21, 2025
本周最热173

Kling-Omni Technical Report

Kling Team, Jialu Chen, Yuanzheng Ci +65 authors

Kling-Omni is an end-to-end generative framework that unifies video generation, editing, and reasoning from multimodal inputs to produce high-fidelity video content.

generative frameworkmultimodal visual language inputsvideo generationvideo editingHF ↗arXiv ↗

50 篇论文 · 按点赞排序

02

Memory in the Age of AI Agents

Yuyang Hu, Shichun Liu, Yanwei Yue +44 authors

This survey provides an updated overview of agent memory research, distinguishing its forms, functions, and dynamics, and highlights emerging research directions.

160agent memoryLLM memoryHF ↗arXiv ↗
03

Step-GUI Technical Report

Haolong Yan, Jia Wang, Xin Huang +94 authors

A self-evolving training pipeline with the Calibrated Step Reward System and GUI-MCP protocol improve GUI automation efficiency, accuracy, and privacy in real-world scenarios.

134multimodal large language modelsGUI automationHF ↗arXiv ↗
05

MMGR: Multi-Modal Generative Reasoning

Zefan Cai, Haoyi Qiu, Tianyi Ma +9 authors

MMGR evaluates video and image models across reasoning abilities in multiple domains, revealing performance gaps and highlighting limitations in perceptual data and causal correctness.

121Frechet Video Distance (FVD)MMGRHF ↗arXiv ↗
07

Adaptation of Agentic AI

Pengcheng Jiang, Jiacheng Lin, Zhiyi Shi +31 authors

This paper presents a framework for agent and tool adaptation in agentic AI systems, clarifying design strategies and identifying open challenges for improving AI capabilities.

111agentic AI systemsfoundation modelsHF ↗arXiv ↗
08

Towards Scalable Pre-training of Visual Tokenizers for Generation

Jingfeng Yao, Yuda Song, Yucong Zhou +1 authors

A unified visual tokenizer pre-training framework (VTP) improves generative performance by optimizing image-text contrastive, self-supervised, and reconstruction losses, leading to better scaling properties and higher zero-shot accuracy and faster convergence.

108latent spacevisual tokenizersHF ↗arXiv ↗
10

Next-Embedding Prediction Makes Strong Vision Learners

Sihan Xu, Ziqiao Ma, Wenhao Chai +5 authors

Generative pretraining using next embedding prediction outperforms traditional self-supervised methods in visual learning tasks, achieving high accuracy on ImageNet and effective transfer to semantic segmentation.

91generative pretrainingpredictive tasksHF ↗arXiv ↗
11

LLaDA2.0: Scaling Up Diffusion Language Models to 100B

Tiwei Bie, Maosong Cao, Kun Chen +28 authors

LLaDA2.0 converts auto-regressive models into discrete diffusion large language models with a novel training scheme, achieving superior performance and efficiency at scale.

89discrete diffusion large language modelsdLLMHF ↗arXiv ↗
19

DentalGPT: Incentivizing Multimodal Complex Reasoning in Dentistry

Zhenyang Cai, Jiaming Zhang, Junjie Zhao +21 authors

DentalGPT, a specialized multimodal large language model for dentistry, achieves superior performance in disease classification and dental visual question answering through high-quality domain data and staged adaptation techniques.

46multimodal large language modelsdental visual question answeringHF ↗arXiv ↗
20

DEER: Draft with Diffusion, Verify with Autoregressive Models

Zicong Cheng, Guo-Wei Yang, Jia Li +3 authors

Diffusion large language model drafters address limitations of autoregressive decoding in speculative frameworks through parallel processing and improved probabilistic modeling, achieving significant speedup in code generation tasks.

45speculative decodingautoregressive decodingHF ↗arXiv ↗
21

Universal Reasoning Model

Zitian Gao, Lynx Chen, Yihao Xiao +5 authors

The Universal Reasoning Model enhances Universal Transformers with short convolution and truncated backpropagation to improve reasoning performance on ARC-AGI tasks.

44Universal TransformersARC-AGIHF ↗arXiv ↗
22

Fast and Accurate Causal Parallel Decoding using Jacobi Forcing

Lanxiang Hu, Siqi Kou, Yichao Fu +5 authors

Parallel decoding for transformers is improved through a progressive distillation method that maintains causal inference properties while achieving significant speedup, combined with multi-block decoding that further reduces inference latency.

44transformer-based large modelsparallel decodingHF ↗arXiv ↗
23

KlingAvatar 2.0 Technical Report

Kling Team, Jialu Chen, Yikang Ding +25 authors

KlingAvatar 2.0 addresses inefficiencies in generating long-duration, high-resolution videos by using a spatio-temporal cascade framework with a Co-Reasoning Director and Negative Director for improved multimodal instruction alignment.

44spatio-temporal cascadeblueprint video keyframesHF ↗arXiv ↗
1 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号