TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Oct 27 – Nov 2, 2025
本周最热236

Scaling Latent Reasoning via Looped Language Models

Rui-Jie Zhu, Zixuan Wang, Kai Hua +30 authors

LoopLM, a family of pre-trained Looped Language Models, enhances reasoning by integrating iterative computation and entropy regularization during pre-training, achieving superior performance with better knowledge manipulation.

Looped Language ModelsLoopLMiterative computationlatent spaceHF ↗arXiv ↗

50 篇论文 · 按点赞排序

06

Emu3.5: Native Multimodal Models are World Learners

Yufeng Cui, Honghao Chen, Haoge Deng +20 authors

Emu3.5, a large-scale multimodal world model, predicts next states in vision and language, enhanced with reinforcement learning and Discrete Diffusion Adaptation for efficient inference, achieving strong performance in various multimodal tasks.

117multimodal world modelnext-token predictionHF ↗arXiv ↗
07

Tongyi DeepResearch Technical Report

Tongyi DeepResearch Team, Baixuan Li, Bo Zhang +53 authors

Tongyi DeepResearch, a large language model with agentic capabilities, achieves top performance in various deep research tasks through an end-to-end training framework and automatic data synthesis.

105agentic large language modelend-to-end training frameworkHF ↗arXiv ↗
08

DeepAgent: A General Reasoning Agent with Scalable Toolsets

Xiaoxi Li, Wenxiang Jiao, Jiarui Jin +8 authors

DeepAgent, an end-to-end deep reasoning agent, autonomously performs thinking, tool discovery, and action execution using memory folding and reinforcement learning, outperforming baselines in various tool-use and application tasks.

103DeepAgentautonomous thinkingHF ↗arXiv ↗
14

The Principles of Diffusion Models

Chieh-Hsin Lai, Yang Song, Dongjun Kim +2 authors

Diffusion models are explored through variational, score-based, and flow-based perspectives, focusing on their mathematical foundations and applications in controllable generation and efficient sampling.

64diffusion modelsvariational viewHF ↗arXiv ↗
15

RoboOmni: Proactive Robot Manipulation in Omni-modal Context

Siyin Wang, Jinlan Fu, Feihong Liu +11 authors

RoboOmni, a Perceiver-Thinker-Talker-Executor framework using end-to-end omni-modal LLMs, improves robotic manipulation by inferring user intentions from spoken dialogue, environmental sounds, and visual cues.

62Multimodal Large Language ModelsVision-Language-Action modelsHF ↗arXiv ↗
16

FARMER: Flow AutoRegressive Transformer over Pixels

Guangting Zheng, Qinyu Zhao, Tao Yang +6 authors

FARMER, a unified generative framework combining Normalizing Flows and Autoregressive models, achieves competitive image synthesis performance with exact likelihoods and scalable training.

59Normalizing FlowsAutoregressive modelsHF ↗arXiv ↗
18

Video-As-Prompt: Unified Semantic Control for Video Generation

Yuxuan Bian, Xin Chen, Zenan Li +4 authors

Video-As-Prompt (VAP) uses a reference video to guide a frozen Video Diffusion Transformer via a Mixture-of-Transformers expert, achieving state-of-the-art results in semantic-controlled video generation with strong zero-shot generalization.

50Video-As-PromptVideo Diffusion TransformerHF ↗arXiv ↗
23

Uniform Discrete Diffusion with Metric Path for Video Generation

Haoge Deng, Ting Pan, Fan Zhang +8 authors

URSA, a discrete generative model, bridges the gap with continuous approaches in video generation by using iterative refinement, linearized metric paths, and resolution-dependent timestep shifting, achieving performance comparable to state-of-the-art continuous methods.

43discrete generative modelingUniform discRete diffuSion with metric pAth (URSA)HF ↗arXiv ↗
24

Reasoning-Aware GRPO using Process Mining

Taekhyun Park, Yongjae Lee, Hyerim Bae

PM4GRPO, a reasoning-aware Group Relative Policy Optimization, enhances policy models by incorporating process mining to align reasoning with a teacher model, outperforming existing methods.

42reinforcement learningmulti-step reasoningHF ↗arXiv ↗
27

WorldGrow: Generating Infinite 3D World

Sikuang Li, Chen Yang, Jiemin Fang +6 authors

WorldGrow, a hierarchical framework, generates large, continuous 3D environments with coherent geometry and realistic appearance using pre-trained 3D models and a coarse-to-fine generation strategy.

423D implicit representations3D foundation modelsHF ↗arXiv ↗
29

LongCat-Video Technical Report

Meituan LongCat Team, Xunliang Cai, Qilong Huang +8 authors

LongCat-Video, a 13.6B parameter video generation model based on the Diffusion Transformer framework, excels in efficient and high-quality long video generation across multiple tasks using unified architecture, coarse-to-fine generation, and block sparse attention.

41Diffusion TransformerText-to-VideoHF ↗arXiv ↗
1 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号