TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Jan 26 – Feb 1, 2026

50 篇论文 · 按点赞排序

31

Shaping capabilities with token-level data filtering

Neil Rathi, Alec Radford

Token filtering during pretraining effectively reduces unwanted language model capabilities while maintaining alignment, becoming more effective at larger scales and tolerating noisy labels with sufficient compute.

32language modelspretrainingHF ↗arXiv ↗
32

Masked Depth Modeling for Spatial Perception

Bin Tan, Changjiang Sun, Xiage Qin +8 authors

LingBot-Depth is a depth completion model that uses visual context to refine depth maps through masked depth modeling and automated data curation for improved spatial perception in robotics and autonomous systems.

31depth completionmasked depth modelingHF ↗arXiv ↗
34

Post-LayerNorm Is Back: Stable, ExpressivE, and Deep

Chen Chen, Lai Wei

A novel Post-LayerNorm Transformer architecture called Keel addresses training instability in extremely deep networks by replacing residual connections with Highway-style connections, enabling stable training beyond 1000 layers.

27Post-LayerNormPre-LayerNormHF ↗arXiv ↗
35

VIBEVOICE-ASR Technical Report

Zhiliang Peng, Jianwei Yu, Yaoyao Chang +21 authors

VibeVoice-ASR is a unified end-to-end speech understanding framework that processes long-form audio without chunking, supports multiple languages and code-switching, and uses prompt-based context injection for improved domain-specific accuracy.

25speech understanding frameworkVibeVoiceHF ↗arXiv ↗
38

Self-Refining Video Sampling

Sangwon Jang, Taekyung Ki, Jaehyeong Jo +3 authors

Self-refining video sampling improves motion coherence and physics alignment by using a pre-trained video generator as its own denoising autoencoder for iterative refinement with uncertainty-aware region selection.

25video generatorsdenoising autoencoderHF ↗arXiv ↗
40

Exploring Reasoning Reward Model for Agents

Kaixuan Fan, Kaituo Feng, Manyuan Zhang +7 authors

Agent-RRM, a multi-faceted reward model, provides structured feedback for agentic trajectories through reasoning traces, critiques, and performance scores, with unified feedback integration showing superior performance across diverse benchmarks.

24Agentic Reinforcement Learningreward modelHF ↗arXiv ↗
42

LoL: Longer than Longer, Scaling Video Generation to Hour

Justin Cui, Jie Wu, Ming Li +6 authors

Researchers developed a method to overcome sink-collapse in autoregressive video generation by addressing the conflict between Rotary Position Embedding and multi-head attention mechanisms, enabling real-time streaming of videos up to 12 hours long.

23Rotary Position Embeddingmulti-head attentionHF ↗arXiv ↗
43

Discovering Hidden Gems in Model Repositories

Jonathan Kahana, Eliahu Horwitz, Yedid Hoshen

Hidden superior models exist in public repositories but are overlooked due to inefficient discovery methods; a multi-armed bandit approach using shared query sets and aggressive elimination significantly accelerates identification of top-performing models.

22multi-armed banditsequential halvingHF ↗arXiv ↗
44

Agentic Very Long Video Understanding

Aniket Rege, Arka Sadhu, Yuliang Li +5 authors

An agentic framework using entity scene graphs enables long-horizon video understanding with structured search, temporal reasoning, and cross-modal capabilities for extended visual and audio interpretation.

22entity scene graphsagentic frameworkHF ↗arXiv ↗
49

Endless Terminals: Scaling RL Environments for Terminal Agents

Kanishk Gandhi, Shivam Garg, Noah D. Goodman +1 authors

Endless Terminals introduces an autonomous pipeline for generating procedural terminal tasks that significantly improves agent performance on both synthetic and human-curated benchmarks through scalable reinforcement learning environments.

20reinforcement learningPPOHF ↗arXiv ↗
2 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号