TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Nov 3 – Nov 9, 2025

50 篇论文 · 按点赞排序

31

LongCat-Flash-Omni Technical Report

Meituan LongCat Team, Bairui Wang, Bayan +129 authors

LongCat-Flash-Omni, a 560 billion parameter omni-modal model, achieves real-time audio-visual interaction through curriculum-inspired training and modality-decoupled parallelism.

27curriculum-inspired progressive trainingShortcut-connected Mixture-of-Experts (MoE)HF ↗arXiv ↗
35

The Collaboration Gap

Tim R. Davidson, Adam Fourney, Saleema Amershi +3 authors

Evaluation of agent-based systems reveals a collaboration gap where solo-performing models degrade in pairings, suggesting the need for collaboration-aware evaluation and training strategies.

23agent-based systemscollaborative maze-solving benchmarkHF ↗arXiv ↗
37

Revisiting Multimodal Positional Encoding in Vision-Language Models

Jie Huang, Xuejing Liu, Sibo Song +4 authors

A comprehensive analysis of multimodal Rotary Positional Embedding (RoPE) leads to the proposal of Multi-Head RoPE (MHRoPE) and MRoPE-Interleave (MRoPE-I), which improve multimodal understanding in vision-language models.

23multimodal position encodingRotary Positional Embedding (RoPE)HF ↗arXiv ↗
38

OpenSIR: Open-Ended Self-Improving Reasoner

Wai-Chung Kwan, Joshua Ong Jun Leang, Pavlos Vougiouklis +3 authors

OpenSIR is a self-play framework that enables large language models to improve their reasoning abilities through open-ended problem generation and solving without external supervision.

21large language modelreinforcement learningHF ↗arXiv ↗
48

Orion-MSP: Multi-Scale Sparse Attention for Tabular In-Context Learning

Mohamed Bouadi, Pratinav Seth, Aditya Tanna +1 authors

Orion-MSP, a tabular in-context learning architecture, addresses limitations in current models by incorporating multi-scale processing, block-sparse attention, and a Perceiver-style memory, achieving state-of-the-art performance on diverse benchmarks.

15tabular in-context learningTabPFNHF ↗arXiv ↗
49

Higher-order Linear Attention

Yifan Zhang, Zhen Qin, Quanquan Gu

Higher-order Linear Attention (HLA) is a scalable, causal, and efficient mechanism for long-context autoregressive language models, combining attention-like mixing with recurrent architecture efficiency.

15scaled dot-product attentionlinear-time attentionHF ↗arXiv ↗
2 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号