TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Mar 16 – Mar 22, 2026

50 篇论文 · 按点赞排序

34

Can Vision-Language Models Solve the Shell Game?

Tiedong Liu, Wee Sun Lee

Vision-Language Models exhibit poor performance on visual entity tracking due to reliance on static features; a proposed method using spatiotemporal grounded chain-of-thought achieves high accuracy by generating object trajectories as intermediate states.

39Vision-Language Modelsspatiotemporal continuityHF ↗arXiv ↗
36

Cheers: Decoupling Patch Details from Semantic Representations Enables Unified Multimodal Comprehension and Generation

Yichen Zhang, Da Peng, Zonghao Guo +19 authors

Cheers is a unified multimodal model that decouples visual details from semantic representations using a vision tokenizer, LLM-based Transformer, and cascaded flow matching head to achieve efficient joint optimization for both visual understanding and generation tasks.

38unified multimodal modelvision tokenizerHF ↗arXiv ↗
38

Complementary Reinforcement Learning

Dilxat Muhtar, Jiashun Liu, Wei Gao +8 authors

Complementary RL enables efficient agent learning by synchronizing experience extraction with policy optimization through dual objectives that evolve together during training.

37Reinforcement LearningLLM-based agentsHF ↗arXiv ↗
41

Effective Distillation to Hybrid xLSTM Architectures

Lukas Hauzenberger, Niklas Schmidinger, Thomas Schmied +7 authors

There have been numerous attempts to distill quadratic attention-based large language models (LLMs) into sub-quadratic linearized architectures. However, despite extensive research, such distilled models often fail to match the performance of their teacher LLMs on various downstream tasks. We set out the goal of lossless distillation, which we define in terms of tolerance-corrected Win-and-Tie rates between student and teacher on sets of tasks. To this end, we introduce an effective distillation pipeline for xLSTM-based students. We propose an additional merging stage, where individually linearized experts are combined into a single model. We show the effectiveness of this pipeline by distilling base and instruction-tuned models from the Llama, Qwen, and Olmo families. In many settings, our xLSTM-based students recover most of the teacher's performance, and even exceed it on some downstream tasks. Our contributions are an important step towards more energy-efficient and cost-effective replacements for transformer-based LLMs.

34HF ↗arXiv ↗
44

LoST: Level of Semantics Tokenization for 3D Shapes

Niladri Shekhar Dutt, Zifan Shi, Paul Guerrero +4 authors

Level-of-Semantics Tokenization (LoST) improves 3D shape generation by ordering tokens based on semantic salience and using a novel relational alignment loss for better reconstruction and efficiency.

32tokenizationautoregressive modelsHF ↗arXiv ↗
45

When AI Navigates the Fog of War

Ming Li, Xirui Li, Tianyi Zhou

Large language models demonstrate varying capabilities in reasoning about unfolding geopolitical conflicts, showing strategic realism in structured settings but inconsistent performance in complex political environments.

32large language modelsgeopolitical predictionHF ↗arXiv ↗
47

OmniForcing: Unleashing Real-time Joint Audio-Visual Generation

Yaofeng Su, Yuming Li, Zeyue Xue +7 authors

OmniForcing distills a dual-stream bidirectional diffusion model into a streaming autoregressive generator while addressing training instability and synchronization issues through asymmetric alignment and specialized token mechanisms.

31diffusion modelsbidirectional attentionHF ↗arXiv ↗
48

daVinci-Env: Open SWE Environment Synthesis at Scale

Dayuan Fu, Shenyu Wu, Yunze Wu +11 authors

OpenSWE presents the largest transparent framework for software engineering agent training, featuring 45,320 executable environments and achieving state-of-the-art performance on SWE-bench Verified.

30software engineering agentsexecutable environmentsHF ↗arXiv ↗
2 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号