TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Sep 22 – Sep 28, 2025
本周最热155

Qwen3-Omni Technical Report

Jin Xu, Zhifang Guo, Hangrui Hu +35 authors

Qwen3-Omni, a multimodal model, achieves state-of-the-art performance across text, image, audio, and video, using a Thinker-Talker MoE architecture and a lightweight causal ConvNet for efficient streaming synthesis.

multimodal modelThinker-Talker MoE architectureaudio tasksaudio-visual benchmarksHF ↗arXiv ↗

50 篇论文 · 按点赞排序

02

Baseer: A Vision-Language Model for Arabic Document-to-Markdown OCR

Khalil Hennara, Muhammad Hreden, Mohamed Motasim Hamed +4 authors

Baseer, a vision-language model fine-tuned for Arabic document OCR, achieves state-of-the-art performance using a decoder-only strategy and a large-scale dataset, outperforming existing solutions with a WER of 0.25.

135Multimodal Large Language Modelsvision-language modelHF ↗arXiv ↗
06

LIMI: Less is More for Agency

Yang Xiao, Mohan Jiang, Jie Sun +18 authors

LIMI demonstrates that sophisticated agentic intelligence can emerge from minimal, strategically curated demonstrations, outperforming data-intensive models on agency benchmarks.

104Agencyautonomous agentsHF ↗arXiv ↗
08

Video models are zero-shot learners and reasoners

Thaddäus Wiedemer, Yuxuan Li, Paul Vicol +6 authors

Veo 3, a generative video model, exhibits zero-shot capabilities across various visual tasks, suggesting a trajectory towards becoming a unified, generalist vision foundation model.

101Large Language ModelsLLMsHF ↗arXiv ↗
09

Tree Search for LLM Agent Reinforcement Learning

Yuxiang Ji, Ziyu Ma, Yong Wang +3 authors

Tree-based Group Relative Policy Optimization (Tree-GRPO) enhances reinforcement learning for large language models by using tree search to improve rollouts and estimate grouped relative advantages, outperforming chain-based methods.

92reinforcement learninglarge language modelsHF ↗arXiv ↗
10

Seedream 4.0: Toward Next-generation Multimodal Image Generation

Team Seedream, Yunpeng Chen, Yu Gao +47 authors

Seedream 4.0 is a high-performance multimodal image generation system that integrates text-to-image synthesis, image editing, and multi-image composition using a diffusion transformer and VAE, achieving state-of-the-art results with efficient training and inference.

89diffusion transformerVAEHF ↗arXiv ↗
11

Reinforcement Learning on Pre-Training Data

Siheng Li, Kejiao Li, Zenan Xu +33 authors

Reinforcement Learning on Pre-Training data (RLPT) optimizes large language models by autonomously exploring meaningful trajectories in pre-training data, improving generalizable reasoning skills without human annotation.

67Reinforcement Learning on Pre-Training dataRLPTHF ↗arXiv ↗
15

EmbeddingGemma: Powerful and Lightweight Text Representations

Henrique Schechter Vera, Sahil Dua, Biao Zhang +85 authors

EmbeddingGemma, a lightweight text embedding model based on Gemma 3, achieves state-of-the-art performance with fewer parameters through encoder-decoder initialization, geometric embedding distillation, and spread-out regularization.

51Gemma 3encoder-decoder initializationHF ↗arXiv ↗
16

Do You Need Proprioceptive States in Visuomotor Policies?

Juntu Zhao, Wenbo Lu, Di Zhang +10 authors

A state-free policy using only visual observations achieves better spatial generalization and data efficiency in robot manipulation tasks compared to state-based policies.

50imitation-learning-based visuomotor policiesproprioceptive state inputHF ↗arXiv ↗
18

SIM-CoT: Supervised Implicit Chain-of-Thought

Xilin Wei, Xiaoran Liu, Yuhang Zang +5 authors

SIM-CoT, a plug-and-play training module, introduces step-level supervision to stabilize and enrich the latent reasoning space of implicit Chain-of-Thought methods, enhancing their performance and efficiency.

43implicit Chain-of-Thoughtexplicit Chain-of-ThoughtHF ↗arXiv ↗
20

AutoIntent: AutoML for Text Classification

Ilya Alekseev, Roman Solomatin, Darina Rustamova +1 authors

AutoIntent is an automated machine learning tool for text classification that offers end-to-end automation, including embedding model selection, classifier optimization, and decision threshold tuning, and supports multi-label classification and out-of-scope detection.

37embedding model selectionclassifier optimizationHF ↗arXiv ↗
21

ARE: Scaling Up Agent Environments and Evaluations

Pierre Andrews, Amine Benhalloum, Gerard Moreno-Torres Bertran +21 authors

Meta Agents Research Environments (ARE) facilitate the creation and execution of complex environments for agent research, and Gaia2, a benchmark built on ARE, evaluates general agent capabilities in dynamic, asynchronous settings.

36Meta Agents Research EnvironmentsAREHF ↗arXiv ↗
27

SPATIALGEN: Layout-guided 3D Indoor Scene Generation

Chuan Fang, Heng Li, Yixun Liang +6 authors

SpatialGen, a multi-view multi-modal diffusion model, generates realistic and semantically consistent 3D indoor scenes using a large synthetic dataset, outperforming previous methods.

28diffusion model3D indoor scenesHF ↗arXiv ↗
28

MAPO: Mixed Advantage Policy Optimization

Wenke Huang, Quan Zhang, Yiyang Fang +11 authors

Mixed Advantage Policy Optimization (MAPO) dynamically reweights the advantage function to improve trajectory ranking in reinforcement learning for foundation models.

27Group Relative Policy Optimization (GRPO)advantage functionHF ↗arXiv ↗
1 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号