TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Jun 9 – Jun 15, 2025
本周最热265

Reinforcement Pre-Training

Qingxiu Dong, Li Dong, Yao Tang +4 authors

Reinforcement Pre-Training (RPT) improves language model accuracy through reinforcement learning and offers a scalable method for leveraging text data for general-purpose RL.

Reinforcement Pre-Training (RPT)next-token predictionreasoning taskreinforcement learning (RL)HF ↗arXiv ↗

50 篇论文 · 按点赞排序

06

MiniCPM4: Ultra-Efficient LLMs on End Devices

MiniCPM Team, Chaojun Xiao, Yuxuan Li +72 authors

MiniCPM4, a highly efficient large language model for end-side devices, achieves superior performance using innovations in sparse attention, pre-training datasets, training algorithms, and inference systems.

104InfLLM v2sparse attention mechanismHF ↗arXiv ↗
11

Magistral

Mistral-AI, Abhinav Rastogi, Albert Q. Jiang +97 authors

Magistral, a scalable reinforcement learning pipeline, demonstrates that RL can enhance multimodal understanding and instruction following in large language models without requiring existing RL traces.

68reinforcement learningRLHF ↗arXiv ↗
17

Text-Aware Image Restoration with Diffusion Models

Jaewon Min, Jin Hyeon Kim, Paul Hyunbin Cho +6 authors

The proposed Text-Aware Image Restoration (TAIR) system integrates a multi-task diffusion framework with a text-spotting module to enhance both image recovery and textual fidelity, outperforming existing diffusion-based methods.

45diffusion-based restorationtext-image hallucinationHF ↗arXiv ↗
22

PlayerOne: Egocentric World Simulator

Yuanpeng Tu, Hao Luo, Xi Chen +3 authors

PlayerOne is an egocentric realistic world simulator that constructs and generates videos from user-captured images, using a coarse-to-fine training pipeline and advanced motion injection and reconstruction frameworks.

34egocentric realistic world simulatorcoarse-to-fine pipelineHF ↗arXiv ↗
26

Discrete Audio Tokens: More Than a Survey!

Pooneh Mousavi, Gallil Maimon, Adel Moumen +18 authors

A systematic review and benchmark of discrete audio tokenizers across speech, music, and general audio domains is presented, covering their taxonomy, evaluation metrics, and limitations.

31discrete audio tokensperceptual qualityHF ↗arXiv ↗
27

Ming-Omni: A Unified Multimodal Model for Perception and Generation

Inclusion AI, Biao Gong, Cheng Zou +55 authors

Ming-Omni is a unified multimodal model with dedicated encoders and modality-specific routers that can process images, text, audio, and video, and performs tasks like speech and image generation, context-aware chatting, and versatile image editing.

31multimodal modelencodersHF ↗arXiv ↗
1 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号