TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

本月最热346

Qwen3 Technical Report

An Yang, Anfeng Li, Baosong Yang +57 authors

Qwen3, a unified series of large language models, integrates thinking and non-thinking modes, reduces computational resources, and achieves state-of-the-art performance across various tasks and languages.

large language modelsdense architectureMixture-of-Expertthinking modeHF ↗arXiv ↗

50 篇论文 · 按点赞排序

05

Seed1.5-VL Technical Report

Dong Guo, Faming Wu, Feida Zhu +194 authors

Seed1.5-VL, a vision-language foundation model combining a vision encoder and a large MoE LLM, achieves state-of-the-art performance across various benchmarks and excels in multimodal reasoning tasks such as visual puzzles.

157vision-language foundation modelvision encoderHF ↗arXiv ↗
08

Emerging Properties in Unified Multimodal Pretraining

Chaorui Deng, Deyao Zhu, Kunchang Li +9 authors

BAGEL, an open-source foundational model trained on diverse multimodal data, significantly outperforms existing models in both generation and understanding tasks.

136multimodal understandingmultimodal generationHF ↗arXiv ↗
10

NovelSeek: When Agent Becomes the Scientist -- Building Closed-Loop System from Hypothesis to Verification

NovelSeek Team, Bo Zhang, Shiyang Feng +22 authors

Artificial Intelligence (AI) is accelerating the transformation of scientific research paradigms, not only enhancing research efficiency but also driving innovation. We introduce NovelSeek, a unified closed-loop multi-agent framework to conduct Autonomous Scientific Research (ASR) across various scientific research fields, enabling researchers to tackle complicated problems in these fields with unprecedented speed and precision. NovelSeek highlights three key advantages: 1) Scalability: NovelSeek has demonstrated its versatility across 12 scientific research tasks, capable of generating innovative ideas to enhance the performance of baseline code. 2) Interactivity: NovelSeek provides an interface for human expert feedback and multi-agent interaction in automated end-to-end processes, allowing for the seamless integration of domain expert knowledge. 3) Efficiency: NovelSeek has achieved promising performance gains in several scientific fields with significantly less time cost compared to human efforts. For instance, in reaction yield prediction, it increased from 27.6% to 35.4% in just 12 hours; in enhancer activity prediction, accuracy rose from 0.52 to 0.79 with only 4 hours of processing; and in 2D semantic segmentation, precision advanced from 78.8% to 81.0% in a mere 30 hours.

121HF ↗arXiv ↗
11

Chain-of-Model Learning for Language Model

Kaitao Song, Xiaohua Wang, Xu Tan +14 authors

A novel Chain-of-Model framework introduces hierarchical hidden state chains in Transformers to improve scaling efficiency and inference flexibility for language models.

121Chain-of-Model (CoM)Chain-of-Representation (CoR)HF ↗arXiv ↗
15

Web-Shepherd: Advancing PRMs for Reinforcing Web Agents

Hyungjoo Chae, Sunghwan Kim, Junhee Cho +18 authors

The paper introduces Web-Shepherd, a process reward model for web navigation, which improves accuracy and cost-effectiveness in step-level trajectory assessment compared to existing multimodal large language models.

104multimodal large language modelprocess reward modelHF ↗arXiv ↗
17

MMaDA: Multimodal Large Diffusion Language Models

Ling Yang, Ye Tian, Bowen Li +4 authors

MMaDA, a multimodal diffusion foundation model, achieves superior performance through a unified architecture, mixed long chain-of-thought fine-tuning, and a unified policy-gradient-based RL algorithm.

99multimodal diffusion foundation modelsunified diffusion architectureHF ↗arXiv ↗
21

Table-R1: Inference-Time Scaling for Table Reasoning

Zheyuan Yang, Lyuhao Chen, Arman Cohan +1 authors

Two post-training strategies, distillation and RLVR, enable inference-time scaling in table reasoning tasks, resulting in a model (Table-R1-Zero) that matches GPT-4.1's performance using fewer parameters and shows strong generalization.

93distillationreinforcement learningHF ↗arXiv ↗
23

Flow-GRPO: Training Flow Matching Models via Online RL

Jie Liu, Gongye Liu, Jiajun Liang +6 authors

Flow-GRPO combines online reinforcement learning with flow matching models through an ODE-to-SDE conversion and denoising reduction, improving sampling efficiency and performance across text-to-image tasks.

89Flow-GRPOonline reinforcement learningHF ↗arXiv ↗
30

Parallel Scaling Law for Language Models

Mouxiang Chen, Binyuan Hui, Zeyu Cui +5 authors

Parallel scaling (ParScale) improves inference efficiency by reusing existing parameters and executing multiple transformations in parallel, offering superior performance with reduced memory and latency compared to parameter scaling.

83parallel computationparallel scalingHF ↗arXiv ↗
1 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号