TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Oct 20 – Oct 26, 2025
本周最热151

A Theoretical Study on Bridging Internal Probability and Self-Consistency for LLM Reasoning

Zhi Zhou, Yuhao Tan, Zenan Li +4 authors

A theoretical framework for sampling-based test-time scaling in large language models reveals limitations in self-consistency and perplexity, and introduces RPC to improve reasoning performance and reduce sampling costs.

sampling-based test-time scalinglarge language modelsself-consistencyperplexityHF ↗arXiv ↗

50 篇论文 · 按点赞排序

06

DeepSeek-OCR: Contexts Optical Compression

Haoran Wei, Yaofeng Sun, Yukun Li

DeepSeek-OCR uses optical 2D mapping to compress long contexts, achieving high OCR precision with reduced vision tokens and demonstrating practical value in document processing.

95DeepSeek-OCRDeepEncoderHF ↗arXiv ↗
09

FineVision: Open Data Is All You Need

Luis Wiedmann, Orr Zohar, Amir Mahla +6 authors

FineVision, a large-scale and curated dataset, enhances vision-language models through rigorous data collection, de-duplication, and human oversight, leading to improved performance.

81vision-language modelsFineVisionHF ↗arXiv ↗
10

World-in-World: World Models in a Closed-Loop World

Jiahan Zhang, Muqing Jiang, Nanru Dai +14 authors

World-in-World evaluates generative world models in closed-loop environments, emphasizing task success over visual quality and revealing insights into controllability, data scaling, and compute allocation.

78generative world modelsWMHF ↗arXiv ↗
13

Language Models are Injective and Hence Invertible

Giorgos Nikolaou, Tommaso Mencattini, Donato Crisostomi +3 authors

Transformer language models are proven to be injective, allowing exact input reconstruction from hidden activations, which has implications for transparency and safety.

70transformer componentsnon-linear activationsHF ↗arXiv ↗
22

Chem-R: Learning to Reason as a Chemist

Weida Wang, Benteng Chen, Di Zhang +14 authors

Chem-R, a three-phase trained Chemical Reasoning model, achieves superior performance on chemical tasks by integrating core knowledge, expert reasoning, and multi-task optimization.

53Chem-RChemical Foundation TrainingHF ↗arXiv ↗
25

Attention Sinks in Diffusion Language Models

Maximo Eduardo Rulli, Simone Petruzzi, Edoardo Michielon +3 authors

Empirical analysis of Masked Diffusion Language Models (DLMs) reveals distinct attention sinking phenomena and robustness compared to Autoregressive Models (ARMs).

50Masked Diffusion Language ModelsDLMsHF ↗arXiv ↗
26

Latent Diffusion Model without Variational Autoencoder

Minglei Shi, Haolin Wang, Wenzhao Zheng +6 authors

SVG, a novel latent diffusion model without VAEs, uses self-supervised representations to enable efficient training, few-step sampling, and high-quality visual generation with semantic and discriminative capabilities.

50latent diffusion modelsvariational autoencodersHF ↗arXiv ↗
27

RL makes MLLMs see better than SFT

Junha Song, Sangdoo Yun, Dongyoon Han +2 authors

Reinforcement Learning enhances vision encoders in Multimodal Language Models, leading to better visual representations and performance compared to Supervised Fine-tuning.

49Multimodal Language ModelLLM backboneHF ↗arXiv ↗
1 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号