TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Jun 30 – Jul 6, 2025
本周最热257

GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning

Wenyi Hong, Wenmeng Yu, Xiaotao Gu +74 authors

A vision-language model (VLM) named GLM-4.1V-Thinking, developed with a reasoning-centric training framework, achieves state-of-the-art performance across various tasks, including STEM problem solving, video understanding, and long document understanding, outperforming larger models on many benchmarks.

vision-language modelVLMreasoning-centric training frameworklarge-scale pre-trainingHF ↗arXiv ↗

50 篇论文 · 按点赞排序

02

Kwai Keye-VL Technical Report

Kwai Keye Team, Biao Yang, Bin Wen +57 authors

Kwai Keye-VL, an 8-billion-parameter multimodal model, excels in short-video understanding and general vision-language tasks through a comprehensive pre-training and post-training process, including a five-mode data mixture and reinforcement learning.

133Multimodal Large Language ModelsKwai Keye-VLHF ↗arXiv ↗
07

Energy-Based Transformers are Scalable Learners and Thinkers

Alexi Gladstone, Ganesh Nanduru, Md Mofijul Islam +7 authors

Energy-Based Transformers (EBTs) improve model performance and scalability across modalities by learning to verify predictions through unsupervised learning and energy minimization.

71Energy-Based TransformersEnergy-Based ModelsHF ↗arXiv ↗
09

Ovis-U1 Technical Report

Guo-Hua Wang, Shanshan Zhao, Xinjie Zhang +9 authors

Ovis-U1, a 3-billion-parameter unified model, integrates multimodal understanding, text-to-image generation, and image editing using a diffusion-based visual decoder and bidirectional token refiner, achieving state-of-the-art performance across various benchmarks.

63diffusion-based visual decoderbidirectional token refinerHF ↗arXiv ↗
13

Depth Anything at Any Condition

Boyuan Sun, Modi Jin, Bowen Yin +1 authors

DepthAnything-AC is a monocular depth estimation model that uses unsupervised consistency regularization and spatial distance constraints to handle complex environmental conditions and achieve zero-shot performance across various benchmarks.

49monocular depth estimationunsupervised consistency regularizationHF ↗arXiv ↗
18

Calligrapher: Freestyle Text Image Customization

Yue Ma, Qingyan Bai, Hao Ouyang +8 authors

Calligrapher, a diffusion-based framework, integrates advanced text customization with artistic typography using self-distillation, a localized style injection framework, and in-context generation to achieve high-quality, visually consistent typography.

37diffusion-based frameworkself-distillation mechanismHF ↗arXiv ↗
27

Fast and Simplex: 2-Simplicial Attention in Triton

Aurko Roy, Timothy Chou, Sai Surya Duvvuri +5 authors

The 2-simplicial Transformer improves token efficiency over standard Transformers, offering better performance on knowledge and reasoning tasks with a fixed token budget.

252-simplicial Transformertrilinear functionsHF ↗arXiv ↗
28

Thinking Beyond Tokens: From Brain-Inspired Intelligence to Cognitive Foundations for Artificial General Intelligence and its Societal Impact

Rizwan Qureshi, Ranjan Sapkota, Abbas Shah +17 authors

The paper explores the development of Artificial General Intelligence by integrating insights from various fields, focusing on modular reasoning, memory, and multi-agent coordination, and highlights the challenges in achieving true intelligence.

24Artificial General Intelligence (AGI)token-level predictionHF ↗arXiv ↗
30

Listener-Rewarded Thinking in VLMs for Image Preferences

Alexander Gambashidze, Li Pengyi, Matvey Skripkin +5 authors

A listener-augmented GRPO framework improves the accuracy and generalization of reward models for aligning text-to-image and text-to-video generative models with human preferences.

22reinforcement learningGroup Relative Policy OptimizationHF ↗arXiv ↗
1 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号