TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Dec 15 – Dec 21, 2025

50 篇论文 · 按点赞排序

33

Olmo 3

Team Olmo, Allyson Ettinger, Amanda Bertsch +65 authors

Olmo 3, a family of state-of-the-art fully-open language models at 7B and 32B parameter scales, excels in long-context reasoning, function calling, coding, instruction following, general chat, and knowledge recall.

37fully-open language modelslong-context reasoningHF ↗arXiv ↗
35

PuzzleCraft: Exploration-Aware Curriculum Learning for Puzzle-Based RLVR in VLMs

Ahmadreza Jeddi, Hakki Can Karaimer, Hue Nguyen +8 authors

PuzzleCraft presents a supervision-free framework for scaling vision-centric RLVR using lightweight puzzle environments with built-in verification, incorporating an exploration-aware curriculum and a new consistency metric to improve reasoning robustness and downstream performance.

36RL post-trainingvision-language modelsHF ↗arXiv ↗
37

JustRL: Scaling a 1.5B LLM with a Simple RL Recipe

Bingxiang He, Zekai Qu, Zeyuan Liu +9 authors

JustRL achieves state-of-the-art performance on reasoning models with minimal complexity, using single-stage training and fixed hyperparameters, outperforming sophisticated approaches in terms of compute and stability.

34reinforcement learninglarge language modelsHF ↗arXiv ↗
40

V-RGBX: Video Editing with Accurate Controls over Intrinsic Properties

Ye Fang, Tong Wu, Valentin Deschaintre +6 authors

V-RGBX presents an end-to-end framework for intrinsic-aware video editing that combines video inverse rendering, photorealistic video synthesis, and keyframe-based editing with physically grounded intrinsic channel manipulation.

30video inverse renderingphotorealistic video synthesisHF ↗arXiv ↗
41

Sharp Monocular View Synthesis in Less Than a Second

Lars Mescheder, Wei Dong, Shiwei Li +10 authors

SHARP synthesizes photorealistic views from a single image using a 3D Gaussian representation, achieving state-of-the-art results with rapid processing.

30photorealistic view synthesis3D Gaussian representationHF ↗arXiv ↗
45

TimeLens: Rethinking Video Temporal Grounding with Multimodal LLMs

Jun Zhang, Teng Wang, Yuying Ge +4 authors

TimeLens establishes a robust baseline for video temporal grounding by improving benchmark quality, addressing noisy training data, and developing efficient algorithmic design principles for multimodal large language models.

27multimodal large language modelsvideo temporal groundingHF ↗arXiv ↗
2 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号