TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Mar 17 – Mar 23, 2025
本周最热174

SmolDocling: An ultra-compact vision-language model for end-to-end multi-modal document conversion

Ahmed Nassar, Andres Marafioti, Matteo Omenetti +10 authors

SmolDocling is a compact vision-language model that performs end-to-end document conversion with robust performance across various document types using 256M parameters and a new markup format.

vision-language modelDocTagsend-to-end document conversionuniversal markup formatHF ↗arXiv ↗

50 篇论文 · 按点赞排序

02

RWKV-7 "Goose" with Expressive Dynamic State Evolution

Bo Peng, Ruichong Zhang, Daniel Goldstein +12 authors

RWKV-7 "Goose" achieves state-of-the-art performance in multilingual tasks with optimal memory and inference efficiency, exceeding Transformer capabilities in complexity.

154sequence modeling architecturedelta ruleHF ↗arXiv ↗
03

DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Qiying Yu, Zheng Zhang, Ruofei Zhu +32 authors

The DAPO algorithm, which includes decoupled clip and dynamic sampling policy optimization, enables open-source, high-performance reinforcement learning training for large-scale language models, enhancing reproducibility in the field.

148LLMsreinforcement learningHF ↗arXiv ↗
04

ReCamMaster: Camera-Controlled Generative Rendering from A Single Video

Jianhong Bai, Menghan Xia, Xiao Fu +8 authors

ReCamMaster uses pre-trained text-to-video models to render dynamic scenes of input videos from novel camera trajectories, leveraging a unique video conditioning mechanism and a custom dataset created with Unreal Engine 5.

148camera-controlled generative video re-renderingpre-trained text-to-video modelsHF ↗arXiv ↗
06

Survey on Evaluation of LLM-based Agents

Asaf Yehudai, Lilach Eden, Alan Li +5 authors

This survey analyzes evaluation methodologies for large language model-based agents, covering fundamental capabilities, application-specific benchmarks, and generalist agents, highlighting trends and gaps in the field.

97LLM-based agentsautonomous systemsHF ↗arXiv ↗
11

Impossible Videos

Zechen Bai, Hai Ci, Mike Zheng Shou

IPV-Bench evaluates video generation and understanding models on creating and interpreting impossible videos, highlighting their limitations and guiding future advancements.

61video generation modelsprompt followingHF ↗arXiv ↗
12

Inside-Out: Hidden Factual Knowledge in LLMs

Zorik Gekhman, Eyal Ben David, Hadas Orgad +5 authors

LLMs encode more internal factual knowledge than they express externally, with some knowledge so deeply hidden that it is never generated, despite repeated sampling.

56large language modelsLLMsHF ↗arXiv ↗
14

Cosmos-Reason1: From Physical Common Sense To Embodied Reasoning

NVIDIA, Alisson Azzolini, Hannah Brandon +42 authors

Cosmos-Reason1 models, using hierarchical and two-dimensional ontologies for physical common sense and embodied reasoning, generate embodied decisions through multimodal large language models trained in vision and Physical AI stages.

52Physical AIreasoningHF ↗arXiv ↗
16

Why Do Multi-Agent LLM Systems Fail?

Mert Cemri, Melissa Z. Pan, Shuyi Yang +10 authors

A comprehensive study identifies 14 unique failure modes in Multi-Agent Systems, categorizing them into specification issues, inter-agent conflicts, and task verification problems, and proposes interventions for preventing these failures.

49Multi-Agent Systems (MAS)LLMHF ↗arXiv ↗
20

Personalize Anything for Free with Diffusion Transformer

Haoran Feng, Zehuan Huang, Lin Li +2 authors

Personalize Anything is a training-free framework for personalized image generation using diffusion transformers, enabling subject consistency and flexibility through timestep-adaptive token replacement and patch perturbation.

44denoising tokensdiffusion transformersHF ↗arXiv ↗
23

Scale-wise Distillation of Diffusion Models

Nikita Starodubcev, Denis Kuznedelev, Artem Babenko +1 authors

A scale-wise distillation framework for diffusion models reduces computational costs and improves inference times by incorporating next-scale predictions and enhancing distribution matching methods.

42diffusion modelsscale-wise distillationHF ↗arXiv ↗
25

VGGT: Visual Geometry Grounded Transformer

Jianyuan Wang, Minghao Chen, Nikita Karaev +3 authors

VGGT, a feed-forward neural network, efficiently infers multiple 3D attributes from single or multiple views, outperforming alternatives and enhancing downstream tasks without post-processing.

40feed-forward neural networkcamera parametersHF ↗arXiv ↗
29

SPIN-Bench: How Well Do LLMs Plan Strategically and Reason Socially?

Jianzhu Yao, Kevin Wang, Ryan Hsieh +5 authors

SPIN-Bench evaluates strategic planning and social reasoning in AI through a unified framework combining PDDL tasks, board games, card games, and negotiation scenarios, highlighting challenges in deep multi-hop reasoning and social coordination.

34PDDL taskscompetitive board gamesHF ↗arXiv ↗
1 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号