TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

50 篇论文 · 按点赞排序

34

Wan: Open and Advanced Large-Scale Video Generative Models

WanTeam, Ang Wang, Baole Ai +59 authors

Wan, a comprehensive suite of video foundation models built on the diffusion transformer paradigm, advannces video generation by introducing a novel VAE, scalable pre-training strategies, and large-scale data curation, offering superior performance and versatility across various applications with both large and efficient models.

71diffusion transformerVAEHF ↗arXiv ↗
39

Impossible Videos

Zechen Bai, Hai Ci, Mike Zheng Shou

IPV-Bench evaluates video generation and understanding models on creating and interpreting impossible videos, highlighting their limitations and guiding future advancements.

61video generation modelsprompt followingHF ↗arXiv ↗
41

Gemma 3 Technical Report

Gemma Team, Aishwarya Kamath, Johan Ferret +213 authors

Gemma 3 introduces vision capabilities, broader language coverage, and extended context length, featuring an optimized architecture and post-training enhancements to outperform previous versions.

59multimodal modelsvision understandingHF ↗arXiv ↗
42

VACE: All-in-One Video Creation and Editing

Zeyinzi Jiang, Zhen Han, Chaojie Mao +3 authors

VACE, an all-in-one framework for video creation and editing, integrates multiple tasks within a unified model using a Video Condition Unit and Context Adapter for flexible and consistent video synthesis.

58diffusion transformervideo synthesisHF ↗arXiv ↗
46

Inside-Out: Hidden Factual Knowledge in LLMs

Zorik Gekhman, Eyal Ben David, Hadas Orgad +5 authors

LLMs encode more internal factual knowledge than they express externally, with some knowledge so deeply hidden that it is never generated, despite repeated sampling.

56large language modelsLLMsHF ↗arXiv ↗
50

Cosmos-Reason1: From Physical Common Sense To Embodied Reasoning

NVIDIA, Alisson Azzolini, Hannah Brandon +42 authors

Cosmos-Reason1 models, using hierarchical and two-dimensional ontologies for physical common sense and embodied reasoning, generate embodied decisions through multimodal large language models trained in vision and Physical AI stages.

52Physical AIreasoningHF ↗arXiv ↗
2 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号