TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Dec 1 – Dec 7, 2025

50 篇论文 · 按点赞排序

34

Geometrically-Constrained Agent for Spatial Reasoning

Zeren Chen, Xiaoya Lu, Zhijie Zheng +6 authors

Geometrically-Constrained Agent (GCA) addresses the semantic-to-geometric gap in vision language models by decoupling semantic analysis and task solving with formal constraints, achieving state-of-the-art performance in spatial reasoning.

41Vision Language Modelssemantic-to-geometric gapHF ↗arXiv ↗
36

PixelDiT: Pixel Diffusion Transformers for Image Generation

Yongsheng Yu, Wei Xiong, Weili Nie +3 authors

PixelDiT is a single-stage, end-to-end diffusion model that operates directly in pixel space, overcoming the limitations of latent-space modeling by using a dual-level transformer architecture and achieving competitive performance in image and text-to-image generation.

39Latent-space modelingDiffusion TransformersHF ↗arXiv ↗
38

DynamicVerse: A Physically-Aware Multimodal Framework for 4D World Modeling

Kairun Wen, Yuzhi Huang, Runyu Chen +16 authors

DynamicVerse is a framework that models dynamic real-world videos by integrating large vision, geometric, and multimodal models to produce a comprehensive 4D multimodal dataset, achieving superior performance in video depth estimation, camera pose estimation, and camera intrinsics estimation.

37DynamicVersemultimodal 4D world modelingHF ↗arXiv ↗
42

OneThinker: All-in-one Reasoning Model for Image and Video

Kaituo Feng, Manyuan Zhang, Hongyu Li +11 authors

OneThinker, an all-in-one multimodal reasoning model, unifies image and video understanding across various tasks using RL and demonstrates strong performance and knowledge transfer.

35Reinforcement learningMultimodal Large Language ModelsHF ↗arXiv ↗
45

DiP: Taming Diffusion Models in Pixel Space

Zhennan Chen, Junwei Zhu, Xu Chen +6 authors

DiP, a pixel space diffusion framework, combines a Diffusion Transformer and a Patch Detailer Head to achieve computational efficiency and high-quality image generation without using VAEs.

31diffusion modelslatent diffusion models (LDMs)HF ↗arXiv ↗
47

ViDiC: Video Difference Captioning

Jiangtao Wu, Shihao Li, Zhaozhou Bian +7 authors

The ViDiC task and ViDiC-1K dataset evaluate Multimodal Large Language Models' ability to describe differences between video pairs, addressing limitations in capturing motion continuity and event evolution.

29Image Difference CaptioningVideo Difference CaptioningHF ↗arXiv ↗
2 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号