TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

686 篇论文 · 按点赞排序

511

Articulated Object Reconstruction from Rest-State Observation

Daeun Lee, Jaeah Lee, Woosung Kim +2 authors

A rest-state framework reconstructs articulated objects from a single closed configuration by fusing vision-language outputs into consistent part meshes and validating synthesized motion hypotheses via geometric consistency.

46digital twinsarticulated object reconstructionHF ↗arXiv ↗
513

GST-Bench: Can VLMs Develop Global Spatial Awareness from Video?

Qifeng Zhang, Kaixiang Huang, Heng Dong +6 authors

The study introduces a video benchmark requiring global spatial reasoning across long videos and reveals that vision-language models struggle to build consistent global scene representations despite strong local perception.

46global spatial intelligenceVQA benchmarkHF ↗arXiv ↗
515

T1: Terminal Agent Reinforcement Learning for Long-Horizon Tasks

Junyao Yang, Yucheng Shi, Zhongzhi Li +4 authors

T1 is a 122B Mixture-of-Experts model trained with reinforcement learning to execute long-horizon terminal tasks in a cloud sandbox, achieving state-of-the-art results through stable actor-critic optimization and out-of-distribution training.

46Mixture-of-Expertsreinforcement learningHF ↗arXiv ↗
517

Vision as Unified Multimodal Generation

Xiaoyang Han, Jianhua Li, Kewang Deng +14 authors

A unified multimodal model formulates computer vision tasks as generation problems using natural language and visual prompts, achieving performance comparable to specialized systems across diverse vision tasks.

44multimodal generationunified multimodal modelHF ↗arXiv ↗
518

Vision Pretraining for Dense Spatial Perception

Zelin Fu, Bin Tan, Changjiang Sun +6 authors

Boundary modeling enables dense spatial perception by learning sub-pixel representations that enhance depth estimation and support embodied AI applications.

43boundary modelingmasked boundary modelingHF ↗arXiv ↗
519

Infinite Worlds with Versatile Interactions

Zelin Gao, Qiuyu Wang, Jiapeng Zhu +17 authors

An advanced world modeling system with extended interaction capabilities, real-time processing, diverse interactive elements, and multi-agent behavior control for collaborative virtual environments.

43causal pretraining paradigmreal-time variantHF ↗arXiv ↗
520

Motif 3: Technical Report

Junghwan Lim, Joon Son Chung, Sungmin Lee +24 authors

Motif 3 is a large sparse mixture-of-experts language model using grouped differential latent attention and specialized training techniques to achieve strong reasoning, coding, and long-context performance.

43Mixture-of-Expertssparse MoEHF ↗arXiv ↗
523

Editable Visual Design

Junyan Ye, Wei Liu, Dongzhi Jiang +9 authors

A coding agent guided by a vision-language model generates editable layered designs by synthesizing isolated visual assets and iteratively refining native HTML/CSS layouts.

42diffusion base modelsVLMHF ↗arXiv ↗
530

Multi-Block Diffusion Language Models

Yijie Jin, Jiajun Xu, Yuxuan Liu +8 authors

Multi-Block Diffusion Language Models extend single-block diffusion to concurrent block decoding with improved training strategies and optimized decoding algorithms.

40Block Diffusion Language ModelsMulti-Block DiffusionHF ↗arXiv ↗
536

CADENA: Stepwise CAD Reverse Engineering

Soslan Kabisov, Gennadiy Savrasov, Maksim Elistratov +9 authors

CADENA reconstructs 3D meshes into parametric CAD programs step-by-step with intermediate geometry checks and introduces a benchmark for mechanical part reverse engineering.

39parametric CAD programreverse-engineeringHF ↗arXiv ↗
537

ChronoVision: Temporal Reasoning via Latent State Reconstruction

Yifan Shen, Jian Xu, Boyi Li +6 authors

ChronoVision improves visual temporal reasoning by aligning latent imagery with logic through reconstructive prediction, ROI attention, and reinforcement learning, achieving strong results on video reasoning benchmarks.

39multimodal large language modelsReconstructive Visual HeadHF ↗arXiv ↗
18 / 23

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号