TensorX
返回文献探索

Paper · arXiv 2608.00486

DreamTraj: Generating 6-DoF Object Trajectories by Reading Unrendered Video Diffusion Latents

Tongsheng Ding, Zhen Luo, Yixuan Yang, Boyu Wang, Luyang Xie, Jinyu Yang, Feng Zheng

15 upvotesAugust 1, 2026arXiv 预印本
AI 摘要

DreamTraj predicts 6-DoF object trajectories from a single image and instruction by decoding motion directly from intermediate representations of a frozen image-to-video diffusion model, achieving state-of-the-art results without privileged inputs.

6-DoF trajectory predictionegocentric trajectoriesimage-to-video diffusion modeldenoisingflow-matchingquery-key attention trackshidden statesrelative 6-DoF poses

Abstract

Accurate prediction of object trajectories during manipulation is essential for closing the perception-action loop. Progress is limited on two fronts: available datasets lack fine-grained language-to-motion annotations, and existing predictors either rely on privileged inputs such as video, depth, or CAD models, or recover motion from fully generated videos through costly, error-prone perception pipelines. We close the supervision gap with the MOVE dataset, 5,038 object-centric egocentric trajectories, each paired with a fine-grained natural-language instruction rather than a coarse verb-noun label. We further propose DreamTraj, which predicts a 6-DoF object trajectory from a single RGB image and a task instruction, requiring no video, depth, or CAD model at inference: rather than generating a video, it reads motion from the internal representations of a frozen image-to-video diffusion model at an early denoising step. A lightweight flow-matching Reader decodes query-key attention tracks and pooled hidden states into relative 6-DoF poses. To our knowledge, this is the first approach to directly decode object 6-DoF trajectories from intermediate video diffusion representations rather than generated pixels. DreamTraj sets a new state of the art on both translation and rotation against forecasters that consume multi-frame or privileged inputs, and runs 4.6x faster than generate-then-extract pipelines.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号