TensorX
返回文献探索

Paper · arXiv 2312.03913

Controllable Human-Object Interaction Synthesis

Jiaman Li, Alexander Clegg, Roozbeh Mottaghi, Jiajun Wu, Xavier Puig, C. Karen Liu

23 upvotesDecember 6, 2023arXiv 预印本
AI 摘要

CHOIS uses a conditional diffusion model to generate synchronized human and object motion in 3D scenes, guided by language descriptions and waypoints, with improvements for object geometry alignment and contact constraints.

conditional diffusion modelobject motionhuman motionlanguage descriptionsinitial object and human statessparse object waypointsobject geometry lossguidance termscontact constraints

Abstract

Synthesizing semantic-aware, long-horizon, human-object interaction is critical to simulate realistic human behaviors. In this work, we address the challenging problem of generating synchronized object motion and human motion guided by language descriptions in 3D scenes. We propose Controllable Human-Object Interaction Synthesis (CHOIS), an approach that generates object motion and human motion simultaneously using a conditional diffusion model given a language description, initial object and human states, and sparse object waypoints. While language descriptions inform style and intent, waypoints ground the motion in the scene and can be effectively extracted using high-level planning methods. Naively applying a diffusion model fails to predict object motion aligned with the input waypoints and cannot ensure the realism of interactions that require precise hand-object contact and appropriate contact grounded by the floor. To overcome these problems, we introduce an object geometry loss as additional supervision to improve the matching between generated object motion and input object waypoints. In addition, we design guidance terms to enforce contact constraints during the sampling process of the trained diffusion model.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号