TensorX
返回文献探索

Paper · arXiv 2505.02094

SkillMimic-V2: Learning Robust and Generalizable Interaction Skills from Sparse and Noisy Demonstrations

Runyi Yu, Yinhuai Wang, Qihan Zhao, Hok Wai Tsui, Jingbo Wang, Ping Tan, Qifeng Chen

19 upvotesMay 4, 2025arXiv 预印本
AI 摘要

The approach uses data augmentation techniques and adaptive sampling to improve robust skill acquisition and generalization in Reinforcement Learning from Interaction Demonstration despite noisy and sparse demonstrations.

Reinforcement Learning from Interaction Demonstration (RLID)demonstration noisecoverage limitationssparse trajectoriesskill variationstransitionsphysically feasible trajectoriesStitched Trajectory Graph (STG)State Transition Field (STF)Adaptive Trajectory Sampling (ATS)dynamic curriculum generationmemory-dependent skill learningconvergence stabilitygeneralization capabilityrecovery robustness

Abstract

We address a fundamental challenge in Reinforcement Learning from Interaction Demonstration (RLID): demonstration noise and coverage limitations. While existing data collection approaches provide valuable interaction demonstrations, they often yield sparse, disconnected, and noisy trajectories that fail to capture the full spectrum of possible skill variations and transitions. Our key insight is that despite noisy and sparse demonstrations, there exist infinite physically feasible trajectories that naturally bridge between demonstrated skills or emerge from their neighboring states, forming a continuous space of possible skill variations and transitions. Building upon this insight, we present two data augmentation techniques: a Stitched Trajectory Graph (STG) that discovers potential transitions between demonstration skills, and a State Transition Field (STF) that establishes unique connections for arbitrary states within the demonstration neighborhood. To enable effective RLID with augmented data, we develop an Adaptive Trajectory Sampling (ATS) strategy for dynamic curriculum generation and a historical encoding mechanism for memory-dependent skill learning. Our approach enables robust skill acquisition that significantly generalizes beyond the reference demonstrations. Extensive experiments across diverse interaction tasks demonstrate substantial improvements over state-of-the-art methods in terms of convergence stability, generalization capability, and recovery robustness.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
SkillMimic-V2: Learning Robust and Generalizable Interaction Skills from Sparse and Noisy Demonstrations | TensorX