TensorX
返回文献探索

Paper · arXiv 2410.18977

MotionCLR: Motion Generation and Training-free Editing via Understanding Attention Mechanisms

Ling-Hao Chen, Wenxun Dai, Xuan Ju, Shunlin Lu, Lei Zhang

14 upvotesOctober 24, 2024arXiv 预印本
AI 摘要

MotionCLR, an attention-based motion diffusion model, enhances interactive motion editing by explicitly modeling word-level text-motion correspondence and improving explainability through self-attention and cross-attention mechanisms.

attention-based motion diffusion modelMotionCLRCLeaR modelingself-attentioncross-attentiontext-motion correspondencemotion editingexplainabilityattention mapsmotion (de-)emphasizingin-place motion replacementexample-based motion generationaction-countinggrounded motion generation

Abstract

This research delves into the problem of interactive editing of human motion generation. Previous motion diffusion models lack explicit modeling of the word-level text-motion correspondence and good explainability, hence restricting their fine-grained editing ability. To address this issue, we propose an attention-based motion diffusion model, namely MotionCLR, with CLeaR modeling of attention mechanisms. Technically, MotionCLR models the in-modality and cross-modality interactions with self-attention and cross-attention, respectively. More specifically, the self-attention mechanism aims to measure the sequential similarity between frames and impacts the order of motion features. By contrast, the cross-attention mechanism works to find the fine-grained word-sequence correspondence and activate the corresponding timesteps in the motion sequence. Based on these key properties, we develop a versatile set of simple yet effective motion editing methods via manipulating attention maps, such as motion (de-)emphasizing, in-place motion replacement, and example-based motion generation, etc. For further verification of the explainability of the attention mechanism, we additionally explore the potential of action-counting and grounded motion generation ability via attention maps. Our experimental results show that our method enjoys good generation and editing ability with good explainability.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
MotionCLR: Motion Generation and Training-free Editing via Understanding Attention Mechanisms | TensorX