TensorX
返回文献探索

Paper · arXiv 2510.22201

ACG: Action Coherence Guidance for Flow-based VLA models

Minho Park, Kinam Kim, Junha Hyung, Hyojin Jang, Hoiyeong Jin, Jooyeol Yun, Hojoon Lee, Jaegul Choo

37 upvotesOctober 25, 2025arXiv 预印本
AI 摘要

Action Coherence Guidance (ACG) improves action coherence in Vision-Language-Action (VLA) models during test time, enhancing performance in diverse manipulation tasks.

diffusion modelsflow matching modelsVision-Language-Action (VLA) modelsimitation learningaction coherencetrajectory driftRoboCasaDexMimicGenSO-101 tasksAction Coherence Guidance (ACG)

Abstract

Diffusion and flow matching models have emerged as powerful robot policies, enabling Vision-Language-Action (VLA) models to generalize across diverse scenes and instructions. Yet, when trained via imitation learning, their high generative capacity makes them sensitive to noise in human demonstrations: jerks, pauses, and jitter which reduce action coherence. Reduced action coherence causes instability and trajectory drift during deployment, failures that are catastrophic in fine-grained manipulation where precision is crucial. In this paper, we present Action Coherence Guidance (ACG) for VLA models, a training-free test-time guidance algorithm that improves action coherence and thereby yields performance gains. Evaluated on RoboCasa, DexMimicGen, and real-world SO-101 tasks, ACG consistently improves action coherence and boosts success rates across diverse manipulation tasks. Code and project page are available at https://github.com/DAVIAN-Robotics/ACG and https://DAVIAN-Robotics.github.io/ACG , respectively.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
ACG: Action Coherence Guidance for Flow-based VLA models | TensorX