TensorX
返回文献探索

Paper · arXiv 2511.20614

The Consistency Critic: Correcting Inconsistencies in Generated Images via Reference-Guided Attentive Alignment

Ziheng Ouyang, Yiren Song, Yaoli Liu, Shihao Zhu, Qibin Hou, Ming-Ming Cheng, Mike Zheng Shou

38 upvotesNovember 25, 2025arXiv 预印本
AI 摘要

ImageCritic addresses detail inconsistency in image generation through reference-guided post-editing, using attention alignment loss and a detail encoder.

reference-guided post-editingVLM-based selectionattention mechanismsattention alignment lossdetail encodermulti-round editinglocal editing

Abstract

Previous works have explored various customized generation tasks given a reference image, but they still face limitations in generating consistent fine-grained details. In this paper, our aim is to solve the inconsistency problem of generated images by applying a reference-guided post-editing approach and present our ImageCritic. We first construct a dataset of reference-degraded-target triplets obtained via VLM-based selection and explicit degradation, which effectively simulates the common inaccuracies or inconsistencies observed in existing generation models. Furthermore, building on a thorough examination of the model's attention mechanisms and intrinsic representations, we accordingly devise an attention alignment loss and a detail encoder to precisely rectify inconsistencies. ImageCritic can be integrated into an agent framework to automatically detect inconsistencies and correct them with multi-round and local editing in complex scenarios. Extensive experiments demonstrate that ImageCritic can effectively resolve detail-related issues in various customized generation scenarios, providing significant improvements over existing methods.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
The Consistency Critic: Correcting Inconsistencies in Generated Images via Reference-Guided Attentive Alignment | TensorX