TensorX
返回文献探索

Paper · arXiv 2509.15185

Understand Before You Generate: Self-Guided Training for Autoregressive Image Generation

Xiaoyu Yue, Zidong Wang, Yuqing Wang, Wenlong Zhang, Xihui Liu, Wanli Ouyang, Lei Bai, Luping Zhou

29 upvotesSeptember 18, 2025arXiv 预印本
AI 摘要

Self-guided Training for AutoRegressive models (ST-AR) enhances image understanding and generation quality in autoregressive models by addressing key visual semantics challenges through self-supervised objectives.

autoregressive modelsnext-token predictionhigh-level visual semanticslocal and conditional dependenceinter-step semantic inconsistencyspatial invariance deficiencyself-supervised objectivesSelf-guided Training for AutoRegressive modelsST-ARFID improvementLlamaGen-LLlamaGen-XL

Abstract

Recent studies have demonstrated the importance of high-quality visual representations in image generation and have highlighted the limitations of generative models in image understanding. As a generative paradigm originally designed for natural language, autoregressive models face similar challenges. In this work, we present the first systematic investigation into the mechanisms of applying the next-token prediction paradigm to the visual domain. We identify three key properties that hinder the learning of high-level visual semantics: local and conditional dependence, inter-step semantic inconsistency, and spatial invariance deficiency. We show that these issues can be effectively addressed by introducing self-supervised objectives during training, leading to a novel training framework, Self-guided Training for AutoRegressive models (ST-AR). Without relying on pre-trained representation models, ST-AR significantly enhances the image understanding ability of autoregressive models and leads to improved generation quality. Specifically, ST-AR brings approximately 42% FID improvement for LlamaGen-L and 49% FID improvement for LlamaGen-XL, while maintaining the same sampling strategy.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Understand Before You Generate: Self-Guided Training for Autoregressive Image Generation | TensorX