TensorX
返回文献探索

Paper · arXiv 2409.04410

Open-MAGVIT2: An Open-Source Project Toward Democratizing Auto-regressive Visual Generation

Zhuoyan Luo, Fengyuan Shi, Yixiao Ge, Yujiu Yang, Limin Wang, Ying Shan

24 upvotesSeptember 6, 2024arXiv 预印本
AI 摘要

Open-MAGVIT2 models, ranging from 300M to 1.5B parameters, achieve state-of-the-art image reconstruction and exploration in auto-regressive models with super-large token vocabularies.

auto-regressive image generationOpen-MAGVIT2MAGVIT-v2 tokenizersuper-large codebookreconstruction performanceasymmetric token factorizationnext sub-token prediction

Abstract

We present Open-MAGVIT2, a family of auto-regressive image generation models ranging from 300M to 1.5B. The Open-MAGVIT2 project produces an open-source replication of Google's MAGVIT-v2 tokenizer, a tokenizer with a super-large codebook (i.e., 2^{18} codes), and achieves the state-of-the-art reconstruction performance (1.17 rFID) on ImageNet 256 times 256. Furthermore, we explore its application in plain auto-regressive models and validate scalability properties. To assist auto-regressive models in predicting with a super-large vocabulary, we factorize it into two sub-vocabulary of different sizes by asymmetric token factorization, and further introduce "next sub-token prediction" to enhance sub-token interaction for better generation quality. We release all models and codes to foster innovation and creativity in the field of auto-regressive visual generation.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Open-MAGVIT2: An Open-Source Project Toward Democratizing Auto-regressive Visual Generation | TensorX