TensorX
返回文献探索

Paper · arXiv 2504.08003

Have we unified image generation and understanding yet? An empirical study of GPT-4o's image generation ability

Ning Li, Jingran Zhang, Justin Cui

49 upvotesApril 9, 2025arXiv 预印本
AI 摘要

Evaluation of GPT-4o shows limitations in its ability to integrate world knowledge, apply contextual reasoning, and adhere to complex instructions in image generation and editing.

multimodal GPT-4osemantic synthesisglobal instruction adherencefine-grained editing precisionpost-generation reasoningknowledge constraintsconditional reasoningcontext-aware generationreasoning-grounded generation

Abstract

OpenAI's multimodal GPT-4o has demonstrated remarkable capabilities in image generation and editing, yet its ability to achieve world knowledge-informed semantic synthesis--seamlessly integrating domain knowledge, contextual reasoning, and instruction adherence--remains unproven. In this study, we systematically evaluate these capabilities across three critical dimensions: (1) Global Instruction Adherence, (2) Fine-Grained Editing Precision, and (3) Post-Generation Reasoning. While existing benchmarks highlight GPT-4o's strong capabilities in image generation and editing, our evaluation reveals GPT-4o's persistent limitations: the model frequently defaults to literal interpretations of instructions, inconsistently applies knowledge constraints, and struggles with conditional reasoning tasks. These findings challenge prevailing assumptions about GPT-4o's unified understanding and generation capabilities, exposing significant gaps in its dynamic knowledge integration. Our study calls for the development of more robust benchmarks and training strategies that go beyond surface-level alignment, emphasizing context-aware and reasoning-grounded multimodal generation.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Have we unified image generation and understanding yet? An empirical study of GPT-4o's image generation ability | TensorX