UniG2U-Bench: Do Unified Models Advance Multimodal Understanding?
Zimo Wen, Boxiu Li, Wanbo Zhang +11 authors
Unified multimodal models show mixed performance in generation-to-understanding tasks, with specific subtasks benefiting from enhanced spatial and reasoning capabilities while overall performance lags behind specialized vision-language models.