TensorX
返回文献探索

Paper · arXiv 2607.12752

Hallo4D: Multi-Modal Hallucination Mitigation for Consistent Spatio-Temporal Generation

Hongbo Wang, Huaibo Huang, Jie Cao, Jin Liu, Haoyang Tong, Ran He

21 upvotesJuly 15, 2026arXiv 预印本
AI 摘要

Hallo4D reduces spatial and temporal inconsistencies in 3D and 4D generation by using multimodal language models to detect errors and guide consensus-based image optimization without retraining.

3D generation4D generation2D diffusion supervisiongeometric consistencyspatial hallucinationslarge multimodal language modelsmulti-view renderingmulti-frame renderingconsensus-driven optimizationmulti-model votingmotion-aware keyframe samplingappearance alignmentexposure-aware optimizationvisibility pruning

Abstract

While recent advances in 3D generation have enabled impressive visual synthesis, existing methods often rely on 2D diffusion supervision without explicit mechanisms for geometric consistency, leading to spatial hallucinations such as duplicated structures and misaligned geometry. These issues become more severe in 4D generation, where maintaining consistency across viewpoints and temporal evolution introduces additional challenges, including jitter, identity flicker, and structural drift. We present Hallo4D, a unified and model-agnostic framework for mitigating spatiotemporal hallucinations in 3D and 4D content generation. Hallo4D introduces a generation-detection-correction paradigm that leverages large multimodal language models (LMMs) to identify and summarize spatial and temporal inconsistencies from multi-view and multi-frame renderings. These insights guide a consensus-driven image-space consistency optimization, where an LMM-based selector evaluates candidate corrections through multi-model voting, without requiring retraining or architectural modifications. To further improve temporal consistency and optimization efficiency, Hallo4D incorporates motion-aware keyframe sampling, LMM-guided initialization, and appearance alignment. We additionally introduce exposure-aware optimization and visibility pruning to enhance robustness under challenging viewpoints. Extensive experiments demonstrate that Hallo4D consistently outperforms strong baselines across diverse 3D and 4D generation settings, providing a scalable and generalizable solution for consistency-aware content generation.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Hallo4D: Multi-Modal Hallucination Mitigation for Consistent Spatio-Temporal Generation | TensorX