TensorX
返回文献探索

Paper · arXiv 2608.02437

InfiniSplat: Implicit Gaussian Decoding for Large-Baseline Monocular View Synthesis

Jiawei Wang, Hao Yu, Yongzhen Hu, Xinyi Yang, Tao Ni, Xin Zhan, Junbo Chen, Xiaowei Zhou, Ruizhen Hu, Sida Peng

72 upvotesAugust 3, 2026arXiv 预印本
AI 摘要

InfiniSplat improves single-image 3D Gaussian Splatting by aligning Gaussian primitives to scene surfaces via geometry-guided sampling and implicit decoding, yielding more coherent renderings under large viewpoint changes.

3D Gaussian Splattingfeed-forwardpixel-aligned representationsurface-aligned representationgeometry-guided samplingimplicit decodercross-dataset NVSzero-shot generalization

Abstract

Single-image feed-forward 3D Gaussian Splatting (3DGS) aims to directly generate a renderable 3D scene representation from one input image, avoiding the cost of multi-view capture and per-scene optimization. However, existing methods are often constrained by a pixel-aligned representation, where Gaussians are predicted from fixed image-grid locations. Such pixel-aligned primitives can produce promising nearby-view renderings, but they remain weakly coupled to underlying scene surfaces and struggle to preserve coherent structures under large viewpoint shifts. We present InfiniSplat, a feed-forward single-image 3DGS framework that moves from a pixel-aligned representation toward a surface-aligned representation. InfiniSplat constructs this representation by first using geometry-guided sampling to place 2D supports according to depth-induced local surface structure, and then applying a query-conditioned implicit decoder to predict Gaussian attributes from the image features queried at these supports.By grounding support locations in geometry while decoupling Gaussian prediction from fixed pixel centers, InfiniSplat produces Gaussian layouts that better follow scene surfaces and reduce scattered primitives caused by grid discretization.Across multiple cross-dataset NVS evaluations, InfiniSplat achieves state-of-the-art performance compared with single-image feed-forward baselines, and demonstrates zero-shot generalization from Hypersim indoor synthetic training to complex open-world scenes.Project page: https://zju3dv.github.io/InfiniSplat.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
InfiniSplat: Implicit Gaussian Decoding for Large-Baseline Monocular View Synthesis | TensorX