TensorX
返回文献探索

Paper · arXiv 2507.08776

CLiFT: Compressive Light-Field Tokens for Compute-Efficient and Adaptive Neural Rendering

Zhengqing Wang, Yuefan Wu, Jiacheng Chen, Fuyang Zhang, Yasutaka Furukawa

55 upvotesJuly 11, 2025arXiv 预印本
AI 摘要

A neural rendering approach using compressed light-field tokens (CLiFTs) achieves efficient rendering with adjustable token numbers, maintaining high quality and speed.

neural renderingcompressed light-field tokensCLiFTsmulti-view encoderlatent-space K-meanscondensercompute-adaptive rendererRealEstate10KDL3DV datasets

Abstract

This paper proposes a neural rendering approach that represents a scene as "compressed light-field tokens (CLiFTs)", retaining rich appearance and geometric information of a scene. CLiFT enables compute-efficient rendering by compressed tokens, while being capable of changing the number of tokens to represent a scene or render a novel view with one trained network. Concretely, given a set of images, multi-view encoder tokenizes the images with the camera poses. Latent-space K-means selects a reduced set of rays as cluster centroids using the tokens. The multi-view ``condenser'' compresses the information of all the tokens into the centroid tokens to construct CLiFTs. At test time, given a target view and a compute budget (i.e., the number of CLiFTs), the system collects the specified number of nearby tokens and synthesizes a novel view using a compute-adaptive renderer. Extensive experiments on RealEstate10K and DL3DV datasets quantitatively and qualitatively validate our approach, achieving significant data reduction with comparable rendering quality and the highest overall rendering score, while providing trade-offs of data size, rendering quality, and rendering speed.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
CLiFT: Compressive Light-Field Tokens for Compute-Efficient and Adaptive Neural Rendering | TensorX