Abstract
Reconstructing 3D scenes using 3D Gaussian Splatting (3DGS) from sparse views is an ill-posed problem due to insufficient information, often resulting in noticeable artifacts. While recent approaches have sought to leverage generative priors to complete information for under-constrained regions, they struggle to generate content that remains consistent with input observations. To address this challenge, we propose GSFixer, a novel framework designed to improve the quality of 3DGS representations reconstructed from sparse inputs. The core of our approach is the reference-guided video restoration model, built upon a DiT-based video diffusion model trained on paired artifact 3DGS renders and clean frames with additional reference-based conditions. Considering the input sparse views as references, our model integrates both 2D semantic features and 3D geometric features of reference views extracted from the visual geometry foundation model, enhancing the semantic coherence and 3D consistency when fixing artifact novel views. Furthermore, considering the lack of suitable benchmarks for 3DGS artifact restoration evaluation, we present DL3DV-Res which contains artifact frames rendered using low-quality 3DGS. Extensive experiments demonstrate our GSFixer outperforms current state-of-the-art methods in 3DGS artifact restoration and sparse-view 3D reconstruction. Project page: this https URL.
Abstract (translated)
使用3D高斯点(3D Gaussian Splatting,简称3DGS)从稀疏视角重构三维场景是一个病态问题,因为信息不足常常导致明显的伪影。虽然最近的方法试图利用生成先验来完成欠约束区域的信息,但它们难以产生与输入观察一致的内容。为了解决这一挑战,我们提出了GSFixer,这是一种旨在提高从稀疏输入重建的3DGS表示质量的新框架。 我们的方法的核心是参考引导视频恢复模型,该模型基于DiT(Diffusion Transformer)基础视频扩散模型训练而成,并且在配对的带有伪影的3DGS渲染图和干净帧的基础上进行额外参考条件的训练。通过将输入稀疏视图视为参考,我们的模型整合了从视觉几何基础模型中提取的参考视图的二维语义特征和三维几何特征,从而在修复具有伪影的新视图时增强语义一致性和3D一致性。 此外,考虑到缺少适合评估3DGS伪影恢复效果的基准数据集,我们提出了DL3DV-Res,其中包含使用低质量的3DGS渲染生成的带有伪影的帧。广泛的实验表明,我们的GSFixer在3DGS伪影修复和稀疏视图三维重建方面优于当前最先进方法。 项目主页:[这个链接](https://this-url.com)(请将“this https URL”替换为实际项目页面URL)。
URL
https://arxiv.org/abs/2508.09667