Paper Reading AI Learner

GSFixer: Improving 3D Gaussian Splatting with Reference-Guided Video Diffusion Priors

2025-08-13 09:56:28
Xingyilang Yin, Qi Zhang, Jiahao Chang, Ying Feng, Qingnan Fan, Xi Yang, Chi-Man Pun, Huaqi Zhang, Xiaodong Cun

Abstract

Reconstructing 3D scenes using 3D Gaussian Splatting (3DGS) from sparse views is an ill-posed problem due to insufficient information, often resulting in noticeable artifacts. While recent approaches have sought to leverage generative priors to complete information for under-constrained regions, they struggle to generate content that remains consistent with input observations. To address this challenge, we propose GSFixer, a novel framework designed to improve the quality of 3DGS representations reconstructed from sparse inputs. The core of our approach is the reference-guided video restoration model, built upon a DiT-based video diffusion model trained on paired artifact 3DGS renders and clean frames with additional reference-based conditions. Considering the input sparse views as references, our model integrates both 2D semantic features and 3D geometric features of reference views extracted from the visual geometry foundation model, enhancing the semantic coherence and 3D consistency when fixing artifact novel views. Furthermore, considering the lack of suitable benchmarks for 3DGS artifact restoration evaluation, we present DL3DV-Res which contains artifact frames rendered using low-quality 3DGS. Extensive experiments demonstrate our GSFixer outperforms current state-of-the-art methods in 3DGS artifact restoration and sparse-view 3D reconstruction. Project page: this https URL.

Abstract (translated)

使用3D高斯点(3D Gaussian Splatting,简称3DGS)从稀疏视角重构三维场景是一个病态问题,因为信息不足常常导致明显的伪影。虽然最近的方法试图利用生成先验来完成欠约束区域的信息,但它们难以产生与输入观察一致的内容。为了解决这一挑战,我们提出了GSFixer,这是一种旨在提高从稀疏输入重建的3DGS表示质量的新框架。 我们的方法的核心是参考引导视频恢复模型,该模型基于DiT(Diffusion Transformer)基础视频扩散模型训练而成,并且在配对的带有伪影的3DGS渲染图和干净帧的基础上进行额外参考条件的训练。通过将输入稀疏视图视为参考,我们的模型整合了从视觉几何基础模型中提取的参考视图的二维语义特征和三维几何特征,从而在修复具有伪影的新视图时增强语义一致性和3D一致性。 此外,考虑到缺少适合评估3DGS伪影恢复效果的基准数据集,我们提出了DL3DV-Res,其中包含使用低质量的3DGS渲染生成的带有伪影的帧。广泛的实验表明,我们的GSFixer在3DGS伪影修复和稀疏视图三维重建方面优于当前最先进方法。 项目主页:[这个链接](https://this-url.com)(请将“this https URL”替换为实际项目页面URL)。

URL

https://arxiv.org/abs/2508.09667

PDF

https://arxiv.org/pdf/2508.09667.pdf


Tags
3D Action Action_Localization Action_Recognition Activity Adversarial Agent Attention Autonomous Bert Boundary_Detection Caption Chat Classification CNN Compressive_Sensing Contour Contrastive_Learning Deep_Learning Denoising Detection Dialog Diffusion Drone Dynamic_Memory_Network Edge_Detection Embedding Embodied Emotion Enhancement Face Face_Detection Face_Recognition Facial_Landmark Few-Shot Gait_Recognition GAN Gaze_Estimation Gesture Gradient_Descent Handwriting Human_Parsing Image_Caption Image_Classification Image_Compression Image_Enhancement Image_Generation Image_Matting Image_Retrieval Inference Inpainting Intelligent_Chip Knowledge Knowledge_Graph Language_Model LLM Matching Medical Memory_Networks Multi_Modal Multi_Task NAS NMT Object_Detection Object_Tracking OCR Ontology Optical_Character Optical_Flow Optimization Person_Re-identification Point_Cloud Portrait_Generation Pose Pose_Estimation Prediction QA Quantitative Quantitative_Finance Quantization Re-identification Recognition Recommendation Reconstruction Regularization Reinforcement_Learning Relation Relation_Extraction Represenation Represenation_Learning Restoration Review RNN Robot Salient Scene_Classification Scene_Generation Scene_Parsing Scene_Text Segmentation Self-Supervised Semantic_Instance_Segmentation Semantic_Segmentation Semi_Global Semi_Supervised Sence_graph Sentiment Sentiment_Classification Sketch SLAM Sparse Speech Speech_Recognition Style_Transfer Summarization Super_Resolution Surveillance Survey Text_Classification Text_Generation Time_Series Tracking Transfer_Learning Transformer Unsupervised Video_Caption Video_Classification Video_Indexing Video_Prediction Video_Retrieval Visual_Relation VQA Weakly_Supervised Zero-Shot