Abstract
Monocular depth and pose estimation play an important role in the development of colonoscopy-assisted navigation, as they enable improved screening by reducing blind spots, minimizing the risk of missed or recurrent lesions, and lowering the likelihood of incomplete examinations. However, this task remains challenging due to the presence of texture-less surfaces, complex illumination patterns, deformation, and a lack of in-vivo datasets with reliable ground truth. In this paper, we propose **PRISM** (Pose-Refinement with Intrinsic Shading and edge Maps), a self-supervised learning framework that leverages anatomical and illumination priors to guide geometric learning. Our approach uniquely incorporates edge detection and luminance decoupling for structural guidance. Specifically, edge maps are derived using a learning-based edge detector (e.g., DexiNed or HED) trained to capture thin and high-frequency boundaries, while luminance decoupling is obtained through an intrinsic decomposition module that separates shading and reflectance, enabling the model to exploit shading cues for depth estimation. Experimental results on multiple real and synthetic datasets demonstrate state-of-the-art performance. We further conduct a thorough ablation study on training data selection to establish best practices for pose and depth estimation in colonoscopy. This analysis yields two practical insights: (1) self-supervised training on real-world data outperforms supervised training on realistic phantom data, underscoring the superiority of domain realism over ground truth availability; and (2) video frame rate is an extremely important factor for model performance, where dataset-specific video frame sampling is necessary for generating high quality training data.
Abstract (translated)
单目深度和姿态估计在结肠镜辅助导航的发展中扮演着重要角色,因为它们通过减少盲点、降低漏诊或复发病变的风险以及降低检查不完全的可能性来提高筛查效果。然而,由于无纹理表面、复杂的光照模式、变形及缺乏可靠地面真实数据的体内数据集等原因,该任务仍然具有挑战性。本文提出了一种新的方法**PRISM**(基于内在阴影和边缘图的姿态精炼),这是一种自监督学习框架,利用解剖学和照明先验指导几何学习。我们的方法独特地结合了边缘检测和亮度分离以提供结构引导。具体而言,通过使用训练有素的边缘探测器(如DexiNed或HED)来捕捉薄且高频边界,可以得到边缘图;同时通过内在分解模块获得亮度分离,该模块能够将阴影与反射分离开来,使模型利用阴影线索进行深度估计。在多个真实和合成数据集上的实验结果表明了其优越的性能表现。我们进一步进行了详尽的数据选择消融研究,以建立结肠镜检查中姿态和深度估算的最佳实践。这项分析提供了两个实用见解:(1)基于现实世界数据的自监督训练优于基于逼真仿真数据的监督训练,突显了领域真实性的优越性而非地面真相的可用性;(2)视频帧率是模型性能的一个极其重要的因素,在生成高质量训练数据时需要针对特定的数据集进行视频帧采样。
URL
https://arxiv.org/abs/2602.17785