Paper Reading AI Learner

Multi-Modal Monocular Endoscopic Depth and Pose Estimation with Edge-Guided Self-Supervision

2026-02-19 19:38:11
Xinwei Ju, Rema Daher, Danail Stoyanov, Sophia Bano, Francisco Vasconcelos

Abstract

Monocular depth and pose estimation play an important role in the development of colonoscopy-assisted navigation, as they enable improved screening by reducing blind spots, minimizing the risk of missed or recurrent lesions, and lowering the likelihood of incomplete examinations. However, this task remains challenging due to the presence of texture-less surfaces, complex illumination patterns, deformation, and a lack of in-vivo datasets with reliable ground truth. In this paper, we propose **PRISM** (Pose-Refinement with Intrinsic Shading and edge Maps), a self-supervised learning framework that leverages anatomical and illumination priors to guide geometric learning. Our approach uniquely incorporates edge detection and luminance decoupling for structural guidance. Specifically, edge maps are derived using a learning-based edge detector (e.g., DexiNed or HED) trained to capture thin and high-frequency boundaries, while luminance decoupling is obtained through an intrinsic decomposition module that separates shading and reflectance, enabling the model to exploit shading cues for depth estimation. Experimental results on multiple real and synthetic datasets demonstrate state-of-the-art performance. We further conduct a thorough ablation study on training data selection to establish best practices for pose and depth estimation in colonoscopy. This analysis yields two practical insights: (1) self-supervised training on real-world data outperforms supervised training on realistic phantom data, underscoring the superiority of domain realism over ground truth availability; and (2) video frame rate is an extremely important factor for model performance, where dataset-specific video frame sampling is necessary for generating high quality training data.

Abstract (translated)

单目深度和姿态估计在结肠镜辅助导航的发展中扮演着重要角色,因为它们通过减少盲点、降低漏诊或复发病变的风险以及降低检查不完全的可能性来提高筛查效果。然而,由于无纹理表面、复杂的光照模式、变形及缺乏可靠地面真实数据的体内数据集等原因,该任务仍然具有挑战性。本文提出了一种新的方法**PRISM**(基于内在阴影和边缘图的姿态精炼),这是一种自监督学习框架,利用解剖学和照明先验指导几何学习。我们的方法独特地结合了边缘检测和亮度分离以提供结构引导。具体而言,通过使用训练有素的边缘探测器(如DexiNed或HED)来捕捉薄且高频边界,可以得到边缘图;同时通过内在分解模块获得亮度分离,该模块能够将阴影与反射分离开来,使模型利用阴影线索进行深度估计。在多个真实和合成数据集上的实验结果表明了其优越的性能表现。我们进一步进行了详尽的数据选择消融研究,以建立结肠镜检查中姿态和深度估算的最佳实践。这项分析提供了两个实用见解:(1)基于现实世界数据的自监督训练优于基于逼真仿真数据的监督训练,突显了领域真实性的优越性而非地面真相的可用性;(2)视频帧率是模型性能的一个极其重要的因素,在生成高质量训练数据时需要针对特定的数据集进行视频帧采样。

URL

https://arxiv.org/abs/2602.17785

PDF

https://arxiv.org/pdf/2602.17785.pdf


Tags
3D Action Action_Localization Action_Recognition Activity Adversarial Agent Attention Autonomous Bert Boundary_Detection Caption Chat Classification CNN Compressive_Sensing Contour Contrastive_Learning Deep_Learning Denoising Detection Dialog Diffusion Drone Dynamic_Memory_Network Edge_Detection Embedding Embodied Emotion Enhancement Face Face_Detection Face_Recognition Facial_Landmark Few-Shot Gait_Recognition GAN Gaze_Estimation Gesture Gradient_Descent Handwriting Human_Parsing Image_Caption Image_Classification Image_Compression Image_Enhancement Image_Generation Image_Matting Image_Retrieval Inference Inpainting Intelligent_Chip Knowledge Knowledge_Graph Language_Model LLM Matching Medical Memory_Networks Multi_Modal Multi_Task NAS NMT Object_Detection Object_Tracking OCR Ontology Optical_Character Optical_Flow Optimization Person_Re-identification Point_Cloud Portrait_Generation Pose Pose_Estimation Prediction QA Quantitative Quantitative_Finance Quantization Re-identification Recognition Recommendation Reconstruction Regularization Reinforcement_Learning Relation Relation_Extraction Represenation Represenation_Learning Restoration Review RNN Robot Salient Scene_Classification Scene_Generation Scene_Parsing Scene_Text Segmentation Self-Supervised Semantic_Instance_Segmentation Semantic_Segmentation Semi_Global Semi_Supervised Sence_graph Sentiment Sentiment_Classification Sketch SLAM Sparse Speech Speech_Recognition Style_Transfer Summarization Super_Resolution Surveillance Survey Text_Classification Text_Generation Time_Series Tracking Transfer_Learning Transformer Unsupervised Video_Caption Video_Classification Video_Indexing Video_Prediction Video_Retrieval Visual_Relation VQA Weakly_Supervised Zero-Shot