Paper Reading AI Learner

Microsurgical Instrument Segmentation for Robot-Assisted Surgery

2025-09-15 09:29:27
Tae Kyeong Jeong, Garam Kim, Juyoun Park

Abstract

Accurate segmentation of thin structures is critical for microsurgical scene understanding but remains challenging due to resolution loss, low contrast, and class imbalance. We propose Microsurgery Instrument Segmentation for Robotic Assistance(MISRA), a segmentation framework that augments RGB input with luminance channels, integrates skip attention to preserve elongated features, and employs an Iterative Feedback Module(IFM) for continuity restoration across multiple passes. In addition, we introduce a dedicated microsurgical dataset with fine-grained annotations of surgical instruments including thin objects, providing a benchmark for robust evaluation Dataset available at this https URL. Experiments demonstrate that MISRA achieves competitive performance, improving the mean class IoU by 5.37% over competing methods, while delivering more stable predictions at instrument contacts and overlaps. These results position MISRA as a promising step toward reliable scene parsing for computer-assisted and robotic microsurgery.

Abstract (translated)

精确分割微细结构对于显微手术场景的理解至关重要,但由于分辨率损失、对比度低和类别不平衡等问题,这一任务仍然具有挑战性。我们提出了一个名为“机器人辅助显微外科器械分割”(Microsurgery Instrument Segmentation for Robotic Assistance, MISRA)的分割框架。该框架通过增加RGB输入中的亮度通道来增强图像信息,并集成跳跃注意力机制以保留拉长特征,同时采用迭代反馈模块(IFM)在多次传递中恢复连续性。此外,我们还引入了一个专门针对显微手术的数据集,其中包含了对手术器械(包括细小物体)的精细标注,为稳健评估提供了基准数据集。该数据集可在提供的网址获取。 实验结果表明,MISRA实现了与竞争方法相媲美的性能,在平均类别交并比(mIoU)上提高了5.37%,并且在器械接触和重叠时能提供更稳定的预测。这些成果将MISRA定位为向计算机辅助和机器人显微手术中可靠场景解析迈出的有希望的一步。

URL

https://arxiv.org/abs/2509.11727

PDF

https://arxiv.org/pdf/2509.11727.pdf


Tags
3D Action Action_Localization Action_Recognition Activity Adversarial Agent Attention Autonomous Bert Boundary_Detection Caption Chat Classification CNN Compressive_Sensing Contour Contrastive_Learning Deep_Learning Denoising Detection Dialog Diffusion Drone Dynamic_Memory_Network Edge_Detection Embedding Embodied Emotion Enhancement Face Face_Detection Face_Recognition Facial_Landmark Few-Shot Gait_Recognition GAN Gaze_Estimation Gesture Gradient_Descent Handwriting Human_Parsing Image_Caption Image_Classification Image_Compression Image_Enhancement Image_Generation Image_Matting Image_Retrieval Inference Inpainting Intelligent_Chip Knowledge Knowledge_Graph Language_Model LLM Matching Medical Memory_Networks Multi_Modal Multi_Task NAS NMT Object_Detection Object_Tracking OCR Ontology Optical_Character Optical_Flow Optimization Person_Re-identification Point_Cloud Portrait_Generation Pose Pose_Estimation Prediction QA Quantitative Quantitative_Finance Quantization Re-identification Recognition Recommendation Reconstruction Regularization Reinforcement_Learning Relation Relation_Extraction Represenation Represenation_Learning Restoration Review RNN Robot Salient Scene_Classification Scene_Generation Scene_Parsing Scene_Text Segmentation Self-Supervised Semantic_Instance_Segmentation Semantic_Segmentation Semi_Global Semi_Supervised Sence_graph Sentiment Sentiment_Classification Sketch SLAM Sparse Speech Speech_Recognition Style_Transfer Summarization Super_Resolution Surveillance Survey Text_Classification Text_Generation Time_Series Tracking Transfer_Learning Transformer Unsupervised Video_Caption Video_Classification Video_Indexing Video_Prediction Video_Retrieval Visual_Relation VQA Weakly_Supervised Zero-Shot