Paper Reading AI Learner

Towards Comprehensive Cellular Characterisation of H&E slides

2025-08-13 16:24:15
Benjamin Adjadj (Owkin), Pierre-Antoine Bannier (Owkin), Guillaume Horent (Owkin), Sebastien Mandela (Owkin), Aurore Lyon (Owkin), Kathryn Schutte (Owkin), Ulysse Marteau (Owkin), Valentin Gaury (Owkin), Laura Dumont (Owkin), Thomas Mathieu (Owkin), Reda Belbahri (Owkin), Beno\^it Schmauch (Owkin), Eric Durand (Owkin), Katharina Von Loga (Owkin), Lucie Gillet (Owkin)

Abstract

Cell detection, segmentation and classification are essential for analyzing tumor microenvironments (TME) on hematoxylin and eosin (H&E) slides. Existing methods suffer from poor performance on understudied cell types (rare or not present in public datasets) and limited cross-domain generalization. To address these shortcomings, we introduce HistoPLUS, a state-of-the-art model for cell analysis, trained on a novel curated pan-cancer dataset of 108,722 nuclei covering 13 cell types. In external validation across 4 independent cohorts, HistoPLUS outperforms current state-of-the-art models in detection quality by 5.2% and overall F1 classification score by 23.7%, while using 5x fewer parameters. Notably, HistoPLUS unlocks the study of 7 understudied cell types and brings significant improvements on 8 of 13 cell types. Moreover, we show that HistoPLUS robustly transfers to two oncology indications unseen during training. To support broader TME biomarker research, we release the model weights and inference code at this https URL.

Abstract (translated)

细胞检测、分割和分类对于分析肿瘤微环境(TME)在苏木精-伊红(H&E)切片上的情况至关重要。现有的方法在处理研究不足的细胞类型(罕见或未出现在公共数据集中的细胞类型)以及跨领域泛化方面表现不佳。为了解决这些问题,我们引入了HistoPLUS,这是一种先进的细胞分析模型,它是在一个新的精心策划的、包含108,722个核(覆盖13种细胞类型)的泛癌症数据集上训练出来的。在外部验证中,HistoPLUS在四个独立队列中的检测质量比目前最先进的模型高出5.2%,整体F1分类得分提高了23.7%,同时使用的参数减少了5倍。值得注意的是,HistoPLUS解锁了对七种研究不足的细胞类型的研究,并且在这十三种细胞类型中有八种取得了显著改进。此外,我们还展示了HistoPLUS能够稳健地转移到训练期间未见过的两种肿瘤学指征上。为了支持更广泛的TME生物标志物研究,我们在[这里](https://example.com)发布了模型权重和推理代码。

URL

https://arxiv.org/abs/2508.09926

PDF

https://arxiv.org/pdf/2508.09926.pdf


Tags
3D Action Action_Localization Action_Recognition Activity Adversarial Agent Attention Autonomous Bert Boundary_Detection Caption Chat Classification CNN Compressive_Sensing Contour Contrastive_Learning Deep_Learning Denoising Detection Dialog Diffusion Drone Dynamic_Memory_Network Edge_Detection Embedding Embodied Emotion Enhancement Face Face_Detection Face_Recognition Facial_Landmark Few-Shot Gait_Recognition GAN Gaze_Estimation Gesture Gradient_Descent Handwriting Human_Parsing Image_Caption Image_Classification Image_Compression Image_Enhancement Image_Generation Image_Matting Image_Retrieval Inference Inpainting Intelligent_Chip Knowledge Knowledge_Graph Language_Model LLM Matching Medical Memory_Networks Multi_Modal Multi_Task NAS NMT Object_Detection Object_Tracking OCR Ontology Optical_Character Optical_Flow Optimization Person_Re-identification Point_Cloud Portrait_Generation Pose Pose_Estimation Prediction QA Quantitative Quantitative_Finance Quantization Re-identification Recognition Recommendation Reconstruction Regularization Reinforcement_Learning Relation Relation_Extraction Represenation Represenation_Learning Restoration Review RNN Robot Salient Scene_Classification Scene_Generation Scene_Parsing Scene_Text Segmentation Self-Supervised Semantic_Instance_Segmentation Semantic_Segmentation Semi_Global Semi_Supervised Sence_graph Sentiment Sentiment_Classification Sketch SLAM Sparse Speech Speech_Recognition Style_Transfer Summarization Super_Resolution Surveillance Survey Text_Classification Text_Generation Time_Series Tracking Transfer_Learning Transformer Unsupervised Video_Caption Video_Classification Video_Indexing Video_Prediction Video_Retrieval Visual_Relation VQA Weakly_Supervised Zero-Shot