Paper Reading AI Learner

RDFace: A Benchmark Dataset for Rare Disease Facial Image Analysis under Extreme Data Scarcity and Phenotype-Aware Synthetic Generation

2026-04-03 20:56:30
Ganlin Feng, Yuxi Long, Hafsa Ali, Erin Lou, Fahad Butt, Qian Liu, Yang Wang, Pingzhao Hu

Abstract

Rare diseases often manifest with distinctive facial phenotypes in children, offering valuable diagnostic cues for clinicians and AI-assisted screening systems. However, progress in this field is severely limited by the scarcity of curated, ethically sourced facial data and the high similarity among phenotypes across different conditions. To address these challenges, we introduce RDFace, a curated benchmark dataset comprising 456 pediatric facial images spanning 103 rare genetic conditions (average 4.4 samples per condition). Each ethically verified image is paired with standardized metadata. RDFace enables the development and evaluation of data-efficient AI models for rare disease diagnosis under real-world low-data constraints. We benchmark multiple pretrained vision backbones using cross-validation and explore synthetic augmentation with DreamBooth and FastGAN. Generated images are filtered via facial landmark similarity to maintain phenotype fidelity and merged with real data, improving diagnostic accuracy by up to 13.7% in ultra-low-data regimes. To assess semantic validity, phenotype descriptions generated by a vision-language model from real and synthetic images achieve a report similarity score of 0.84. RDFace establishes a transparent, benchmark-ready dataset for equitable rare disease AI research and presents a scalable framework for evaluating both diagnostic performance and the integrity of synthetic medical imagery.

Abstract (translated)

罕见病在儿童中常表现为独特的 facial phenotypes(面部表型),为临床医生和人工智能辅助筛查系统提供了宝贵的诊断线索。然而,该领域的发展严重受限于:符合伦理规范的面部数据稀缺,且不同疾病间的表型高度相似。为应对这些挑战,我们推出了 RDFace——一个经过精心整理的基准数据集,包含 456 张儿科面部图像,涵盖 103 种罕见遗传病(平均每种疾病 4.4 个样本)。每张符合伦理验证的图像均配有标准化元数据。RDFace 使得在真实世界低数据约束下,开发和评估数据高效的人工智能模型用于罕见病诊断成为可能。我们通过交叉验证对多个预训练视觉骨干网络进行了基准测试,并探索了使用 DreamBooth 和 FastGAN 进行合成数据增强。生成的图像通过面部关键点相似度进行筛选以保持表型真实性,并与真实数据合并,在极低数据场景下将诊断准确率最高提升了 13.7%。为评估语义有效性,视觉-语言模型从真实和合成图像生成的表型描述,其报告相似度评分达到 0.84。RDFace 建立了一个透明、即用型基准数据集,以促进公平的罕见病人工智能研究,并提出了一种可扩展的框架,用于评估诊断性能及合成医学图像的完整性。

URL

https://arxiv.org/abs/2604.03454

PDF

https://arxiv.org/pdf/2604.03454.pdf


Tags
3D Action Action_Localization Action_Recognition Activity Adversarial Agent Attention Autonomous Bert Boundary_Detection Caption Chat Classification CNN Compressive_Sensing Contour Contrastive_Learning Deep_Learning Denoising Detection Dialog Diffusion Drone Dynamic_Memory_Network Edge_Detection Embedding Embodied Emotion Enhancement Face Face_Detection Face_Recognition Facial_Landmark Few-Shot Gait_Recognition GAN Gaze_Estimation Gesture Gradient_Descent Handwriting Human_Parsing Image_Caption Image_Classification Image_Compression Image_Enhancement Image_Generation Image_Matting Image_Retrieval Inference Inpainting Intelligent_Chip Knowledge Knowledge_Graph Language_Model LLM Matching Medical Memory_Networks Multi_Modal Multi_Task NAS NMT Object_Detection Object_Tracking OCR Ontology Optical_Character Optical_Flow Optimization Person_Re-identification Point_Cloud Portrait_Generation Pose Pose_Estimation Prediction QA Quantitative Quantitative_Finance Quantization Re-identification Recognition Recommendation Reconstruction Regularization Reinforcement_Learning Relation Relation_Extraction Represenation Represenation_Learning Restoration Review RNN Robot Salient Scene_Classification Scene_Generation Scene_Parsing Scene_Text Segmentation Self-Supervised Semantic_Instance_Segmentation Semantic_Segmentation Semi_Global Semi_Supervised Sence_graph Sentiment Sentiment_Classification Sketch SLAM Sparse Speech Speech_Recognition Style_Transfer Summarization Super_Resolution Surveillance Survey Text_Classification Text_Generation Time_Series Tracking Transfer_Learning Transformer Unsupervised Video_Caption Video_Classification Video_Indexing Video_Prediction Video_Retrieval Visual_Relation VQA Weakly_Supervised Zero-Shot