Paper Reading AI Learner

RTGMFF: Enhanced fMRI-based Brain Disorder Diagnosis via ROI-driven Text Generation and Multimodal Feature Fusion

2025-09-03 11:05:57
Junhao Jia, Yifei Sun, Yunyou Liu, Cheng Yang, Changmiao Wang, Feiwei Qin, Yong Peng, Wenwen Min

Abstract

Functional magnetic resonance imaging (fMRI) is a powerful tool for probing brain function, yet reliable clinical diagnosis is hampered by low signal-to-noise ratios, inter-subject variability, and the limited frequency awareness of prevailing CNN- and Transformer-based models. Moreover, most fMRI datasets lack textual annotations that could contextualize regional activation and connectivity patterns. We introduce RTGMFF, a framework that unifies automatic ROI-level text generation with multimodal feature fusion for brain-disorder diagnosis. RTGMFF consists of three components: (i) ROI-driven fMRI text generation deterministically condenses each subject's activation, connectivity, age, and sex into reproducible text tokens; (ii) Hybrid frequency-spatial encoder fuses a hierarchical wavelet-mamba branch with a cross-scale Transformer encoder to capture frequency-domain structure alongside long-range spatial dependencies; and (iii) Adaptive semantic alignment module embeds the ROI token sequence and visual features in a shared space, using a regularized cosine-similarity loss to narrow the modality gap. Extensive experiments on the ADHD-200 and ABIDE benchmarks show that RTGMFF surpasses current methods in diagnostic accuracy, achieving notable gains in sensitivity, specificity, and area under the ROC curve. Code is available at this https URL.

Abstract (translated)

功能磁共振成像(fMRI)是一种强大的工具,用于探究大脑的功能,然而由于信号与噪声比率低、个体间差异大以及现有基于卷积神经网络(CNN)和变换器模型对频率感知有限的问题,可靠临床诊断的实现受到了阻碍。此外,大多数fMRI数据集缺乏能为区域激活及连接模式提供上下文信息的文字注释。我们引入了RTGMFF框架,该框架整合了自动ROI级别的文本生成与多模态特征融合,用于大脑疾病的诊断。 RTGMFF包含三个组成部分: (i) ROI驱动的fMRI文本生成:确定性地将每个受试者的激活情况、连接性、年龄和性别浓缩为可重复的文字标记; (ii) 混合频率-空间编码器:结合层次小波-蟒蛇分支与跨尺度变换器编码器,捕捉频域结构以及长程空间依赖关系; (iii) 自适应语义对齐模块:在共享空间中嵌入ROI令牌序列和视觉特征,并使用正则化的余弦相似度损失来缩小模态差距。 在ADHD-200和ABIDE基准上的广泛实验表明,RTGMFF超越了现有方法,在诊断准确性方面取得了显著的进步,具体体现在敏感性、特异性及ROC曲线下的面积等指标上。代码可在提供的网址获取。

URL

https://arxiv.org/abs/2509.03214

PDF

https://arxiv.org/pdf/2509.03214.pdf


Tags
3D Action Action_Localization Action_Recognition Activity Adversarial Agent Attention Autonomous Bert Boundary_Detection Caption Chat Classification CNN Compressive_Sensing Contour Contrastive_Learning Deep_Learning Denoising Detection Dialog Diffusion Drone Dynamic_Memory_Network Edge_Detection Embedding Embodied Emotion Enhancement Face Face_Detection Face_Recognition Facial_Landmark Few-Shot Gait_Recognition GAN Gaze_Estimation Gesture Gradient_Descent Handwriting Human_Parsing Image_Caption Image_Classification Image_Compression Image_Enhancement Image_Generation Image_Matting Image_Retrieval Inference Inpainting Intelligent_Chip Knowledge Knowledge_Graph Language_Model LLM Matching Medical Memory_Networks Multi_Modal Multi_Task NAS NMT Object_Detection Object_Tracking OCR Ontology Optical_Character Optical_Flow Optimization Person_Re-identification Point_Cloud Portrait_Generation Pose Pose_Estimation Prediction QA Quantitative Quantitative_Finance Quantization Re-identification Recognition Recommendation Reconstruction Regularization Reinforcement_Learning Relation Relation_Extraction Represenation Represenation_Learning Restoration Review RNN Robot Salient Scene_Classification Scene_Generation Scene_Parsing Scene_Text Segmentation Self-Supervised Semantic_Instance_Segmentation Semantic_Segmentation Semi_Global Semi_Supervised Sence_graph Sentiment Sentiment_Classification Sketch SLAM Sparse Speech Speech_Recognition Style_Transfer Summarization Super_Resolution Surveillance Survey Text_Classification Text_Generation Time_Series Tracking Transfer_Learning Transformer Unsupervised Video_Caption Video_Classification Video_Indexing Video_Prediction Video_Retrieval Visual_Relation VQA Weakly_Supervised Zero-Shot