Abstract
Functional magnetic resonance imaging (fMRI) is a powerful tool for probing brain function, yet reliable clinical diagnosis is hampered by low signal-to-noise ratios, inter-subject variability, and the limited frequency awareness of prevailing CNN- and Transformer-based models. Moreover, most fMRI datasets lack textual annotations that could contextualize regional activation and connectivity patterns. We introduce RTGMFF, a framework that unifies automatic ROI-level text generation with multimodal feature fusion for brain-disorder diagnosis. RTGMFF consists of three components: (i) ROI-driven fMRI text generation deterministically condenses each subject's activation, connectivity, age, and sex into reproducible text tokens; (ii) Hybrid frequency-spatial encoder fuses a hierarchical wavelet-mamba branch with a cross-scale Transformer encoder to capture frequency-domain structure alongside long-range spatial dependencies; and (iii) Adaptive semantic alignment module embeds the ROI token sequence and visual features in a shared space, using a regularized cosine-similarity loss to narrow the modality gap. Extensive experiments on the ADHD-200 and ABIDE benchmarks show that RTGMFF surpasses current methods in diagnostic accuracy, achieving notable gains in sensitivity, specificity, and area under the ROC curve. Code is available at this https URL.
Abstract (translated)
功能磁共振成像(fMRI)是一种强大的工具,用于探究大脑的功能,然而由于信号与噪声比率低、个体间差异大以及现有基于卷积神经网络(CNN)和变换器模型对频率感知有限的问题,可靠临床诊断的实现受到了阻碍。此外,大多数fMRI数据集缺乏能为区域激活及连接模式提供上下文信息的文字注释。我们引入了RTGMFF框架,该框架整合了自动ROI级别的文本生成与多模态特征融合,用于大脑疾病的诊断。 RTGMFF包含三个组成部分: (i) ROI驱动的fMRI文本生成:确定性地将每个受试者的激活情况、连接性、年龄和性别浓缩为可重复的文字标记; (ii) 混合频率-空间编码器:结合层次小波-蟒蛇分支与跨尺度变换器编码器,捕捉频域结构以及长程空间依赖关系; (iii) 自适应语义对齐模块:在共享空间中嵌入ROI令牌序列和视觉特征,并使用正则化的余弦相似度损失来缩小模态差距。 在ADHD-200和ABIDE基准上的广泛实验表明,RTGMFF超越了现有方法,在诊断准确性方面取得了显著的进步,具体体现在敏感性、特异性及ROC曲线下的面积等指标上。代码可在提供的网址获取。
URL
https://arxiv.org/abs/2509.03214