Paper Reading AI Learner

Attn-Adapter: Attention Is All You Need for Online Few-shot Learner of Vision-Language Model

2025-09-04 05:42:02
Phuoc-Nguyen Bui, Khanh-Binh Nguyen, Hyunseung Choo

Abstract

Contrastive vision-language models excel in zero-shot image recognition but face challenges in few-shot scenarios due to computationally intensive offline fine-tuning using prompt learning, which risks overfitting. To overcome these limitations, we propose Attn-Adapter, a novel online few-shot learning framework that enhances CLIP's adaptability via a dual attention mechanism. Our design incorporates dataset-specific information through two components: the Memory Attn-Adapter, which refines category embeddings using support examples, and the Local-Global Attn-Adapter, which enriches image embeddings by integrating local and global features. This architecture enables dynamic adaptation from a few labeled samples without retraining the base model. Attn-Adapter outperforms state-of-the-art methods in cross-category and cross-dataset generalization, maintaining efficient inference and scaling across CLIP backbones.

Abstract (translated)

对比视觉语言模型在零样本图像识别方面表现出色,但在少量样本场景中却面临挑战。由于使用提示学习进行离线微调计算成本高昂,并且存在过拟合的风险,这些问题限制了它们的应用范围。为了克服这些局限性,我们提出了一种新颖的在线少量样本学习框架——Attn-Adapter。该框架通过双注意力机制增强了CLIP(对比语言-图像预训练模型)的适应能力。 我们的设计包括两个组件,旨在利用特定数据集的信息:记忆Attn-Adapter使用支持示例来优化类别嵌入;全局-局部Attn-Adapter则通过融合局部和全局特征来丰富图像嵌入。这种架构使得从少量标注样本进行动态适应成为可能,而无需重新训练基础模型。 实验结果显示,Attn-Adapter在跨类别的泛化能力和跨数据集的迁移学习上超越了当前最先进的方法,并且能够在保持高效推理的同时,在不同的CLIP骨干网络中实现规模扩展。

URL

https://arxiv.org/abs/2509.03895

PDF

https://arxiv.org/pdf/2509.03895.pdf


Tags
3D Action Action_Localization Action_Recognition Activity Adversarial Agent Attention Autonomous Bert Boundary_Detection Caption Chat Classification CNN Compressive_Sensing Contour Contrastive_Learning Deep_Learning Denoising Detection Dialog Diffusion Drone Dynamic_Memory_Network Edge_Detection Embedding Embodied Emotion Enhancement Face Face_Detection Face_Recognition Facial_Landmark Few-Shot Gait_Recognition GAN Gaze_Estimation Gesture Gradient_Descent Handwriting Human_Parsing Image_Caption Image_Classification Image_Compression Image_Enhancement Image_Generation Image_Matting Image_Retrieval Inference Inpainting Intelligent_Chip Knowledge Knowledge_Graph Language_Model LLM Matching Medical Memory_Networks Multi_Modal Multi_Task NAS NMT Object_Detection Object_Tracking OCR Ontology Optical_Character Optical_Flow Optimization Person_Re-identification Point_Cloud Portrait_Generation Pose Pose_Estimation Prediction QA Quantitative Quantitative_Finance Quantization Re-identification Recognition Recommendation Reconstruction Regularization Reinforcement_Learning Relation Relation_Extraction Represenation Represenation_Learning Restoration Review RNN Robot Salient Scene_Classification Scene_Generation Scene_Parsing Scene_Text Segmentation Self-Supervised Semantic_Instance_Segmentation Semantic_Segmentation Semi_Global Semi_Supervised Sence_graph Sentiment Sentiment_Classification Sketch SLAM Sparse Speech Speech_Recognition Style_Transfer Summarization Super_Resolution Surveillance Survey Text_Classification Text_Generation Time_Series Tracking Transfer_Learning Transformer Unsupervised Video_Caption Video_Classification Video_Indexing Video_Prediction Video_Retrieval Visual_Relation VQA Weakly_Supervised Zero-Shot