Paper Reading AI Learner

PointAD+: Learning Hierarchical Representations for Zero-shot 3D Anomaly Detection

2025-09-03 12:53:40
Qihang Zhou, Shibo He, Jiangtao Yan, Wenchao Meng, Jiming Chen

Abstract

In this paper, we aim to transfer CLIP's robust 2D generalization capabilities to identify 3D anomalies across unseen objects of highly diverse class semantics. To this end, we propose a unified framework to comprehensively detect and segment 3D anomalies by leveraging both point- and pixel-level information. We first design PointAD, which leverages point-pixel correspondence to represent 3D anomalies through their associated rendering pixel representations. This approach is referred to as implicit 3D representation, as it focuses solely on rendering pixel anomalies but neglects the inherent spatial relationships within point clouds. Then, we propose PointAD+ to further broaden the interpretation of 3D anomalies by introducing explicit 3D representation, emphasizing spatial abnormality to uncover abnormal spatial relationships. Hence, we propose G-aggregation to involve geometry information to enable the aggregated point representations spatially aware. To simultaneously capture rendering and spatial abnormality, PointAD+ proposes hierarchical representation learning, incorporating implicit and explicit anomaly semantics into hierarchical text prompts: rendering prompts for the rendering layer and geometry prompts for the geometry layer. A cross-hierarchy contrastive alignment is further introduced to promote the interaction between the rendering and geometry layers, facilitating mutual anomaly learning. Finally, PointAD+ integrates anomaly semantics from both layers to capture the generalized anomaly semantics. During the test, PointAD+ can integrate RGB information in a plug-and-play manner and further improve its detection performance. Extensive experiments demonstrate the superiority of PointAD+ in ZS 3D anomaly detection across unseen objects with highly diverse class semantics, achieving a holistic understanding of abnormality.

Abstract (translated)

在这篇论文中,我们的目标是将CLIP强大的二维泛化能力转移到三维异常检测上,以识别不同类别的未知物体中的三维异常。为此,我们提出了一种统一的框架,利用点和像素级别的信息全面地检测并分割三维异常。 首先,我们设计了PointAD,它通过利用点-像素对应关系来表示3D异常,这些异常是通过它们相关的渲染像素表现出来的。这种表示方法被称为隐式3D表示法,因为它仅关注于渲染像素的异常,并忽略了点云内部的空间关系。然后,为了进一步扩展对三维异常的理解,我们提出了PointAD+,引入了显式的三维表示形式,强调空间异常以揭示异常的空间关系。 为使聚合后的点表示具有空间感知能力,我们提出了一种G-聚集方法,将几何信息纳入考量。为了同时捕捉渲染和空间异常,PointAD+ 提出了分层表示学习,将隐式和显式的异常语义融入层次化的文本提示中:渲染层的渲染提示和几何层的几何提示。 此外,还引入了跨层级对比对齐,以促进渲染层与几何层之间的交互作用,从而实现互惠的异常学习。最终,PointAD+ 将来自两个层次的异常语义整合起来,捕捉到泛化的异常语义。 在测试过程中,PointAD+ 可以通过即插即用的方式集成RGB信息,并进一步提高其检测性能。广泛的实验表明,PointAD+ 在处理具有高度多样类别语义未知物体的零样本(ZS)三维异常检测方面表现出优越性,实现了对异常性的整体理解。

URL

https://arxiv.org/abs/2509.03277

PDF

https://arxiv.org/pdf/2509.03277.pdf


Tags
3D Action Action_Localization Action_Recognition Activity Adversarial Agent Attention Autonomous Bert Boundary_Detection Caption Chat Classification CNN Compressive_Sensing Contour Contrastive_Learning Deep_Learning Denoising Detection Dialog Diffusion Drone Dynamic_Memory_Network Edge_Detection Embedding Embodied Emotion Enhancement Face Face_Detection Face_Recognition Facial_Landmark Few-Shot Gait_Recognition GAN Gaze_Estimation Gesture Gradient_Descent Handwriting Human_Parsing Image_Caption Image_Classification Image_Compression Image_Enhancement Image_Generation Image_Matting Image_Retrieval Inference Inpainting Intelligent_Chip Knowledge Knowledge_Graph Language_Model LLM Matching Medical Memory_Networks Multi_Modal Multi_Task NAS NMT Object_Detection Object_Tracking OCR Ontology Optical_Character Optical_Flow Optimization Person_Re-identification Point_Cloud Portrait_Generation Pose Pose_Estimation Prediction QA Quantitative Quantitative_Finance Quantization Re-identification Recognition Recommendation Reconstruction Regularization Reinforcement_Learning Relation Relation_Extraction Represenation Represenation_Learning Restoration Review RNN Robot Salient Scene_Classification Scene_Generation Scene_Parsing Scene_Text Segmentation Self-Supervised Semantic_Instance_Segmentation Semantic_Segmentation Semi_Global Semi_Supervised Sence_graph Sentiment Sentiment_Classification Sketch SLAM Sparse Speech Speech_Recognition Style_Transfer Summarization Super_Resolution Surveillance Survey Text_Classification Text_Generation Time_Series Tracking Transfer_Learning Transformer Unsupervised Video_Caption Video_Classification Video_Indexing Video_Prediction Video_Retrieval Visual_Relation VQA Weakly_Supervised Zero-Shot