Abstract
In this paper, we aim to transfer CLIP's robust 2D generalization capabilities to identify 3D anomalies across unseen objects of highly diverse class semantics. To this end, we propose a unified framework to comprehensively detect and segment 3D anomalies by leveraging both point- and pixel-level information. We first design PointAD, which leverages point-pixel correspondence to represent 3D anomalies through their associated rendering pixel representations. This approach is referred to as implicit 3D representation, as it focuses solely on rendering pixel anomalies but neglects the inherent spatial relationships within point clouds. Then, we propose PointAD+ to further broaden the interpretation of 3D anomalies by introducing explicit 3D representation, emphasizing spatial abnormality to uncover abnormal spatial relationships. Hence, we propose G-aggregation to involve geometry information to enable the aggregated point representations spatially aware. To simultaneously capture rendering and spatial abnormality, PointAD+ proposes hierarchical representation learning, incorporating implicit and explicit anomaly semantics into hierarchical text prompts: rendering prompts for the rendering layer and geometry prompts for the geometry layer. A cross-hierarchy contrastive alignment is further introduced to promote the interaction between the rendering and geometry layers, facilitating mutual anomaly learning. Finally, PointAD+ integrates anomaly semantics from both layers to capture the generalized anomaly semantics. During the test, PointAD+ can integrate RGB information in a plug-and-play manner and further improve its detection performance. Extensive experiments demonstrate the superiority of PointAD+ in ZS 3D anomaly detection across unseen objects with highly diverse class semantics, achieving a holistic understanding of abnormality.
Abstract (translated)
在这篇论文中,我们的目标是将CLIP强大的二维泛化能力转移到三维异常检测上,以识别不同类别的未知物体中的三维异常。为此,我们提出了一种统一的框架,利用点和像素级别的信息全面地检测并分割三维异常。 首先,我们设计了PointAD,它通过利用点-像素对应关系来表示3D异常,这些异常是通过它们相关的渲染像素表现出来的。这种表示方法被称为隐式3D表示法,因为它仅关注于渲染像素的异常,并忽略了点云内部的空间关系。然后,为了进一步扩展对三维异常的理解,我们提出了PointAD+,引入了显式的三维表示形式,强调空间异常以揭示异常的空间关系。 为使聚合后的点表示具有空间感知能力,我们提出了一种G-聚集方法,将几何信息纳入考量。为了同时捕捉渲染和空间异常,PointAD+ 提出了分层表示学习,将隐式和显式的异常语义融入层次化的文本提示中:渲染层的渲染提示和几何层的几何提示。 此外,还引入了跨层级对比对齐,以促进渲染层与几何层之间的交互作用,从而实现互惠的异常学习。最终,PointAD+ 将来自两个层次的异常语义整合起来,捕捉到泛化的异常语义。 在测试过程中,PointAD+ 可以通过即插即用的方式集成RGB信息,并进一步提高其检测性能。广泛的实验表明,PointAD+ 在处理具有高度多样类别语义未知物体的零样本(ZS)三维异常检测方面表现出优越性,实现了对异常性的整体理解。
URL
https://arxiv.org/abs/2509.03277