A Unified Game-Theoretic Interpretation of Adversarial Robustness

2021-11-05 14:57:49

Jie Ren, Die Zhang, Yisen Wang, Lu Chen, Zhanpeng Zhou, Yiting Chen, Xu Cheng, Xin Wang, Meng Zhou, Jie Shi, Quanshi Zhang

arXiv_AI

arXiv_AI Adversarial Action

Abstract
Abstract (translated)
URL
PDF

Abstract

This paper provides a unified view to explain different adversarial attacks and defense methods, \emph{i.e.} the view of multi-order interactions between input variables of DNNs. Based on the multi-order interaction, we discover that adversarial attacks mainly affect high-order interactions to fool the DNN. Furthermore, we find that the robustness of adversarially trained DNNs comes from category-specific low-order interactions. Our findings provide a potential method to unify adversarial perturbations and robustness, which can explain the existing defense methods in a principle way. Besides, our findings also make a revision of previous inaccurate understanding of the shape bias of adversarially learned features.

Abstract (translated)

URL

https://arxiv.org/abs/2111.03536

PDF

https://arxiv.org/pdf/2111.03536.pdf