Beyond Model Extraction: Imitation Attack for Black-Box NLP APIs

2021-08-29 10:52:04

Qiongkai Xu, Xuanli He, Lingjuan Lyu, Lizhen Qu, Gholamreza Haffari

arXiv_CL

arXiv_CL Unsupervised Action

Abstract
Abstract (translated)
URL
PDF

Abstract

Machine-learning-as-a-service (MLaaS) has attracted millions of users to their outperforming sophisticated models. Although published as black-box APIs, the valuable models behind these services are still vulnerable to imitation attacks. Recently, a series of works have demonstrated that attackers manage to steal or extract the victim models. Nonetheless, none of the previous stolen models can outperform the original black-box APIs. In this work, we take the first step of showing that attackers could potentially surpass victims via unsupervised domain adaptation and multi-victim ensemble. Extensive experiments on benchmark datasets and real-world APIs validate that the imitators can succeed in outperforming the original black-box models. We consider this as a milestone in the research of imitation attack, especially on NLP APIs, as the superior performance could influence the defense or even publishing strategy of API providers.

Abstract (translated)

URL

https://arxiv.org/abs/2108.13873

PDF

https://arxiv.org/pdf/2108.13873.pdf