Paper Reading AI Learner

All in One: Unifying Deepfake Detection, Tampering Localization, and Source Tracing with a Robust Landmark-Identity Watermark

2026-02-26 21:57:15
Junjiang Wu, Liejun Wang, Zhiqing Guo

Abstract

With the rapid advancement of deepfake technology, malicious face manipulations pose a significant threat to personal privacy and social security. However, existing proactive forensics methods typically treat deepfake detection, tampering localization, and source tracing as independent tasks, lacking a unified framework to address them jointly. To bridge this gap, we propose a unified proactive forensics framework that jointly addresses these three core tasks. Our core framework adopts an innovative 152-dimensional landmark-identity watermark termed LIDMark, which structurally interweaves facial landmarks with a unique source identifier. To robustly extract the LIDMark, we design a novel Factorized-Head Decoder (FHD). Its architecture factorizes the shared backbone features into two specialized heads (i.e., regression and classification), robustly reconstructing the embedded landmarks and identifier, respectively, even when subjected to severe distortion or tampering. This design realizes an "all-in-one" trifunctional forensic solution: the regression head underlies an "intrinsic-extrinsic" consistency check for detection and localization, while the classification head robustly decodes the source identifier for tracing. Extensive experiments show that the proposed LIDMark framework provides a unified, robust, and imperceptible solution for the detection, localization, and tracing of deepfake content. The code is available at this https URL.

Abstract (translated)

随着深度伪造技术的迅速发展,恶意的脸部篡改对个人隐私和社会安全构成了重大威胁。然而,现有的主动取证方法通常将深度伪造检测、篡改定位和源头追溯视为独立的任务,并缺乏一个统一框架来共同解决这些问题。为弥补这一空白,我们提出了一种统一的主动取证框架,该框架能够同时应对这三个核心任务。 我们的核心框架采用了一种创新性的152维地标-身份水印(LIDMark),这种水印将面部特征点与唯一来源标识符结构化地交织在一起。为了稳健地提取LIDMark,我们设计了一个新颖的因子头部解码器(Factorized-Head Decoder, FHD)。FHD架构将其共享骨干特征分解为两个专门化的头部(即回归和分类),即使在遭受严重扭曲或篡改的情况下,也能分别稳健地重构嵌入的地标和标识符。这种设计实现了“三位一体”的多功能取证解决方案:回归头部用于深度伪造检测和定位的内在-外在一致性检查,而分类头部则通过解码来源标识符来进行稳健的溯源。 广泛的实验表明,所提出的LIDMark框架为深度伪造内容的检测、定位和追踪提供了一种统一、鲁棒且不可感知的解决方案。代码可在[此处](https://这个URL)获取。

URL

https://arxiv.org/abs/2602.23523

PDF

https://arxiv.org/pdf/2602.23523.pdf


Tags
3D Action Action_Localization Action_Recognition Activity Adversarial Agent Attention Autonomous Bert Boundary_Detection Caption Chat Classification CNN Compressive_Sensing Contour Contrastive_Learning Deep_Learning Denoising Detection Dialog Diffusion Drone Dynamic_Memory_Network Edge_Detection Embedding Embodied Emotion Enhancement Face Face_Detection Face_Recognition Facial_Landmark Few-Shot Gait_Recognition GAN Gaze_Estimation Gesture Gradient_Descent Handwriting Human_Parsing Image_Caption Image_Classification Image_Compression Image_Enhancement Image_Generation Image_Matting Image_Retrieval Inference Inpainting Intelligent_Chip Knowledge Knowledge_Graph Language_Model LLM Matching Medical Memory_Networks Multi_Modal Multi_Task NAS NMT Object_Detection Object_Tracking OCR Ontology Optical_Character Optical_Flow Optimization Person_Re-identification Point_Cloud Portrait_Generation Pose Pose_Estimation Prediction QA Quantitative Quantitative_Finance Quantization Re-identification Recognition Recommendation Reconstruction Regularization Reinforcement_Learning Relation Relation_Extraction Represenation Represenation_Learning Restoration Review RNN Robot Salient Scene_Classification Scene_Generation Scene_Parsing Scene_Text Segmentation Self-Supervised Semantic_Instance_Segmentation Semantic_Segmentation Semi_Global Semi_Supervised Sence_graph Sentiment Sentiment_Classification Sketch SLAM Sparse Speech Speech_Recognition Style_Transfer Summarization Super_Resolution Surveillance Survey Text_Classification Text_Generation Time_Series Tracking Transfer_Learning Transformer Unsupervised Video_Caption Video_Classification Video_Indexing Video_Prediction Video_Retrieval Visual_Relation VQA Weakly_Supervised Zero-Shot