Paper Reading AI Learner

Deep Adaptive Attention for Joint Facial Action Unit Detection and Face Alignment

2018-07-24 09:05:07
Zhiwen Shao, Zhilei Liu, Jianfei Cai, Lizhuang Ma

Abstract

Facial action unit (AU) detection and face alignment are two highly correlated tasks since facial landmarks can provide precise AU locations to facilitate the extraction of meaningful local features for AU detection. Most existing AU detection works often treat face alignment as a preprocessing and handle the two tasks independently. In this paper, we propose a novel end-to-end deep learning framework for joint AU detection and face alignment, which has not been explored before. In particular, multi-scale shared features are learned firstly, and high-level features of face alignment are fed into AU detection. Moreover, to extract precise local features, we propose an adaptive attention learning module to refine the attention map of each AU adaptively. Finally, the assembled local features are integrated with face alignment features and global features for AU detection. Experiments on BP4D and DISFA benchmarks demonstrate that our framework significantly outperforms the state-of-the-art methods for AU detection.

Abstract (translated)

面部动作单元(AU)检测和面部对齐是两个高度相关的任务,因为面部地标可以提供精确的AU位置以便于提取用于AU检测的有意义的局部特征。大多数现有的AU检测工作通常将面部对齐视为预处理并独立处理这两个任务。在本文中,我们提出了一种新的端对端深度学习框架,用于联合AU检测和面部对齐,这在以前尚未探索过。特别地,首先学习多尺度共享特征,并且将面部对齐的高级特征馈送到AU检测中。此外,为了提取精确的局部特征,我们提出了一种自适应注意力学习模块,以自适应地细化每个AU的注意力图。最后,组装的局部特征与面部对齐特征和用于AU检测的全局特征集成在一起。 BP4D和DISFA基准测试的实验表明,我们的框架明显优于AU检测的最先进方法。

URL

https://arxiv.org/abs/1803.05588

PDF

https://arxiv.org/pdf/1803.05588.pdf


Tags
3D Action Action_Localization Action_Recognition Activity Adversarial Agent Attention Autonomous Bert Boundary_Detection Caption Chat Classification CNN Compressive_Sensing Contour Contrastive_Learning Deep_Learning Denoising Detection Dialog Diffusion Drone Dynamic_Memory_Network Edge_Detection Embedding Embodied Emotion Enhancement Face Face_Detection Face_Recognition Facial_Landmark Few-Shot Gait_Recognition GAN Gaze_Estimation Gesture Gradient_Descent Handwriting Human_Parsing Image_Caption Image_Classification Image_Compression Image_Enhancement Image_Generation Image_Matting Image_Retrieval Inference Inpainting Intelligent_Chip Knowledge Knowledge_Graph Language_Model Matching Medical Memory_Networks Multi_Modal Multi_Task NAS NMT Object_Detection Object_Tracking OCR Ontology Optical_Character Optical_Flow Optimization Person_Re-identification Point_Cloud Portrait_Generation Pose Pose_Estimation Prediction QA Quantitative Quantitative_Finance Quantization Re-identification Recognition Recommendation Reconstruction Regularization Reinforcement_Learning Relation Relation_Extraction Represenation Represenation_Learning Restoration Review RNN Salient Scene_Classification Scene_Generation Scene_Parsing Scene_Text Segmentation Self-Supervised Semantic_Instance_Segmentation Semantic_Segmentation Semi_Global Semi_Supervised Sence_graph Sentiment Sentiment_Classification Sketch SLAM Sparse Speech Speech_Recognition Style_Transfer Summarization Super_Resolution Surveillance Survey Text_Classification Text_Generation Tracking Transfer_Learning Transformer Unsupervised Video_Caption Video_Classification Video_Indexing Video_Prediction Video_Retrieval Visual_Relation VQA Weakly_Supervised Zero-Shot