Abstract
High-precision facial landmark detection (FLD) relies on high-resolution deep feature representations. However, low-resolution face images or the compression (via pooling or strided convolution) of originally high-resolution images hinder the learning of such features, thereby reducing FLD accuracy. Moreover, insufficient training data and imprecise annotations further degrade performance. To address these challenges, we propose a weakly-supervised framework called Supervision-by-Hallucination-and-Transfer (SHT) for more robust and precise FLD. SHT contains two novel mutually enhanced modules: Dual Hallucination Learning Network (DHLN) and Facial Pose Transfer Network (FPTN). By incorporating FLD and face hallucination tasks, DHLN is able to learn high-resolution representations with low-resolution inputs for recovering both facial structures and local details and generating more effective landmark heatmaps. Then, by transforming faces from one pose to another, FPTN can further improve landmark heatmaps and faces hallucinated by DHLN for detecting more accurate landmarks. To the best of our knowledge, this is the first study to explore weakly-supervised FLD by integrating face hallucination and facial pose transfer tasks. Experimental results of both face hallucination and FLD demonstrate that our method surpasses state-of-the-art techniques.
Abstract (translated)
高精度面部标志检测(FLD)依赖于高质量的深度特征表示。然而,低分辨率的人脸图像或对原本高分辨率图像通过池化或步进卷积进行压缩会妨碍这些特征的学习,从而降低FLD的准确性。此外,训练数据不足和标注不准确进一步降低了性能。为了解决这些问题,我们提出了一种弱监督框架,称为幻觉与迁移监督(SHT),以实现更稳健且精确的FLD。SHT包含两个新颖的相互增强模块:双幻觉学习网络(DHLN)和面部姿态转换网络(FPTN)。通过结合FLD和人脸幻觉任务,DHLN能够使用低分辨率输入来学习高分辨率表示,从而恢复面部结构和局部细节,并生成更有效的标志热图。随后,通过将面孔从一种姿态变换为另一种姿态,FPTN可以进一步改进由DHLN产生的面部标志热图及幻觉出的人脸,以检测到更加准确的标志点。据我们所知,这是首次探索结合人脸幻觉和面部姿态转换任务的弱监督FLD的研究。实验结果表明,在人脸幻觉和FLD方面,我们的方法超越了现有的先进技术。
URL
https://arxiv.org/abs/2601.12919