TokenPose: Learning Keypoint Tokens for Human Pose Estimation

2021-04-08 05:12:38

Yanjie Li, Shoukui Zhang, Zhicheng Wang, Sen Yang, Wankou Yang, Shu-Tao Xia, Erjin Zhou

arXiv_CV

Abstract
Abstract (translated)
URL
PDF

Abstract

Human pose estimation deeply relies on visual clues and anatomical constraints between parts to locate keypoints. Most existing CNN-based methods do well in visual representation, however, lacking in the ability to explicitly learn the constraint relationships between keypoints. In this paper, we propose a novel approach based on Token representation for human Pose estimation~(TokenPose). In detail, each keypoint is explicitly embedded as a token to simultaneously learn constraint relationships and appearance cues from images. Extensive experiments show that the small and large TokenPose models are on par with state-of-the-art CNN-based counterparts while being more lightweight. Specifically, our TokenPose-S and TokenPose-L achieve 72.5 AP and 75.8 AP on COCO validation dataset respectively, with significant reduction in parameters (\textcolor{red}{ $\downarrow80.6\%$} ; \textcolor{red}{$\downarrow$ $56.8\%$}) and GFLOPs (\textcolor{red}{$\downarrow$$ 75.3\%$}; \textcolor{red}{$\downarrow$ $24.7\%$}).

Abstract (translated)

URL

https://arxiv.org/abs/2104.03516

PDF

https://arxiv.org/pdf/2104.03516.pdf