Abstract Recent advancements in deep learning have significantly improved the accuracy of multi-person pose estimation from RGB images. However, these deep learning methods typically rely on a large number of deep refinement modules to refine the features of body joints and limbs, which hugely reduce the run-time speed and therefore limit the application domain. In this paper, we propose a feature transfer framework to capture the concurrent correlations between body joint and limb features. The concurrent correlations of these features form a complementary structural relationship, which mutually strengthens the network’s inferences and reduces the needs of refinement modules. The transfer sub-network is implemented with multiple convolutional layers, and is merged with the body part detection network to form an end-to-end system. The transfer relationship is automatically learned from ground-truth data instead of being manually encoded, resulting in a more general and efficient design. The proposed framework is validated on the multiple popular multi-person pose estimation benchmarks - MPII, COCO 2018 and PoseTrack 2017 and 2018. Experimental results show that our method not only significantly increases the inference speed to 73.8 frame per second (FPS), but also attains comparable state-of-the-art performance.
[1]
Kyoung Mu Lee,et al.
2D-3D Pose Consistency-based Conditional Random Fields for 3D Human Pose Estimation
,
2017,
Comput. Vis. Image Underst..
[2]
Meng Wang,et al.
Multimodal Deep Autoencoder for Human Pose Recovery
,
2015,
IEEE Transactions on Image Processing.
[3]
Chaoqun Hong,et al.
Hypergraph regularized autoencoder for image-based 3D human pose recovery
,
2016,
Signal Process..
[4]
Rich Caruana,et al.
Multitask Learning
,
1997,
Machine Learning.
[5]
Ming-Hsuan Yang,et al.
Ensemble convolutional neural networks for pose estimation
,
2018,
Comput. Vis. Image Underst..
[6]
Trevor Darrell,et al.
Caffe: Convolutional Architecture for Fast Feature Embedding
,
2014,
ACM Multimedia.
[7]
Meng Wang,et al.
Image-Based Three-Dimensional Human Pose Recovery by Multiview Locality-Sensitive Sparse Retrieval
,
2015,
IEEE Transactions on Industrial Electronics.