A cross-feature interaction network for 3D human pose estimation

被引:0
|
作者
Peng, Jihua [1 ]
Zhou, Yanghong [3 ]
Mok, P. Y. [1 ,2 ,4 ,5 ]
机构
[1] Hong Kong Polytech Univ, Sch Fash & Text, Hong Kong, Peoples R China
[2] Lab Artificial Intelligence Design, Hong Kong, Peoples R China
[3] Hong Kong Polytech Univ, Res Ctr Text Future Fash, Hong Kong, Peoples R China
[4] Hong Kong Polytech Univ, Res Inst Sports Sci & Technol, Hong Kong, Peoples R China
[5] Hong Kong Univ Sci & Technol, Div Integrat Syst & Design, Hong Kong, Peoples R China
关键词
3D human pose estimation; graph convolutional network (GCN); self-attention; cross-attention;
D O I
10.1016/j.patrec.2025.01.016
中图分类号
TP18 [人工智能理论];
学科分类号
081104 ; 0812 ; 0835 ; 1405 ;
摘要
The task of estimating 3D human poses from single monocular images is challenging because, unlike video sequences, single images can hardly provide any temporal information for the prediction. Most existing methods attempt to predict 3D poses by modeling the spatial dependencies inherent in the anatomical structure of the human skeleton, yet these methods fail to capture the complex local and global relationships that exist among various joints. To solve this problem, we propose a novel Cross-Feature Interaction Network to effectively model spatial correlations between body joints. Specifically, we exploit graph convolutional networks (GCNs) to learn the local features between neighboring joints and the self-attention structure to learn the global features among all joints. We then design a cross-feature interaction (CFI) module to facilitate cross-feature communications among the three different features, namely the local features, global features, and initial 2D pose features, aggregating them to form enhanced spatial representations of human pose. Furthermore, a novel graph-enhanced module (GraMLP) with parallel GCN and multi-layer perceptron is introduced to inject the skeletal knowledge of the human body into the final representation of 3D pose. Extensive experiments on two datasets (Human3.6M (Ionescu et al., 2013) and MPI-INF-3DHP (Mehta et al., 2017)) show the superior performance of our method in comparison to existing state-of-the-art (SOTA) models. The code and data are shared at https://github.com/JihuaPeng/CFI-3DHPE
引用
收藏
页码:175 / 181
页数:7
相关论文
共 50 条
  • [41] A survey on monocular 3D human pose estimation
    Ji X.
    Fang Q.
    Dong J.
    Shuai Q.
    Jiang W.
    Zhou X.
    Virtual Reality and Intelligent Hardware, 2020, 2 (06): : 471 - 500
  • [42] Precise 3D Pose Estimation of Human Faces
    Pernek, Akos
    Hajder, Levente
    PROCEEDINGS OF THE 2014 9TH INTERNATIONAL CONFERENCE ON COMPUTER VISION, THEORY AND APPLICATIONS (VISAPP 2014), VOL 3, 2014, : 618 - 625
  • [43] A survey on deep 3D human pose estimation
    Neupane, Rama Bastola
    Li, Kan
    Boka, Tesfaye Fenta
    ARTIFICIAL INTELLIGENCE REVIEW, 2024, 58 (01)
  • [44] Deep 3D human pose estimation: A review
    Wang, Jinbao
    Tan, Shujie
    Zhen, Xiantong
    Xu, Shuo
    Zheng, Feng
    He, Zhenyu
    Shao, Ling
    COMPUTER VISION AND IMAGE UNDERSTANDING, 2021, 210
  • [45] Augmented Reality with Human Body Interaction Based on Monocular 3D Pose Estimation
    Lin, Huei-Yung
    Chen, Ting-Wen
    ADVANCED CONCEPTS FOR INTELLIGENT VISION SYSTEMS, PT I, 2010, 6474 : 321 - 331
  • [46] Human 3D Pose Estimation with a Tilting Camera for Social Mobile Robot Interaction
    Garcia-Salguero, Mercedes
    Gonzalez-Jimenez, Javier
    Moreno, Francisco-Angel
    SENSORS, 2019, 19 (22)
  • [47] 3D human pose estimation by depth map
    Wu, Jianzhai
    Hu, Dewen
    Xiang, Fengtao
    Yuan, Xingsheng
    Su, Jiongming
    VISUAL COMPUTER, 2020, 36 (07): : 1401 - 1410
  • [48] View Invariant 3D Human Pose Estimation
    Wei, Guoqiang
    Lan, Cuiling
    Zeng, Wenjun
    Chen, Zhibo
    IEEE TRANSACTIONS ON CIRCUITS AND SYSTEMS FOR VIDEO TECHNOLOGY, 2020, 30 (12) : 4601 - 4610
  • [49] 3D Human Pose Estimation With Adversarial Learning
    Meng, Wenming
    Hu, Tao
    Shuai, Li
    2019 INTERNATIONAL CONFERENCE ON VIRTUAL REALITY AND VISUALIZATION (ICVRV), 2019, : 93 - 99
  • [50] MONOCULAR 3D HUMAN POSE ESTIMATION BY CLASSIFICATION
    Greif, Thomas
    Lienhart, Rainer
    Sengupta, Debabrata
    2011 IEEE INTERNATIONAL CONFERENCE ON MULTIMEDIA AND EXPO (ICME), 2011,