EMHIFormer: An Enhanced Multi-Hypothesis Interaction Transformer for 3D human estimation in video✩

被引:2
|
作者
Xiang, Xuezhi [1 ,2 ]
Zhang, Kaixu [1 ]
Qiao, Yulong [1 ,2 ]
El Saddik, Abdulmotaleb [3 ]
机构
[1] Harbin Engn Univ, Sch Informat & Commun Engn, Harbin 150001, Peoples R China
[2] Minist Ind & Informat Technol, Key Lab Adv Marine Commun & Informat Technol, Harbin 150001, Peoples R China
[3] Univ Ottawa, Sch Elect Engn & Comp Sci, Ottawa, ON K1N 6N5, Canada
基金
黑龙江省自然科学基金; 中国国家自然科学基金;
关键词
3D human pose estimation; Transformer; Cross-hypothesis; Enhanced regression head;
D O I
10.1016/j.jvcir.2023.103890
中图分类号
TP [自动化技术、计算机技术];
学科分类号
0812 ;
摘要
Monocular 3D human pose estimation is a challenging task because of depth ambiguity and occlusion. Recent methods exploit spatio-temporal information and generate different hypotheses for simulating diverse solutions to alleviate these problems. However, these methods do not fully extract spatial and temporal information and the relationship of each hypothesis. To ease these limitations, we propose EMHIFormer (Enhanced Multi-Hypothesis Interaction Transformer) to model 3D human pose with better performance. In detail, we build connections between different Transformer layers so that our model is able to integrate spatio-temporal information from the previous layer and establish more comprehensive hypotheses. Furthermore, a cross-hypothesis model consisting of a parallel Transformer is proposed to strengthen the relationship between various hypotheses. We also design an enhanced regression head which adaptively adjusts the channel weights to export the final 3D human pose. Extensive experiments are conducted on two challenging datasets: Human3.6M and MPI-INF-3DHP to evaluate our EMHIFormer. The results show that EMHIFormer achieves competitive performance on Human3.6M and state-of-the-art performance on MPI-INF-3DHP. Compared with the closest counterpart, MHFormer, our model outperforms it by 0.6% P-MPJPE and 0.5% MPJPE on Human3.6M dataset and 46.0% MPJPE on MPI-INF-3DHP.
引用
下载
收藏
页数:11
相关论文
共 50 条
  • [1] MHFormer: Multi-Hypothesis Transformer for 3D Human Pose Estimation
    Li, Wenhao
    Liu, Hong
    Tang, Hao
    Wang, Pichao
    Van Gool, Luc
    2022 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR), 2022, : 13137 - 13146
  • [2] DBMHT: A double-branch multi-hypothesis transformer for 3D human pose estimation in video
    Xiang, Xuezhi
    Li, Xiaoheng
    Bao, Weijie
    Qiaoa, Yulong
    El Saddik, Abdulmotaleb
    COMPUTER VISION AND IMAGE UNDERSTANDING, 2024, 249
  • [3] Multi-hypothesis representation learning for transformer-based 3D human pose estimation
    Li, Wenhao
    Liu, Hong
    Tang, Hao
    Wang, Pichao
    PATTERN RECOGNITION, 2023, 141
  • [4] MHCanonNet: Multi-Hypothesis Canonical lifting Network for 3D human estimation in the wild video
    Kim, Hyun-Woo
    Lee, Gun-Hee
    Nam, Woo-Jeoung
    Jin, Kyung-Min
    Kang, Tae-Kyung
    Yang, Geon-Jun
    Lee, Seong-Whan
    PATTERN RECOGNITION, 2024, 145
  • [5] Diffusion-Based 3D Human Pose Estimation with Multi-Hypothesis Aggregation
    Shan, Wenkang
    Liu, Zhenhua
    Zhang, Xinfeng
    Wang, Zhao
    Han, Kai
    Wang, Shanshe
    Ma, Siwei
    Gao, Wen
    2023 IEEE/CVF INTERNATIONAL CONFERENCE ON COMPUTER VISION (ICCV 2023), 2023, : 14715 - 14725
  • [6] PoSynDA: Multi-Hypothesis Pose Synthesis Domain Adaptation for Robust 3D Human Pose Estimation
    Liu, Hanbing
    He, Jun-Yan
    Cheng, Zhi-Qi
    Xiang, Wangmeng
    Yang, Qize
    Chai, Wenhao
    Wang, Gaoang
    Bao, Xu
    Luo, Bin
    Geng, Yifeng
    Xie, Xuansong
    PROCEEDINGS OF THE 31ST ACM INTERNATIONAL CONFERENCE ON MULTIMEDIA, MM 2023, 2023, : 5542 - 5551
  • [7] 3D Human Pose Estimation in Video with Temporal and Spatial Transformer
    Peng, Sha
    Hu, Jiwei
    Proceedings of SPIE - The International Society for Optical Engineering, 2023, 12707
  • [8] Shape Estimation of a 3D Printed Soft Sensor Using Multi-Hypothesis Extended Kalman Filter
    Tan, Kaige
    Ji, Qinglei
    Feng, Lei
    Torngren, Martin
    IEEE ROBOTICS AND AUTOMATION LETTERS, 2022, 7 (03) : 8383 - 8390
  • [9] LOCAL TO GLOBAL TRANSFORMER FOR VIDEO BASED 3D HUMAN POSE ESTIMATION
    Ma, Haifeng
    Ke Lu
    Xue, Jian
    Niu, Zehai
    Gao, Pengcheng
    2022 IEEE INTERNATIONAL CONFERENCE ON MULTIMEDIA AND EXPO WORKSHOPS (IEEE ICMEW 2022), 2022,
  • [10] A multi-granular joint tracing transformer for video-based 3D human pose estimation
    Yingying Hou
    Zhenhua Huang
    Wentao Zhu
    Signal, Image and Video Processing, 2025, 19 (1)