DBMHT: A double-branch multi-hypothesis transformer for 3D human pose estimation in video

被引：0

作者：

Xiang, Xuezhi ^{[1
,2
]}

Li, Xiaoheng ^{[1
]}

Bao, Weijie ^{[1
]}

Qiaoa, Yulong ^{[1
,3
]}

El Saddik, Abdulmotaleb ^{[3
]}

机构：

[1] Harbin Engn Univ, Sch Informat & Commun Engn, Harbin 150001, Peoples R China

[2] Minist Ind & Informat Technol, Key Lab Adv Marine Commun & Informat Technol, Harbin 150001, Peoples R China

[3] Univ Ottawa, Sch Elect Engn & Comp Sci, Ottawa, ON K1N 6N5, Canada

来源：

COMPUTER VISION AND IMAGE UNDERSTANDING | 2024年 / 249卷

基金：

黑龙江省自然科学基金; 中国国家自然科学基金;

关键词：

3D human pose estimation; Transformer; Dual-branch; Cross-hypothesis;

D O I：

10.1016/j.cviu.2024.104147

中图分类号：

TP18 [人工智能理论];

学科分类号：

081104 ; 0812 ; 0835 ; 1405 ;

摘要：

The estimation of 3D human poses from monocular videos presents a significant challenge. The existing methods face the problems of deep ambiguity and self-occlusion. To overcome these problems, we propose a Double-Branch Multi-Hypothesis Transformer (DBMHT). In detail, we utilize a Double-Branch architecture to capture temporal and spatial information and generate multiple hypotheses. To merge these hypotheses, we adopt a lightweight module to integrate spatial and temporal representations. The DBMHT can not only capture spatial information from each joint in the human body and temporal information from each frame in the video but also merge multiple hypotheses that have different spatio-temporal information. Comprehensive evaluation on two challenging datasets (i.e. Human3.6M and MPI-INF-3DHP) demonstrates the superior performance of DBMHT, marking it as a robust and efficient approach for accurate 3D HPE in dynamic scenarios. The results show that our model surpasses the state-of-the-art approach by 1.9% MPJPE with ground truth 2D keypoints as input.

引用

页数：8

共 50 条

[1] MHFormer: Multi-Hypothesis Transformer for 3D Human Pose Estimation
Li, Wenhao
Liu, Hong
Tang, Hao
Wang, Pichao
Van Gool, Luc
2022 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR), 2022, : 13137 - 13146
[2] EMHIFormer: An Enhanced Multi-Hypothesis Interaction Transformer for 3D human estimation in video✩
Xiang, Xuezhi
Zhang, Kaixu
Qiao, Yulong
El Saddik, Abdulmotaleb
JOURNAL OF VISUAL COMMUNICATION AND IMAGE REPRESENTATION, 2023, 95
[3] Multi-hypothesis representation learning for transformer-based 3D human pose estimation
Li, Wenhao
Liu, Hong
Tang, Hao
Wang, Pichao
PATTERN RECOGNITION, 2023, 141
[4] Diffusion-Based 3D Human Pose Estimation with Multi-Hypothesis Aggregation
Shan, Wenkang
Liu, Zhenhua
Zhang, Xinfeng
Wang, Zhao
Han, Kai
Wang, Shanshe
Ma, Siwei
Gao, Wen
2023 IEEE/CVF INTERNATIONAL CONFERENCE ON COMPUTER VISION (ICCV 2023), 2023, : 14715 - 14725
[5] PoSynDA: Multi-Hypothesis Pose Synthesis Domain Adaptation for Robust 3D Human Pose Estimation
Liu, Hanbing
He, Jun-Yan
Cheng, Zhi-Qi
Xiang, Wangmeng
Yang, Qize
Chai, Wenhao
Wang, Gaoang
Bao, Xu
Luo, Bin
Geng, Yifeng
Xie, Xuansong
PROCEEDINGS OF THE 31ST ACM INTERNATIONAL CONFERENCE ON MULTIMEDIA, MM 2023, 2023, : 5542 - 5551
[6] MHCanonNet: Multi-Hypothesis Canonical lifting Network for 3D human estimation in the wild video
Kim, Hyun-Woo
Lee, Gun-Hee
Nam, Woo-Jeoung
Jin, Kyung-Min
Kang, Tae-Kyung
Yang, Geon-Jun
Lee, Seong-Whan
PATTERN RECOGNITION, 2024, 145
[7] 3D Human Pose Estimation in Video with Temporal and Spatial Transformer
Peng, Sha
Hu, Jiwei
Proceedings of SPIE - The International Society for Optical Engineering, 2023, 12707
[8] LOCAL TO GLOBAL TRANSFORMER FOR VIDEO BASED 3D HUMAN POSE ESTIMATION
Ma, Haifeng
Ke Lu
Xue, Jian
Niu, Zehai
Gao, Pengcheng
2022 IEEE INTERNATIONAL CONFERENCE ON MULTIMEDIA AND EXPO WORKSHOPS (IEEE ICMEW 2022), 2022,
[9] Unsupervised 3D Human Pose Estimation in Multi-view-multi-pose Video
Sun, Cheng
Thomas, Diego
Kawasaki, Hiroshi
2020 25TH INTERNATIONAL CONFERENCE ON PATTERN RECOGNITION (ICPR), 2021, : 5959 - 5964
[10] 3D human pose estimation with multi-hypotheses gated transformer
Dong, Xiena
Zhang, Jian
Yu, Jun
Yu, Ting
MULTIMEDIA SYSTEMS, 2024, 30 (06)

← 1 2 3 4 5 →