Category-Level Pose Estimation and Iterative Refinement for Monocular RGB-D Image

被引:0
|
作者
Bao, Yongtang [1 ]
Qi, Yutong [2 ]
Su, Chunjian [1 ]
Geng, Yanbing [3 ]
Li, Haojie [1 ]
机构
[1] Shandong Univ Sci & Technol, Coll Comp Sci & Engn, Qingdao, Peoples R China
[2] Univ Toronto, Dept Comp & Math Sci, Scarborough, ON, Canada
[3] North Univ China, Sch Data Sci & Technol, Taiyuan, Peoples R China
基金
中国国家自然科学基金;
关键词
Deep learning; category-level pose estimation; scene understanding; transformer; TRANSFORMER;
D O I
10.1145/3695877
中图分类号
TP [自动化技术、计算机技术];
学科分类号
0812 ;
摘要
Category-level pose estimation is proposed to predict the 6D pose of objects under a specific category and has wide applications in fields such as robotics, virtual reality, and autonomous driving. With the development of VR/AR technology, pose estimation has gradually become a research hotspot in 3D scene understanding. However, most methods fail to fully utilize geometric and color information to solve intra-class shape variations, which leads to inaccurate prediction results. To solve the above problems, we propose a novel pose estimation and iterative refinement network, use an attention mechanism to fuse multi-modal information to obtain color features after a coordinate transformation, and design iterative modules to ensure the accuracy of object geometric features. Specifically, we use an encoder-decoder architecture to implicitly generate a coarse-grained initial pose and refine it through an iterative refinement module. In addition, due to the differences between rotation and position estimation, we design a multi-head pose decoder that utilizes the local geometry and global features. Finally, we design a transformer-based coordinate transformation attention module to extract pose-sensitive features from RGB images and supervise color information by correlating point cloud features in different coordinate systems. We train and test our network on the synthetic dataset CAMERA25 and the real dataset REAL275. Experimental results show that our method achieves state-of-the-art performance on multiple evaluation metrics.
引用
收藏
页数:20
相关论文
共 50 条
  • [41] GeoReF: Geometric Alignment Across Shape Variation for Category-level Object Pose Refinement
    Zheng, Linfang
    Tse, Tze Ho Elden
    Wang, Chen
    Sun, Yinghan
    Chen, Hua
    Leonardis, Ales
    Zhang, Wei
    Chang, Hyung Jin
    2024 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR), 2024, : 10693 - 10703
  • [42] GarmentTracking: Category-Level Garment Pose Tracking
    Xue, Han
    Xu, Wenqiang
    Zhang, Jieyi
    Tang, Tutian
    Li, Yutong
    Du, Wenxin
    Ye, Ruolin
    Lu, Cewu
    2023 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR), 2023, : 21233 - 21242
  • [43] GS-Pose: Category-Level Object Pose Estimation via Geometric and Semantic Correspondence
    Wang, Pengyuan
    Ikeda, Takuya
    Lee, Robert
    Nishiwaki, Koichi
    COMPUTER VISION - ECCV 2024, PT XXVII, 2025, 15085 : 108 - 126
  • [44] HS-Pose: Hybrid Scope Feature Extraction for Category-level Object Pose Estimation
    Zheng, Linfang
    Wang, Chen
    Sun, Yinghan
    Dasgupta, Esha
    Chen, Hua
    Leonardis, Ales
    Zhang, Wei
    Chang, Hyung Jin
    2023 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR), 2023, : 17163 - 17173
  • [45] GenPose: Generative Category-level Object Pose Estimation via Diffusion Models
    Zhang, Jiyao
    Wu, Mingdong
    Dong, Hao
    ADVANCES IN NEURAL INFORMATION PROCESSING SYSTEMS 36 (NEURIPS 2023), 2023,
  • [46] GSNet: Model Reconstruction Network for Category-level 6D Object Pose and Size Estimation
    Liu, Penglei
    Zhang, Qieshi
    Cheng, Jun
    2023 IEEE INTERNATIONAL CONFERENCE ON ROBOTICS AND AUTOMATION, ICRA, 2023, : 2898 - 2904
  • [47] Multiple-Hand 2D Pose Estimation From a Monocular RGB Image
    Mishra, Purnendu
    Sarawadekar, Kishor
    IEEE ACCESS, 2024, 12 : 40722 - 40735
  • [48] RGB-D Hand Pose Estimation Using Fourier Descriptor
    Rong, Zihao
    Kong, Dehui
    Wang, Shaofan
    Yin, Baocai
    2018 7TH INTERNATIONAL CONFERENCE ON DIGITAL HOME (ICDH 2018), 2018, : 50 - 56
  • [49] HUMAN POSE ESTIMATION USING TWO RGB-D SENSORS
    Xu, Wanxin
    Su, Po-chang
    Cheung, Sen-ching S.
    2016 IEEE INTERNATIONAL CONFERENCE ON IMAGE PROCESSING (ICIP), 2016, : 1279 - 1283
  • [50] A HYBRID METRIC FOR CAMERA POSE ESTIMATION IN RGB-D RECONSTRUCTION
    Guo, Fei
    He, Yifeng
    Guan, Ling
    2016 IEEE INTERNATIONAL CONFERENCE ON MULTIMEDIA & EXPO WORKSHOPS (ICMEW), 2016,