Rethinking Masked Representation Learning for 3D Point Cloud Understanding

被引:0
|
作者
Wang, Chuxin [1 ,2 ]
Zha, Yixin [1 ,2 ]
He, Jianfeng [1 ,2 ]
Yang, Wenfei [1 ,2 ]
Zhang, Tianzhu [1 ,2 ]
机构
[1] Univ Sci & Technol China, Sch Informat Sci & Technol, Hefei 230027, Peoples R China
[2] Univ Sci & Technol China, Deep Space Explorat Lab, Hefei 230027, Peoples R China
关键词
Point cloud compression; Semantics; Feature extraction; Three-dimensional displays; Representation learning; Solid modeling; Prototypes; Shape; Nearest neighbor methods; Image reconstruction; Self-supervised point cloud representation learning; optimal transport; and part modeling; NETWORK;
D O I
10.1109/TIP.2024.3520008
中图分类号
TP18 [人工智能理论];
学科分类号
081104 ; 0812 ; 0835 ; 1405 ;
摘要
Self-supervised point cloud representation learning aims to acquire robust and general feature representations from unlabeled data. Recently, masked point modeling-based methods have shown significant performance improvements for point cloud understanding, yet these methods rely on overlapping grouping strategies (k-nearest neighbor algorithm) resulting in early leakage of structural information of mask groups, and overlook the semantic modeling of object components resulting in parts with the same semantics having obvious feature differences due to position differences. In this work, we rethink grouping strategies and pretext tasks that are more suitable for self-supervised point cloud representation learning and propose a novel hierarchical masked representation learning method, including an optimal transport-based hierarchical grouping strategy, a prototype-based part modeling module, and a hierarchical attention encoder. The proposed method enjoys several merits. First, the proposed grouping strategy partitions the point cloud into non-overlapping groups, eliminating the early leakage of structural information in the masked groups. Second, the proposed prototype-based part modeling module dynamically models different object components, ensuring feature consistency on parts with the same semantics. Extensive experiments on four downstream tasks demonstrate that our method surpasses state-of-the-art 3D representation learning methods. Furthermore, Comprehensive ablation studies and visualizations demonstrate the effectiveness of the proposed modules.
引用
收藏
页码:247 / 262
页数:16
相关论文
共 50 条
  • [31] Learning multiview 3D point cloud registration
    Gojcic, Zan
    Zhou, Caifa
    Wegner, Jan D.
    Guibas, Leonidas J.
    Birdal, Tolga
    2020 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR), 2020, : 1756 - 1766
  • [32] Interpolated Convolutional Networks for 3D Point Cloud Understanding
    Mao, Jiageng
    Wang, Xiaogang
    Li, Hongsheng
    2019 IEEE/CVF INTERNATIONAL CONFERENCE ON COMPUTER VISION (ICCV 2019), 2019, : 1578 - 1587
  • [33] Advancing 3D point cloud understanding through deep transfer learning: A comprehensive survey
    Sohail, Shahab Saquib
    Himeur, Yassine
    Kheddar, Hamza
    Amira, Abbes
    Fadli, Fodil
    Atalla, Shadi
    Copiaco, Igail
    INFORMATION FUSION, 2025, 113
  • [34] Deep Learning Approach to Point Cloud Scene Understanding for Automated Scan to 3D Reconstruction
    Chen, Jingdao
    Kira, Zsolt
    Cho, Yong K.
    JOURNAL OF COMPUTING IN CIVIL ENGINEERING, 2019, 33 (04)
  • [35] Open-Set Semi-Supervised Learning for 3D Point Cloud Understanding
    Shi, Xian
    Xu, Xun
    Zhang, Wanyue
    Zhu, Xiatian
    Foo, Chuan Sheng
    Jia, Kui
    2022 26TH INTERNATIONAL CONFERENCE ON PATTERN RECOGNITION (ICPR), 2022, : 5045 - 5051
  • [36] Learning Progressive Point Embeddings for 3D Point Cloud Generation
    Wen, Cheng
    Yu, Baosheng
    Tao, Dacheng
    2021 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION, CVPR 2021, 2021, : 10261 - 10270
  • [37] Rethinking Few-shot 3D Point Cloud Semantic Segmentation
    An, Zhaochong
    Sun, Guolei
    Liu, Yun
    Liu, Fayao
    Wu, Zongwei
    Wang, Dan
    Van Gool, Luc
    Belongie, Serge
    2024 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION, CVPR 2024, 2024, : 3996 - 4006
  • [38] Interpreting Representation Quality of DNNs for 3D Point Cloud Processing
    Shen, Wen
    Ren, Qihan
    Liu, Dongrui
    Zhang, Quanshi
    ADVANCES IN NEURAL INFORMATION PROCESSING SYSTEMS 34 (NEURIPS 2021), 2021, 34
  • [39] 3D Point Cloud Registration based on the Vector Field Representation
    Van Tung Nguyen
    Trung-Thien Tran
    Van-Toan Cao
    Laurendeau, Denis
    2013 SECOND IAPR ASIAN CONFERENCE ON PATTERN RECOGNITION (ACPR 2013), 2013, : 491 - 495
  • [40] Point-BERT: Pre-training 3D Point Cloud Transformers with Masked Point Modeling
    Yu, Xumin
    Tang, Lulu
    Rao, Yongming
    Huang, Tiejun
    Zhou, Jie
    Lu, Jiwen
    2022 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR 2022), 2022, : 19291 - 19300