Rethinking Masked Representation Learning for 3D Point Cloud Understanding

被引:0
|
作者
Wang, Chuxin [1 ,2 ]
Zha, Yixin [1 ,2 ]
He, Jianfeng [1 ,2 ]
Yang, Wenfei [1 ,2 ]
Zhang, Tianzhu [1 ,2 ]
机构
[1] Univ Sci & Technol China, Sch Informat Sci & Technol, Hefei 230027, Peoples R China
[2] Univ Sci & Technol China, Deep Space Explorat Lab, Hefei 230027, Peoples R China
关键词
Point cloud compression; Semantics; Feature extraction; Three-dimensional displays; Representation learning; Solid modeling; Prototypes; Shape; Nearest neighbor methods; Image reconstruction; Self-supervised point cloud representation learning; optimal transport; and part modeling; NETWORK;
D O I
10.1109/TIP.2024.3520008
中图分类号
TP18 [人工智能理论];
学科分类号
081104 ; 0812 ; 0835 ; 1405 ;
摘要
Self-supervised point cloud representation learning aims to acquire robust and general feature representations from unlabeled data. Recently, masked point modeling-based methods have shown significant performance improvements for point cloud understanding, yet these methods rely on overlapping grouping strategies (k-nearest neighbor algorithm) resulting in early leakage of structural information of mask groups, and overlook the semantic modeling of object components resulting in parts with the same semantics having obvious feature differences due to position differences. In this work, we rethink grouping strategies and pretext tasks that are more suitable for self-supervised point cloud representation learning and propose a novel hierarchical masked representation learning method, including an optimal transport-based hierarchical grouping strategy, a prototype-based part modeling module, and a hierarchical attention encoder. The proposed method enjoys several merits. First, the proposed grouping strategy partitions the point cloud into non-overlapping groups, eliminating the early leakage of structural information in the masked groups. Second, the proposed prototype-based part modeling module dynamically models different object components, ensuring feature consistency on parts with the same semantics. Extensive experiments on four downstream tasks demonstrate that our method surpasses state-of-the-art 3D representation learning methods. Furthermore, Comprehensive ablation studies and visualizations demonstrate the effectiveness of the proposed modules.
引用
收藏
页码:247 / 262
页数:16
相关论文
共 50 条
  • [41] Masked Scene Contrast: A Scalable Framework for Unsupervised 3D Representation Learning
    Wu, Xiaoyang
    Wen, Xin
    Liu, Xihui
    Zhao, Hengshuang
    2023 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR), 2023, : 9415 - 9424
  • [42] Learning 3D Face Representation with Vision Transformer for Masked Face Recognition
    Wang, Yuan
    Yang, Zhen
    Zhang, Zhiqiang
    Zang, Huaijuan
    Zhu, Qiang
    Zhan, Shu
    2022 ASIA CONFERENCE ON ALGORITHMS, COMPUTING AND MACHINE LEARNING (CACML 2022), 2022, : 505 - 511
  • [43] T-MAE : Temporal Masked Autoencoders for Point Cloud Representation Learning
    Wei, Weijie
    Nejadasl, Fatemeh Karimi
    Gevers, Theo
    Oswald, Martin R.
    COMPUTER VISION - ECCV 2024, PT XI, 2025, 15069 : 178 - 195
  • [44] Point Cloud Domain Adaptation via Masked Local 3D Structure Prediction
    Liang, Hanxue
    Fan, Hehe
    Fan, Zhiwen
    Wang, Yi
    Chen, Tianlong
    Cheng, Yu
    Wang, Zhangyang
    COMPUTER VISION - ECCV 2022, PT III, 2022, 13663 : 156 - 172
  • [45] PiMAE: Point Cloud and Image Interactive Masked Autoencoders for 3D Object Detection
    Chen, Anthony
    Zhang, Kevin
    Zhang, Renrui
    Wang, Zihan
    Lu, Yuheng
    Guo, Yandong
    Zhang, Shanghang
    2023 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION, CVPR, 2023, : 5291 - 5301
  • [46] Masked Autoencoder for Pre-Training on 3D Point Cloud Object Detection
    Xie, Guangda
    Li, Yang
    Qu, Hongquan
    Sun, Zaiming
    MATHEMATICS, 2022, 10 (19)
  • [47] DCCN: A dual-cross contrastive neural network for 3D point cloud representation learning
    Wu, Xiaopeng
    Shi, Guangsi
    Zhao, Zexing
    Li, Mingjie
    Gao, Xiaojun
    Yan, Xiaoli
    EXPERT SYSTEMS WITH APPLICATIONS, 2024, 249
  • [48] Progressive Framework of Learning 3D Object Classes and Orientations from Deep Point Cloud Representation
    Lee, Sukhan
    Cheng, Wencan
    PROCEEDINGS OF THE 2020 14TH INTERNATIONAL CONFERENCE ON UBIQUITOUS INFORMATION MANAGEMENT AND COMMUNICATION (IMCOM), 2020,
  • [49] Point Cloud Representation of 3D Shape for Laser-Plasma Scanning 3D Display
    Ishikawa, Hiroyo
    Saito, Hideo
    IECON 2008: 34TH ANNUAL CONFERENCE OF THE IEEE INDUSTRIAL ELECTRONICS SOCIETY, VOLS 1-5, PROCEEDINGS, 2008, : 1849 - 1854
  • [50] LinK3D: Linear Keypoints Representation for 3D LiDAR Point Cloud
    Cui, Yunge
    Zhang, Yinlong
    Dong, Jiahua
    Sun, Haibo
    Chen, Xieyuanli
    Zhu, Feng
    IEEE ROBOTICS AND AUTOMATION LETTERS, 2024, 9 (03) : 2128 - 2135