Rethinking Masked Representation Learning for 3D Point Cloud Understanding

被引:0
|
作者
Wang, Chuxin [1 ,2 ]
Zha, Yixin [1 ,2 ]
He, Jianfeng [1 ,2 ]
Yang, Wenfei [1 ,2 ]
Zhang, Tianzhu [1 ,2 ]
机构
[1] Univ Sci & Technol China, Sch Informat Sci & Technol, Hefei 230027, Peoples R China
[2] Univ Sci & Technol China, Deep Space Explorat Lab, Hefei 230027, Peoples R China
关键词
Point cloud compression; Semantics; Feature extraction; Three-dimensional displays; Representation learning; Solid modeling; Prototypes; Shape; Nearest neighbor methods; Image reconstruction; Self-supervised point cloud representation learning; optimal transport; and part modeling; NETWORK;
D O I
10.1109/TIP.2024.3520008
中图分类号
TP18 [人工智能理论];
学科分类号
081104 ; 0812 ; 0835 ; 1405 ;
摘要
Self-supervised point cloud representation learning aims to acquire robust and general feature representations from unlabeled data. Recently, masked point modeling-based methods have shown significant performance improvements for point cloud understanding, yet these methods rely on overlapping grouping strategies (k-nearest neighbor algorithm) resulting in early leakage of structural information of mask groups, and overlook the semantic modeling of object components resulting in parts with the same semantics having obvious feature differences due to position differences. In this work, we rethink grouping strategies and pretext tasks that are more suitable for self-supervised point cloud representation learning and propose a novel hierarchical masked representation learning method, including an optimal transport-based hierarchical grouping strategy, a prototype-based part modeling module, and a hierarchical attention encoder. The proposed method enjoys several merits. First, the proposed grouping strategy partitions the point cloud into non-overlapping groups, eliminating the early leakage of structural information in the masked groups. Second, the proposed prototype-based part modeling module dynamically models different object components, ensuring feature consistency on parts with the same semantics. Extensive experiments on four downstream tasks demonstrate that our method surpasses state-of-the-art 3D representation learning methods. Furthermore, Comprehensive ablation studies and visualizations demonstrate the effectiveness of the proposed modules.
引用
收藏
页码:247 / 262
页数:16
相关论文
共 50 条
  • [21] Quadratic Terms Based Point-to-Surface 3D Representation for Deep Learning of Point Cloud
    Sun, Tiecheng
    Liu, Guanghui
    Li, Ru
    Liu, Shuaicheng
    Zhu, Shuyuan
    Zeng, Bing
    IEEE TRANSACTIONS ON CIRCUITS AND SYSTEMS FOR VIDEO TECHNOLOGY, 2022, 32 (05) : 2705 - 2718
  • [22] ULIP: Learning a Unified Representation of Language, Images, and Point Clouds for 3D Understanding
    Xue, Le
    Gao, Mingfei
    Xing, Chen
    Martin-Martin, Roberto
    Wu, Jiajun
    Xiong, Caiming
    Xu, Ran
    Niebles, Juan Carlos
    Savarese, Silvio
    2023 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION, CVPR, 2023, : 1179 - 1189
  • [23] Bioinspired point cloud representation: 3D object tracking
    Sergio Orts-Escolano
    Jose Garcia-Rodriguez
    Miguel Cazorla
    Vicente Morell
    Jorge Azorin
    Marcelo Saval
    Alberto Garcia-Garcia
    Victor Villena
    Neural Computing and Applications, 2018, 29 : 663 - 672
  • [24] Bioinspired point cloud representation: 3D object tracking
    Orts-Escolano, Sergio
    Garcia-Rodriguez, Jose
    Cazorla, Miguel
    Morell, Vicente
    Azorin, Jorge
    Saval, Marcelo
    Garcia-Garcia, Alberto
    Villena, Victor
    NEURAL COMPUTING & APPLICATIONS, 2018, 29 (09): : 663 - 672
  • [25] Rethinking Design and Evaluation of 3D Point Cloud Segmentation Models
    Zoumpekas, Thanasis
    Salamo, Maria
    Puig, Anna
    REMOTE SENSING, 2022, 14 (23)
  • [26] Unsupervised 3D Point Cloud Representation Learning by Triangle Constrained Contrast for Autonomous Driving
    Pang, Bo
    Xia, Hongchi
    Lu, Cewu
    2023 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION, CVPR, 2023, : 5229 - 5239
  • [27] Rotation-Invariant Local-to-Global Representation Learning for 3D Point Cloud
    Kim, Seohyun
    Park, Jaeyoo
    Han, Bohyung
    ADVANCES IN NEURAL INFORMATION PROCESSING SYSTEMS 33, NEURIPS 2020, 2020, 33
  • [28] UMA-Net: an unsupervised representation learning network for 3D point cloud classification
    Liu, Jie
    Tian, Yu
    Geng, Guohua
    Wang, Haolin
    Song, Da
    Li, Kang
    Zhou, Mingquan
    Cao, Xin
    JOURNAL OF THE OPTICAL SOCIETY OF AMERICA A-OPTICS IMAGE SCIENCE AND VISION, 2022, 39 (06) : 1085 - 1094
  • [29] Learning Interpretable Representation for 3D Point Clouds
    Su, Feng-Guang
    Lin, Ci-Siang
    Wang, Yu-Chiang Frank
    2020 25TH INTERNATIONAL CONFERENCE ON PATTERN RECOGNITION (ICPR), 2021, : 7470 - 7477
  • [30] Learning from 3D (Point Cloud) Data
    Hsu, Winston H.
    PROCEEDINGS OF THE 27TH ACM INTERNATIONAL CONFERENCE ON MULTIMEDIA (MM'19), 2019, : 2697 - 2698