Mask2Former with Improved Query for Semantic Segmentation in Remote-Sensing Images

被引:1
|
作者
Guo, Shichen [1 ,2 ]
Yang, Qi [2 ,3 ]
Xiang, Shiming [2 ,3 ]
Wang, Shuwen [4 ]
Wang, Xuezhi [1 ]
机构
[1] Chinese Acad Sci, Comp Network Informat Ctr, Beijing 100083, Peoples R China
[2] Univ Chinese Acad Sci, Beijing 100049, Peoples R China
[3] Chinese Acad Sci, Inst Automat, State Key Lab Multimodal Artificial Intelligence S, Beijing 100190, Peoples R China
[4] Portland State Univ, Dept Comp Sci, Portland, OR 97201 USA
关键词
semantic segmentation; remote-sensing image; transformer; Mask2Former; query;
D O I
10.3390/math12050765
中图分类号
O1 [数学];
学科分类号
0701 ; 070101 ;
摘要
Semantic segmentation of remote sensing (RS) images is vital in various practical applications, including urban construction planning, natural disaster monitoring, and land resources investigation. However, RS images are captured by airplanes or satellites at high altitudes and long distances, resulting in ground objects of the same category being scattered in various corners of the image. Moreover, objects of different sizes appear simultaneously in RS images. For example, some objects occupy a large area in urban scenes, while others only have small regions. Technically, the above two universal situations pose significant challenges to the segmentation with a high quality for RS images. Based on these observations, this paper proposes a Mask2Former with an improved query (IQ2Former) for this task. The fundamental motivation behind the IQ2Former is to enhance the capability of the query of Mask2Former by exploiting the characteristics of RS images well. First, we propose the Query Scenario Module (QSM), which aims to learn and group the queries from feature maps, allowing the selection of distinct scenarios such as the urban and rural areas, building clusters, and parking lots. Second, we design the query position module (QPM), which is developed to assign the image position information to each query without increasing the number of parameters, thereby enhancing the model's sensitivity to small targets in complex scenarios. Finally, we propose the query attention module (QAM), which is constructed to leverage the characteristics of query attention to extract valuable features from the preceding queries. Being positioned between the duplicated transformer decoder layers, QAM ensures the comprehensive utilization of the supervisory information and the exploitation of those fine-grained details. Architecturally, the QSM, QPM, and QAM as well as an end-to-end model are assembled to achieve high-quality semantic segmentation. In comparison to the classical or state-of-the-art models (FCN, PSPNet, DeepLabV3+, OCRNet, UPerNet, MaskFormer, Mask2Former), IQ2Former has demonstrated exceptional performance across three publicly challenging remote-sensing image datasets, 83.59 mIoU on the Vaihingen dataset, 87.89 mIoU on Potsdam dataset, and 56.31 mIoU on LoveDA dataset. Additionally, overall accuracy, ablation experiment, and visualization segmentation results all indicate IQ2Former validity.
引用
收藏
页数:24
相关论文
共 50 条
  • [1] SeMask-Mask2Former: A Semantic Segmentation Model for High Resolution Remote Sensing Images
    Qiao, Yicheng
    Liu, Wei
    Liang, Bin
    Wang, Pengyun
    Zhang, Haopeng
    Yang, Junli
    [J]. 2023 IEEE AEROSPACE CONFERENCE, 2023,
  • [2] Edge Detection Guide Network for Semantic Segmentation of Remote-Sensing Images
    Jin, Jianhui
    Zhou, Wujie
    Yang, Rongwang
    Ye, Lv
    Yu, Lu
    [J]. IEEE GEOSCIENCE AND REMOTE SENSING LETTERS, 2023, 20
  • [3] Edge Detection Guide Network for Semantic Segmentation of Remote-Sensing Images
    Jin, Jianhui
    Zhou, Wujie
    Yang, Rongwang
    Ye, Lv
    Yu, Lu
    [J]. IEEE GEOSCIENCE AND REMOTE SENSING LETTERS, 2023, 20
  • [4] Learnable Gated Convolutional Neural Network for Semantic Segmentation in Remote-Sensing Images
    Guo, Shichen
    Jin, Qizhao
    Wang, Hongzhen
    Wang, Xuezhi
    Wang, Yangang
    Xiang, Shiming
    [J]. REMOTE SENSING, 2019, 11 (16)
  • [5] Dynamic High-Resolution Network for Semantic Segmentation in Remote-Sensing Images
    Guo, Shichen
    Yang, Qi
    Xiang, Shiming
    Wang, Pengfei
    Wang, Xuezhi
    [J]. REMOTE SENSING, 2023, 15 (09)
  • [6] Progressive Guidance Edge Perception Network for Semantic Segmentation of Remote-Sensing Images
    Pan, Shaoming
    Tao, Yulong
    Chen, Xiaoshu
    Chong, Yanwen
    [J]. IEEE GEOSCIENCE AND REMOTE SENSING LETTERS, 2022, 19
  • [7] Sea-Land Segmentation of Remote-Sensing Images with Prompt Mask-Attention
    Ji, Yingjie
    Wu, Weiguo
    Nie, Shiqiang
    Wang, Jinyu
    Liu, Song
    [J]. REMOTE SENSING, 2024, 16 (18)
  • [8] SDFCNv2: An Improved FCN Framework for Remote Sensing Images Semantic Segmentation
    Chen, Guanzhou
    Tan, Xiaoliang
    Guo, Beibei
    Zhu, Kun
    Liao, Puyun
    Wang, Tong
    Wang, Qing
    Zhang, Xiaodong
    [J]. REMOTE SENSING, 2021, 13 (23)
  • [9] Semantic Segmentation of Remote-Sensing Images Based on Multiscale Feature Fusion and Attention Refinement
    He, Xin
    Zhou, Yong
    Zhao, Jiaqi
    Zhang, Man
    Yao, Rui
    Liu, Bing
    Li, Haichao
    [J]. IEEE Geoscience and Remote Sensing Letters, 2022, 19
  • [10] Unsupervised Domain Adaptation Semantic Segmentation for Remote-Sensing Images via Covariance Attention
    Liu, Yikun
    Kang, Xudong
    Huang, Yuwen
    Wang, Kuikui
    Yang, Gongping
    [J]. IEEE GEOSCIENCE AND REMOTE SENSING LETTERS, 2022, 19