Improving RGB-D Salient Object Detection via Modality-Aware Decoder

被引:0
|
作者
Song, Mengke [1 ,2 ]
Song, Wenfeng [3 ]
Yang, Guowei [4 ]
Chen, Chenglizhao [1 ,2 ]
机构
[1] China Univ Petr East China, Coll Comp Sci & Technol, Qingdao 266580, Peoples R China
[2] China Univ Petr East China, Qingdao Inst Software, Qingdao 266580, Peoples R China
[3] Beijing Informat Sci & Technol Univ, Comp Sch, Beijing 100192, Peoples R China
[4] Qingdao Univ, Sch Elect Informat, Qingdao 266071, Peoples R China
基金
中国国家自然科学基金;
关键词
Decoding; Object detection; Training; Task analysis; Saliency detection; Image segmentation; Feature extraction; RGB-D salient object detection; modality-aware fusion; deep learning; GRAPH CONVOLUTION NETWORK; IMAGE; ATTENTION; FUSION;
D O I
10.1109/TIP.2022.3205747
中图分类号
TP18 [人工智能理论];
学科分类号
081104 ; 0812 ; 0835 ; 1405 ;
摘要
Most existing RGB-D salient object detection (SOD) methods are primarily focusing on cross-modal and cross-level saliency fusion, which has been proved to be efficient and effective. However, these methods still have a critical limitation, i.e., their fusion patterns - typically the combination of selective characteristics and its variations, are too highly dependent on the network's non-linear adaptability. In such methods, the balances between RGB and D (Depth) are formulated individually considering the intermediate feature slices, but the relation at the modality level may not be learned properly. The optimal RGB-D combinations differ depending on the RGB-D scenarios, and the exact complementary status is frequently determined by multiple modality-level factors, such as D quality, the complexity of the RGB scene, and degree of harmony between them. Therefore, given the existing approaches, it may be difficult for them to achieve further performance breakthroughs, as their methodologies belong to some methods that are somewhat less modality sensitive. To conquer this problem, this paper presents the Modality-aware Decoder (MaD). The critical technical innovations include a series of feature embedding, modality reasoning, and feature back-projecting and collecting strategies, all of which upgrade the widely-used multi-scale and multi-level decoding process to be modality-aware. Our MaD achieves competitive performance over other state-of-the-art (SOTA) models without using any fancy tricks in the decoder's design. Codes and results will be publicly available at https://github.com/MengkeSong/MaD.
引用
收藏
页码:6124 / 6138
页数:15
相关论文
共 50 条
  • [21] Depth-aware inverted refinement network for RGB-D salient object detection
    Gao, Lina
    Liu, Bing
    Fu, Ping
    Xu, Mingzhu
    [J]. NEUROCOMPUTING, 2023, 518 : 507 - 522
  • [22] HiDAnet: RGB-D Salient Object Detection via Hierarchical Depth Awareness
    Wu, Zongwei
    Allibert, Guillaume
    Meriaudeau, Fabrice
    Ma, Chao
    Demonceaux, Cedric
    [J]. IEEE TRANSACTIONS ON IMAGE PROCESSING, 2023, 32 : 2160 - 2173
  • [23] RGB-D salient object detection via deep fusion of semantics and details
    Zhao, Shimin
    Chen, Miaomiao
    Wang, Pengjie
    Cao, Ying
    Zhang, Pingping
    Yang, Xin
    [J]. COMPUTER ANIMATION AND VIRTUAL WORLDS, 2020, 31 (4-5)
  • [24] MULTI-MODALITY DIVERSITY FUSION NETWORK WITH SWINTRANSFORMER FOR RGB-D SALIENT OBJECT DETECTION
    Duan, Songsong
    Xia, Chenxing
    Gao, Xiuju
    Ge, Bin
    Zhang, Hanling
    Li, Kuan-Ching
    [J]. 2022 IEEE INTERNATIONAL CONFERENCE ON IMAGE PROCESSING, ICIP, 2022, : 1076 - 1080
  • [25] Multi-modality information refinement fusion network for RGB-D salient object detection
    Bao, Hua
    Fan, Bo
    [J]. VISUAL COMPUTER, 2024, 40 (06): : 4183 - 4199
  • [26] Double cross-modality progressively guided network for RGB-D salient object detection
    Yao, Cuili
    Feng, Lin
    Kong, Yuqiu
    Li, Shengming
    Li, Hang
    [J]. IMAGE AND VISION COMPUTING, 2022, 117
  • [27] Object Discovery on RGB-D Data via Salient Object Proposals
    Li, Wanyi
    Wang, Peng
    Qiao, Hong
    Fan, Naiji
    Zhou, Hai
    Jing, Feng
    [J]. 2015 CHINESE AUTOMATION CONGRESS (CAC), 2015, : 737 - 739
  • [28] AirSOD: A Lightweight Network for RGB-D Salient Object Detection
    Zeng, Zhihong
    Liu, Haijun
    Chen, Fenglei
    Tan, Xiaoheng
    [J]. IEEE TRANSACTIONS ON CIRCUITS AND SYSTEMS FOR VIDEO TECHNOLOGY, 2024, 34 (03) : 1656 - 1669
  • [29] Three-Stream Attention-Aware Network for RGB-D Salient Object Detection
    Chen, Hao
    Li, Youfu
    [J]. IEEE TRANSACTIONS ON IMAGE PROCESSING, 2019, 28 (06) : 2825 - 2835
  • [30] Aggregate interactive learning for RGB-D salient object detection
    Wu, Jingyu
    Sun, Fuming
    Xu, Rui
    Meng, Jie
    Wang, Fasheng
    [J]. EXPERT SYSTEMS WITH APPLICATIONS, 2022, 195