Unified Information Fusion Network for Multi-Modal RGB-D and RGB-T Salient Object Detection

被引：116

作者：

Gao, Wei ^{[1
,2
]}

Liao, Guibiao ^{[1
,2
]}

Ma, Siwei ^{[3
]}

Li, Ge ^{[1
,2
]}

Liang, Yongsheng ^{[4
]}

Lin, Weisi ^{[5
]}

机构：

[1] Peking Univ, Sch Elect & Comp Engn, Shenzhen Grad Sch, Shenzhen 518055, Peoples R China

[2] Peng Cheng Lab, Shenzhen 518066, Peoples R China

[3] Peking Univ, Inst Digital Media, Beijing 100871, Peoples R China

[4] Harbin Inst Technol, Sch Elect & Informat Engn, Shenzhen 518055, Peoples R China

[5] Nanyang Technol Univ, Sch Comp Sci & Engn, Singapore 639798, Singapore

来源：

IEEE TRANSACTIONS ON CIRCUITS AND SYSTEMS FOR VIDEO TECHNOLOGY | 2022年 / 32卷 / 04期

关键词：

Dynamic cross-modal guided mechanism; RGB-D/RGB-T multi-modal data; information fusion; salient object detection; VISUAL-ATTENTION; COLOR-VISION; IMAGE; SEGMENTATION; MECHANISMS; MODEL;

D O I：

10.1109/TCSVT.2021.3082939

中图分类号：

TM [电工技术]; TN [电子技术、通信技术];

学科分类号：

0808 ; 0809 ;

摘要：

The use of complementary information, namely depth or thermal information, has shown its benefits to salient object detection (SOD) during recent years. However, the RGB-D or RGB-T SOD problems are currently only solved independently, and most of them directly extract and fuse raw features from backbones. Such methods can he easily restricted by low-quality modality data and redundant cross-modal features. In this work, a unified end-to-end framework is designed to simultaneously analyze RCB-D and RGB-T SOD tasks. Specifically, to effectively tackle multi-modal features, we propose a novel multi-stage and multi-scale fusion network (MMNet), which consists of a cross-modal multi-stage fusion module (CMFM) and a bi-directional multi-scale decoder (BMD). Similar to the visual color stage doctrine in the human visual system (HVS), the proposed CMFM aims to explore important feature representations in feature response stage, and integrate them into cross-modal features in adversarial combination stage. Moreover, the proposed BMD learns the combination of multilevel cross-modal fused features to capture both local and global information of salient objects, and can further boost the multimodal SOD performance. The proposed unified cross-modality feature analysis framework based on two-stage and multi-scale information fusion can be used for diverse multi-modal SOD tasks. Comprehensive experiments (similar to 92K image-pairs) demonstrate that the proposed method consistently outperforms the other 21 state-of-the-art methods on nine benchmark datasets. This validates that our proposed method can work well on diverse multi-modal SOD tasks with good generalization and robustness, and provides a good multi-modal SOD benchmark.

引用

页码：2091 / 2106

页数：16

共 50 条

[21] FEATURE ENHANCEMENT AND FUSION FOR RGB-T SALIENT OBJECT DETECTION
Sun, Fengming
Zhang, Kang
Yuan, Xia
Zhao, Chunxia
2023 IEEE INTERNATIONAL CONFERENCE ON IMAGE PROCESSING, ICIP, 2023, : 1300 - 1304
[22] Unsupervised RGB-T object tracking with attentional multi-modal feature fusion
Shenglan Li
Rui Yao
Yong Zhou
Hancheng Zhu
Bing Liu
Jiaqi Zhao
Zhiwen Shao
Multimedia Tools and Applications, 2023, 82 : 23595 - 23613
[23] Revisiting Feature Fusion for RGB-T Salient Object Detection
Zhang, Qiang
Xiao, Tonglin
Huang, Nianchang
Zhang, Dingwen
Han, Jungong
IEEE TRANSACTIONS ON CIRCUITS AND SYSTEMS FOR VIDEO TECHNOLOGY, 2021, 31 (05) : 1804 - 1818
[24] Cross-Modal Fusion and Progressive Decoding Network for RGB-D Salient Object Detection
Hu, Xihang
Sun, Fuming
Sun, Jing
Wang, Fasheng
Li, Haojie
INTERNATIONAL JOURNAL OF COMPUTER VISION, 2024, 132 (08) : 3067 - 3085
[25] Progressive multi-scale fusion network for RGB-D salient object detection
Ren, Guangyu
Xie, Yanchun
Dai, Tianhong
Stathaki, Tania
COMPUTER VISION AND IMAGE UNDERSTANDING, 2022, 223
[26] RGB-D salient object detection with asymmetric cross-modal fusion
Yu M.
Xing Z.-H.
Liu Y.
Kongzhi yu Juece/Control and Decision, 2023, 38 (09): : 2487 - 2495
[27] An adaptive guidance fusion network for RGB-D salient object detection
Sun, Haodong
Wang, Yu
Ma, Xinpeng
SIGNAL IMAGE AND VIDEO PROCESSING, 2024, 18 (02) : 1683 - 1693
[28] Unsupervised RGB-T object tracking with attentional multi-modal feature fusion
Li, Shenglan
Yao, Rui
Zhou, Yong
Zhu, Hancheng
Liu, Bing
Zhao, Jiaqi
Shao, Zhiwen
MULTIMEDIA TOOLS AND APPLICATIONS, 2023, 82 (15) : 23595 - 23613
[29] Scale Adaptive Fusion Network for RGB-D Salient Object Detection
Kong, Yuqiu
Zheng, Yushuo
Yao, Cuili
Liu, Yang
Wang, He
COMPUTER VISION - ACCV 2022, PT III, 2023, 13843 : 608 - 625
[30] An adaptive guidance fusion network for RGB-D salient object detection
Haodong Sun
Yu Wang
Xinpeng Ma
Signal, Image and Video Processing, 2024, 18 : 1683 - 1693

← 1 2 3 4 5 →