TRANSFORMER-BASED MULTI-MODAL LEARNING FOR MULTI-LABEL REMOTE SENSING IMAGE CLASSIFICATION

被引:0
|
作者
Hoffmann, David Sebastian [1 ]
Clasen, Kai Norman [1 ]
Demir, Begum [1 ,2 ]
机构
[1] Tech Univ Berlin, Fac Elect Engn & Comp Sci, Berlin, Germany
[2] BIFOLD Berlin Inst Fdn Learning & Data, Berlin, Germany
基金
欧洲研究理事会;
关键词
Multi-modal fusion; multi-label image classification; deep learning; transformer; remote sensing;
D O I
10.1109/IGARSS52108.2023.10281927
中图分类号
P [天文学、地球科学];
学科分类号
07 ;
摘要
In this paper, we introduce a novel Synchronized Class Token Fusion (SCT Fusion) architecture in the framework of multi-modal multi-label classification (MLC) of remote sensing (RS) images. The proposed architecture leverages modality-specific attention-based transformer encoders to process varying input modalities, while exchanging information across modalities by synchronizing the special class tokens after each transformer encoder block. The synchronization involves fusing the class tokens with a trainable fusion transformation, resulting in a synchronized class token that contains information from all modalities. As the fusion transformation is trainable, it allows to reach an accurate representation of the shared features among different modalities. Experimental results show the effectiveness of the proposed architecture over single-modality architectures and an early fusion multi-modal architecture when evaluated on a multi-modal MLC dataset. The code of the proposed architecture is publicly available at https://git.tu- berlin.de/rsim/sct- fusion.
引用
收藏
页码:4891 / 4894
页数:4
相关论文
共 50 条
  • [41] Multi-modal, Multi-task and Multi-label for Music Genre Classification and Emotion Regression
    Pandeya, Yagya Raj
    You, Jie
    Bhattarai, Bhuwan
    Lee, Joonwhoan
    [J]. 12TH INTERNATIONAL CONFERENCE ON ICT CONVERGENCE (ICTC 2021): BEYOND THE PANDEMIC ERA WITH ICT CONVERGENCE INNOVATION, 2021, : 1042 - 1045
  • [42] Multi-modality multi-label ocular abnormalities detection with transformer-based semantic dictionary learning
    Siswadi, Anneke Annassia Putri
    Bricq, Stephanie
    Meriaudeau, Fabrice
    [J]. MEDICAL & BIOLOGICAL ENGINEERING & COMPUTING, 2024, 62 (11) : 3433 - 3444
  • [43] Multi-modal Multi-label Emotion Detection with Modality and Label Dependence
    Dong Zhang
    Ju, Xincheng
    Li, Junhui
    Li, Shoushan
    Zhu, Qiaoming
    Zhou, Guodong
    [J]. PROCEEDINGS OF THE 2020 CONFERENCE ON EMPIRICAL METHODS IN NATURAL LANGUAGE PROCESSING (EMNLP), 2020, : 3584 - 3593
  • [44] Pyramidal Cross-Modal Transformer with Sustained Visual Guidance for Multi-Label Image Classification
    Li, Zhuohua
    Wang, Ruyun
    Zhu, Fuqing
    Han, Jizhong
    Hu, Songlin
    [J]. PROCEEDINGS OF THE 4TH ANNUAL ACM INTERNATIONAL CONFERENCE ON MULTIMEDIA RETRIEVAL, ICMR 2024, 2024, : 740 - 748
  • [45] Deep Feature Correlation Learning for Multi-Modal Remote Sensing Image Registration
    Quan, Dou
    Wang, Shuang
    Gu, Yu
    Lei, Ruiqi
    Yang, Bowu
    Wei, Shaowei
    Hou, Biao
    Jiao, Licheng
    [J]. IEEE TRANSACTIONS ON GEOSCIENCE AND REMOTE SENSING, 2022, 60
  • [46] Graph Attention Transformer Network for Multi-label Image Classification
    Yuan, Jin
    Chen, Shikai
    Zhang, Yao
    Shi, Zhongchao
    Geng, Xin
    Fan, Jianping
    Rui, Yong
    [J]. ACM TRANSACTIONS ON MULTIMEDIA COMPUTING COMMUNICATIONS AND APPLICATIONS, 2023, 19 (04)
  • [47] MMDL-Net: Multi-Band Multi-Label Remote Sensing Image Classification Model
    Cheng, Xiaohui
    Li, Bingwu
    Deng, Yun
    Tang, Jian
    Shi, Yuanyuan
    Zhao, Junyu
    [J]. APPLIED SCIENCES-BASEL, 2024, 14 (06):
  • [48] Open-Vocabulary Multi-Label Classification via Multi-Modal Knowledge Transfer
    He, Sunan
    Guo, Taian
    Dai, Tao
    Qiao, Ruizhi
    Shu, Xiujun
    Ren, Bo
    Xia, Shu-Tao
    [J]. THIRTY-SEVENTH AAAI CONFERENCE ON ARTIFICIAL INTELLIGENCE, VOL 37 NO 1, 2023, : 808 - 816
  • [49] Two-Stream Transformer for Multi-Label Image Classification
    Zhu, Xuelin
    Cao, Jiuxin
    Ge, Jiawei
    Liu, Weijia
    Liu, Bo
    [J]. PROCEEDINGS OF THE 30TH ACM INTERNATIONAL CONFERENCE ON MULTIMEDIA, MM 2022, 2022, : 3598 - 3607
  • [50] DATran: Dual Attention Transformer for Multi-Label Image Classification
    Zhou, Wei
    Zheng, Zhijie
    Su, Tao
    Hu, Haifeng
    [J]. IEEE TRANSACTIONS ON CIRCUITS AND SYSTEMS FOR VIDEO TECHNOLOGY, 2024, 34 (01) : 342 - 356