EFCMF: A Multimodal Robustness Enhancement Framework for Fine-Grained Recognition

被引:0
|
作者
Zou, Rongping [1 ,2 ,3 ]
Zhu, Bin [1 ,2 ,3 ]
Chen, Yi [1 ,2 ,3 ]
Xie, Bo [1 ,2 ,3 ]
Shao, Bin [1 ,2 ,3 ]
机构
[1] Natl Univ Def Technol, Coll Elect Engn, Hefei 230037, Peoples R China
[2] State Key Lab Pulsed Power Laser Technol, Hefei 230037, Peoples R China
[3] Key Lab Infrared & Low Temp Plasma Anhui Prov, Hefei 230037, Peoples R China
来源
APPLIED SCIENCES-BASEL | 2023年 / 13卷 / 03期
基金
美国国家科学基金会;
关键词
fine-grained recognition; multimodal; modal missing; adversarial examples;
D O I
10.3390/app13031640
中图分类号
O6 [化学];
学科分类号
0703 ;
摘要
Fine-grained recognition has many applications in many fields and aims to identify targets from subcategories. This is a highly challenging task due to the minor differences between subcategories. Both modal missing and adversarial sample attacks are easily encountered in fine-grained recognition tasks based on multimodal data. These situations can easily lead to the model needing to be fixed. An Enhanced Framework for the Complementarity of Multimodal Features (EFCMF) is proposed in this study to solve this problem. The model's learning of multimodal data complementarity is enhanced by randomly deactivating modal features in the constructed multimodal fine-grained recognition model. The results show that the model gains the ability to handle modal missing without additional training of the model and can achieve 91.14% and 99.31% accuracy on Birds and Flowers datasets. The average accuracy of EFCMF on the two datasets is 52.85%, which is 27.13% higher than that of Bi-modal PMA when facing four adversarial example attacks, namely FGSM, BIM, PGD and C&W. In the face of missing modal cases, the average accuracy of EFCMF is 76.33% on both datasets respectively, which is 32.63% higher than that of Bi-modal PMA. Compared with existing methods, EFCMF is robust in the face of modal missing and adversarial example attacks in multimodal fine-grained recognition tasks. The source code is available at https://github.com/RPZ97/EFCMF (accessed on 8 January 2023).
引用
收藏
页数:15
相关论文
共 50 条
  • [1] FACESEC: A Fine-grained Robustness Evaluation Framework for Face Recognition Systems
    Tong, Liang
    Chen, Zhengzhang
    Ni, Jingchao
    Cheng, Wei
    Song, Dongjin
    Chen, Haifeng
    Vorobeychik, Yevgeniy
    [J]. 2021 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION, CVPR 2021, 2021, : 13249 - 13258
  • [2] Fine-Grained Multimodal Named Entity Recognition and Grounding with a Generative Framework
    Wang, Jieming
    Li, Ziyan
    Yu, Jianfei
    Yang, Li
    Xia, Rui
    [J]. PROCEEDINGS OF THE 31ST ACM INTERNATIONAL CONFERENCE ON MULTIMEDIA, MM 2023, 2023, : 3934 - 3943
  • [3] Fine-Grained Grounding for Multimodal Speech Recognition
    Srinivasan, Tejas
    Sanabria, Ramon
    Metze, Florian
    Elliott, Desmond
    [J]. FINDINGS OF THE ASSOCIATION FOR COMPUTATIONAL LINGUISTICS, EMNLP 2020, 2020, : 2667 - 2677
  • [4] Fine-Grained Crowdsourcing for Fine-Grained Recognition
    Jia Deng
    Krause, Jonathan
    Li Fei-Fei
    [J]. 2013 IEEE CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR), 2013, : 580 - 587
  • [5] Multimodal Fine-Grained Transformer Model for Pest Recognition
    Zhang, Yinshuo
    Chen, Lei
    Yuan, Yuan
    [J]. ELECTRONICS, 2023, 12 (12)
  • [6] Dynamic Perception Framework for Fine-Grained Recognition
    Ding, Yao
    Han, Zhenjun
    Zhou, Yanzhao
    Zhu, Yi
    Chen, Jie
    Ye, Qixiang
    Jiao, Jianbin
    [J]. IEEE TRANSACTIONS ON CIRCUITS AND SYSTEMS FOR VIDEO TECHNOLOGY, 2022, 32 (03) : 1353 - 1365
  • [7] Multimodal Wearable Sensing for Fine-Grained Activity Recognition in Healthcare
    De, Debraj
    Bharti, Pratool
    Das, Sajal K.
    Chellappan, Sriram
    [J]. IEEE INTERNET COMPUTING, 2015, 19 (05) : 26 - 35
  • [8] A Semantic-driven Image Scene Fine-grained Enhancement Recognition
    Qu, Dongyang
    Li, Yaling
    Luo, Xiaoyan
    Shi, Xiaofeng
    [J]. SEVENTH ASIA PACIFIC CONFERENCE ON OPTICS MANUFACTURE (APCOM 2021), 2022, 12166
  • [9] Multimodal fine-grained grocery product recognition using image and OCR text
    Pettersson, Tobias
    Riveiro, Maria
    Lofstrom, Tuwe
    [J]. MACHINE VISION AND APPLICATIONS, 2024, 35 (04)
  • [10] A progressive deep learning framework for fine-grained primate behavior recognition
    Feng, Jiangfan
    Luo, Hongxin
    Fang, Dongxu
    [J]. APPLIED ANIMAL BEHAVIOUR SCIENCE, 2023, 269