High-Order Interaction Learning for Image Captioning

被引:68
|
作者
Wang, Yanhui [1 ]
Xu, Ning [1 ]
Liu, An-An [1 ]
Li, Wenhui [1 ]
Zhang, Yongdong [2 ]
机构
[1] Tianjin Univ, Sch Elect & Informat Engn, Tianjin 300072, Peoples R China
[2] Univ Sci & Technol China, Sch Informat Sci & Technol, Hefei 230052, Peoples R China
基金
中国国家自然科学基金;
关键词
Visualization; Semantics; Feature extraction; Decoding; Task analysis; Ions; Encoding; Image captioning; high-order interaction; encoder-decoder framework;
D O I
10.1109/TCSVT.2021.3121062
中图分类号
TM [电工技术]; TN [电子技术、通信技术];
学科分类号
0808 ; 0809 ;
摘要
Image captioning aims at understanding various semantic concepts (e.g., objects and relationships) from an image and integrating them in a sentence-level description. Hence, it is necessary to learn the interaction among these concepts. If we define the context of the interaction to be involved in the subject-predicate-object triplet, most current methods only focus on the single triplet for the first-order interaction to generate sentences. Intuitively, we humans are able to perceive the high-order interaction among concepts from two or more triplets to describe an image. For example, when we see the triplets man-cutting-sandwich and man-with-knife, it is natural to integrate and predict the sentence man cutting sandwich with knife. This depends on the high-order interaction between cutting and knife in different triplets. Therefore, exploiting high-order interaction is expected to benefit image captioning and focus on reasoning. In this paper, we introduce the novel high-order interaction learning method over detected objects and relationships for image captioning under the umbrella of the encoder-decoder framework. We first extract a set of object and relationship features in an image. During the encoding stage, the interactive refining network is proposed to learn high-order representations by modeling intra- and inter-object feature interaction in the self-attention fashion. During the decoding stage, the interactive fusion network is proposed to integrate object and relationship information by strengthening their high-order interaction based on language context for sentence generation. In this way, we learn the object-relationship dependencies in different stages, which can provide abundant cues for both visual understanding and caption generation. Extensive experiments show that the proposed method can achieve competitive performances against the state-of-the-art methods on MSCOCO dataset. Additional ablation studies further validate its effectiveness.
引用
收藏
页码:4417 / 4430
页数:14
相关论文
共 50 条
  • [21] Learning to Guide Decoding for Image Captioning
    Jiang, Wenhao
    Ma, Lin
    Chen, Xinpeng
    Zhang, Hanwang
    Liu, Wei
    THIRTY-SECOND AAAI CONFERENCE ON ARTIFICIAL INTELLIGENCE / THIRTIETH INNOVATIVE APPLICATIONS OF ARTIFICIAL INTELLIGENCE CONFERENCE / EIGHTH AAAI SYMPOSIUM ON EDUCATIONAL ADVANCES IN ARTIFICIAL INTELLIGENCE, 2018, : 6959 - 6966
  • [22] Image Captioning using Deep Learning
    Jain, Yukti Sanjay
    Dhopeshwar, Tanisha
    Chadha, Supreet Kaur
    Pagire, Vrushali
    2021 INTERNATIONAL CONFERENCE ON COMPUTATIONAL PERFORMANCE EVALUATION (COMPE-2021), 2021,
  • [23] Learning Transferable Perturbations for Image Captioning
    Wu, Hanjie
    Liu, Yongtuo
    Cai, Hongmin
    He, Shengfeng
    ACM Transactions on Multimedia Computing, Communications and Applications, 2022, 18 (02)
  • [24] Image Captioning Using Deep Learning
    Adithya, Paluvayi Veera
    Kalidindi, Mourya Viswanadh
    Swaroop, Nallani Jyothi
    Vishwas, H. N.
    ADVANCED NETWORK TECHNOLOGIES AND INTELLIGENT COMPUTING, ANTIC 2023, PT III, 2024, 2092 : 42 - 58
  • [25] Efficient high-order image subsampling using FANNs
    Dumitras, A
    Kossentini, F
    2000 INTERNATIONAL CONFERENCE ON IMAGE PROCESSING, VOL III, PROCEEDINGS, 2000, : 316 - 319
  • [26] Deep High-order Supervised Hashing for Image Retrieval
    Cheng, Jingdong
    Sun, Qiule
    Zhang, Jianxin
    Wei, Xiaopeng
    Zhang, Qiang
    2018 24TH INTERNATIONAL CONFERENCE ON PATTERN RECOGNITION (ICPR), 2018, : 2693 - 2698
  • [27] On the accuracy bounds of high-order image correlation spectroscopy
    Katoozi, Delaram
    Clayton, Andrew H. A.
    Moss, David J.
    Chon, James W. M.
    OPTICS EXPRESS, 2024, 32 (13): : 22095 - 22109
  • [28] Distributed image classification based on high-order features
    Liu Qi
    Liang Peng
    Zhang Haitao
    Zhou Jianxiong
    Zhou Yishu
    PROCEEDINGS OF 2015 IEEE 12TH INTERNATIONAL CONFERENCE ON ELECTRONIC MEASUREMENT & INSTRUMENTS (ICEMI), VOL. 3, 2015, : 1122 - 1125
  • [29] High-order iterative learning controller with initial state learning
    Chen, Yangquan
    Wen, Changyun
    Sun, Mingxuan
    IMA Journal of Mathematical Control and Information, 2000, 17 (02) : 111 - 121
  • [30] Interaction of high-order solitons with external dispersive waves
    Oreshnikov, I.
    Driben, R.
    Yulin, A. V.
    OPTICS LETTERS, 2015, 40 (23) : 5554 - 5557