Enhancing Multi-modal Multi-hop Question Answering via Structured Knowledge and Unified Retrieval-Generation

被引:1
|
作者
Yang, Qian [1 ]
Chen, Qian
Wang, Wen
Hu, Baotian [1 ]
Zhang, Min [1 ]
机构
[1] Harbin Inst Technol, Shenzhen, Peoples R China
关键词
Question Answering; Cross-modal Reasoning; Multi-modal Retrieval;
D O I
10.1145/3581783.3611964
中图分类号
TP18 [人工智能理论];
学科分类号
081104 ; 0812 ; 0835 ; 1405 ;
摘要
Multi-modal multi-hop question answering involves answering a question by reasoning over multiple input sources from different modalities. Existing methods often retrieve evidences separately and then use a language model to generate an answer based on the retrieved evidences, and thus do not adequately connect candidates and are unable to model the interdependent relations during retrieval. Moreover, the pipelined approaches of retrieval and generation might result in poor generation performance when retrieval performance is low. To address these issues, we propose a Structured Knowledge and Unified Retrieval-Generation (SKURG) approach. SKURG employs an Entity-centered Fusion Encoder to align sources from different modalities using shared entities. It then uses a unified Retrieval-Generation Decoder to integrate intermediate retrieval results for answer generation and also adaptively determine the number of retrieval steps. Extensive experiments on two representative multi-modal multi-hop QA datasets MultimodalQA and WebQA demonstrate that SKURG outperforms the state-of-the-art models in both source retrieval and answer generation performance with fewer parameters(1).
引用
收藏
页码:5223 / 5234
页数:12
相关论文
共 50 条
  • [1] Multi-Modal Alignment of Visual Question Answering Based on Multi-Hop Attention Mechanism
    Xia, Qihao
    Yu, Chao
    Hou, Yinong
    Peng, Pingping
    Zheng, Zhengqi
    Chen, Wen
    [J]. ELECTRONICS, 2022, 11 (11)
  • [2] Unsupervised Multi-hop Question Answering by Question Generation
    Pan, Liangming
    Chen, Wenhu
    Xiong, Wenhan
    Kan, Min-Yen
    Wang, William Yang
    [J]. 2021 CONFERENCE OF THE NORTH AMERICAN CHAPTER OF THE ASSOCIATION FOR COMPUTATIONAL LINGUISTICS: HUMAN LANGUAGE TECHNOLOGIES (NAACL-HLT 2021), 2021, : 5866 - 5880
  • [3] Subgraph Retrieval Enhanced Model for Multi-hop Knowledge Base Question Answering
    Zhang, Jing
    Zhang, Xiaokang
    Yu, Jifan
    Tang, Jian
    Tang, Jie
    Li, Cuiping
    Chen, Hong
    [J]. PROCEEDINGS OF THE 60TH ANNUAL MEETING OF THE ASSOCIATION FOR COMPUTATIONAL LINGUISTICS (ACL 2022), VOL 1: (LONG PAPERS), 2022, : 5773 - 5784
  • [4] Multi-hop Question Answering
    Mavi, Vaibhav
    Jangra, Anubhav
    Jatowt, Adam
    [J]. FOUNDATIONS AND TRENDS IN INFORMATION RETRIEVAL, 2023, 17 (05): : 457 - 586
  • [5] Discovering Multimodal Hierarchical Structures with Graph Neural Networks for Multi-modal and Multi-hop Question Answering
    Zhang, Qing
    Lv, Haocheng
    Liu, Jie
    Chen, Zhiyun
    Duan, Jianyong
    Xv, Mingying
    Wang, Hao
    [J]. PATTERN RECOGNITION AND COMPUTER VISION, PRCV 2023, PT I, 2024, 14425 : 383 - 394
  • [6] Ask to Understand: Question Generation for Multi-hop Question Answering
    Li, Jiawei
    Ren, Mucheng
    Gao, Yang
    Yang, Yizhe
    [J]. CHINESE COMPUTATIONAL LINGUISTICS, CCL 2023, 2023, 14232 : 19 - 36
  • [7] Multi-Hop Reasoning for Question Answering with Knowledge Graph
    Zhang, Jiayuan
    Cai, Yifei
    Zhang, Qian
    Cao, Zehao
    Cheng, Zhenrong
    Li, Dongmei
    Meng, Xianghao
    [J]. 2021 IEEE/ACIS 20TH INTERNATIONAL CONFERENCE ON COMPUTER AND INFORMATION SCIENCE (ICIS 2021-SUMMER), 2021, : 121 - 125
  • [8] Multi-modal multi-hop interaction network for dialogue response generation
    Zhou, Jie
    Tian, Junfeng
    Wang, Rui
    Wu, Yuanbin
    Yan, Ming
    He, Liang
    Huang, Xuanjing
    [J]. EXPERT SYSTEMS WITH APPLICATIONS, 2023, 227
  • [9] Multi-Hop Paragraph Retrieval for Open-Domain Question Answering
    Feldman, Yair
    El-Yaniv, Ran
    [J]. 57TH ANNUAL MEETING OF THE ASSOCIATION FOR COMPUTATIONAL LINGUISTICS (ACL 2019), 2019, : 2296 - 2309
  • [10] Translational relation embeddings for multi-hop knowledge base question answering
    Li, Ziyan
    Wang, Haofen
    Zhang, Wenqiang
    [J]. JOURNAL OF WEB SEMANTICS, 2022, 74