Hybrid explainable image caption generation using image processing and natural language processing

被引:0
|
作者
Mishra, Atul [1 ]
Agrawal, Anubhav [1 ]
Bhasker, Shailendra [2 ]
机构
[1] BML Munjal Univ, Gurgaon, India
[2] Harcourt Butler Tech Univ, Kanpur, India
关键词
NLP; Image caption generation; CNN; LSTM; InceptionV3; HYPERPARAMETER OPTIMIZATION;
D O I
10.1007/s13198-024-02495-5
中图分类号
T [工业技术];
学科分类号
08 ;
摘要
Image caption generation is among the most rapidly growing research areas that combine image processing methodologies with natural language processing (NLP) technique(s). The effectiveness of the combination of image processing and NLP techniques can revolutionaries the areas of content creation, media analysis, and accessibility. The study proposed a novel model to generate automatic image captions by consuming visual and linguistic features. Visual image features are extracted by applying Convolutional Neural Network and linguistic features by Long Short-Term Memory (LSTM) to generate text. Microsoft Common Objects in Context dataset with over 330,000 images having corresponding captions is used to train the proposed model. A comprehensive evaluation of various models, including VGGNet + LSTM, ResNet + LSTM, GoogleNet + LSTM, VGGNet + RNN, AlexNet + RNN, and AlexNet + LSTM, was conducted based on different batch sizes and learning rates. The assessment was performed using metrics such as BLEU-2 Score, METEOR Score, ROUGE-L Score, and CIDEr. The proposed method demonstrated competitive performance, suggesting its potential for further exploration and refinement. These findings underscore the importance of careful parameter tuning and model selection in image captioning tasks.
引用
收藏
页码:4874 / 4884
页数:11
相关论文
共 50 条
  • [1] CAPTION: Caption Analysis with Proposed Terms, Image of Objects, and Natural Language Processing
    Ferreira L.A.
    De Rizzo Meneghetti D.
    Lopes M.
    Santos P.E.
    SN Computer Science, 3 (5)
  • [2] Natural Language Model for Image Caption
    Guo, Chunjie
    Yang, Weiwei
    Peng, Huan
    2020 4TH INTERNATIONAL CONFERENCE ON NATURAL LANGUAGE PROCESSING AND INFORMATION RETRIEVAL, NLPIR 2020, 2020, : 125 - 130
  • [3] Explainable Natural Language Processing
    Zhang, Zihao
    NATURAL LANGUAGE ENGINEERING, 2023, 30 (04) : 882 - 885
  • [4] Image and Natural Language Processing for Multimedia Information Retrieval
    Lapata, Mirella
    ADVANCES IN INFORMATION RETRIEVAL, PROCEEDINGS, 2010, 5993 : 12 - 12
  • [5] Editorial: Explainable AI in Natural Language Processing
    Banerjee, Somnath
    Tomas, David
    FRONTIERS IN ARTIFICIAL INTELLIGENCE, 2024, 7
  • [6] Generation of Oracles using Natural Language Processing
    Leong, Iat Tou
    Barbosa, Raul
    2021 28TH ASIA-PACIFIC SOFTWARE ENGINEERING CONFERENCE WORKSHOPS (APSECW 2021), 2021, : 25 - 31
  • [7] ICON: Instagram Profile Classification Using Image and Natural Language Processing Methods
    Guven, Ebu Yusuf
    Boyaci, Ali
    Saritemur, Fatma Nur
    Turk, Zehra
    Sutcu, Gizem
    Turna, Ozgur Can
    IEEE TRANSACTIONS ON COMPUTATIONAL SOCIAL SYSTEMS, 2024, 11 (02): : 2776 - 2783
  • [8] Underwater Image Processing using Hybrid Techniques
    Krishnapriya, T. S.
    Kunju, Nissan
    PROCEEDINGS OF 2019 1ST INTERNATIONAL CONFERENCE ON INNOVATIONS IN INFORMATION AND COMMUNICATION TECHNOLOGY (ICIICT 2019), 2019,
  • [9] Image Caption Generation Using A Deep Architecture
    Hani, Ansar
    Tagougui, Najiba
    Kherallah, Monji
    2019 INTERNATIONAL ARAB CONFERENCE ON INFORMATION TECHNOLOGY (ACIT), 2019, : 246 - 251
  • [10] Image Caption Generation Using Attention Model
    Ramalakshmi, Eliganti
    Jain, Moksh Sailesh
    Uddin, Mohammed Ameer
    INNOVATIVE DATA COMMUNICATION TECHNOLOGIES AND APPLICATION, ICIDCA 2021, 2022, 96 : 1009 - 1017