Hybrid explainable image caption generation using image processing and natural language processing

被引:0
|
作者
Mishra, Atul [1 ]
Agrawal, Anubhav [1 ]
Bhasker, Shailendra [2 ]
机构
[1] BML Munjal Univ, Gurgaon, India
[2] Harcourt Butler Tech Univ, Kanpur, India
关键词
NLP; Image caption generation; CNN; LSTM; InceptionV3; HYPERPARAMETER OPTIMIZATION;
D O I
10.1007/s13198-024-02495-5
中图分类号
T [工业技术];
学科分类号
08 ;
摘要
Image caption generation is among the most rapidly growing research areas that combine image processing methodologies with natural language processing (NLP) technique(s). The effectiveness of the combination of image processing and NLP techniques can revolutionaries the areas of content creation, media analysis, and accessibility. The study proposed a novel model to generate automatic image captions by consuming visual and linguistic features. Visual image features are extracted by applying Convolutional Neural Network and linguistic features by Long Short-Term Memory (LSTM) to generate text. Microsoft Common Objects in Context dataset with over 330,000 images having corresponding captions is used to train the proposed model. A comprehensive evaluation of various models, including VGGNet + LSTM, ResNet + LSTM, GoogleNet + LSTM, VGGNet + RNN, AlexNet + RNN, and AlexNet + LSTM, was conducted based on different batch sizes and learning rates. The assessment was performed using metrics such as BLEU-2 Score, METEOR Score, ROUGE-L Score, and CIDEr. The proposed method demonstrated competitive performance, suggesting its potential for further exploration and refinement. These findings underscore the importance of careful parameter tuning and model selection in image captioning tasks.
引用
收藏
页码:4874 / 4884
页数:11
相关论文
共 50 条
  • [31] Natural Language Processing Based Approach for Identification of Problems in Medical Image Management Using PACS
    Yagahara, Ayako
    Tanikawa, Takumi
    Fukuda, Akihisa
    Ando, Daisuke
    Suzuki, Tastuya
    Harada, Kohei
    Karata, Shuichi
    Uesugi, Masahito
    MEDINFO 2019: HEALTH AND WELLBEING E-NETWORKS FOR ALL, 2019, 264 : 1815 - 1816
  • [32] WaveICA: a Hybrid Tool for Image Processing
    Coltuc, Daniela
    Udroiu, Iulian
    Angelescu, Nicoleta
    2010 3RD INTERNATIONAL SYMPOSIUM ON ELECTRICAL AND ELECTRONICS ENGINEERING (ISEEE), 2010, : 234 - 237
  • [33] PROCESSING IMAGE DATA BY HYBRID TECHNIQUES
    RAO, KR
    NARASIMHAN, MA
    GORZINSKI, WJ
    IEEE TRANSACTIONS ON SYSTEMS MAN AND CYBERNETICS, 1977, 7 (10): : 728 - 734
  • [34] Image processing to automate mesh generation
    Hattangady, NV
    Fridy, JM
    Vemuri, KR
    Dick, RE
    ENGINEERING WITH COMPUTERS, 1999, 15 (02) : 127 - 136
  • [35] An Efficient Deep Learning based Hybrid Model Image Caption Generation for
    Kaur, Mehzabeen
    Kaur, Harpreet
    INTERNATIONAL JOURNAL OF ADVANCED COMPUTER SCIENCE AND APPLICATIONS, 2023, 14 (03) : 231 - 237
  • [36] Hand Sign Language Feature Extraction Using Image Processing
    Chabchoub, Abdelkader
    Hamouda, Ali
    Al-Ahmadi, Saleh
    Barkouti, Wahid
    Cherif, Adnen
    PROCEEDINGS OF THE FUTURE TECHNOLOGIES CONFERENCE (FTC) 2019, VOL 2, 2020, 1070 : 122 - 131
  • [37] HYBRID IMAGE-PROCESSING WITH FEEDBACK
    HAUSLER, G
    LOHMANN, A
    OPTICS COMMUNICATIONS, 1977, 21 (03) : 365 - 368
  • [38] SUGGESTIONS FOR HYBRID IMAGE-PROCESSING
    LOHMANN, AW
    OPTICS COMMUNICATIONS, 1977, 22 (02) : 165 - 168
  • [39] Image Processing to Automate Mesh Generation
    N. V. Hattangady
    J. M. Fridy
    K. Rao Vemuri
    R. E. Dick
    Engineering with Computers, 1999, 15 : 127 - 136
  • [40] Objective Type Question Generation using Natural Language Processing
    Deena, G.
    Raja, K.
    INTERNATIONAL JOURNAL OF ADVANCED COMPUTER SCIENCE AND APPLICATIONS, 2022, 13 (02) : 539 - 548