Attribute-Driven Filtering: A new attributes predicting approach for fine-grained image captioning

被引：0

作者：

Hossen, Md. Bipul ^{[1
]}

Ye, Zhongfu ^{[1
]}

Abdussalam, Amr ^{[1
]}

Ul Hassan, Shabih ^{[1
]}

机构：

[1] Univ Sci & Technol China, Sch Informat Sci & Technol, Hefei 230027, Anhui, Peoples R China

来源：

ENGINEERING APPLICATIONS OF ARTIFICIAL INTELLIGENCE | 2024年 / 137卷

关键词：

Fine-grained captioning; Fusion mechanism; Encoder-decoder architecture; Attribute predictor module; ATTENTION;

D O I：

10.1016/j.engappai.2024.109134

中图分类号：

TP [自动化技术、计算机技术];

学科分类号：

0812 ;

摘要：

Fine-grained image captioning with attribute information has garnered significant attention in the realms of computer vision and natural language processing, demanding precise and contextually relevant descriptions of visual content. While previous attribute-driven image captioning models have shown improvements, challenges remain, such as the independence of attribute predictors and caption generators and the semantic gap between images and attributes. Another common issue is the inclusion of all attributes at every time step, despite most attributes being irrelevant to the word currently being generated. This can divert the model's attention toward erroneous semantic details, resulting in a performance decline. To address these issues, we propose a novel Attribute-Driven Filtering (ADF) captioning network designed to provide rich and nuanced descriptions. This model incorporates a unique Attribute Predictor Module (APM) that dynamically predicts the most pertinent attributes in accordance with the textual context, utilizing different attributes at various time steps. The novelty of this approach lies in recognizing that not all attributes hold equal relevance at each time step, and the APM filters out irrelevant attributes to generate precise and contextually relevant captions. Furthermore, this model features a fusion mechanism that integrates visual information from a conventional attention module with attribute information predicted by the APM, aiming to reduce the visual semantic gap between images and attributes. Extensive experimentation demonstrates that the ADF model outperforms advanced models, achieving impressive CIDEr-D scores of 72.0 (Flickr30K) and 123.3 (MS-COCO) through reinforcement learning optimization. It consistently surpasses baseline models across diverse evaluation metrics, highlighting its effectiveness and robustness.

引用

页数：17

共 50 条

[41] Unlocking Efficiency in Fine-Grained Compositional Image Synthesis: A Single-Generator Approach
Wang, Zongtao
Liu, Zhiming
[J]. APPLIED SCIENCES-BASEL, 2023, 13 (13):
[42] Mining Fine-Grained Image-Text Alignment for Zero-Shot Captioning via Text-Only Training
Qiu, Longtian
Ning, Shan
He, Xuming
[J]. THIRTY-EIGHTH AAAI CONFERENCE ON ARTIFICIAL INTELLIGENCE, VOL 38 NO 5, 2024, : 4605 - 4613
[43] New Constructions of Hierarchical Attribute-Based Encryption for Fine-Grained Access Control in Cloud Computing
Zhang, Leyou
Hu, Yupu
[J]. KSII TRANSACTIONS ON INTERNET AND INFORMATION SYSTEMS, 2013, 7 (05): : 1343 - 1356
[44] A Motion-Driven Approach for Fine-Grained Temporal Segmentation of User-Generated Videos
Apostolidis, Konstantinos
Apostolidis, Evlampios
Mezaris, Vasileios
[J]. MULTIMEDIA MODELING, MMM 2018, PT I, 2018, 10704 : 29 - 41
[45] Compositional controls on early diagenetic pathways in fine-grained sedimentary rocks: Implications for predicting unconventional reservoir attributes of mudstones
Macquaker, Joe H. S.
Taylor, Kevin G.
Keller, Margaret
Polya, David
[J]. AAPG BULLETIN, 2014, 98 (03) : 587 - 603
[46] Text-Enhanced Attribute-Based Attention for Generalized Zero-Shot Fine-Grained Image Classification
Chen, Yan-He
Yeh, Mei-Chen
[J]. PROCEEDINGS OF THE 2021 INTERNATIONAL CONFERENCE ON MULTIMEDIA RETRIEVAL (ICMR '21), 2021, : 447 - 450
[47] Attribute-Aware Deep Hashing With Self-Consistency for Large-Scale Fine-Grained Image Retrieval
Wei, Xiu-Shen
Shen, Yang
Sun, Xuhao
Wang, Peng
Peng, Yuxin
[J]. IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE, 2023, 45 (11) : 13904 - 13920
[48] A2-NET: Learning Attribute-Aware Hash Codes for Large-Scale Fine-Grained Image Retrieval
Wei, Xiu-Shen
Shen, Yang
Sun, Xuhao
Ye, Han-Jia
Yang, Jian
[J]. ADVANCES IN NEURAL INFORMATION PROCESSING SYSTEMS 34 (NEURIPS 2021), 2021, 34
[49] Integrating Multiple Models Using Image-as-Documents Approach for Recognizing Fine-Grained Home Contexts
Chen, Sinan
Saiki, Sachio
Nakamura, Masahide
[J]. SENSORS, 2020, 20 (03)
[50] A Fine-Grained System Driven of Attacks Over Several New Representation Techniques Using Machine Learning
Al Ghamdi, Mohammed A.
[J]. IEEE ACCESS, 2023, 11 : 96615 - 96625

← 1 2 3 4 5 →