Fine-Grained Visual Textual Alignment for Cross-Modal Retrieval Using Transformer Encoders

被引：73

作者：

Messina, Nicola ^{[1
]}

Amato, Giuseppe ^{[1
]}

Esuli, Andrea ^{[1
]}

Falchi, Fabrizio ^{[1
]}

Gennaro, Claudio ^{[1
]}

Marchand-Maillet, Stephane ^{[2
]}

机构：

[1] ISTI CNR, Pisa, Italy

[2] Univ Geneva, VIPER Grp, Geneva, Switzerland

来源：

ACM TRANSACTIONS ON MULTIMEDIA COMPUTING COMMUNICATIONS AND APPLICATIONS | 2021年 / 17卷 / 04期

基金：

欧盟地平线“2020”;

关键词：

Deep learning; cross-modal retrieval; multi-modal matching; computer vision; natural language processing; LANGUAGE; GENOME;

D O I：

10.1145/3451390

中图分类号：

TP [自动化技术、计算机技术];

学科分类号：

0812 ;

摘要：

Despite the evolution of deep-learning-based visual-textual processing systems, precise multi-modal matching remains a challenging task. In this work, we tackle the task of cross-modal retrieval through image-sentence matching based on word-region alignments, using supervision only at the global image-sentence level. Specifically, we present a novel approach called Transformer Encoder Reasoning and Alignment Network (TERAN). TERAN enforces a fine-grained match between the underlying components of images and sentences (i.e., image regions and words, respectively) to preserve the informative richness of both modalities. TERAN obtains state-of-the-art results on the image retrieval task on both MS-COCO and Flickr30k datasets. Moreover, on MS-COCO, it also outperforms current approaches on the sentence retrieval task. 000Focusing on scalable cross-modal information retrieval, TERAN is designed to keep the visual and textual data pipelines well separated. Cross-attention links invalidate any chance to separately extract visual and textual features needed for the online search and the offline indexing steps in large-scale retrieval systems. In this respect, TERAN merges the information from the two domains only during the final alignment phase, immediately before the loss computation. We argue that the fine-grained alignments produced by TERAN pave the way toward the research for effective and efficient methods for large-scale cross-modal information retrieval. We compare the effectiveness of our approach against relevant state-of-the-art methods. On the MS-COCO 1K test set, we obtain an improvement of 5.7% and 3.5% respectively on the image and the sentence retrieval tasks on the Recall@1 metric.

引用

页数：23

共 50 条

[41] Cross-modal knowledge learning with scene text for fine-grained image classification
Xiong, Li
Mao, Yingchi
Wang, Zicheng
Nie, Bingbing
Li, Chang
IET IMAGE PROCESSING, 2024, 18 (06) : 1447 - 1459
[42] Fine-grained sentiment Feature Extraction Method for Cross-modal Sentiment Analysis
Sun, Ye
Jin, Guozhe
Zhao, Yahui
Cui, Rongyi
2024 16TH INTERNATIONAL CONFERENCE ON MACHINE LEARNING AND COMPUTING, ICMLC 2024, 2024, : 602 - 608
[43] Token Embeddings Alignment for Cross-Modal Retrieval
Xie, Chen-Wei
Wu, Jianmin
Zheng, Yun
Pan, Pan
Hua, Xian-Sheng
PROCEEDINGS OF THE 30TH ACM INTERNATIONAL CONFERENCE ON MULTIMEDIA, MM 2022, 2022, : 4555 - 4563
[44] Robust cross-modal retrieval with alignment refurbishment
Guo, Jinyi
Ding, Jieyu
FRONTIERS OF INFORMATION TECHNOLOGY & ELECTRONIC ENGINEERING, 2023, 24 (10) : 1403 - 1415
[45] Adequate alignment and interaction for cross-modal retrieval
Mingkang WANG
Min MENG
Jigang LIU
Jigang WU
虚拟现实与智能硬件(中英文), 2023, 5 (06) : 509 - 522
[46] Fine-grained Image Classification and Retrieval by Combining Visual and Locally Pooled Textual Features
Mafla, Andres
Dey, Sounak
Biten, Ali Furkan
Gomez, Lluis
Karatzas, Dimosthenis
2020 IEEE WINTER CONFERENCE ON APPLICATIONS OF COMPUTER VISION (WACV), 2020, : 2939 - 2948
[47] Multimodal Encoders for Food-Oriented Cross-Modal Retrieval
Chen, Ying
Zhou, Dong
Li, Lin
Han, Jun-mei
WEB AND BIG DATA, APWEB-WAIM 2021, PT II, 2021, 12859 : 253 - 266
[48] Semantic Modeling of Textual Relationships in Cross-modal Retrieval
Yu, Jing
Yang, Chenghao
Qin, Zengchang
Yang, Zhuoqian
Hu, Yue
Shi, Zhiguo
KNOWLEDGE SCIENCE, ENGINEERING AND MANAGEMENT, KSEM 2019, PT I, 2019, 11775 : 24 - 32
[49] Fine-Grained Visual-Textual Representation Learning
He, Xiangteng
Peng, Yuxin
IEEE TRANSACTIONS ON CIRCUITS AND SYSTEMS FOR VIDEO TECHNOLOGY, 2020, 30 (02) : 520 - 531
[50] Fine-grained Image-text Matching by Cross-modal Hard Aligning Network
Pan, Zhengxin
Wu, Fangyu
Zhang, Bailing
2023 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR), 2023, : 19275 - 19284

← 1 2 3 4 5 →