Toward Robust Referring Image Segmentation

被引：6

作者：

Wu, Jianzong ^{[1
]}

Li, Xiangtai ^{[1
]}

Li, Xia ^{[2
]}

Ding, Henghui ^{[3
]}

Tong, Yunhai ^{[1
]}

Tao, Dacheng ^{[4
,5
]}

机构：

[1] Peking Univ, Sch Intelligence Sci & Technol, Natl Key Lab Gen Artificial Intelligence, Beijing 100871, Peoples R China

[2] Swiss Fed Inst Technol, Dept Comp Sci, CH-8092 Zurich, Switzerland

[3] Swiss Fed Inst Technol, Dept Informat Technol & Elect Engn, CH-8092 Zurich, Switzerland

[4] Univ Sydney, Camperdown, NSW 2050, Australia

[5] Nanyang Technol Univ, Sch Comp Sci & Engn SCSE, Singapore 639798, Singapore

来源：

IEEE TRANSACTIONS ON IMAGE PROCESSING | 2024年 / 33卷

关键词：

Computer vision; image segmentation; natural language processing;

D O I：

10.1109/TIP.2024.3371348

中图分类号：

TP18 [人工智能理论];

学科分类号：

081104 ; 0812 ; 0835 ; 1405 ;

摘要：

Referring Image Segmentation (RIS) is a fundamental vision-language task that outputs object masks based on text descriptions. Many works have achieved considerable progress for RIS, including different fusion method designs. In this work, we explore an essential question, "What if the text description is wrong or misleading?" For example, the described objects are not in the image. We term such a sentence as a negative sentence. However, existing solutions for RIS cannot handle such a setting. To this end, we propose a new formulation of RIS, named Robust Referring Image Segmentation (R-RIS). It considers the negative sentence inputs besides the regular positive text inputs. To facilitate this new task, we create three R-RIS datasets by augmenting existing RIS datasets with negative sentences and propose new metrics to evaluate both types of inputs in a unified manner. Furthermore, we propose a new transformer-based model, called RefSegformer, with a token-based vision and language fusion module. Our design can be easily extended to our R-RIS setting by adding extra blank tokens. Our proposed RefSegformer achieves state-of-the-art results on both RIS and R-RIS datasets, establishing a solid baseline for both settings. Our project page is at https://github.com/jianzongwu/robust-ref-seg.

引用

页码：1782 / 1794

页数：13

共 50 条

[41] See-Through-Text Grouping for Referring Image Segmentation
Chen, Ding-Jie
Jia, Songhao
Lo, Yi-Chen
Chen, Hwann-Tzong
Liu, Tyng-Luh
2019 IEEE/CVF INTERNATIONAL CONFERENCE ON COMPUTER VISION (ICCV 2019), 2019, : 7453 - 7462
[42] Comprehensive Multi-Modal Interactions for Referring Image Segmentation
Jain, Kanishk
Gandhi, Vineet
FINDINGS OF THE ASSOCIATION FOR COMPUTATIONAL LINGUISTICS (ACL 2022), 2022, : 3427 - 3435
[43] Vision-Aware Language Reasoning for Referring Image Segmentation
Fayou Xu
Bing Luo
Chao Zhang
Li Xu
Mingxing Pu
Bo Li
Neural Processing Letters, 2023, 55 : 11313 - 11331
[44] Bidirectional Relationship Inferring Network for Referring Image Localization and Segmentation
Feng, Guang
Hu, Zhiwei
Zhang, Lihe
Sun, Jiayu
Lu, Huchuan
IEEE TRANSACTIONS ON NEURAL NETWORKS AND LEARNING SYSTEMS, 2023, 34 (05) : 2246 - 2258
[45] Global Selection and Local Attention Network for Referring Image Segmentation
Ding, Haixin
Zhang, Shengchuan
Cao, Liujuan
PATTERN RECOGNITION AND COMPUTER VISION, PRCV 2023, PT VII, 2024, 14431 : 284 - 295
[46] Bottom-Up Shift and Reasoning for Referring Image Segmentation
Yang, Sibei
Xia, Meng
Li, Guanbin
Zhou, Hong-Yu
Yu, Yizhou
2021 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION, CVPR 2021, 2021, : 11261 - 11270
[47] De-noising mask transformer for referring image segmentation
Wang, Yehui
Lei, Fang
Wang, Baoyan
Zhang, Qiang
Zhen, Xiantong
Zhang, Lei
IMAGE AND VISION COMPUTING, 2025, 154
[48] Text-Vision Relationship Alignment for Referring Image Segmentation
Mingxing Pu
Bing Luo
Chao Zhang
Li Xu
Fayou Xu
Mingming Kong
Neural Processing Letters, 56
[49] Shatter and Gather: Learning Referring Image Segmentation with Text Supervision
Kim, Dongwon
Kim, Namyup
Lan, Cuiling
Kwak, Suha
2023 IEEE/CVF INTERNATIONAL CONFERENCE ON COMPUTER VISION (ICCV 2023), 2023, : 15501 - 15511
[50] Toward Robust Image Classification
Alshemali, Basemah
Graham, Alta
Kalita, Jugal
INTELLIGENT SYSTEMS AND APPLICATIONS, VOL 2, 2020, 1038 : 483 - 489

← 1 2 3 4 5 →