Adversarial text-to-image synthesis: A review

被引:58
|
作者
Frolov, Stanislav [1 ,2 ]
Hinz, Tobias [3 ,4 ]
Raue, Federico [2 ]
Hees, Joern [2 ]
Dengel, Andreas [1 ,2 ]
机构
[1] Tech Univ Kaiserslautern, Kaiserslautern, Germany
[2] Deutsch Forschungszentrum Kunstliche Intelligenz, Kaiserslautern, Germany
[3] Univ Hamburg, Hamburg, Germany
[4] Adobe Res, San Jose, CA USA
关键词
Text-to-image synthesis; Generative adversarial networks; GENERATION; NETWORKS;
D O I
10.1016/j.neunet.2021.07.019
中图分类号
TP18 [人工智能理论];
学科分类号
081104 ; 0812 ; 0835 ; 1405 ;
摘要
With the advent of generative adversarial networks, synthesizing images from text descriptions has recently become an active research area. It is a flexible and intuitive way for conditional image generation with significant progress in the last years regarding visual realism, diversity, and semantic alignment. However, the field still faces several challenges that require further research efforts such as enabling the generation of high-resolution images with multiple objects, and developing suitable and reliable evaluation metrics that correlate with human judgement. In this review, we contextualize the state of the art of adversarial text-to-image synthesis models, their development since their inception five years ago, and propose a taxonomy based on the level of supervision. We critically examine current strategies to evaluate text-to-image synthesis models, highlight shortcomings, and identify new areas of research, ranging from the development of better datasets and evaluation metrics to possible improvements in architectural design and model training. This review complements previous surveys on generative adversarial networks with a focus on text-to-image synthesis which we believe will help researchers to further advance the field. @2021 Published Elsevier Ltd
引用
收藏
页码:187 / 209
页数:23
相关论文
共 50 条
  • [1] Dual Adversarial Inference for Text-to-Image Synthesis
    Lao, Qicheng
    Havaei, Mohammad
    Pesaranghader, Ahmad
    Dutil, Francis
    Di Jorio, Lisa
    Fevens, Thomas
    [J]. 2019 IEEE/CVF INTERNATIONAL CONFERENCE ON COMPUTER VISION (ICCV 2019), 2019, : 7566 - 7575
  • [2] Perceptual Pyramid Adversarial Networks for Text-to-Image Synthesis
    Gao, Lianli
    Chen, Daiyuan
    Song, Jingkuan
    Xu, Xing
    Zhang, Dongxiang
    Shen, Heng Tao
    [J]. THIRTY-THIRD AAAI CONFERENCE ON ARTIFICIAL INTELLIGENCE / THIRTY-FIRST INNOVATIVE APPLICATIONS OF ARTIFICIAL INTELLIGENCE CONFERENCE / NINTH AAAI SYMPOSIUM ON EDUCATIONAL ADVANCES IN ARTIFICIAL INTELLIGENCE, 2019, : 8312 - 8319
  • [3] Advancements in adversarial generative text-to-image models: a review
    Zaghloul, Rawan
    Rawashdeh, Enas
    Bani-Ata, Tomader
    [J]. IMAGING SCIENCE JOURNAL, 2024,
  • [4] GALIP: Generative Adversarial CLIPs for Text-to-Image Synthesis
    Tao, Ming
    Bao, Bing-Kun
    Tang, Hao
    Xu, Changsheng
    [J]. 2023 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR), 2023, : 14214 - 14223
  • [5] Semantic Distance Adversarial Learning for Text-to-Image Synthesis
    Yuan, Bowen
    Sheng, Yefei
    Bao, Bing-Kun
    Chen, Yi-Ping Phoebe
    Xu, Changsheng
    [J]. IEEE TRANSACTIONS ON MULTIMEDIA, 2024, 26 : 1255 - 1266
  • [6] ADVERSARIAL NETS WITH PERCEPTUAL LOSSES FOR TEXT-TO-IMAGE SYNTHESIS
    Cha, Miriam
    Gwon, Youngjune
    Kung, H. T.
    [J]. 2017 IEEE 27TH INTERNATIONAL WORKSHOP ON MACHINE LEARNING FOR SIGNAL PROCESSING, 2017,
  • [7] A survey of generative adversarial networks and their application in text-to-image synthesis
    Zeng, Wu
    Zhu, Heng-liang
    Lin, Chuan
    Xiao, Zheng-ying
    [J]. ELECTRONIC RESEARCH ARCHIVE, 2023, 31 (12): : 7142 - 7181
  • [8] Semantics-enhanced Adversarial Nets for Text-to-Image Synthesis
    Tan, Hongchen
    Liu, Xiuping
    Li, Xin
    Zhang, Yi
    Yin, Baocai
    [J]. 2019 IEEE/CVF INTERNATIONAL CONFERENCE ON COMPUTER VISION (ICCV 2019), 2019, : 10500 - 10509
  • [9] Survey About Generative Adversarial Network and Text-to-Image Synthesis
    Lai, Lina
    Mi, Yu
    Zhou, Longlong
    Rao, Jiyong
    Xu, Tianyang
    Song, Xiaoning
    [J]. Computer Engineering and Applications, 2023, 59 (19): : 21 - 39
  • [10] TextControlGAN: Text-to-Image Synthesis with Controllable Generative Adversarial Networks
    Ku, Hyeeun
    Lee, Minhyeok
    [J]. APPLIED SCIENCES-BASEL, 2023, 13 (08):