Dual Adversarial Inference for Text-to-Image Synthesis

被引:19
|
作者
Lao, Qicheng [1 ,2 ]
Havaei, Mohammad [1 ]
Pesaranghader, Ahmad [1 ,3 ]
Dutil, Francis [1 ]
Di Jorio, Lisa [1 ]
Fevens, Thomas [2 ]
机构
[1] Imagia Inc, Montreal, PQ, Canada
[2] Concordia Univ, Montreal, PQ, Canada
[3] Dalhousie Univ, Halifax, NS, Canada
关键词
D O I
10.1109/ICCV.2019.00766
中图分类号
TP18 [人工智能理论];
学科分类号
081104 ; 0812 ; 0835 ; 1405 ;
摘要
Synthesizing images from a given text description involves engaging two types of information: the content, which includes information explicitly described in the text (e.g., color, composition, etc.), and the style, which is usually not well described in the text (e.g., location, quantity, size, etc.). However, in previous works, it is typically treated as a process of generating images only from the content, i.e., without considering learning meaningful style representations. In this paper, we aim to learn two variables that are disentangled in the latent space, representing content and style respectively. We achieve this by augmenting current text-to-image synthesis frameworks with a dual adversarial inference mechanism. Through extensive experiments, we show that our model learns, in an unsupervised manner, style representations corresponding to certain meaningful information present in the image that are not well described in the text. The new framework also improves the quality of synthesized images when evaluated on Oxford-102, CUB and COCO datasets.
引用
收藏
页码:7566 / 7575
页数:10
相关论文
共 50 条
  • [1] Adversarial text-to-image synthesis: A review
    Frolov, Stanislav
    Hinz, Tobias
    Raue, Federico
    Hees, Joern
    Dengel, Andreas
    [J]. NEURAL NETWORKS, 2021, 144 : 187 - 209
  • [2] Perceptual Pyramid Adversarial Networks for Text-to-Image Synthesis
    Gao, Lianli
    Chen, Daiyuan
    Song, Jingkuan
    Xu, Xing
    Zhang, Dongxiang
    Shen, Heng Tao
    [J]. THIRTY-THIRD AAAI CONFERENCE ON ARTIFICIAL INTELLIGENCE / THIRTY-FIRST INNOVATIVE APPLICATIONS OF ARTIFICIAL INTELLIGENCE CONFERENCE / NINTH AAAI SYMPOSIUM ON EDUCATIONAL ADVANCES IN ARTIFICIAL INTELLIGENCE, 2019, : 8312 - 8319
  • [3] Weakly Supervised Attention Inference Generative Adversarial Network for Text-to-Image
    Mei, Lingrui
    Ran, Xuming
    Hu, Jin
    [J]. 2019 IEEE SYMPOSIUM SERIES ON COMPUTATIONAL INTELLIGENCE (IEEE SSCI 2019), 2019, : 1574 - 1578
  • [4] GALIP: Generative Adversarial CLIPs for Text-to-Image Synthesis
    Tao, Ming
    Bao, Bing-Kun
    Tang, Hao
    Xu, Changsheng
    [J]. 2023 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR), 2023, : 14214 - 14223
  • [5] Semantic Distance Adversarial Learning for Text-to-Image Synthesis
    Yuan, Bowen
    Sheng, Yefei
    Bao, Bing-Kun
    Chen, Yi-Ping Phoebe
    Xu, Changsheng
    [J]. IEEE TRANSACTIONS ON MULTIMEDIA, 2024, 26 : 1255 - 1266
  • [6] ADVERSARIAL NETS WITH PERCEPTUAL LOSSES FOR TEXT-TO-IMAGE SYNTHESIS
    Cha, Miriam
    Gwon, Youngjune
    Kung, H. T.
    [J]. 2017 IEEE 27TH INTERNATIONAL WORKSHOP ON MACHINE LEARNING FOR SIGNAL PROCESSING, 2017,
  • [7] A survey of generative adversarial networks and their application in text-to-image synthesis
    Zeng, Wu
    Zhu, Heng-liang
    Lin, Chuan
    Xiao, Zheng-ying
    [J]. ELECTRONIC RESEARCH ARCHIVE, 2023, 31 (12): : 7142 - 7181
  • [8] Semantics-enhanced Adversarial Nets for Text-to-Image Synthesis
    Tan, Hongchen
    Liu, Xiuping
    Li, Xin
    Zhang, Yi
    Yin, Baocai
    [J]. 2019 IEEE/CVF INTERNATIONAL CONFERENCE ON COMPUTER VISION (ICCV 2019), 2019, : 10500 - 10509
  • [9] Survey About Generative Adversarial Network and Text-to-Image Synthesis
    Lai, Lina
    Mi, Yu
    Zhou, Longlong
    Rao, Jiyong
    Xu, Tianyang
    Song, Xiaoning
    [J]. Computer Engineering and Applications, 2023, 59 (19): : 21 - 39
  • [10] TextControlGAN: Text-to-Image Synthesis with Controllable Generative Adversarial Networks
    Ku, Hyeeun
    Lee, Minhyeok
    [J]. APPLIED SCIENCES-BASEL, 2023, 13 (08):