Dual Attention GANs for Semantic Image Synthesis

被引:34
|
作者
Tang, Hao [1 ]
Bai, Song [2 ]
Sebe, Nicu [1 ,3 ]
机构
[1] Univ Trento, DISI, Trento, Italy
[2] Univ Oxford, Dept Engn Sci, Oxford, England
[3] Huawei Res Ireland, Dublin, Ireland
关键词
Generative Adversarial Networks (GANs); Semantic Image Synthesis; Spatial Attention; Channel Attention;
D O I
10.1145/3394171.3416270
中图分类号
TP18 [人工智能理论];
学科分类号
081104 ; 0812 ; 0835 ; 1405 ;
摘要
In this paper, we focus on the semantic image synthesis task that aims at transferring semantic label maps to photo-realistic images. Existing methods lack effective semantic constraints to preserve the semantic information and ignore the structural correlations in both spatial and channel dimensions, leading to unsatisfactory blurry and artifact-prone results. To address these limitations, we propose a novel Dual Attention GAN (DAGAN) to synthesize photo-realistic and semantically-consistent images with fine details from the input layouts without imposing extra training overhead or modifying the network architectures of existing methods. We also propose two novel modules, i.e., position-wise Spatial Attention Module (SAM) and scale-wise Channel Attention Module (CAM), to capture semantic structure attention in spatial and channel dimensions, respectively. Specifically, SAM selectively correlates the pixels at each position by a spatial attention map, leading to pixels with the same semantic label being related to each other regardless of their spatial distances. Meanwhile, CAM selectively emphasizes the scalewise features at each channel by a channel attention map, which integrates associated features among all channel maps regardless of their scales. We finally sum the outputs of SAM and CAM to further improve feature representation. Extensive experiments on four challenging datasets show that DAGAN achieves remarkably better results than state-of-the-art methods, while using fewer model parameters. The source code and trained models are available at https://github.com/Ha0Tang/DAGAN.
引用
收藏
页码:1994 / 2002
页数:9
相关论文
共 50 条
  • [31] CCASinGAN: Cascaded Channel Attention Guided Single-Image GANs
    Wang, Xueqin
    Jiang, Wenzong
    Zhao, Lifei
    Liu, Baodi
    Wang, Yanjiang
    [J]. 2022 16TH IEEE INTERNATIONAL CONFERENCE ON SIGNAL PROCESSING (ICSP2022), VOL 1, 2022, : 61 - 65
  • [32] Text to Image GANs with RoBERTa and Fine-grained Attention Networks
    Siddharth, M.
    Aarthi, R.
    [J]. INTERNATIONAL JOURNAL OF ADVANCED COMPUTER SCIENCE AND APPLICATIONS, 2021, 12 (12) : 947 - 955
  • [33] SEMANTIC-FUSION GANS FOR SEMI-SUPERVISED SATELLITE IMAGE CLASSIFICATION
    Roy, Subhankar
    Sangineto, Enver
    Demir, Begum
    Sebe, Nicu
    [J]. 2018 25TH IEEE INTERNATIONAL CONFERENCE ON IMAGE PROCESSING (ICIP), 2018, : 684 - 688
  • [34] Dual-attention-transformer-based semantic reranking for large-scale image localization
    Xiao, Yilin
    Du, Siliang
    Chen, Xu
    Liu, Mingzhong
    Sun, Mingwei
    [J]. APPLIED INTELLIGENCE, 2024, 54 (9-10) : 6946 - 6958
  • [35] Scaling up GANs for Text-to-Image Synthesis
    Kang, Minguk
    Zhu, Jun-Yan
    Zhang, Richard
    Park, Jaesik
    Shechtman, Eli
    Paris, Sylvain
    Park, Taesung
    [J]. 2023 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR), 2023, : 10124 - 10134
  • [36] The Application of Attention Mechanism in Semantic Image Segmentation
    Qin, Qi
    Hu, Xijing
    [J]. PROCEEDINGS OF 2020 IEEE 4TH INFORMATION TECHNOLOGY, NETWORKING, ELECTRONIC AND AUTOMATION CONTROL CONFERENCE (ITNEC 2020), 2020, : 1573 - 1580
  • [37] Image Captioning Based on Visual and Semantic Attention
    Wei, Haiyang
    Li, Zhixin
    Zhang, Canlong
    [J]. MULTIMEDIA MODELING (MMM 2020), PT I, 2020, 11961 : 151 - 162
  • [38] Adaptive Semantic Bayesian Framework for Image Attention
    Zhang, Wei
    Wu, Q. M. Jonathan
    Wang, Guanghui
    [J]. 19TH INTERNATIONAL CONFERENCE ON PATTERN RECOGNITION, VOLS 1-6, 2008, : 2092 - 2095
  • [39] GAN for Semantic Image Synthesis with Laplacian Pyramid and Multi-Scale Channel Attention
    Dong, Xinhua
    Li, Chuang
    Xu, Zhigang
    Han, Hongmu
    Jiang, Lifeng
    [J]. IEEE Access, 2024, 12 : 178010 - 178021
  • [40] DEANet: A Real-Time Image Semantic Segmentation Method Based on Dual Efficient Attention Mechanism
    Liu, Xu
    Liu, Rui
    Dong, Jing
    Yi, Pengfei
    Zhou, Dongsheng
    [J]. WIRELESS ALGORITHMS, SYSTEMS, AND APPLICATIONS (WASA 2022), PT II, 2022, 13472 : 193 - 205