Text Data Augmentation for Deep Learning

被引:187
|
作者
Shorten, Connor [1 ]
Khoshgoftaar, Taghi M. [1 ]
Furht, Borko [1 ]
机构
[1] Florida Atlantic Univ, 777 Glades Rd, Boca Raton, FL 33431 USA
关键词
Data Augmentation; Natural Language Processing; Overfitting; Big Data; NLP; Text Data;
D O I
10.1186/s40537-021-00492-0
中图分类号
TP301 [理论、方法];
学科分类号
081202 ;
摘要
Natural Language Processing (NLP) is one of the most captivating applications of Deep Learning. In this survey, we consider how the Data Augmentation training strategy can aid in its development. We begin with the major motifs of Data Augmentation summarized into strengthening local decision boundaries, brute force training, causality and counterfactual examples, and the distinction between meaning and form. We follow these motifs with a concrete list of augmentation frameworks that have been developed for text data. Deep Learning generally struggles with the measurement of generalization and characterization of overfitting. We highlight studies that cover how augmentations can construct test sets for generalization. NLP is at an early stage in applying Data Augmentation compared to Computer Vision. We highlight the key differences and promising ideas that have yet to be tested in NLP. For the sake of practical implementation, we describe tools that facilitate Data Augmentation such as the use of consistency regularization, controllers, and offline and online augmentation pipelines, to preview a few. Finally, we discuss interesting topics around Data Augmentation in NLP such as task-specific augmentations, the use of prior knowledge in self-supervised learning versus Data Augmentation, intersections with transfer and multi-task learning, and ideas for AI-GAs (AI-Generating Algorithms). We hope this paper inspires further research interest in Text Data Augmentation.
引用
收藏
页数:34
相关论文
共 50 条
  • [41] Performance Benchmarking of Data Augmentation and Deep Learning for Tornado Prediction
    Barajas, Carlos A.
    Gobbert, Matthias K.
    Wang, Jianwu
    [J]. 2019 IEEE INTERNATIONAL CONFERENCE ON BIG DATA (BIG DATA), 2019, : 3607 - 3615
  • [42] A deep learning approach with data augmentation for median filtering forensics
    Dong, Wanli
    Zeng, Hui
    Peng, Yong
    Gao, Xiaoming
    Peng, Anjie
    [J]. MULTIMEDIA TOOLS AND APPLICATIONS, 2022, 81 (08) : 11087 - 11105
  • [43] Deep Learning for Neuroimaging Segmentation with a Novel Data Augmentation Strategy
    Wu, Wenshan
    Lu, Yuhao
    Mane, Ravikiran
    Guan, Cuntai
    [J]. 42ND ANNUAL INTERNATIONAL CONFERENCES OF THE IEEE ENGINEERING IN MEDICINE AND BIOLOGY SOCIETY: ENABLING INNOVATIVE TECHNOLOGIES FOR GLOBAL HEALTHCARE EMBC'20, 2020, : 1516 - 1519
  • [44] Sparse Signal Models for Data Augmentation in Deep Learning ATR
    Agarwal, Tushar
    Sugavanam, Nithin
    Ertin, Emre
    [J]. REMOTE SENSING, 2023, 15 (16)
  • [45] Deep Generative Models for Data Synthesis and Augmentation in Machine Learning
    Adavala, Kiran Mayee
    Vhatkar, Sangeeta
    Ruprah, Taranpreet Singh
    Bhatia, Sukhwinder Kaur
    Kumar, Vipin
    Sharma, Dharmendra
    Praveen, B. Shyam
    [J]. JOURNAL OF ELECTRICAL SYSTEMS, 2024, 20 (03) : 1242 - 1249
  • [46] Molecular communication data augmentation and deep learning based detection
    Scazzoli, Davide
    Vakilipoor, Fardad
    Magarini, Maurizio
    [J]. NANO COMMUNICATION NETWORKS, 2024, 40
  • [47] Training data augmentation for deep learning radio frequency systems
    Clark, William H.
    Hauser, Steven
    Headley, William C.
    Michaels, Alan J.
    [J]. JOURNAL OF DEFENSE MODELING AND SIMULATION-APPLICATIONS METHODOLOGY TECHNOLOGY-JDMS, 2021, 18 (03): : 217 - 237
  • [48] A deep learning approach with data augmentation for median filtering forensics
    Wanli Dong
    Hui Zeng
    Yong Peng
    Xiaoming Gao
    Anjie Peng
    [J]. Multimedia Tools and Applications, 2022, 81 : 11087 - 11105
  • [49] A Preliminary Study on Data Augmentation of Deep Learning for Image Classification
    Lei, Cheng
    Hu, Benlin
    Wang, Dong
    Zhang, Shu
    Chen, Zhenyu
    [J]. 11TH ASIA-PACIFIC SYMPOSIUM ON INTERNETWARE (INTERNETWARE 2019), 2019,
  • [50] ECG Quality Assessment via Deep Learning and Data Augmentation
    Huerta, Alvaro
    Martinez-Rodrigo, Arturo
    Rieta, Jose J.
    Alcaraz, Raul
    [J]. 2021 COMPUTING IN CARDIOLOGY (CINC), 2021,