Performance vs. hardware requirements in state-of-the-art automatic speech recognition

被引：13

作者：

Georgescu, Alexandru-Lucian ^{[1
]}

Pappalardo, Alessandro ^{[2
]}

Cucu, Horia ^{[1
]}

Blott, Michaela ^{[2
]}

机构：

[1] Univ Politehn Bucuresti, Speech & Dialogue Res Lab, Bucharest, Romania

[2] Xilinx, Res Labs, Dublin, Ireland

来源：

EURASIP JOURNAL ON AUDIO SPEECH AND MUSIC PROCESSING | 2021年 / 2021卷 / 01期

关键词：

Automatic speech recognition; Survey; End-to-end ASR systems; Deep learning; Performance analysis; DEEP NEURAL-NETWORKS; MODELS; SCALE;

D O I：

10.1186/s13636-021-00217-4

中图分类号：

O42 [声学];

学科分类号：

070206 ; 082403 ;

摘要：

The last decade brought significant advances in automatic speech recognition (ASR) thanks to the evolution of deep learning methods. ASR systems evolved from pipeline-based systems, that modeled hand-crafted speech features with probabilistic frameworks and generated phone posteriors, to end-to-end (E2E) systems, that translate the raw waveform directly into words using one deep neural network (DNN). The transcription accuracy greatly increased, leading to ASR technology being integrated into many commercial applications. However, few of the existing ASR technologies are suitable for integration in embedded applications, due to their hard constrains related to computing power and memory usage. This overview paper serves as a guided tour through the recent literature on speech recognition and compares the most popular ASR implementations. The comparison emphasizes the trade-off between ASR performance and hardware requirements, to further serve decision makers in choosing the system which fits best their embedded application. To the best of our knowledge, this is the first study to provide this kind of trade-off analysis for state-of-the-art ASR systems.

引用

页数：30

共 50 条

[1] Performance vs. hardware requirements in state-of-the-art automatic speech recognition
Alexandru-Lucian Georgescu
Alessandro Pappalardo
Horia Cucu
Michaela Blott
[J]. EURASIP Journal on Audio, Speech, and Music Processing, 2021
[2] State-of-the-Art Review on Recent Trends in Automatic Speech Recognition
Kandji, Abdou Karim
Ba, Cheikh
Ndiaye, Samba
[J]. EMERGING TECHNOLOGIES FOR DEVELOPING COUNTRIES, AFRICATEK 2023, 2024, 520 : 185 - 203
[3] THE STATE-OF-THE-ART IN SPEECH RECOGNITION
BISIANI, R
[J]. TRENDS IN NEUROSCIENCES, 1985, 8 (01) : 9 - 11
[4] Automatic Speech Recognition System for Tonal Languages: State-of-the-Art Survey
Kaur, Jaspreet
Singh, Amitoj
Kadyan, Virender
[J]. ARCHIVES OF COMPUTATIONAL METHODS IN ENGINEERING, 2021, 28 (03) : 1039 - 1068
[5] Automatic Speech Recognition System for Tonal Languages: State-of-the-Art Survey
Jaspreet Kaur
Amitoj Singh
Virender Kadyan
[J]. Archives of Computational Methods in Engineering, 2021, 28 : 1039 - 1068
[6] STATE-OF-THE-ART IN CONTINUOUS SPEECH RECOGNITION
MAKHOUL, J
SCHWARTZ, R
[J]. PROCEEDINGS OF THE NATIONAL ACADEMY OF SCIENCES OF THE UNITED STATES OF AMERICA, 1995, 92 (22) : 9956 - 9963
[7] AUTOMATIC TARGET RECOGNITION - STATE-OF-THE-ART SURVEY
BHANU, B
[J]. IEEE TRANSACTIONS ON AEROSPACE AND ELECTRONIC SYSTEMS, 1986, 22 (04) : 364 - 379
[8] State-of-the-art of SAR automatic target recognition
Novak, LM
[J]. RECORD OF THE IEEE 2000 INTERNATIONAL RADAR CONFERENCE, 2000, : 836 - 843
[9] State-of-the-art of SAR automatic target recognition
Novak, L.M.
[J]. IEEE National Radar Conference - Proceedings, 2000, : 836 - 843
[10] An automatic speech recognition system in Indian and foreign languages: A state-of-the-art review analysis
Gupta, Astha
Kumar, Rakesh
Kumar, Yogesh
[J]. Intelligent Decision Technologies, 2023, 17 (02) : 505 - 526

← 1 2 3 4 5 →