TIRec: Transformer-based Invoice Text Recognition

被引：0

作者：

Chen, Yanlan ^{[1
]}

机构：

[1] Beijing Univ Posts & Telecommun, Sch Artificial Intelligence, Beijing, Peoples R China

来源：

2023 2ND ASIA CONFERENCE ON ALGORITHMS, COMPUTING AND MACHINE LEARNING, CACML 2023 | 2023年

关键词：

Text recognition; Invoice; Convolutional Vision Transformer;

D O I：

10.1145/3590003.3590034

中图分类号：

TP18 [人工智能理论];

学科分类号：

081104 ; 0812 ; 0835 ; 1405 ;

摘要：

A novel invoice text recognition model is proposed. In the past few years, researchers have explored text recognition methods with RNN-like structures to model semantic information. However, RNN-based approaches have some obvious drawbacks, such as the level-by-level decoding approach and the one-way serial transmission of semantic information, which greatly limit semantic information's effectiveness and computational efficiency. In contrast, invoice text has obvious contextual relationships due to its fixed text pattern, the text font in the invoice is more fixed and the complexity of the background is much lower than that of natural scenes. To further exploit these contextual relationships and adapt to the characteristics of invoice text, we propose a new text recognition framework inspired by Transformer [1]. Self-attention-based architectures, in particular Transformer, have been successful in natural language processing (NLP). It has demonstrated powerful semantic information modeling capabilities in NLP. Inspired by its success, we try to apply Transformer to invoice text recognition. Unlike the RNN-based approach, we reduce the parameters of the vision network used to extract image features, use the Convolutional Vision Transformer Attention module to capture the semantic information, and use the Transformer decoding module to decode all characters in parallel. We hope that this Transformer-based architecture can better model the semantic information in invoices while remaining lightweight. Meanwhile, we collected text images of more than 40,000 train invoices, VAT invoices, rolled invoices, and cab invoices. Experiments on the collected invoice text recognition dataset show that our approach outperforms previous methods in terms of accuracy and speed.

引用

页码：175 / 180

页数：6

共 50 条

[31] Development of a Text Classification Framework using Transformer-based Embeddings
Yeasmin, Sumona
Afrin, Nazia
Saif, Kashfia
Huq, Mohammad Rezwanul
PROCEEDINGS OF THE 11TH INTERNATIONAL CONFERENCE ON DATA SCIENCE, TECHNOLOGY AND APPLICATIONS (DATA), 2022, : 74 - 82
[32] Mention Flags (MF): Constraining Transformer-based Text Generators
Wang, Yufei
Wood, Ian D.
Wan, Stephen
Dras, Mark
Johnson, Mark
59TH ANNUAL MEETING OF THE ASSOCIATION FOR COMPUTATIONAL LINGUISTICS AND THE 11TH INTERNATIONAL JOINT CONFERENCE ON NATURAL LANGUAGE PROCESSING, VOL 1 (ACL-IJCNLP 2021), 2021, : 103 - 113
[33] Transformer-Based Automatic Speech Recognition with Auxiliary Input of Source Language Text Toward Transcribing Simultaneous Interpretation
Taniguchi, Shuta
Kato, Tsuneo
Tamura, Akihiro
Yasuda, Keiji
INTERSPEECH 2022, 2022, : 2813 - 2817
[34] Lightweight Scene Text Recognition Based on Transformer
Luan, Xin
Zhang, Jinwei
Xu, Miaomiao
Silamu, Wushouer
Li, Yanbing
SENSORS, 2023, 23 (09)
[35] Transformer-based Human Action Recognition with Dynamic Feature Selection
Lamghari, Soufiane
Bilodeau, Guillaume-Alexandre
Saunier, Nicolas
2023 20TH CONFERENCE ON ROBOTS AND VISION, CRV, 2023, : 129 - 136
[36] Transformer-based Unified Recognition of Two Hands Manipulating Objects
Cho, Hoseong
Kim, Chanwoo
Kim, Jihyeon
Lee, Seongyeong
Ismayilzada, Elkhan
Baek, Seungryul
2023 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION, CVPR, 2023, : 4769 - 4778
[37] UNTIED POSITIONAL ENCODINGS FOR EFFICIENT TRANSFORMER-BASED SPEECH RECOGNITION
Samarakoon, Lahiru
Fung, Ivan
2022 IEEE SPOKEN LANGUAGE TECHNOLOGY WORKSHOP, SLT, 2022, : 108 - 114
[38] Transformer-based network with temporal depthwise convolutions for sEMG recognition
Wang, Zefeng
Yao, Junfeng
Xu, Meiyan
Jiang, Min
Su, Jinsong
PATTERN RECOGNITION, 2024, 145
[39] ERTNet: an interpretable transformer-based framework for EEG emotion recognition
Liu, Ruixiang
Chao, Yihu
Ma, Xuerui
Sha, Xianzheng
Sun, Limin
Li, Shuo
Chang, Shijie
FRONTIERS IN NEUROSCIENCE, 2024, 18
[40] Transformer-Based Approaches for Legal Text ProcessingJNLP Team - COLIEE 2021
Ha-Thanh Nguyen
Minh-Phuong Nguyen
Thi-Hai-Yen Vuong
Minh-Quan Bui
Minh-Chau Nguyen
Tran-Binh Dang
Vu Tran
Le-Minh Nguyen
Ken Satoh
The Review of Socionetwork Strategies, 2022, 16 : 135 - 155

← 1 2 3 4 5 →