Arabic Handwritten Documents Segmentation into Text-lines and Words using Deep Learning

被引:22
|
作者
Neche, Chemseddine [1 ]
Belaid, Abdel [1 ]
Kacem-Echi, Afef [2 ]
机构
[1] Univ Lorraine, LORIA, F-54500 Vandoeuvre Les Nancy, France
[2] Univ Tunis, ENSIT LaTICE, 5 Ave Taha Hussein, Tunis 1008, Tunisia
关键词
D O I
10.1109/ICDARW.2019.50110
中图分类号
TP18 [人工智能理论];
学科分类号
081104 ; 0812 ; 0835 ; 1405 ;
摘要
One of the most important steps in a handwriting recognition system is text-line and word segmentation. But, this step is made difficult by the differences in handwriting styles, problems of skewness, overlapping and touching of text and the fluctuations of text-lines. It is even more difficult for ancient and calligraphic writings, as in Arabic manuscripts, due to the cursive connection in Arabic text, the erroneous position of diacritic marks, the presence of ascending and descending letters, etc. In this work, we propose an effective segmentation of Arabic handwritten text into text-lines and words, using deep learning. For text-line segmentation, we used an RU-net which allows a pixel-wise classification to separate text-lines pixels from the background ones. For word segmentation, we resorted to the text-line transcription, as we have not got a ground truth at word level. A BLSTM-CTC (Bidirectional Long Short Term Memory followed by a Connectionist Temporal Classification) is then used to perform the mapping between the transcription and text-line image, avoiding the need of the input segmentation. A CNN (Convolutional Neural Network) precedes the BLST-CTC to extract the features and to feed the BLSTM with the essential of the text-line image. Tested on the standard KHATT Arabic database, the experimental results confirm a segmentation success rate of no less than 96.7% for text-lines and 80.1% for words.
引用
收藏
页码:19 / 24
页数:6
相关论文
共 50 条
  • [1] Segmentation of Arabic Handwritten Documents into Text Lines using Watershed Transform
    Souhar, A.
    Boulid, Y.
    Ameur, ElB.
    Ouagague, Mly. M.
    INTERNATIONAL JOURNAL OF INTERACTIVE MULTIMEDIA AND ARTIFICIAL INTELLIGENCE, 2017, 4 (06): : 96 - 102
  • [2] Handwritten Arabic Documents Segmentation into Text Lines using Seam Carving
    Daldali, M.
    Souhar, A.
    INTERNATIONAL JOURNAL OF INTERACTIVE MULTIMEDIA AND ARTIFICIAL INTELLIGENCE, 2019, 5 (05): : 89 - 96
  • [3] Segmentation of Text-Lines and Words from JPEG Compressed Printed Text Documents Using DCT Coefficients
    Rajesh, Bulla
    Javed, Mohammed
    Nagabhushan, P.
    Osamu, Watanabe
    2020 DATA COMPRESSION CONFERENCE (DCC 2020), 2020, : 389 - 389
  • [4] Segmentation of Arabic handwritten text to lines
    Mokhtari, Younes
    Yousfi, Abdellah
    INTERNATIONAL CONFERENCE ON ADVANCED WIRELESS INFORMATION AND COMMUNICATION TECHNOLOGIES (AWICT 2015), 2015, 73 : 115 - 121
  • [5] Unconstrained Handwritten Arabic Text-lines Segmentation based on AR2U-Net
    Gader, Takwa Ben Aicha
    Echi, Afef Kacem
    2020 17TH INTERNATIONAL CONFERENCE ON FRONTIERS IN HANDWRITING RECOGNITION (ICFHR 2020), 2020, : 349 - 354
  • [6] Segmentation of Historical Handwritten Documents into Text Zones and Text Lines
    Gatos, Basilis
    Louloudis, Georgios
    Stamatopoulos, Nikolaos
    2014 14TH INTERNATIONAL CONFERENCE ON FRONTIERS IN HANDWRITING RECOGNITION (ICFHR), 2014, : 464 - 469
  • [7] Lines segmentation and word extraction of Arabic handwritten text
    Lamsaf, Asmae
    Aitkerroum, Mounir
    Boulaknadel, Siham
    Fakhri, Youssef
    PROCEEDINGS OF THE 3RD INTERNATIONAL CONFERENCE ON SMART CITY APPLICATIONS (SCA'18), 2018,
  • [8] Deep Learning-Based Segmentation of Connected Components in Arabic Handwritten Documents
    Gader, Takwa Ben Aïcha
    Echi, Afef Kacem
    Communications in Computer and Information Science, 2022, 1589 CCIS : 93 - 106
  • [9] Handwritten document image segmentation into text lines and words
    Papavassiliou, Vassilis
    Stafylakis, Themos
    Katsouros, Vassilis
    Carayannis, George
    PATTERN RECOGNITION, 2010, 43 (01) : 369 - 377
  • [10] Detection and Segmentation of Lines and Words in Gurmukhi Handwritten Text
    Kumar, Rajiv
    Singh, Amardeep
    2010 IEEE 2ND INTERNATIONAL ADVANCE COMPUTING CONFERENCE, 2010, : 353 - +