Strokelets: A Learned Multi-Scale Mid-Level Representation for Scene Text Recognition

被引:71
|
作者
Bai, Xiang [1 ]
Yao, Cong [1 ]
Liu, Wenyu [1 ]
机构
[1] Huazhong Univ Sci & Technol, Sch Elect Informat & Commun, Wuhan 430074, Peoples R China
基金
中国国家自然科学基金;
关键词
Scene text recognition; scene text detection; mid-level representation; multi-scale representation; natural images; OBJECT DETECTION; DESCRIPTOR; VISION; MODEL;
D O I
10.1109/TIP.2016.2555080
中图分类号
TP18 [人工智能理论];
学科分类号
081104 ; 0812 ; 0835 ; 1405 ;
摘要
In this paper, we are concerned with the problem of automatic scene text recognition, which involves localizing and reading characters in natural images. We investigate this problem from the perspective of representation and propose a novel multi-scale representation, which leads to accurate, robust character identification and recognition. This representation consists of a set of mid-level primitives, termed strokelets, which capture the underlying substructures of characters at different granularities. The Strokelets possess four distinctive advantages: 1) usability: automatically learned from character level annotations; 2) robustness: insensitive to interference factors; 3) generality: applicable to variant languages; and 4) expressivity: effective at describing characters. Extensive experiments on standard benchmarks verify the advantages of the strokelets and demonstrate the effectiveness of the text recognition algorithm built upon the strokelets. Moreover, we show the method to incorporate the strokelets to improve the performance of scene text detection.
引用
收藏
页码:2789 / 2802
页数:14
相关论文
共 50 条
  • [31] Action Recognition with Discriminative Mid-Level Features
    Liu, Cuiwei
    Kong, Yu
    Wu, Xinxiao
    Jia, Yunde
    2012 21ST INTERNATIONAL CONFERENCE ON PATTERN RECOGNITION (ICPR 2012), 2012, : 3366 - 3369
  • [32] Multi-Scale Deep Representation Aggregation for Vein Recognition
    Pan, Zaiyu
    Wang, Jun
    Wang, Guoqing
    Zhu, Jihong
    IEEE TRANSACTIONS ON INFORMATION FORENSICS AND SECURITY, 2021, 16 : 1 - 15
  • [33] SCENE TEXT DETECTION BASED ON MULTI-SCALE SWT AND EDGE FILTERING
    Feng, Yuanyuan
    Song, Yonghong
    YualinZhang
    2016 23RD INTERNATIONAL CONFERENCE ON PATTERN RECOGNITION (ICPR), 2016, : 645 - 650
  • [34] Pyrboxes: An efficient multi-scale scene text detector with feature pyramids
    Sheng, Fenfen
    Chen, Zhineng
    Zhang, Wei
    Xu, Bo
    PATTERN RECOGNITION LETTERS, 2019, 125 : 228 - 234
  • [35] Multi-Scale Scene Text Detection Based on Convolutional Neural Network
    Lu, Yan-Feng
    Zhang, Ai-Xuan
    Li, Yi
    Yu, Qian-Hui
    Qiao, Hong
    2019 CHINESE AUTOMATION CONGRESS (CAC2019), 2019, : 583 - 587
  • [36] A generic mid-level representation for semantic video analysis
    Tang, Q
    Lim, JH
    Jin, JS
    Sun, HP
    Tian, Q
    ICIP: 2004 INTERNATIONAL CONFERENCE ON IMAGE PROCESSING, VOLS 1- 5, 2004, : 629 - 632
  • [37] Grey level image components for multi-scale representation
    Ramella, G
    di Baja, GS
    PROGRESS IN PATTERN RECOGNITION, IMAGE ANALYSIS AND APPLICATIONS, 2004, 3287 : 574 - 581
  • [38] Pitch Contours as a Mid-Level Representation for Music Informatics
    Bittner, Rachel M.
    Salamon, Justin
    Bosch, Juan J.
    Bello, Juan P.
    2017 AES INTERNATIONAL CONFERENCE ON SEMANTIC AUDIO, 2017,
  • [39] Supervised Mid-Level Features for Word Image Representation
    Gordo, Albert
    2015 IEEE CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR), 2015, : 2956 - 2964
  • [40] A Mid-Level Representation of Visual Structures for Video Compression
    Georgiadis, Georgios
    Soatto, Stefano
    2016 IEEE WINTER CONFERENCE ON APPLICATIONS OF COMPUTER VISION (WACV 2016), 2016,