Please use this identifier to cite or link to this item:
https://scholarhub.balamand.edu.lb/handle/uob/393
Title: | Arabic handwritten document preprocessing and recognition | Authors: | Chammas, Edgar Mokbel, Chafic Likforman-Sulem, Laurence |
Affiliations: | Department of Electrical Engineering | Keywords: | Hidden Markov models Image recognition Optical imaging Optical reflection Text recognition |
Subjects: | Image segmentation Writing |
Issue Date: | 2015 | Part of: | 2015 13th International Conference on Document Analysis and Recognition (ICDAR) | Start page: | 451 | End page: | 455 | Conference: | International Conference on Document Analysis and Recognition (ICDAR) (13th : 23-26 Aug 2015 : Tunisia) | Abstract: | Arabic handwritten documents present specific challenges due to the cursive nature of the writing and the presence of diacritical marks. Moreover, one of the largest labeled database of Arabic handwritten documents, the OpenHart-NIST database includes specific noise, namely guidelines, that has to be addressed. We propose several approaches to process these documents. First a guideline detection approach has been developed, based on K-means, that detects the documents that include guidelines. We then propose a series of preprocessing at text-line level to reduce the noise effects. For text-lines including guidelines, a guideline removal preprocessing is described and existing keystroke restoration approaches are assessed. In addition, we propose a preprocessing that combines noise removal and deskewing by removing line fragments from neighboring text lines, while searching for the principal orientation of the text-line. We provide recognition results, showing the significant improvement brought by the proposed processings. |
URI: | https://scholarhub.balamand.edu.lb/handle/uob/393 | Ezproxy URL: | Link to full text | Type: | Conference Paper |
Appears in Collections: | Department of Electrical Engineering |
Show full item record
Google ScholarTM
Check
Items in DSpace are protected by copyright, with all rights reserved, unless otherwise indicated.