Text Normalization Method for Arabic Handwritten Script
Autor: | Bilal Bataineh, Khairuddin Omar, Ashraf Abu-Ein, Waleed Abu-Ain, Siti Norul Huda Sheikh Abdullah, Tarik Abu-Ain |
---|---|
Rok vydání: | 2013 |
Předmět: |
Information Systems and Management
General Computer Science Computer science business.industry Text segmentation Feature extraction Pattern recognition computer.software_genre Set (abstract data type) ComputingMethodologies_DOCUMENTANDTEXTPROCESSING Benchmark (computing) Text normalization Preprocessor Segmentation Artificial intelligence Electrical and Electronic Engineering Line (text file) business computer Natural language processing |
Zdroj: | Journal of ICT Research and Applications. 7:164-175 |
ISSN: | 2338-5499 2337-5787 |
DOI: | 10.5614/itbj.ict.res.appl.2013.7.2.5 |
Popis: | Text normalization is an important technique in document image analysis and recognition. It consists of many preprocessing stages, which include slope correction, text padding, skew correction, and straight the writing line. In this side, text normalization has an important role in many procedures such as text segmentation, feature extraction and characters recognition. In the present article, a new method for text baseline detection, straightening, and slant correction for Arabic handwritten texts is proposed. The method comprises a set of sequential steps: first components segmentation is done followed by components text thinning; then, the direction features of the skeletons are extracted, and the candidate baseline regions are determined. After that, selection of the correct baseline region is done, and finally, the baselines of all components are aligned with the writing line. The experiments are conducted on IFN/ENIT benchmark Arabic dataset. The results show that the proposed method has a promising and encouraging performance. |
Databáze: | OpenAIRE |
Externí odkaz: |