首页> 美国政府科技报告 >Arabic Optical Character Recognition (OCR) Evaluation in Order to Develop a Post-OCR Module

【24h】

Arabic Optical Character Recognition (OCR) Evaluation in Order to Develop a Post-OCR Module

机译：阿拉伯语光学字符识别（OCR）评估，以开发后OCR模块

获取原文

页面导航

摘要
著录项
引文网络
相似文献
相关主题

摘要

Optical character recognition (OCR) is the process of converting an image of a document into text. While progress in OCR research has enabled low error rates for English text in low-noise images, performance is still poor for noisy images and documents in other languages. We intend to create a post-OCR processing module for noisy Arabic documents which can correct OCR errors before passing the resulting Arabic text to a translation system. To this end, we are evaluating an Arabic-script OCR engine on documents with the same content but varying levels of image quality. We have found that OCR text accuracy can be improved with different stages of pre-OCR image processing: (1) filtering out low-contrast images to avoid hallucination of characters, (2) removing marks from images with cleanup software to prevent their misrecognition, and (3) zoning multi-column images with segmentation software to enable recognition of all zones. The specific errors observed in OCR will form the basis of training data for our post-OCR correction module.

著录项

作者
Kjersten, B.;
展开▼
作者单位

展开▼
年度 2011
页码 1-18
总页数 18
原文格式 PDF
正文语种 eng
中图分类工业技术;
关键词
Documents ; Image processing ; Optical character recognition ; Optical images ; Computer programs ; English language ; Errors ; Low rate ; Quality;

机译：文件;图像处理;光学字符识别;光学图像;计算机程序;英语;错误;低速率;质量;

相似文献

外文文献
中文文献
专利

1. A Proposed OCR Algorithm for the Recognition of Handwritten Arabic Characters [J] . Ahmed T. Sahlol, Cheng Y. Suen, Mohammed R. Elbasyouni, Journal of Pattern Recognition and Intelligent Systems . 2014,第1期

机译：一种用于手写阿拉伯字符识别的OCR算法
2. The use of Hartley transform in OCR with application to printed Arabic character recognition [J] . Sabri A. Mahmoud, Ashraf S. Mahmoud Pattern Analysis and Applications . 2009,第4期

机译：Hartley变换在OCR中的应用及其在印刷阿拉伯字符识别中的应用
3. An Arabic optical character recognition system using recognition-based segmentation [J] . Cheung A., Bergmann NW., Bennamoun M. Pattern Recognition: The Journal of the Pattern Recognition Society . 2001,第2期

机译：使用基于识别的分割的阿拉伯光学字符识别系统
4. A survey on Arabic Optical Character Recognition and an isolated handwritten Arabic Character Recognition algorithm using encoded freeman chain code [C] . Hassan Althobaiti, Chao Lu Annual Conference on Information Sciences and Systems . 2017

机译：使用编码的弗里曼链码的阿拉伯语光学字符识别研究和孤立的手写阿拉伯语字符识别算法
5. A multimodal fusion approach for automatic postal address recognition system using Optical Character Recognition (OCR) and Automatic Speech Recognition (ASR) techniques. [D] . Singh, Amriteshwar. 2011

机译：一种使用光学字符识别（OCR）和自动语音识别（ASR）技术的自动邮政地址识别系统的多模式融合方法。
6. The use of Optical Character Recognition (OCR) in the digitisation of herbarium specimen labels [O] . Robyn E. Drinkwater, Robert W. N. Cubey, Elspeth M. Haston 2014

机译：光学字符识别（OCR）在植物标本标签数字化中的使用
7. FAWA: Fast Adversarial Watermark Attack on Optical Character Recognition (OCR) Systems [O] . Lu Chen, Jiao Sun, Wei Xu 2021

机译：Fawa：光学字符识别（OCR）系统的快速逆境水印攻击

Arabic Optical Character Recognition (OCR) Evaluation in Order to Develop a Post-OCR Module

摘要

著录项

引文网络

相似文献

相关主题

期刊订阅