KuroNet: Pre-Modern Japanese Kuzushiji Character Recognition with Deep Learning

机译：KuroNet：带深度学习的近现代日语Kuzushiji字符识别

获取原文

页面导航

摘要
著录项
相似文献
相关主题

摘要

Kuzushiji, a cursive writing style, had been used in Japan for over a thousand years starting from the 8th century. Over 3 millions books on a diverse array of topics, such as literature, science, mathematics and even cooking are preserved. However, following a change to the Japanese writing system in 1900, Kuzushiji has not been included in regular school curricula. Therefore, most Japanese natives nowadays cannot read books written or printed just 150 years ago. Museums and libraries have invested a great deal of effort into creating digital copies of these historical documents as a safeguard against fires, earthquakes and tsunamis. The result has been datasets with hundreds of millions of photographs of historical documents which can only be read by a small number of specially trained experts. Thus there has been a great deal of interest in using Machine Learning to automatically recognize these historical texts and transcribe them into modern Japanese characters. Nevertheless, several challenges in Kuzushiji recognition have made the performance of existing systems extremely poor. To tackle these challenges, we propose KuroNet, a new end-to-end model which jointly recognizes an entire page of text by using a residual U-Net architecture which predicts the location and identity of all characters given a page of text (without any pre-processing). This allows the model to handle long range context, large vocabularies, and non-standardized character layouts. We demonstrate that our system is able to successfully recognize a large fraction of pre-modern Japanese documents, but also explore areas where our system is limited and suggest directions for future work.

机译：从8世纪开始，Kuzushiji是一种草书写作风格，在日本已经使用了1000多年。保留了超过300万本涉及各种主题的书籍，例如文学，科学，数学甚至烹饪。但是，随着1900年日本文字系统的变化，《九十九路》没有被纳入普通学校的课程中。因此，当今大多数日本人都无法阅读150年前的书面或印刷书籍。博物馆和图书馆投入了大量精力来创建这些历史文献的数字副本，以防火灾，地震和海啸。结果是获得了具有数亿张历史文献照片的数据集，而这些照片只能由少数经过特殊培训的专家来阅读。因此，使用机器学习自动识别这些历史文本并将其转录成现代日语字符引起了极大的兴趣。尽管如此，在Kuzushiji识别方面的一些挑战已使现有系统的性能极差。为了应对这些挑战，我们提出了一种新的端到端模型KuroNet，该模型可以通过使用残留的U-Net架构共同识别整个文本页面，该体系结构可以预测给定文本页面的所有字符的位置和标识（不包含任何字符）。预处理）。这使模型可以处理远程上下文，大量词汇和非标准化字符布局。我们证明了我们的系统能够成功识别出大部分的前现代日语文件，而且还探索了我们的系统受限的领域并为未来的工作提供了建议。

著录项

来源
《International Conference on Document Analysis and Recognition》|2019年|607-614|共8页
会议地点
作者
Tarin Clanuwat; Alex Lamb; Asanobu Kitamoto;
展开▼
作者单位

展开▼
会议组织
原文格式 PDF
正文语种
中图分类
关键词
Character recognition; Writing; Task analysis; Printing; Training; Text recognition; Image resolution;

机译：字符识别;书写;任务分析;打印;培训;文本识别;图像分辨率;
入库时间 2022-08-26 14:34:51

相似文献

外文文献
中文文献
专利

1. Real-time Automated Detection and Recognition of Nigerian License Plates via Deep Learning Single Shot Detection and Optical Character Recognition [J] . Kayode David Adedayo, Ayomide Oluwaseyi Agunloye Computer and Information Science . 2021,第4期

机译：通过深度学习单次检测和光学字符识别，实时自动检测和识别尼日利亚牌照
2. Handwritten Urdu character recognition via images using different machine learning and deep learning techniques [J] . M Ameen Chhajro, Hadeeb Khan, Farrukh Khan, Indian Journal of Science and Technology . 2020,第17期

机译：手写Urdu字符识别通过使用不同的机器学习和深度学习技术
3. Learning representation hierarchies by sharing visual features: a computational investigation of Persian character recognition with unsupervised deep learning [J] . Sadeghi Zahra, Testolin Alberto Cognitive processing . 2017,第3期

机译：通过分享视觉特征学习代表层次结构：与无监督深度学习的波斯字符识别的计算调查
4. KuroNet: Pre-Modern Japanese Kuzushiji Character Recognition with Deep Learning [C] . Tarin Clanuwat, Alex Lamb, Asanobu Kitamoto International Conference on Document Analysis and Recognition . 2019

机译：Kuronet：前现代日本库祖史，具有深入学习的人物识别
5. Sequence-to-Sequence Learning using Deep Learning for Optical Character Recognition [D] . Mishra, Vishal Vijayshankar. 2017

机译：使用深度学习进行光学字符识别的序列到序列学习
6. Character Recognition of Components Mounted on Printed Circuit Board Using Deep Learning [O] . Sumyung Gang, Ndayishimiye Fabrice, Daewon Chung, 2021

机译：使用深度学习安装在印刷电路板上的部件的字符识别
7. KuroNet: Pre-Modern Japanese Kuzushiji Character Recognition with Deep Learning [O] . Tarin Clanuwat, Alex Lamb, Asanobu Kitamoto 2019

机译：Kuronet：前现代日本库祖史，具有深入学习的人物识别

KuroNet: Pre-Modern Japanese Kuzushiji Character Recognition with Deep Learning

摘要

著录项

相似文献

相关主题

期刊订阅