Showing posts with label OCR. Show all posts
Showing posts with label OCR. Show all posts

Thursday, April 25, 2024

 



"Funded through two grants from The Andrew W. Mellon Foundation, Phase One of the Open Islamicate Texts Initiative Arabic-script OCR Catalyst Project (OpenITI AOCP) is the first undertaking of its kind to tackle the technical and organizational barriers that historically have stymied the development of Arabic-script OCR and digital text production for Islamicate Studies.
OpenITI AOCP is led by an interdisciplinary team of humanities, computer science, and digital humanities co-principal investigators from Roshan Institute for Persian Studies at the University of Maryland, College Park, Northeastern University’s NULab for Texts, Maps, and Networks, the Aga Khan University’s Institute for the Study of Muslim Civilisations in London, and the Maryland Institute for Technology in the Humanities at the University of Maryland, College Park. We are proud to partner with the SHARIAsource project of the Program in Islamic Law at Harvard Law School and the eScripta project of Université Paris Sciences et Lettres for the technical development portion of the project.

The primary technical goal of the first phase of OpenITI AOCP is to achieve ≥97% character accuracy rates (CARs) for OCR on the most used Persian and Arabic print typefaces. 

The second major deliverable of OpenITI AOCP is an open-source and user-friendly digital text production pipeline for Persian and Arabic texts."


Friday, April 27, 2018

Arabic Scientific Manuscripts of the British Library - OCR crowdsourcing project

 https://fromthepage.com/bldigital/arabic-scientific-manuscripts

Arabic Scientific Manuscripts of the British Library - Collaborative transcription project:
https://fromthepage.com/bldigital/arabic-scientific-manuscripts

"Help advance research into automatic text recognition technologies for historical Arabic handwritten texts. Below you will find selected pages from some of the British Library's most important Arabic Scientific Manuscripts. We need your help to transcribe them in order to create a freely available ground truth datataset for anyone wishing to advance the state-of-the-art in optical character recognition (OCR) technology for handwriting."

Read more on the effort:
http://blogs.bl.uk/digital-scholarship/2018/03/arabic-handwrittten-ocr.html


Thursday, October 6, 2016

Open Source Arabic OCR

Working paper :
Important New Developments in Arabographic Optical Character Recognition (OCR) by
Maxim Romanov, Matthew Thomas Miller, Sarah Bowen Savant, and Benjamin Kiessling.
 Highlights from the paper:

"The OpenITI team—building on the foundational open-source OCR work of the Leipzig University’s (LU) Alexander von Humboldt Chair for Digital Humanities—has achieved Optical Character Recognition (OCR) accuracy rates for classical Arabic-script texts in the high nineties " 

"The specific type of OCR software that we employed in our tests is an
open-source OCR program called Kraken, which was developed by Benjamin
Kiessling at Leipzig University’s Alexander von Humboldt Chair for Digital
Humanities. Unlike more traditional OCR approaches, Kraken relies on a neural
network—which mimics the way we learn—to recognize letters in the images of
entire lines of text without trying first to segment lines into words and then words
into letters."


"The most important advantage of Kraken is that its workflow allows one to train new
models relatively easily, including text-specific ones. In a nutshell, the process of
training requires a transcription of approximately 800 lines (the number will vary
depending on the complexity of the typeface) aligned with images of these lines as
they appear in the printed edition."


 "The two rounds of testing presented here indicate that with a fairly modest amount
of gold standard training data (~800–1,000 lines) Kraken is consistently able to
produce OCR results for Arabic-script documents that achieve accuracy rates in the
high nineties."


"In the long term, we will are also planning to train models for other Islamicate languages (Ottoman Turkish,Urdu, Syriac, etc.). Our hope is that an easy-to-use and effective OCR pipeline will allow us all—collectively—to significantly enrich our collection of digital Islamicate texts and thereby enable us to understand better this fascinating and understudied textual tradition."