Skip to main content

OCR/HTR Workshop for Under-resourced and Under-represented Languages in Digital Humanities

Hosting organisations
Central European University (CEU)
Responsible persons
Alíz Horváth
Start
End

Many digital humanities projects depend on the reliable conversion of printed or handwritten sources into machine-readable text. While OCR and HTR methods are already well established for some widely used languages and scripts, substantial technical and methodological challenges remain for under-represented and under-resourced languages, especially those using non-Latin scripts. The workshop addressed this need by bringing together researchers, library professionals and technical experts to discuss experiences, challenges and possible solutions in the field of text recognition. 

About the Project 

The two-day workshop took place on 3 and 4 October 2025 at Central European University and was funded by CLARIAH-AT and the FWF Cluster of Excellence EurAsian Transformations . It brought together 23 on-site participants from Austria, Germany, France, Israel, Japan and the United States, as well as more than 20 online participants from further countries. 

The workshop focused on OCR and HTR workflows for languages and scripts that are often insufficiently supported by existing digital infrastructures. These included, among others, Devanagari, Hebrew, Kanbun, Sanskrit and Newar, ancient Greek, Armeno-Turkish, Garshuni Malayalam, classical Chinese, Ottoman Turkish and Tibetan. Recurring issues discussed during the event included layout and alignment, character variation, recognition accuracy, error correction, right-to-left writing, vertical alignment, and bi- or multidirectional text structures. 

The programme combined short expert presentations, thematic panels, extensive discussions, a technical roundtable and a forum addressing library and infrastructure perspectives. One important aim was to enable exchange among people from different disciplinary, linguistic and institutional backgrounds, since similar technical problems are often addressed separately within distinct scholarly communities. 

As a tangible outcome, a Zenodo community was created to collect presentation slides and further materials. In addition, a collaborative writing sprint produced a zine entitled “Ten annoying things about digitizing under-resourced languages (And what might help fixing them)”. It summarises key challenges and possible solutions in a concise and accessible format and is intended to raise awareness among different stakeholders. The workshop also strengthened networking among the participants and opened up follow-up opportunities with the SCOOP community (Source Codes of the Past), which is planning a further meeting in Vienna in 2026.