Pilot Corpus for Tokenization and Lemmatization of Pannonian Rusyn – A First Step Toward Open Language Resources
- Hosting organisations
- University of Graz
- Responsible persons
- Marko Simonovic
- Start
- End
Pannonian Rusyn is a regional minority language spoken mainly in parts of Serbia and Croatia. Although the language is taught at all levels of education in Serbia and is also present in the media of the language community, very few basic digital language resources are currently available.
The project addresses this gap by developing the first linguistically annotated pilot corpus for Pannonian Rusyn. The corpus will provide a foundation for further work in natural language processing, corpus linguistics and the Digital Humanities. At the same time, the project aims to connect the language more closely with open European research infrastructures.
About the Project
The pilot corpus is based on entries from the Pannonian Rusyn Wikipedia. The technical workflow was defined during the first project meeting on 31 July 2025 and comprises two main stages:
- automatic tokenisation and preprocessing using the CLASSLA-Stanza pipeline,
- manual correction and linguistic annotation of the processed data.
The planned pilot corpus will comprise around 10,000 tokens. The data will be manually annotated for tokenisation, lemmatisation and parts of speech. A simplified tagset will be used for part-of-speech tagging.
The processed data will be made available in CoNLL-U format. Annotation guidelines and a short methodological report will document the corpus-building process and the linguistic decisions taken during annotation. This will ensure that the results are reproducible, extensible and compatible with existing CLARIN standards.
Ana Rimar Simunović, Aleksander Mudri and Aleksandra Dejanović from the Department of Rusyn Studies in Novi Sad are contributing to the annotation process. Initial project results were presented by Marko Simonović in October 2025 at the CLASSLA-Express 1.0 Plus “Corpora meets AI” event in Graz.
The academic discussion focused, among other topics, on the lemmatisation of grammatically mixed categories—particularly adjectival and adverbial participles—and on their classification within a broader comparison of Slavic languages.
After the project has been completed, the pilot corpus will be published under an open licence and made permanently accessible through CLARIN.SI or a suitable institutional repository. The project will thus provide a first openly reusable resource for the digital processing and study of Pannonian Rusyn.