WB meets MHDBDB: The Wenzelsbibel as a Pilot Model for Interoperable Editions
- Hosting organisations
- University of Salzburg
- Responsible persons
- Julia Hintersteiner
- Start
- End
The project “WB meets MHDBDB” focuses on integrating the Wenzelsbibel, a Bohemian prose Bible translation from 1390, into the Middle High German Conceptual Database (MHDBDB). The aim is to establish the Wenzelsbibel as a pilot model for developing interoperable digital editions. For the first time, the edition will be fully lemmatized, POS-tagged, and semantically annotated in a publicly accessible database.
The Wenzelsbibel comprises approximately 150,000 words across five books (Genesis, Exodus, Leviticus, Numbers, Deuteronomy). Its integration into the MHDBDB adheres to the principles of Linked Open Data (LOD) and the FAIR principles (Findable, Accessible, Interoperable, Reusable). The project opens new possibilities for the semantic modelling of Middle High German terms and their interdisciplinary linkage with other research databases.
About the Project
The project is based on a multi-step process for data preparation and integration:
- Development of a sustainable workflow: An automated workflow is being developed to transform the Wenzelsbibel edition data into an MHDBDB-compatible format without loss. This workflow employs Python scripts and LLM technologies to efficiently transform and annotate the data.
- Semantic harmonization: The semantic structure of the Wenzelsbibel is aligned with the ontologies and controlled vocabularies of the MHDBDB. This includes linking key terms to established LOD vocabularies such as Wikidata, CIDOC CRM, and ICONCLASS.
- Automatic lemmatization and POS tagging: Using LLM technologies, the word forms of the Wenzelsbibel are automatically lemmatized and tagged with part-of-speech information. Ambiguous and unmatched forms are further processed through machine-assisted disambiguation.
- Semantic modelling: The Middle High German terms of the Wenzelsbibel are semantically annotated and integrated into the MHDBDB conceptual network. This enables deeper analysis of conceptual usage, narratology, and intertextual connections.
The project contributes to strengthening digital research infrastructures and provides a methodological foundation for the reuse of complex edition data. The results aim to enhance the reusability of digital editions and open new research perspectives on conceptual usage and intertextual connections.