Skip to main content

Automated Recording and Semantic Analysis of Association Dissolutions in the "Official Gazette" 1938–1940

Hosting organisations
University of Vienna
Responsible persons
Markus Stumpf and Martin Gasteiner
Start
End

Following Austria’s “Anschluss” in 1938, numerous associations, organisations and federations were forcibly dissolved by the so-called Stillhaltekommissar. The corresponding decrees were published in the “Amtliches Nachrichtenblatt”. They constitute an important source for research into the systematic dismantling of Austrian civil society and the persecution and expropriation of organisations, particularly Jewish organisations, under National Socialism. 

These sources have not yet been systematically processed in digital form. The project therefore develops an AI-supported workflow for automatically capturing, structuring and semantically enriching the information on dissolved associations contained in the newsletters. The resulting data will be made sustainably reusable for historical research, remembrance initiatives and digital research infrastructures. 

About the Project 

The project focuses on developing and testing a multi-stage processing chain for the 1938 to 1940 volumes of the “Amtliches Nachrichtenblatt”. It combines Optical Character Recognition (OCR), AI-supported information extraction, quality control and structured data modelling. 

The work includes: 

  • Securing, reviewing and analysing the relevant volumes as well as their formats, layout variants and typical formulations 
  • Applying and optimising OCR methods to prepare the digitised sources 
  • Using AI-supported methods to extract key information such as association names, places, dates of dissolution, stated reasons and information on the disposal of association assets 
  • Reviewing the extracted information and transferring it into a structured data model 
  • Preparing export formats and interfaces for existing research and remembrance infrastructures 

Since the start of the project, the relevant volumes have been secured and examined with regard to their formal and textual structures. On this basis, a processing chain combining OCR, preprocessing, information extraction, quality control and data export has been designed. Options for linking the extracted information to authority data and controlled vocabularies such as the Integrated Authority File (GND) and Wikidata have also been explored. 

As an interim result, a technical prototype covering essential parts of the extraction chain is available. Under manual supervision, it enables the preliminary processing of individual issues of the “Amtliches Nachrichtenblatt”. Initial test datasets, some of which have been manually reviewed and curated, have already been recorded in the intended data model. They confirm the general suitability of the approach for the scalable processing of the complete source collection. 

The configuration and alignment of the individual processing steps are expected to be completed in early 2026. Script-based processing of the complete collection is planned for February to early March 2026. Whether a standalone user interface for end users will also be developed remains open. 

The code, scripts and prompts developed within the project will be documented and published in a GitLab repository. This will make the workflows and prompt structures transparent, reproducible and reusable for comparable projects. The project also plans to establish a structured database that could eventually be linked to existing infrastructures of the National Fund. 

The project thus contributes to the FAIR-oriented preparation of historical sources, the digital study of National Socialist persecution and expropriation policies, and the further development of AI-supported source-processing methods in the Digital Humanities.