ACDH Tool Gallery 12.3
Introduction to Named Entity Recognition (NER)
Wann: Mittwoch, 23. September 2026, 10:00 - 16:00
Wo: Seminarraum 1, Erdgeschoss / Innenhof
Österreichische Akademie der Wissenschaften
Bäckerstraße 13, 1010 Wien, Österreich
Anmeldung: Bitte registrieren Sie sich vorab über das Anmeldeformular .
Vortragssprache: Englisch
Diese Tool-Galerie wurde konzipiert, um Forschende und Fachleute systematisch in die vielfältige Methodologie der Named Entity Recognition (NER) einzuführen. Der Workshop bietet eine umfassende Auseinandersetzung mit klassischen maschinellen Lernverfahren sowie mit modernen, transformerbasierten Modellen wie BERT und den Pipelines von SpaCy. Neben einer fundierten theoretischen Einführung liegt ein zentraler Fokus auf praxisorientierten Übungen, in denen die Teilnehmenden die verschiedenen Ansätze anhand unterschiedlicher Datensätze evaluieren. Dabei werden zentrale Aspekte wie Genauigkeit, Rechenaufwand, domänenspezifische Anpassungsfähigkeit und Anforderungen an Annotationen kritisch analysiert und gegeneinander abgewogen. Darüber hinaus widmet sich das Curriculum fortgeschrittenen Themen, darunter Zero-Shot- und Few-Shot-NER mit großen Sprachmodellen, Strategien zur domänenspezifischen Feinabstimmung sowie der Einsatz von Ensemble-Methoden. Sämtliche Workshop-Materialien, einschließlich annotierter Notebooks und Evaluierungsskripte, werden als Open-Access-Ressourcen bereitgestellt, um die Reproduzierbarkeit der Ergebnisse zu gewährleisten und eine breitere Anwendung im akademischen Kontext zu fördern.
Curriculum:
9:45 – 10:00: Doors Open
10:00 – 10:45: Introduction to NER
10:45 – 11:00: Coffee Break
11:00 – 11:45: Classical Machine Learning Methods – Theory
11:45 – 12:30: Classical Machine Learning Methods – Hands-On
12:30 – 13:30: Lunch Break
13:30 – 14:15: Transformer based Methods – Theory
14:15 – 15:00: Transformer based Methods – Hands-On
15:00 – 15:45: Wrap Up – Open Discussion
Requirements
- Bring your laptop
- Basic experience in working with Jupyter notebooks and Google Colab are useful
- Basic Python
Team
Daniel Elsner
… studied development studies at the University of Vienna and received a Bachelor of Arts degree. Continuing with historical studies at the University of Vienna, the master’s degree program he enrolled in had a focus on digital humanities and global history. His master thesis, with the title „A global perspective: Singapore migration and labor recruitment networks“, applied several digital methods. Among them, GIS is for mapping migration and statistical analysis with empirical census data. The master’s program was concluded by receiving a Master of Arts degree in early 2021.
As a research data and software engineer (RDE/RSE) at the ACDH (then ACDH-CH), he supports research projects with the development and implementation of digital methods for digital editions, corpora, linked (open) data, data curation and machine learning (AI).
He collaborates in projects of the research units DH Research & Infrastructure (2020 – today), Musicology (2021 – 2024), Literary & Print Culture Studies (2024 – today) and Linguistics (2025 – today).
Stefan Resch
… is a Software Architect / NLP Analyst in the research unit DH Research & Infrastructure . His stack includes Python, Bash, Linux, Docker, Pandas, NumPy, spaCy, Django, RDF and SPARQL. In his current project CLSInfra he is a Software Architect, responsible for the implementation of VELD , an architecture enabling stable and reproducable interoperability between heterogenous tools and data sets with a focus on NLP pipelining. As Backend Developer of APIS , he is responsible for the major refactoring of the inner architecture of the APIS app , increasing the flexibility of its business logic towards arbitrary project ontologies, and improving the application’s compliancy with RDF paradigms.
In his former project MARA , as Software Architect and Natural Language Processing Analyst, he supported the researchers in media corpus analysis by providing them with iterative Supervised Machine Learning workflows. For this, he implemented the MARA NLP Suite , encapsulating the entire data cycle and providing the public with reproducability of its data and code. In the project SOLA he worked as Backend Developer and Ontology Designer, adapting the entity database APIS towards the research field of investigating juristical texts of early christianity. As Backend Developer and Ontology Designer of the Jelinek Werkverzeichnis Online , he adapted the entity database APIS towards the extensive bibliography of the Austrian writer Elfriede Jelinek and her professional and artistic connections throughout her career. As Data Analyst and supportive Backend Developer of Parthenos , he assured the harmonization results of data aggregation pipelines on heterogeneous data sets originating from diverse research archives.
ACDH Website