BigScience AI4LAM presentation
Abstract
In this work, we explore whether the recently demonstrated zero-shot abilities of the T0 model extend to Named Entity Recognition for out-of-distribution languages and time periods. Using a historical newspaper corpus in 3 languages as test-bed, we use prompts to extract possible named entities. Our results show that a naive approach for prompt-based zero-shot multilingual Named Entity Recognition is error-prone, but highlights the potential of such an approach for historical languages lacking labeled datasets. Moreover, we also find that T0-like models can be probed to predict the publication date and language of a document, which could be very relevant for the study of historical texts.
Authors 6
-
The University of Western Australia
Affiliation as printed
The University of Western Australia ,
UWA - The University of Western Australia (35 Stirling Highway Perth WA 6009 Australia - Australia)
-
The National Library of Norway
Affiliation as printed
National Library of Norway ,
National Library of Norway (Norway)
-
Centre Inria de Paris · ALMANACH: Modélisation et analyse linguistique automatique et humanités computationnelles
Affiliation as printed
Inria Paris
ALMAnaCH - Automatic Language Modelling and ANAlysis & Computational Humanities (France)
-
Enrique Manjavacas Aachen
Affiliation as printed
Leiden University ,
Universiteit Leiden = Leiden University (Leiden University | 2300 RA Leiden The Netherlands - Netherlands)
-
Affiliation as printed
Bayerische Staatsbibliothek ,
BSB - Bayerische Staatsbibliothek (Bibliothèque de l'État de Bavière à Münich - Germany)
-
Affiliation as printed
British Library
British Library (96 Euston Road London NW1 2DB - United Kingdom)
Cited by 0 stored of 0
No patents citing this paper on Lens.org (checked 2026-10-11).