A

BigScience AI4LAM presentation

arXiv (Cornell University)

Abstract

In this work, we explore whether the recently demonstrated zero-shot abilities of the T0 model extend to Named Entity Recognition for out-of-distribution languages and time periods. Using a historical newspaper corpus in 3 languages as test-bed, we use prompts to extract possible named entities. Our results show that a naive approach for prompt-based zero-shot multilingual Named Entity Recognition is error-prone, but highlights the potential of such an approach for historical languages lacking labeled datasets. Moreover, we also find that T0-like models can be probed to predict the publication date and language of a document, which could be very relevant for the study of historical texts.

Authors 6

  1. The University of Western Australia

    Affiliation as printed

    The University of Western Australia ,

    UWA - The University of Western Australia (35 Stirling Highway Perth WA 6009 Australia - Australia)

  2. The National Library of Norway

    Affiliation as printed

    National Library of Norway ,

    National Library of Norway (Norway)

  3. Centre Inria de Paris · ALMANACH: Modélisation et analyse linguistique automatique et humanités computationnelles

    Affiliation as printed

    Inria Paris

    ALMAnaCH - Automatic Language Modelling and ANAlysis & Computational Humanities (France)

  4. Leiden University

    Affiliation as printed

    Leiden University ,

    Universiteit Leiden = Leiden University (Leiden University | 2300 RA Leiden The Netherlands - Netherlands)

  5. Bavarian State Library

    Affiliation as printed

    Bayerische Staatsbibliothek ,

    BSB - Bayerische Staatsbibliothek (Bibliothèque de l'État de Bavière à Münich - Germany)

  6. British Library

    Affiliation as printed

    British Library

    British Library (96 Euston Road London NW1 2DB - United Kingdom)

Cited by 0 stored of 0

No patents citing this paper on Lens.org (checked 2026-10-11).

References 0