A

Sequence-Level Unsupervised Training in Speech Recognition: A Theoretical Study

IEEE International Conference on Acoustics Speech and Signal Processing, pp. 2571–2575

Abstract

Unsupervised speech recognition is a task of training a speech recognition model with unpaired data. To determine when and how unsupervised speech recognition can succeed, and how classification error relates to candidate training objectives, we develop a theoretical framework for unsupervised speech recognition grounded in classification error bounds. We introduce two conditions under which unsupervised speech recognition is possible. The necessity of these conditions are also discussed. Under these conditions, we derive a classification error bound for unsupervised speech recognition and validate this bound in simulations. Motivated by this bound, we propose a single-stage sequence-level cross-entropy loss for unsupervised speech recognition.

Authors 4

  1. RWTH Aachen University

    Affiliation as printed

    RWTH Aachen University,Human Language Technology and Pattern Recognition,Computer Science Department,Germany

  2. RWTH Aachen University

    Affiliation as printed

    RWTH Aachen University,Human Language Technology and Pattern Recognition,Computer Science Department,Germany

  3. RWTH Aachen University

    Affiliation as printed

    RWTH Aachen University,Human Language Technology and Pattern Recognition,Computer Science Department,Germany

  4. RWTH Aachen University

    Affiliation as printed

    RWTH Aachen University,Human Language Technology and Pattern Recognition,Computer Science Department,Germany

Cited by 0 stored of 0

No patents citing this paper on Lens.org (checked 2026-10-06).

References 11

11 results