On the Relevance of Phoneme Duration Variability of Synthesized Training Data for Automatic Speech Recognition
IEEE Automatic Speech Recognition and Understanding Workshop (ASRU), pp. 1–8
Abstract
Synthetic data generated by text-to-speech (TTS) systems can be used to improve automatic speech recognition (ASR) systems in low-resource or domain mismatch tasks. It has been shown that TTS-generated outputs still do not have the same qualities as real data. In this work we focus on the temporal structure of synthetic data and its relation to ASR training. By using a novel oracle setup we show how much the degradation of synthetic data quality is influenced by duration modeling in non-autoregressive (NAR) TTS. To get reference phoneme durations we use two common alignment methods, a hidden Markov Gaussian-mixture model (HMM-GMM) aligner and a neural connectionist temporal classification (CTC) aligner. Using a simple algorithm based on random walks we shift phoneme duration distributions of the TTS system closer to real durations, resulting in an improvement of an ASR system using synthetic data in a semi-supervised setting.
Authors 3
-
Affiliation as printed
RWTH Aachen University,Machine Learning and Human Language Technology,Computer Science Departement,Germany
AppTek GmbH, Germany
Computer Science Departement, Machine Learning and Human Language Technology, RWTH Aachen University, Germany
-
Affiliation as printed
RWTH Aachen University,Machine Learning and Human Language Technology,Computer Science Departement,Germany
Computer Science Departement, Machine Learning and Human Language Technology, RWTH Aachen University, Germany
-
Affiliation as printed
AppTek GmbH,Germany
AppTek GmbH, Germany
Cited by 3 stored of 3
3 results
No patents citing this paper on Lens.org (checked 2026-10-06).