Language model integration into acoustic model training
RWTH Publications (RWTH Aachen)
Abstract
To fuel the training of ever more powerful and accurate speech recognition systems a mountain of bimodal training data is necessary, that includes both the acoustic representation of speech along with the textual transcriptions. The amount of available high-quality training data and the cost of producing more of it is currently the main bottleneck in the development of automatic speech recognition systems. Another component in the improvement of automatic speech recognition systems are language models that capture semantic knowledge of a specific language. These require only textual data and can and have been trained on incomprehensible amounts of it. In this thesis we will investigate how language models can be more tightly integrated into the acoustic model training process to improve the overall system performance. Firstly, we consider sequence discriminative training of hybrid acoustic models. We propose a novel sequence discriminative training criterion that we call frame-level maximum mutual information. This is closely related to the state-level Bayes risk criterion but using it in model training should move them towards the true distribution instead of a sharp 0-1-distribution. While we do not see any improvements in word error rate, we do observe a qualitative change in the resulting probability distribution, which is less sharp if trained with the frame-level maximum mutual information criterion. Secondly, we propose to integrate language models into the training process of attention-based encoder-decoder acoustic models via log-linear combination. The sentence-level normalization of the combination gives rise to a sequence discriminative training criterion that we call maximum mutual information, based on its similarity to the equivalent training criterion for hybrid models. We investigate suitable approximations to the normalization term and the required search space size to capture all relevant contributions.The language model used in this combination is tuned extensively to show which properties are optimal for training integration. The new criterion is then compared to the existing expected error criterion and we show that competitive word error rates can be achieved. Thirdly, we aim to ameliorate the increased training time and complexity of the maximum mutual information criterion by using a symbol-level normalization. We show that the same improvement in word error rates can be achieved without significantly increasing the training time over the baseline. Additionally, we show that symbol-dependent scaling exponents, which are learned by back-propagation, can improve the word error rate.These improvements, however, go away once the acoustic model is jointly trained with the scales and language model. Lastly, we investigate the interplay of language model integration in acoustic model training and the internal language model correction approach. We find that most of the improvements of integrating a language model in the training process come from suppressing the implicit language model that is learned by the acoustic model decoder. We therefore show that the word error rate improvements obtained by integrating a language model in training are largely already captured when an internal language model correction model is used during recognition.
Authors 1
-
Affiliation as printed
RWTH Aachen
Cited by 0 stored of 0
No patents citing this paper on Lens.org (checked 2026-10-06).