Chunked Attention-Based Encoder-Decoder Model for Streaming Speech Recognition
IEEE International Conference on Acoustics Speech and Signal Processing, pp. 11331–11335
Abstract
We study a streamable attention-based encoder-decoder model in which either the decoder, or both the encoder and decoder, operate on pre-defined, fixed-size windows called chunks. A special end-of-chunk (EOC) symbol advances from one chunk to the next chunk, effectively replacing the conventional end-of-sequence symbol. This modification, while minor, situates our model as equivalent to a transducer model that operates on chunks instead of frames, where EOC corresponds to the blank symbol. We further explore the remaining differences between a standard transducer and our model. Additionally, we examine relevant aspects such as long-form speech generalization, beam size, and length normalization. Through experiments on Librispeech and TED-LIUM-v2, and by concatenating consecutive sequences for long-form trials, we find that our streamable model maintains competitive performance compared to the non-streamable variant and generalizes very well to long-form speech.
Authors 4
-
Affiliation as printed
RWTH Aachen University,Machine Learning and Human Language Technology,Computer Science Department,Germany
AppTek GmbH, Germany
Computer Science Department, Machine Learning and Human Language Technology, RWTH Aachen University, Germany
-
Affiliation as printed
RWTH Aachen University,Machine Learning and Human Language Technology,Computer Science Department,Germany
AppTek GmbH, Germany
Computer Science Department, Machine Learning and Human Language Technology, RWTH Aachen University, Germany
-
Affiliation as printed
RWTH Aachen University,Machine Learning and Human Language Technology,Computer Science Department,Germany
AppTek GmbH, Germany
Computer Science Department, Machine Learning and Human Language Technology, RWTH Aachen University, Germany
-
Affiliation as printed
RWTH Aachen University,Machine Learning and Human Language Technology,Computer Science Department,Germany
AppTek GmbH, Germany
Computer Science Department, Machine Learning and Human Language Technology, RWTH Aachen University, Germany
Cited by 10 stored of 10
10 results
No patents citing this paper on Lens.org (checked 2026-10-06).
References 69
-
W2962784628details pending0citations
-
W2963382396details pending0citations
-
W3094957294details pending0citations
-
W3097747488details pending0citations
-
W3105532142details pending0citations
-
W3197507772details pending0citations
-
W3198439131details pending0citations
-
W6623517193details pending0citations
-
W6685711979details pending0citations
-
W6714142977details pending0citations
-
W6747158283details pending0citations
-
W6793472422details pending0citations
-
W6679434410details pending0citations