Chunked Attention-Based Encoder-Decoder Model for Streaming Speech Recognition
IEEE International Conference on Acoustics Speech and Signal Processing, pp. 11331–11335
Abstract
We study a streamable attention-based encoder-decoder model in which either the decoder, or both the encoder and decoder, operate on pre-defined, fixed-size windows called chunks. A special end-of-chunk (EOC) symbol advances from one chunk to the next chunk, effectively replacing the conventional end-of-sequence symbol. This modification, while minor, situates our model as equivalent to a transducer model that operates on chunks instead of frames, where EOC corresponds to the blank symbol. We further explore the remaining differences between a standard transducer and our model. Additionally, we examine relevant aspects such as long-form speech generalization, beam size, and length normalization. Through experiments on Librispeech and TED-LIUM-v2, and by concatenating consecutive sequences for long-form trials, we find that our streamable model maintains competitive performance compared to the non-streamable variant and generalizes very well to long-form speech.
Authors 4
-
Affiliation as printed
RWTH Aachen University,Machine Learning and Human Language Technology,Computer Science Department,Germany
AppTek GmbH, Germany
Computer Science Department, Machine Learning and Human Language Technology, RWTH Aachen University, Germany
-
Affiliation as printed
RWTH Aachen University,Machine Learning and Human Language Technology,Computer Science Department,Germany
AppTek GmbH, Germany
Computer Science Department, Machine Learning and Human Language Technology, RWTH Aachen University, Germany
-
Affiliation as printed
RWTH Aachen University,Machine Learning and Human Language Technology,Computer Science Department,Germany
AppTek GmbH, Germany
Computer Science Department, Machine Learning and Human Language Technology, RWTH Aachen University, Germany
-
Affiliation as printed
RWTH Aachen University,Machine Learning and Human Language Technology,Computer Science Department,Germany
AppTek GmbH, Germany
Computer Science Department, Machine Learning and Human Language Technology, RWTH Aachen University, Germany
Cited by 10 stored of 10
10 results
No patents citing this paper on Lens.org (checked 2026-10-06).
References 69
-
W6691770337details pending0citations
-
W4294619417details pending0citations
-
W3008181812details pending0citations
-
W3037698816details pending0citations
-
W3197654132details pending0citations
-
W4293714597details pending0citations
-
W6787040858details pending0citations
-
W6847363464details pending0citations
-
W6850218400details pending0citations