The Conformer Encoder May Reverse the Time Dimension
Abstract
We sometimes observe monotonically decreasing cross-attention weights in our Conformer-based global attention-based encoder-decoder (AED) models, Further investigation shows that the Conformer encoder reverses the sequence in the time dimension. We analyze the initial behavior of the decoder cross-attention mechanism and find that it encourages the Conformer encoder self-attention to build a connection between the initial frames and all other informative frames. Furthermore, we show that, at some point in training, the self-attention module of the Conformer starts dominating the output over the preceding feed-forward module, which then only allows the reversed information to pass through. We propose methods and ideas of how this flipping can be avoided and investigate a novel method to obtain label-frame-position alignments by using the gradients of the label log probabilities w.r.t. the encoder input frames.
Authors 5
-
Affiliation as printed
Human Language Technology and Pattern Recognition , Computer Science Department , RWTH Aachen University , Aachen , Germany
-
Affiliation as printed
Human Language Technology and Pattern Recognition , Computer Science Department , RWTH Aachen University , Aachen , Germany
-
Mohammad Zeineldeen Aachen Computer Science Department Human Language Technology and Pattern Recognition
Affiliation as printed
AppTek GmbH , Aachen , Germany
Human Language Technology and Pattern Recognition , Computer Science Department , RWTH Aachen University , Aachen , Germany
-
Affiliation as printed
Human Language Technology and Pattern Recognition , Computer Science Department , RWTH Aachen University , Aachen , Germany
Cited by 0 stored of 0
No patents citing this paper on Lens.org (checked 2026-10-06).