Skip to content
Apple ML Research · Cloud & Big Tech

Segmental Attention Decoding with Long Form Acoustic Encodings

We address the fundamental incompatibility of attention-based encoder-decoder (AED) models with long-form acoustic encodings. AED models trained on segmented utterances learn to encode absolute frame positions by exploiting limited acoustic context beyond segment boundaries, but fail to generalize when decoding long-fo