Skip to content
arXiv cs.CL · Papers

Generation-Step-Aware Framework for Cross-Modal Representation and Control in Multilingual Speech-Text Models

arXiv:2601.17387v3 Announce Type: replace Abstract: Multilingual speech-text models rely on cross-modal language alignment to transfer knowledge between speech and text, but it remains unclear whether this reflects shared computation for the same language or modality-specific processing. We introduce a generation-step-