arXiv cs.CL
· Papers
Generation-Step-Aware Framework for Cross-Modal Representation and Control in Multilingual Speech-Text Models
arXiv:2601.17387v3 Announce Type: replace Abstract: Multilingual speech-text models rely on cross-modal language alignment to transfer knowledge between speech and text, but it remains unclear whether this reflects shared computation for the same language or modality-specific processing. We introduce a generation-step-