arXiv cs.AI
· Papers
Unlocking Spatial Grounding in Large Audio-Visual Retrieval models
arXiv:2607.24786v1 Announce Type: cross Abstract: Weak supervision sets a practical regime for audio-visual sound source localization as dense spatial annotations are costly to obtain at scale. The task, however, remains challenging, as models must locate sound sources from temporally aligned audio-visual data without