Skip to content
HF Daily Papers · Papers

FATE: Frame-Level Audio-Visual Temporal Embedding

When a dog opens its mouth and barks, humans naturally recognize what the sound is and when it occurs. Building audio-visual models with this same ability requires representations that capture both semantic and temporal alignment. Current approaches fall short on one side or the other: embedding mod