Skip to content
arXiv cs.CV · Papers

SceneBind: Binding What and Where Across Vision, Audio and Language

arXiv:2607.15265v1 Announce Type: new Abstract: We present SceneBind, an omni-modal representation of realistic scenes with joint semantic and 3D spatial understanding across vision, audio and language. Existing omni-modal encoders excel at instance-level semantics (i.e., what is present), but often lack explicit spati