arXiv cs.CV
· Papers
SceneBind: Binding What and Where Across Vision, Audio and Language
arXiv:2607.15265v1 Announce Type: new Abstract: We present SceneBind, an omni-modal representation of realistic scenes with joint semantic and 3D spatial understanding across vision, audio and language. Existing omni-modal encoders excel at instance-level semantics (i.e., what is present), but often lack explicit spati