SceneBind: Binding What and Where Across Vision, Audio and Language
Mingfei Chen, Zijun Cui, Ruoke Zhang, Hyeonggon Ryu, Eli Shlizerman
Read on arXiv →Key claim
SceneBind enables superior scene understanding across modalities.
In plain English
Current omni-modal encoders struggle with understanding spatial relationships in scenes. They can identify what objects are present but often miss where they are located. SceneBind addresses this by creating a representation that combines semantic meaning with spatial attributes, allowing for better scene understanding. This could be particularly valuable for applications that require precise localization and interaction with the environment.
Introduces a new omni-modal representation that integrates spatial understanding.
Presents a novel dataset and training protocol, but lacks extensive baselines.
Deep reliability assessment
The methodology supports the integration of semantic and spatial understanding across modalities, but the claims of state-of-the-art performance may be overclaimed without extensive comparison to all existing methods.
Reproducibility
No open source code or dataset URL is mentioned in the paper.
Key figure
Figure 1 illustrates how SceneBind models scenes by capturing both semantic and 3D spatial attributes across vision, audio, and language using a shared global embedding with object-centric slots.
