← Back to feed
2026-07-16multimodalvisiondata

SceneBind: Binding What and Where Across Vision, Audio and Language

Mingfei Chen, Zijun Cui, Ruoke Zhang, Hyeonggon Ryu, Eli Shlizerman

PDF preview for SceneBind: Binding What and Where Across Vision, Audio and Language
Read on arXiv →

Key claim

SceneBind enables superior scene understanding across modalities.

In plain English

Current omni-modal encoders struggle with understanding spatial relationships in scenes. They can identify what objects are present but often miss where they are located. SceneBind addresses this by creating a representation that combines semantic meaning with spatial attributes, allowing for better scene understanding. This could be particularly valuable for applications that require precise localization and interaction with the environment.

Novelty
8.0/10

Introduces a new omni-modal representation that integrates spatial understanding.

Reliability
7.5/10

Presents a novel dataset and training protocol, but lacks extensive baselines.

Deep reliability assessment

The methodology supports the integration of semantic and spatial understanding across modalities, but the claims of state-of-the-art performance may be overclaimed without extensive comparison to all existing methods.

Reproducibility

No open source code or dataset URL is mentioned in the paper.

Key figure

Figure 1 illustrates how SceneBind models scenes by capturing both semantic and 3D spatial attributes across vision, audio, and language using a shared global embedding with object-centric slots.

Benchmark results

~novel binaural audio-visual datasetaccuracy: 34.7vs not specifiednot specifiedSOTA
~novel binaural audio-visual datasetaccuracy: 38.4vs not specifiednot specifiedSOTA
~novel binaural audio-visual datasetaccuracy: 19.6vs not specifiednot specifiedSOTA