← Back to feed
2026-09-04agentsreasoningvisiondatacode

RoboSPA: Can VLA Models Go Beyond Simple Scenes and Short-Horizon Tasks?

Zhenxuan Fan, Bo Zhang, Yutong Lin, Yuqian Yuan, Juekai Lin, Liang Liang, Zhuoyi Huang, Wenqiao Zhang, Juncheng Li, Siliang Tang, Jun Xiao, Yueting Zhuang

PDF preview for RoboSPA: Can VLA Models Go Beyond Simple Scenes and Short-Horizon Tasks?
Read on arXiv →

Key claim

RoboSPA enhances evaluation of robotic manipulation models.

In plain English

Imagine you're developing a robot that can manipulate objects in a complex environment, like a kitchen. The challenge isn't just getting the robot to pick up a cup; it's about understanding where the cup is in relation to other objects, planning a sequence of actions to reach it, and adapting if something goes wrong. Current benchmarks often test robots in overly simplified scenarios, which means they miss critical failures in real-world situations, such as misjudging distances or forgetting steps in a multi-part task. This is what's called limited evaluation scope.

To address these shortcomings, the authors created RoboSPA, a comprehensive dataset and benchmark designed to evaluate how well vision-language-action models can handle intricate spatial reasoning and long-term planning. RoboSPA includes a variety of tasks that increase in difficulty, allowing for a more nuanced assessment of a robot's capabilities. By introducing new diagnostic metrics beyond just success rates, RoboSPA helps identify specific areas where current models struggle, such as understanding complex spatial relationships and executing detailed plans. For anyone building robots, this benchmark offers a more realistic framework for testing and improving their systems, pushing the boundaries of what these agents can achieve in real-world applications.

Novelty
8.0/10

RoboSPA introduces a new benchmark for evaluating embodied reasoning in VLA models.

Reliability
7.5/10

The dataset is extensive and includes diverse tasks, but lacks comparison with existing benchmarks.

Deep reliability assessment

The methodology supports the evaluation of VLA models under complex spatial and procedural tasks, but the claim of establishing a challenging diagnostic benchmark may be overclaimed without real-world validation.

Reproducibility

yes, the paper mentions that the data and code are available at https://github.com/fanzhenxuan/RoboSPA.

Key figure

Figure 1 provides an overview of RoboSPA, highlighting its focus on fine-grained spatial reasoning and long-horizon procedural planning across multiple task categories and difficulty levels.

GitHub1 repo
fanzhenxuan/RoboSPAOfficial