FurnitureVLA: Learning Long-Horizon Bimanual Furniture Assembly with Vision-Language-Action Model
Chenyang Ma, Yue Yang, Radu Corcodel, Siddarth Jain, Andrew Wu, Chiori Hori, Diego Romeres
Read on arXiv →Key claim
FurnitureVLA boosts assembly success from 48% to 80%.
In plain English
Imagine you're trying to assemble a large piece of furniture, like a dining table, but you want to do it with two robotic arms instead of your hands. The challenge is that most existing systems only work well with small, simple tasks or just one arm, which limits their usefulness in real-world scenarios. When you try to scale up, things can go wrong: the robots might not coordinate properly, leading to mistakes and frustration. This is what's called coordination failure, where the robots struggle to work together effectively over many steps.
To tackle this, the authors developed FurnitureVLA, a system designed specifically for real-scale bimanual furniture assembly. They created a simulation pipeline to generate expert data and a VR system that allows a single operator to control both arms. The key innovation here is that the system not only predicts what actions the robots should take but also tracks their progress through the assembly process. This helps the robots transition between tasks smoothly, reducing errors that can pile up over time.
Compared to previous methods, FurnitureVLA significantly boosts the success rate of assembly tasks, achieving an 80% success rate across different furniture types. This is a big improvement from the 48% success rate seen before. For anyone building robotic systems for furniture assembly, this means you can expect much better performance and reliability, especially in complex tasks that require multiple steps and coordination.
This work introduces a new systematic approach to real-scale bimanual furniture assembly, extending existing methods significantly.
The claims are supported by strong experimental results and a clear evaluation framework.
Deep reliability assessment
The methodology supports the improvement of simulation success rates and real-world validation, but the claim of generalist robot policy for real-scale bimanual furniture assembly may be overclaimed without extensive real-world testing across diverse environments.
Reproducibility
no
Key figure
Figure 1 illustrates real-scale bimanual furniture assembly using Vision-Language-Action models, highlighting the simulation pipeline and VR teleoperation system.
