← Back to feed
2026-07-01agentsvisionscaling

FurnitureVLA: Learning Long-Horizon Bimanual Furniture Assembly with Vision-Language-Action Model

Chenyang Ma, Yue Yang, Radu Corcodel, Siddarth Jain, Andrew Wu, Chiori Hori, Diego Romeres

PDF preview for FurnitureVLA: Learning Long-Horizon Bimanual Furniture Assembly with Vision-Language-Action Model
Read on arXiv →

Key claim

FurnitureVLA boosts assembly success from 48% to 80%.

In plain English

Imagine you're trying to assemble a large piece of furniture, like a dining table, but you want to do it with two robotic arms instead of your hands. The challenge is that most existing systems only work well with small, simple tasks or just one arm, which limits their usefulness in real-world scenarios. When you try to scale up, things can go wrong: the robots might not coordinate properly, leading to mistakes and frustration. This is what's called coordination failure, where the robots struggle to work together effectively over many steps.

To tackle this, the authors developed FurnitureVLA, a system designed specifically for real-scale bimanual furniture assembly. They created a simulation pipeline to generate expert data and a VR system that allows a single operator to control both arms. The key innovation here is that the system not only predicts what actions the robots should take but also tracks their progress through the assembly process. This helps the robots transition between tasks smoothly, reducing errors that can pile up over time.

Compared to previous methods, FurnitureVLA significantly boosts the success rate of assembly tasks, achieving an 80% success rate across different furniture types. This is a big improvement from the 48% success rate seen before. For anyone building robotic systems for furniture assembly, this means you can expect much better performance and reliability, especially in complex tasks that require multiple steps and coordination.

Novelty
8.0/10

This work introduces a new systematic approach to real-scale bimanual furniture assembly, extending existing methods significantly.

Reliability
8.0/10

The claims are supported by strong experimental results and a clear evaluation framework.

Deep reliability assessment

The methodology supports the improvement of simulation success rates and real-world validation, but the claim of generalist robot policy for real-scale bimanual furniture assembly may be overclaimed without extensive real-world testing across diverse environments.

Reproducibility

no

Key figure

Figure 1 illustrates real-scale bimanual furniture assembly using Vision-Language-Action models, highlighting the simulation pipeline and VR teleoperation system.

Benchmark results

~simulation across three furniture typesaverage simulation success: 80vs prior baselines+32%SOTA
FurnitureVLA: Learning Long-Horizon Bimanual Furniture Assembly with Vision-Language-Action Model — Frontier Papers