RoboTTT: Context Scaling for Robot Policies
Yunfan Jiang, Yevgen Chebotar, Ruijie Zheng, Fengyuan Hu, Yunhao Ge, Jimmy Wu, Tianyuan Dai, Scott Reed, Li Fei-Fei, Yuke Zhu, Linxi "Jim" Fan
Read on arXiv →Key claim
RoboTTT scales context length, enhancing robot task performance.
In plain English
Current robot models struggle with limited visuomotor context, which restricts their ability to perform complex tasks. Existing methods typically operate with short histories, leading to suboptimal performance in multi-stage scenarios. RoboTTT changes this by scaling the context length to 8K timesteps, allowing robots to learn from longer sequences and improve their decision-making in real-time. This advancement could enable builders to create more capable and flexible robotic systems that can handle intricate tasks more effectively.
Introduces a significant scaling of context length for robot policies.
Demonstrates strong performance improvements with clear baselines.
Deep reliability assessment
The methodology supports the claim that longer context lengths improve performance in robot policies, but the generalizability to different robotic tasks and environments may be overclaimed without broader testing.
Reproducibility
No open source code or dataset is mentioned in the paper.
Key figure
Figure 1 illustrates RoboTTT, a long-context visuomotor policy integrating Test-Time Training, highlighting capabilities like one-shot imitation and on-the-fly policy improvement.
