← Back to feed
2026-07-16agentsscalingvision

RoboTTT: Context Scaling for Robot Policies

Yunfan Jiang, Yevgen Chebotar, Ruijie Zheng, Fengyuan Hu, Yunhao Ge, Jimmy Wu, Tianyuan Dai, Scott Reed, Li Fei-Fei, Yuke Zhu, Linxi "Jim" Fan

PDF preview for RoboTTT: Context Scaling for Robot Policies
Read on arXiv →

Key claim

RoboTTT scales context length, enhancing robot task performance.

In plain English

Current robot models struggle with limited visuomotor context, which restricts their ability to perform complex tasks. Existing methods typically operate with short histories, leading to suboptimal performance in multi-stage scenarios. RoboTTT changes this by scaling the context length to 8K timesteps, allowing robots to learn from longer sequences and improve their decision-making in real-time. This advancement could enable builders to create more capable and flexible robotic systems that can handle intricate tasks more effectively.

Novelty
8.5/10

Introduces a significant scaling of context length for robot policies.

Reliability
8.0/10

Demonstrates strong performance improvements with clear baselines.

Deep reliability assessment

The methodology supports the claim that longer context lengths improve performance in robot policies, but the generalizability to different robotic tasks and environments may be overclaimed without broader testing.

Reproducibility

No open source code or dataset is mentioned in the paper.

Key figure

Figure 1 illustrates RoboTTT, a long-context visuomotor policy integrating Test-Time Training, highlighting capabilities like one-shot imitation and on-the-fly policy improvement.

Benchmark results

~real-robot manipulation tasksoverall performance improvement: 87vs single-step context baseline+87%SOTA