← Back to feed
2026-07-15agentsreasoningalignmentcode

Do Agent Optimizers Compound? A Continual-Learning Evaluation on Terminal-Bench 2.0

Wenxiao Wang, Priyatham Kattakinda, Soheil Feizi

PDF preview for Do Agent Optimizers Compound? A Continual-Learning Evaluation on Terminal-Bench 2.0
Read on arXiv →

Key claim

RELAI-VCL enables compounding optimization gains in agents.

In plain English

Current agent optimization methods often report one-time gains without considering ongoing task evolution. This paper identifies that many existing methods fail to maintain improvements when faced with new challenges. By introducing a continual learning framework, the authors demonstrate that their approach, RELAI-VCL, can sustain and even enhance performance over time. Builders might find this relevant as it suggests a pathway to developing more resilient agents that adapt to changing environments.

Novelty
8.0/10

Introduces a new approach to continual learning in agent optimization.

Reliability
7.5/10

Evaluates multiple methods under controlled conditions with clear metrics.

Deep reliability assessment

The methodology supports the claim that optimization gains can compound when regression control is integrated into the optimization loop, but it may overclaim the generalizability of these results across different settings without further validation.

Reproducibility

yes, the paper mentions that baseline, Phase 1, and Phase 2 agent artifacts for all three optimizers are released at https://github.com/relai-ai/Continual-Learning-Terminal-Bench.

Key figure

Figure 1 shows the lifelong average pass rate by agent, highlighting RELAI-VCL's superior performance compared to GEPA, Meta Harness, and the unoptimized baseline.

Benchmark results

Terminal-Bench 2.0lifelong average pass rate: 76.4vs unoptimized baseline+17.7%SOTA
GitHub1 repo
relai-ai/Continual-Learning-Terminal-BenchOfficial