Google researchers have introduced Regularized Recursive Self-Improvement (RRSI), a method designed to stop AI agents from memorizing their test tasks during self-optimization. The approach aims to ensure that improvements made to the agent's operational harness translate to better performance on new, unseen benchmarks rather than just inflating scores on training data.

What Happened

Recent progress in AI agents has been driven largely by refinements to the 'harness'—the framework that determines how an agent reads files, recovers from errors, and delivers results—rather than by new model releases. Traditionally, these harnesses were updated manually. Newer methods automate this by using a language model to rewrite the harness based on feedback from test tasks, a process the researchers describe as a practical form of recursive self-improvement.

However, the paper notes that this self-optimization creates a vulnerability: agents working on a limited set of tasks tend to memorize them. This results in higher scores on training tasks but stagnant or declining performance on new tasks. The researchers identified that the search process often favors candidates that score well by chance, memorizes benchmark-specific patterns, and adds unnecessary complexity that doesn't improve actual capability.

RRSI addresses this by regulating both the proposal and selection phases of the optimization loop. It imposes a shrinking budget on the number of independent edits a candidate can bundle, starting with larger rewrites and narrowing to small, traceable changes. A strict critic reviews every proposal, rejecting those that hardcode task names or benchmark-specific tricks. Additionally, the system only accepts higher compute costs if they yield measurable performance gains and removes components that no longer contribute to success.

Why It Matters

The researchers tested RRSI on eight benchmarks covering coding, agentic office work, and engineering design, using a frozen Claude Opus 4.8 model as the underlying engine. According to the paper, RRSI achieved gains of up to 14.1 points on training tasks and up to 4.7 points on five unseen benchmarks. It also reportedly used about 30 percent fewer tokens at runtime compared to the unregularized version.

While other optimization methods improved training scores, many failed to generalize, with two methods performing below the baseline on unseen tasks. RRSI was the only method among the optimized variants to perform significantly above the baseline on new tasks, albeit with the smallest training gain. This tradeoff is intentional, as the guardrails are designed to prioritize generalization over raw training performance.

The study also highlighted that harnesses optimized for one model can benefit weaker models. A coding harness optimized using Gemini 3.5 Flash improved the accuracy of the weaker Gemini 3.1 Flash Lite from 11.2 to 14.6 points without any modifications, suggesting that the mechanisms discovered are not dependent on the specific capabilities of the model used to find them. This contrasts with previous findings, such as those from ARC-AGI-3 tests, where manually designed harnesses failed to generalize to unfamiliar environments.

The Bottom Line

RRSI demonstrates that self-improvement makes AI agents reliably more capable only when repeated feedback is converted into lasting, generalizable changes. The researchers note that their study focuses on harnesses built around frozen models and does not address scenarios where model weights themselves are updated. The code for RRSI is available on GitHub.