---
title: Google researchers aim to stop self improving agents from learning test tasks by heart
url: https://www.elseif.net/google-researchers-aim-to-stop-self-improving-agents-from-learning-test-tasks-by-heart
published: 2026-10-07T13:04:12+00:00
language: en
section: Agents
source: https://the-decoder.de/wie-google-forscher-verhindern-wollen-dass-sich-selbst-verbessernde-ki-agenten-nur-auswendig-lernen/
organizations: Google Cloud AI Research, Google, Claude Opus 4.8, Gemini 3.5 Flash, Gemini 3.1 Flash Lite, RRSI
publisher: elseif
---

# Google researchers aim to stop self improving agents from learning test tasks by heart

Google Cloud AI Research together with several universities has introduced a method to curb self improving agents that rely on memorising test tasks. The approach adds a harness around a fixed language model that controls how the agent reads files, recovers from errors and delivers results. Recent progress in agent performance stems mainly from improvements to this harness rather than from new model releases. The researchers describe the process as practical recursive self improvement where feedback from test tasks is used to reshape the harness and to steer the agent's behaviour. However the method also creates a risk that the agent learns the limited set of test tasks by heart. When the same benchmark is used repeatedly the agent can specialise on patterns that only fit that benchmark and can overlook genuine improvements on unseen tasks. To address this the team proposes Regularized Recursive Self Improvement of Agent Harnesses which restricts how many independent changes can be bundled at once and gradually reduces that allowance.

The system also remembers previous attempts to avoid repeating failed ideas and deliberately explores untouched parts of the harness when progress stalls. A critic evaluates every proposed change and discards suggestions that embed benchmark specific tricks or that embed solution fragments. Only changes that bring a measurable performance gain relative to the cost are kept and components that no longer help are removed. Experiments were run on eight benchmarks covering programming agent office work and engineering design using Claude Opus 4.8 as the underlying model. RRSI achieved up to 14.1 points on training tasks and up to 4.7 points on five never seen benchmarks while using about 30 percent less token consumption than the unregulated variant. On no unseen benchmark did total performance fall below the starting level which commonly happens with agents that merely memorize tasks. The framework also improves performance on tasks outside the training set with the strongest gain of 4.7 points on JobBench.

All methods performed well on training tasks but reversed on new tasks with two methods even dropping below the baseline. RRSI produced the smallest training gain yet it was the only method that exceeded the baseline on unseen tasks by a clear margin. The reduced token usage and fewer steps make the approach attractive for weaker models such as Gemini 3.1 Flash Lite which rose from 11.2 to 14.6 points when its coding harness was optimised. The authors note that their study focuses on a fixed model and does not cover cases where model weights themselves change. They conclude that self improvement becomes reliable only when repeated feedback is permanently incorporated into the harness. The code is available on GitHub.
