- Google Research and Virginia Tech's WikiSkill stores an agent's failed fixes in a wiki, so later improvement cycles do not repeat them.
- In their tests it led every rival method on every model, by 3.3 to 12 points, while keeping the wiki out of production prompts.
- Ask your team whether agent traces and rejected changes are kept, and who validates new instructions. The paper leaves long-running tasks untested.
Self-improving agents forget their own lessons
Many companies want AI agents that get better with use. One approach is to let the agent review its own successes and failures, then rewrite its instructions. Recent frameworks that automate this have a weak point: memory.
Liyan Tang is a Google research scientist and a co-author of the WikiSkill paper. He says many recent skill-evolution frameworks lose the diagnosis once a patch is proposed. That includes the record of which fixes failed validation. The result, in his words, is a system that keeps "rediscovering the same failures." That is paid work done twice.
The idea here is simple. An agent that improves itself needs a record of what it tried and why it failed, not only a record of what worked. Human teams keep incident logs for the same reason. WikiSkill gives agents one.
How WikiSkill works
WikiSkill comes from Google Research and Virginia Tech. It splits an agent's workspace into three layers. The bottom one, the Raw Layer, is a permanent record of what the agent did. It holds every action, the agent's reasoning, each tool call and its output, and the final answer. None of it is edited afterward.
The Wiki Layer turns those traces into structured pages on recurring failures and winning strategies. It also holds a log of how the skills changed over time. A scorecard beside it notes each proposed change and whether performance improved. The Skill Layer holds the short instructions the agent actually uses. Each skill links back to the wiki patterns behind it.
Each improvement cycle has four steps. An Inference Agent runs training tasks. A Wiki Maintainer updates the wiki from sampled successes and failures. A Skill Proposer drafts new or edited skills. A validation set then tests the result. The change stays only if it beats the best score so far.
The researchers gave one example from ALFWorld, a simulated household-task environment. The agent kept picking up objects and putting them back. A broad skill was proposed and rejected. The wiki kept both the behavior and the failed proposal. In the next round the proposer wrote a tighter skill with a concrete rule: do not put an item back where it started. That version improved validation scores and was accepted.
What the tests showed
The team tested five benchmarks: math reasoning, web search, spreadsheets, long-document question answering and household tasks. They used Qwen, Gemma and Gemini models. They compared WikiSkill with Trace2Skill, EvoSkill, SkillOpt and agents with no skills.
On every model, WikiSkill finished first on average. The gap to the closest competing method on each model ran from 3.3 to 12 percentage points.
Skills also closed part of the gap between model sizes. Qwen-3.5-9B using WikiSkill averaged 47.4%. The larger Qwen-3.6-27B with no skills averaged 39.4%. For a buyer weighing a bigger model against better scaffolding, that is a data point worth testing in-house.
The cost design is as important as the score
The wiki never enters the inference agent's prompt. Tang explained why: "Memory wants to be exhaustive and a production prompt wants to be lean." Production pays only for compact skills, around 45 to 129 lines in the team's runs.
The ablation tests, run on Gemini-3.5-Flash across four benchmarks, support this. The default design averaged 63.7%. With no wiki and no Wiki Maintainer, it averaged 48.7%. Giving the inference agent wiki access too lowered it to 60.9%.
The cost does not disappear. It moves to the improvement phase. Each iteration took roughly 10 to 20 reasoning-and-tool-call turns from the Skill Proposer, plus a Wiki Maintainer call.
What the paper does not answer
These are research results, not production evidence. The paper does not test how agents find the right skill as libraries grow. Its gate rejects changes that do not help immediately, even if they could help later. The wiki grows with no automatic pruning. No task ran for hundreds of actions or several hours.
What to ask your team
First, do your agents keep their execution traces? Those traces are the raw material. Second, when an automated change is rejected, is the reason stored anywhere? Third, who or what validates a new instruction before it goes live? WikiSkill's answer is an independent test set.
Fourth, where does the cost sit: in every production call, or in a separate improvement phase? The lesson here is that an agent's memory and its working instructions are different things. Keep the first thorough and the second short.
Produced by the WebPulse Newsroom with AI assistance from the original reporting credited below, and checked against that source by our editorial review. How we use AI.
Original reporting: VentureBeat.





