Most teams improve their AI agents the way people tuned models in the 1990s: by hand. Someone reads a failed conversation, tweaks the prompt, runs a few examples and hopes nothing else broke.
Deep learning left that behind long ago. You define the parameters, define a loss, and let an optimiser follow the gradient. I wanted to know what happens when you treat an agent the same way.
The answer was a 59-point jump in accuracy on unseen tasks, followed by a collapse when the rules changed. The collapse turned out to be the more interesting half.
An agent has weights too
An agent isn't just a model. It's a model plus a configuration, and that configuration behaves a lot like a set of weights:
| Deep learning | Agent |
|---|---|
| Weights | Instructions, tool descriptions, which tools are on, limits such as how many results a search returns, reasoning effort |
| Loss | The tasks it gets wrong |
| Gradient | An LLM that reads the failures and proposes what to change |
| Optimiser step | Apply the proposed edits |
| Validation set | Tasks the optimiser never sees |
You can't take a derivative of a prompt. You can, however, show a second model the agent's failures and ask it which direction to move. That critique works as the gradient.
My agent's configuration had 25 of these "weights". Some are text, some are switches, and some are numbers. Every change the optimiser makes is recorded as a diff, so you can audit exactly what it did.
The training loop
The loop is short:
- Run the agent on a set of training tasks.
- Hand the failures to a reviewer model.
- The reviewer proposes a few candidate edits.
- Test each candidate, and keep an edit only if its gain is clearly larger than run-to-run noise.
- Repeat.
Step 4 matters more than it looks. LLM agents are noisy, so a change that scores 5 points better on one run may score 5 points worse on the next. Without a noise test you end up accepting luck.
Putting the loop to work: a shop that can be rewired
To try this I needed an environment where three things were true: every answer could be checked automatically, the agent had to use its tools rather than guess, and I could change the rules later. A real customer-service deployment fails the first and the third. So I built a small hardware shop, Bolt & Nut Co., in code.
It has 7 products, 10 customers in three countries and 40 orders, and it gives the agent seven tools: look up an order, look up a customer, get a product's current price and weight, convert currency, run a calculation, search the shop's policy documents, and a decoy that returns out-of-date prices. Prices are in dollars, customers pay in pounds, gold-tier customers get a discount, shipping depends on weight and country, VAT goes on at the end, and return windows depend on the customer's tier and whether the item is a power tool.
The questions are the kind a shop assistant actually gets. "How much did I pay in total for order O1001?" "What was the shipping charge on order O1003?" "Today is 10 September. Can I still return order O1002?" A correct order total needs every line of the order, the customer's tier and country, current prices and weights, and five separate policy rules. There is no way to get it right by guessing, and the answer is a number, so a script can mark it.
I made the starting configuration deliberately weak: customer lookup and policy search switched off, order lookup returning only one line, the decoy tool switched on. That gave the optimiser something to find, and 25 values it could move to find it.
After four rounds, the best agent went from 33% to 92% on tasks it had never seen. The whole experiment cost under $3 in API calls.
It looked like a success.
Then I changed the world
The held-out tasks came from the same shop, with the same prices, VAT rate, shipping rules and return windows. So I built a second shop. It had the same tools, customers and questions, but every policy value was different: a new exchange rate, new discounts, new shipping charges, new VAT and new return windows. The tools served the new values, and the correct answers were recomputed.
A good agent shouldn't care about this, because it should look the rules up.
The 92% agent scored 29%. Its accuracy on order totals went from 14 out of 16 to 1 out of 16.
One detail gave it away. For one order, the correct total under the new rules was £36.77, and the agent answered £36.80. That was the correct answer in the old shop.
What the optimiser had actually learned
When I read the edits the optimiser had accepted, the explanation was obvious. The reviewer had rewritten a 17-word instruction into 327 words, with the shop's rules written in:
"…convert the USD subtotal to GBP at 0.79, then apply a 10% discount for gold-tier customers… For UK shipping, charge £4.99 below 2 kg or £8.99 otherwise…"
It had even rewritten the currency tool's description to say "using the rate stated in the instructions", so the tool's own answer would be ignored.
The optimiser hadn't taught the agent to use its tools better. It had memorised the environment. In deep-learning terms, this is overfitting, and the validation set couldn't catch it because the validation tasks shared the training tasks' facts.
Once the problem was visible, the reasons were too:
- The reviewer could see the answer key. Failure reports included the expected answers, and it could read the full policy documents.
- The objective rewarded memorising. I charged for token cost, and switching on a policy document adds tokens to every call. Copying the rules into the instructions was cheaper. One reviewer even explained its choice: "Keep token cost low by leaving pricing/shipping skills off, but give the agent a precise, step-by-step algorithm."
- Nothing in the loop could notice. Validation from the same world can't detect memorised facts from that world.
It isn't inevitable, though. A smaller model acting as reviewer switched the policy documents on instead, and its agent barely moved when the rules changed.
That difference shows up in a check that costs nothing: count the facts in the optimised text.
Regularisation for agents
The fix is the agent version of regularisation: constrain what the optimiser can learn.
- Procedure, not facts. Reject any instruction edit that contains a number, a product or order ID, or a coupon code. The optimiser may write "look up the shipping rules before computing the total", but not "shipping is £4.99".
- Hide the answer key. The reviewer is told only "wrong answer", never the expected value, and sees a one-line summary of each policy document instead of its contents.
- Validate in a different world. Keep a shifted environment aside, and accept an edit only if most of its gain survives there. This catches leakage the first rule misses, such as facts written out in words.
I ran the loop again with the first two rules in place.
The drop under the rule change shrank from 37 points to 13, and accuracy on unseen tasks rose at the same time, from 58% to 83%.
If you're tuning agents, check these
- Hold out a world, not just tasks. Change the environment's facts and re-test. If the gain disappears, it wasn't real.
- Scan your prompts for facts. Count the numbers, IDs and rule values in your optimised instructions. In my runs, the two configurations with facts in their text were the two that collapsed. This check is free.
- Watch your cost term. Penalising retrieval can reward memorisation.
- Don't show the optimiser the answers. The more it sees, the more it copies.
- Treat single runs with suspicion. Optimisation gains are noisy, so repeat runs before you believe them.
Tuning an agent you already run?
We improve AI features already in production, with the guard rails from this experiment built in: a held-out world as well as held-out tasks, an optimiser that never sees the answers, and every accepted change recorded as a diff you can audit.
See how LLM evaluation works