A gpt-6-luna agent climbs from 33% to 92% accuracy over four optimisation rounds, then falls to 29% when the shop's rules are changed.
Four rounds of automatic tuning took a small agent from 33% to 92% on unseen tasks. Re-sampling the shop's rules took it to 29%.

Most teams improve their AI agents the way people tuned models in the 1990s: by hand. Someone reads a failed conversation, tweaks the prompt, runs a few examples and hopes nothing else broke.

Deep learning left that behind long ago. You define the parameters, define a loss, and let an optimiser follow the gradient. I wanted to know what happens when you treat an agent the same way.

The answer was a 59-point jump in accuracy on unseen tasks, followed by a collapse when the rules changed. The collapse turned out to be the more interesting half.

An agent has weights too

An agent isn't just a model. It's a model plus a configuration, and that configuration behaves a lot like a set of weights:

Deep learningAgent
WeightsInstructions, tool descriptions, which tools are on, limits such as how many results a search returns, reasoning effort
LossThe tasks it gets wrong
GradientAn LLM that reads the failures and proposes what to change
Optimiser stepApply the proposed edits
Validation setTasks the optimiser never sees

You can't take a derivative of a prompt. You can, however, show a second model the agent's failures and ask it which direction to move. That critique works as the gradient.

My agent's configuration had 25 of these "weights". Some are text, some are switches, and some are numbers. Every change the optimiser makes is recorded as a diff, so you can audit exactly what it did.

The 25 configuration values: 9 text fields, 12 switches, 4 numbers.
The 25 values the optimiser was allowed to change.

The training loop

The loop is short:

  1. Run the agent on a set of training tasks.
  2. Hand the failures to a reviewer model.
  3. The reviewer proposes a few candidate edits.
  4. Test each candidate, and keep an edit only if its gain is clearly larger than run-to-run noise.
  5. Repeat.
The five-step loop, with each step's deep-learning analogue: forward pass, loss, gradient, acceptance test, optimiser step.
The loop, with each step's deep-learning analogue. Step 4 is the one people skip.

Step 4 matters more than it looks. LLM agents are noisy, so a change that scores 5 points better on one run may score 5 points worse on the next. Without a noise test you end up accepting luck.

Putting the loop to work: a shop that can be rewired

To try this I needed an environment where three things were true: every answer could be checked automatically, the agent had to use its tools rather than guess, and I could change the rules later. A real customer-service deployment fails the first and the third. So I built a small hardware shop, Bolt & Nut Co., in code.

It has 7 products, 10 customers in three countries and 40 orders, and it gives the agent seven tools: look up an order, look up a customer, get a product's current price and weight, convert currency, run a calculation, search the shop's policy documents, and a decoy that returns out-of-date prices. Prices are in dollars, customers pay in pounds, gold-tier customers get a discount, shipping depends on weight and country, VAT goes on at the end, and return windows depend on the customer's tier and whether the item is a power tool.

The questions are the kind a shop assistant actually gets. "How much did I pay in total for order O1001?" "What was the shipping charge on order O1003?" "Today is 10 September. Can I still return order O1002?" A correct order total needs every line of the order, the customer's tier and country, current prices and weights, and five separate policy rules. There is no way to get it right by guessing, and the answer is a number, so a script can mark it.

I made the starting configuration deliberately weak: customer lookup and policy search switched off, order lookup returning only one line, the decoy tool switched on. That gave the optimiser something to find, and 25 values it could move to find it.

After four rounds, the best agent went from 33% to 92% on tasks it had never seen. The whole experiment cost under $3 in API calls.

Held-out accuracy by round for three models: gpt-6-luna 33 to 92, gpt-5-nano 25 to 67, gpt-5-mini 25 to 58.
Accuracy on held-out tasks after each round, for three models acting as both agent and reviewer. Flat segments are rounds where no candidate beat the noise.

It looked like a success.

Then I changed the world

The held-out tasks came from the same shop, with the same prices, VAT rate, shipping rules and return windows. So I built a second shop. It had the same tools, customers and questions, but every policy value was different: a new exchange rate, new discounts, new shipping charges, new VAT and new return windows. The tools served the new values, and the correct answers were recomputed.

A good agent shouldn't care about this, because it should look the rules up.

The 92% agent scored 29%. Its accuracy on order totals went from 14 out of 16 to 1 out of 16.

Accuracy of each optimised agent in the original shop and in the shop with re-sampled rules. gpt-6-luna drops 63 points, gpt-5-mini 37, gpt-5-nano 8. Right: gpt-6-luna's order-total questions, 14 of 16 correct in the same shop, 1 of 16 in the new one.
Left: each optimised agent, tested in the shop it was tuned in and in a shop with different rules. Right: the luna agent's order-total questions.

One detail gave it away. For one order, the correct total under the new rules was £36.77, and the agent answered £36.80. That was the correct answer in the old shop.

What the optimiser had actually learned

When I read the edits the optimiser had accepted, the explanation was obvious. The reviewer had rewritten a 17-word instruction into 327 words, with the shop's rules written in:

"…convert the USD subtotal to GBP at 0.79, then apply a 10% discount for gold-tier customers… For UK shipping, charge £4.99 below 2 kg or £8.99 otherwise…"

It had even rewritten the currency tool's description to say "using the rate stated in the instructions", so the tool's own answer would be ignored.

The optimiser hadn't taught the agent to use its tools better. It had memorised the environment. In deep-learning terms, this is overfitting, and the validation set couldn't catch it because the validation tasks shared the training tasks' facts.

Once the problem was visible, the reasons were too:

  • The reviewer could see the answer key. Failure reports included the expected answers, and it could read the full policy documents.
  • The objective rewarded memorising. I charged for token cost, and switching on a policy document adds tokens to every call. Copying the rules into the instructions was cheaper. One reviewer even explained its choice: "Keep token cost low by leaving pricing/shipping skills off, but give the agent a precise, step-by-step algorithm."
  • Nothing in the loop could notice. Validation from the same world can't detect memorised facts from that world.

It isn't inevitable, though. A smaller model acting as reviewer switched the policy documents on instead, and its agent barely moved when the rules changed.

That difference shows up in a check that costs nothing: count the facts in the optimised text.

Number of shop facts written into each configuration's text against its accuracy drop when the rules change. The two configurations with facts in their text are the two that collapsed; the eight fact-free ones sit within about plus or minus 13 points.
Facts in the configuration text against the accuracy drop under new rules, for ten configurations. Two data points don't make a law, but the order is right.

Regularisation for agents

The fix is the agent version of regularisation: constrain what the optimiser can learn.

  • Procedure, not facts. Reject any instruction edit that contains a number, a product or order ID, or a coupon code. The optimiser may write "look up the shipping rules before computing the total", but not "shipping is £4.99".
  • Hide the answer key. The reviewer is told only "wrong answer", never the expected value, and sees a one-line summary of each policy document instead of its contents.
  • Validate in a different world. Keep a shifted environment aside, and accept an edit only if most of its gain survives there. This catches leakage the first rule misses, such as facts written out in words.

I ran the loop again with the first two rules in place.

The drop under the rule change shrank from 37 points to 13, and accuracy on unseen tasks rose at the same time, from 58% to 83%.

Slope chart from the same shop to the new shop: gpt-6-luna unconstrained falls 92 to 29, gpt-5-mini unconstrained 63 to 26, gpt-5-mini regularised 74 to 61.
The same test as before, with the regularised agent added. The unconstrained agents lose most of their accuracy when the rules change; the regularised one keeps it, and scores more than double either of them in the new shop.

If you're tuning agents, check these

  1. Hold out a world, not just tasks. Change the environment's facts and re-test. If the gain disappears, it wasn't real.
  2. Scan your prompts for facts. Count the numbers, IDs and rule values in your optimised instructions. In my runs, the two configurations with facts in their text were the two that collapsed. This check is free.
  3. Watch your cost term. Penalising retrieval can reward memorisation.
  4. Don't show the optimiser the answers. The more it sees, the more it copies.
  5. Treat single runs with suspicion. Optimisation gains are noisy, so repeat runs before you believe them.

Tuning an agent you already run?

We improve AI features already in production, with the guard rails from this experiment built in: a held-out world as well as held-out tasks, an optimiser that never sees the answers, and every accepted change recorded as a diff you can audit.

See how LLM evaluation works