Is your AI agent better than it was three months ago? Fresh out of a big round of prompt work, most engineering leaders would say yes. But ask how they know, and the confidence starts to wobble. There are opinions, a ‘feeling’ that things improved, and a nagging suspicion that something slipped when the model version changed. What there usually isn’t is a well-evidenced answer.

When conventional software breaks, you know about it. Something goes red, an alert fires, and someone gets paged at 2am. When an AI agent gets worse, it looks exactly the same as it did when it was working: same response times, same format, no errors. It just keeps answering, a little less well each week, until a customer points it out. Or worse, they don’t, and they simply leave.

In this blog, our Technical Director Ali Nadian explains why AI features need a different kind of testing, how to measure something that doesn’t have one right answer, and what good looks like for teams who want to ship with confidence.

Three things are shifting under your feet

If nobody has touched your agent since launch, you might assume it’s behaving exactly as it did the day you shipped it. Unfortunately, three things are moving underneath you all the time.

The first is the model itself. Providers update, retune and retire models constantly. Even if you’ve pinned a version, the infrastructure it runs on keeps changing, and how it handles your particular edge cases can shift without anything you’d recognise as an announcement.

The second is you. A prompt edit to fix a customer complaint here, a retrieval setting tweaked during an unrelated investigation there, a new tool, a tighter guardrail after an incident. Each change is small and perfectly reasonable on its own. Add up a quarter’s worth, though, and you’ve got a system nobody has looked at as a whole since the day it went live.

The third is your users. They find uses you never designed for. A new customer segment arrives with different documents and different phrasing, and the workflow that made up 80% of your traffic is suddenly splitting it 50/50 with another. Your agent was tuned for the first one, and nobody has changed a line of code.

Without measurement, the loudest voice wins

When there are no real numbers to go on, anecdote fills the gap. A complaint from a big customer becomes a prompt change. An impressive demo becomes proof that everything’s fine. Neither is representative, but both end up steering your roadmap.

Some teams do build an evaluation set, which is a great start. The trouble is they tend to build it once, at launch, and then leave it be while what users actually send drifts further and further away from it. The scores stay flat, the real experience gets worse, and everyone feels reassured. That’s arguably worse than having no measurement at all, because at least then you know you’re guessing.

Underneath it all, “it works” has no owner and no proper definition; quality is something everyone cares about but nobody is actually measured on. And with no threshold to stop a release and no number sitting next to uptime in the monthly report, there’s never a moment where someone has to stop and take a proper look.

What is it really costing you?

It’s easy to file this under “nice to have” while there are features to ship. But not knowing whether your AI is working comes with a very real price tag.

You ship slower. When every change to a prompt or model carries a risk you can’t put a number on, changes get batched up, pushed back and reviewed to death. It rarely shows up as one obvious delay; everything just feels sluggish.

You miss out on upgrades you’re already paying for. New model versions are the cheapest quality improvement available to you, but only if you can prove they’re not worse on the cases that matter. Without that proof, a straightforward upgrade becomes a six-week project, because nobody wants to be the person who shipped the regression. Teams with evaluation in place make the switch in days. Teams without it take months (or never bother), and they fall further behind with every release.

You can’t find the cause when something breaks. Was it one of the last forty changes? Which one? Without testing each change, the answer comes from reverting things until the complaints stop. That takes weeks and teaches you nothing.

The damage hits every customer at once. A 10% drop in quality doesn’t affect one user; it affects every interaction, every day, for as long as nobody notices. By the time it shows up in your churn figures, the cause is months old and the post-mortem has no evidence to work with. And when the board asks how good your AI actually is, “customers seem happy” might get you through one quarter, but it won’t get you through two.

How do you measure something with no right answer?

The fair pushback here is: if the output is free text, what exactly are you scoring? There’s no single correct answer to check against. The good news is there are four ways to go about it. The trick is to only climb as high as you need to, because each step up costs more and tells you less precisely what’s gone wrong.

Step one: check what can be checked. More of an AI’s output is testable with plain code than most teams assume. Is the JSON valid? Do the line items add up to the total? Does every document it cites actually exist in what it retrieved? Is the date in range, is it under the word limit, did it avoid mentioning a competitor? These checks are cheap, instant and give the same result every time. They also catch a surprising number of the mistakes that could embarrass you in front of customers.

Step two: pull out the facts, then check those. This is the step most teams skip, and it’s where you’ll get the most value. Rather than asking a model “how good is this answer?”, break the answer down into the specific claims it makes: which policy it cited, what figure it quoted, what it recommended, whether it asked a clarifying question. Then compare each one to what it should have said, using ordinary code.

You can use a model to do the extracting, and that’s fine. “What number did this text state?” is a far easier question than “is this answer good?”, and when the extraction is wrong, you can see it and fix it. Nobody can pick apart a 7 out of 10. What you end up with are numbers that behave like real numbers: citation accuracy 94%, figure accuracy 89%, correct recommendation 81%. You can compare them across runs, trace them back to a specific change and defend them in a meeting. Pick the three or four things that decide whether an answer is actually usable, and accept that they won’t capture everything. A partial measure you trust is worth far more than a complete one you don’t.

Step three: get a second AI to mark the answer, but only for what’s left. Tone, helpfulness, whether an explanation is easy to follow, whether it was right to refuse. These matter, and no amount of parsing will turn them into a neat field. A model marking against a rubric can work here, as long as you know its quirks:

  • Judges are much better at comparing (“is A better than B?”) than scoring out of ten.
  • They lean towards longer answers.
  • They’re swayed by which option comes first.
  • They favour writing that sounds like their own.

So shuffle the order, control for length, and be wary if your judge and your agent come from the same model family.

A few more rules are worth sticking to:

  • Don’t average judge scores. The gap between a 6 and a 7 isn’t the same as the gap between an 8 and a 9, so report the share that cleared the bar instead (“87% met the standard”, not “average 7.4”).
  • Check your judge against human reviewers and share how often they agree. If it’s 70%, everyone should know every number carries that uncertainty.
  • Pin the judge’s version and rubric. A judge that changes over time gives you a trend line that means nothing.

Step four: keep humans involved, permanently. Not as a fallback, but as the check on everything above. A small sample each week, read by someone who knows the domain and marked against the same rubric the judge uses. It’s the only way to catch the problems nobody thought to measure, and to spot when your automated numbers have stopped meaning what you think they mean.

In practice, that looks like:

  • running the cheap checks on every interaction in production
  • using extracted facts as your main quality measure on every change
  • reporting judged scores separately
  • keeping that weekly human sample going

A full evaluation suite can run to thousands of model calls, so save the expensive steps for a representative sample. And whatever you do, don’t roll it all up into one quality score. Blending citation accuracy, tone and speed into a single number throws away exactly the detail you need to fix anything. It’s also the easiest number in the building to game.

What does good look like?

The teams who get this right tend to have a few things in common.

Their evaluation set gates every release. Change a prompt, a model, a retrieval setting or a tool, and the tests run automatically. A failing result blocks the release the same way a failing unit test blocks a deploy. Almost everything else follows from this.

They treat that evaluation set as a living thing. New cases come in from production every week, and old ones retire when a workflow disappears. Someone owns it, because it’s the asset that makes every other decision cheaper.

They watch production as well as testing. Offline tests tell you whether a change is safe; monitoring live traffic tells you whether the real world has moved. That means scoring a sample of real interactions and keeping an eye on signals like retries, escalations, thumbs-downs and people giving up halfway through.

They version everything together. Prompts, context, model version, tools and settings go out as one release with one version number. “Which prompt was live on the 14th?” becomes a quick lookup rather than an investigation.

They report quality like they report uptime. One small, stable number, shared with the business every month. The point isn’t only measurement, it’s accountability: it creates a regular moment where someone has to look.

They’ve written down what happens when a new model comes out. Run the suite, compare, review the differences, test on a small slice of traffic, roll out. Write it down once and it becomes a routine rather than a project.

You can’t afford to guess

It’s tempting to put evaluation on the “later” pile while there are features to ship and targets to hit. But every week without it is another week your AI could be getting worse with nobody any the wiser. So start small if you need to, and sit down with your team to ask a few honest questions:

  • If we changed model tomorrow, how would we know whether quality moved, and how long would it take to find out?
  • What does our quality number actually measure, and how much of it is a model’s opinion?
  • Which parts of a good answer could we check with ordinary code, and why aren’t we?
  • Who reads real output from our AI every week, and what happens to what they find?

If the answers make you wince, you’re far from alone. Most teams are still working this out, and that’s exactly the opportunity. Because once you can measure your AI, everything changes. A new model release stops being a risk and becomes a free upgrade you can adopt in days. A dip in quality gets caught on a Tuesday afternoon, not in next quarter’s churn figures. And when the board asks how good your AI really is, you can give them a number instead of a shrug. Evaluation might feel like a luxury right now, but it’s what turns AI from something you hope works into something you know does.

Did any of those four questions make you wince?

LLM evaluation is the four-week version of everything above, run on one AI feature you already have in production: we define what good means, build the checks and the judged scores, wire them in as a release gate, and leave the dataset, evals and code in your repo.

See how LLM evaluation works