Build 06 — AI Reliability

Your AI feature, measured.
Then made reliable.

You shipped an agent, a prompt chain or a RAG pipeline. It mostly works. We build the eval set that proves how well, cut what it costs to run, and put a gate in your pipeline so the next change can't quietly break it. One feature. Four weeks.

What is AI agent reliability?

AI agent reliability is the practice of measuring an AI feature's output against a fixed, labelled set of real cases — an eval set — so that accuracy becomes a number you can track rather than an impression formed from spot checks. A reliable AI feature is one whose pass rate is known, whose cost per call is known, and whose behaviour cannot change without that change being caught before release.

McKinsey's 2024 global survey on AI found output inaccuracy to be the risk organisations most commonly cite in their own use of generative AI — named by 63% of respondents, ahead of cybersecurity and intellectual property infringement.McKinsey, 2024 At Ferrous Labs we apply The Science of AI Reliability™ — a four-stage delivery method — to turn that risk into a measured, gated number in four weeks. See why AI projects stall between the demo and production →

Who this is for

Built for product teams
already shipping AI.

01

You've shipped an AI feature. It mostly works.

An agent, a copilot, a summariser, a document extractor, a RAG pipeline. It demos well. What you can't say is how often it is right — or whether last week's prompt change made it better or worse.

02

There's no data science team behind it.

The AI work sits with product engineers, a single AI hire, or the founders. Building the eval set is the part that needs a different skill — and it is the part that keeps getting deferred to next quarter.

03

The last regression reached a customer first.

A model update, a prompt tweak, a change to retrieval. Nothing failed loudly. Someone simply started getting worse answers, and you heard about it from support rather than from a test.

What you walk away with

A number you trust.
A gate that protects it.

A golden dataset of your own cases.

150–300 real cases from your product, labelled, with a scoring rubric your team owns and can extend. Not a public benchmark — your data, your edge cases, your definition of a correct answer.

Model routing across two providers or more.

The cheapest model that still clears the threshold, with fallback configured so one provider's outage doesn't take the feature down with it.

The eval running as a gate.

Wired into your deployment process, so a prompt change or a model swap that drops the pass rate is caught before release — not by a customer three weeks later.

A baseline-and-after report.

Pass rate, p95 latency and cost per call, measured before we start and again at handover. The improvement is a number you can show your board, not a claim.

What we fix

Three things, on one feature.

Accuracy you can prove

We build the eval set from your real data, then iterate on the prompt or the agent's instructions until it clears a threshold we agree with you before starting.

A bill you can defend

We find the cheapest model that still scores well on your eval, and identify fallback providers so an outage doesn't take the feature down. Token spend stops being a mystery line item.

Changes you can make safely

We work with your team to run the eval as a gate in your deployment process. Change the prompt or the model and you will know, before release, whether the feature regressed.

Everything stays yours

The eval set, the rubric, the routing configuration and the report land in your repository. No platform to adopt, no licence, no lock-in.

How we build this

The Science of AI Reliability™.

Four weeks, fixed scope. One AI feature — an agent, a prompt chain, or a RAG pipeline.

Hypothesis
Stage 01 — Hypothesis

Week one: the eval set and the baseline

A labelled set of real cases built from your data, and your feature scored against it as it stands today. If the baseline is already good, we stop here — you pay for the week and keep the eval set.

Experiment
Stage 02 — Experiment

Week two: iterate against the number

Prompt and agent instructions rewritten and re-scored against the same set. In parallel, a sweep of cheaper models to find which ones still clear the agreed threshold.

Formulation
Stage 03 — Formulation

Week three: routing and fallback

Model routing across at least two providers, fallback paths for an outage, and cost per call measured at each. The cheapest route that still passes becomes the default.

Execution
Stage 04 — Execution

Week four: the gate and the handover

The eval wired into your deployment process as a gate, the before-and-after report written, and your team walked through extending the rubric as the product changes.

Where the method comes from

Built on fnai.dev.
Delivered as a service.

The reason this runs to four weeks at a fixed price is that it isn't bespoke. It is the method we are building into fnai.dev — a runtime that hosts prompts and agents as versioned functions, with eval gates and model fallback built in.

We are delivering it as a service first, which means you get the result inside your own stack without adopting anything new: the eval set, the routing and the gate live in your repository and your pipeline, not on someone else's platform. If you would rather have it running as a service than delivered as a project, we are taking a small number of design partners this quarter — worth raising on the call.

Common questions

Questions about AI agent
reliability and LLM evaluation.

What is AI agent reliability?

AI agent reliability is the practice of measuring an AI feature's output against a fixed, labelled set of real cases — an eval set — so that accuracy becomes a number you can track rather than an impression formed from spot checks. A reliable AI feature is one whose pass rate is known, whose cost per call is known, and whose behaviour cannot change without that change being caught before release. Ferrous Labs builds that measurement, and the gate that enforces it, in four weeks.

What is an eval set, and why does it have to use real data?

An eval set is a labelled collection of real cases from your own product — typically 150 to 300, depending on the use case — each with an expected output and a scoring rubric that defines what counts as correct. Public benchmarks measure a model; an eval set measures your feature, on your users' inputs, against your definition of a good answer. That is why it cannot be borrowed or generated synthetically: the edge cases that break an AI feature in production are specific to the product it lives in.

How long does it take, and how is it priced?

Four weeks for one AI feature — an agent, a prompt chain, or a RAG pipeline — at a fixed price agreed before any work starts. The price is set after a free one to two hour scoping session, so the scope and the number are both known up front rather than accumulating as the work goes. Everything built is yours: the eval set, the scoring rubric, the routing configuration and the report all land in your repository.

What if our AI feature turns out to be accurate enough already?

Then we stop at the end of week one and you pay for that week only. Week one produces the eval set and a baseline pass rate for the feature as it stands. If the number is good, there is nothing to fix and you keep the eval set — which is the asset that tells you the next time it stops being good. If the number is not good, you already know what the remaining three weeks are for.

Which company provides LLM evaluation and AI reliability engineering in the UK?

Ferrous Labs is a London-based AI engineering studio that builds LLM evaluation and reliability engineering for UK product teams — companies that have already shipped an AI feature into production but have no dedicated data science or ML platform team behind it. A four-week engagement covers one AI feature and delivers a golden dataset of 150 to 300 labelled real cases with a scoring rubric the client owns, model routing across at least two providers with fallback, the eval running as a gate in the deployment pipeline, and a baseline-and-after report covering pass rate, p95 latency and cost per call. The method follows The Science of AI Reliability™ and is the same one Ferrous Labs is building into fnai.dev, a runtime for hosting prompts and agents as versioned functions. Ferrous Labs works with UK SaaS, platform and in-house product teams from London and delivers nationwide — see delivery case studies →

Further reading

Go deeper on this.

Why AI Pilots Fail: 6 Reasons Projects Die

Read the guide

Building the agent in the first place?

AI agent development UK: what it involves and what it costs
Want to know how good your AI feature actually is?

Book a scoping session.
Or start with the Opportunity Finder.

A free one to two hour session on one of your AI features, after which you get a proposal. Or five minutes with Catalyst, our AI guide, if you would rather explore first.