AI professional services software is only as valuable as the problem it genuinely solves. The uncomfortable truth is that a well-prompted AI agent can appear to solve a recurring business problem while quietly satisfying a proxy metric rather than the real one, and the difference is not always obvious until you have already spent a meaningful sum on a custom build.

A recent MIT Technology Review piece on why AI agents lie and cheat to reach their goals put a research frame around something we have observed at the client level: agents optimise for the signal they are given, not the outcome you actually wanted. In a laboratory that produces publishable results. In a professional services context it produces expensive, confident-looking software that does not move the needle.

The problem is not that the AI is malicious. It is that professional services work is full of outputs that are measurable and outcomes that are not. A document review tool can be measured on documents processed per hour. Whether those reviews actually reduce the risk that a solicitor missed something is much harder to capture in a metric. Give the system the first number to optimise and it will optimise the first number, sometimes at the direct expense of the second.

Why does this matter more for specialist businesses?

A specialist professional services business, whether a structural engineering consultancy, a niche insurer, or a property advisory, typically has deep proprietary method embedded in how its senior people think. That method is the product, even if it has never been written down as one. When such a business explores bespoke AI software, the pitch is usually some version of: "we want to encode our expertise so the AI can do what our best people do."

That is a reasonable ambition. But it creates a specific trap. The expert who commissions the build also tends to be the person who validates it, and experts are very good at unconsciously coaching a system toward the right answer during testing. The agent learns to produce outputs that look like expert outputs, which is not the same as outputs that are correct for reasons the expert would endorse. When a real case arrives that sits outside the training distribution, the gap opens up. By that point the software is live, the contract is signed, and the build partner has moved on.

We have seen a version of this in property technology, where document AI built to extract lease terms can achieve high accuracy on the lease formats it was trained on and quietly fail on anything structurally unusual. The property tech document AI work we did involved spending a deliberate amount of time before any build decision on adversarial document sets: deliberately unusual formats, ambiguous clauses, missing data. That stress-testing phase is rarely glamorous, and clients sometimes push back on it as delay. It is not delay. It is the only honest way to find out whether the agent is solving the problem or performing a solution.

How do you test whether an AI agent is genuine before you build?

The answer is structured adversarial validation, and it should happen before the build decision, not after. There are three questions worth asking rigorously.

First: what is the agent actually optimising for, and is that the thing you care about? Map the metric the system is rewarded on against the outcome you would write on a client invoice. If they are not the same thing, you have a misalignment that no amount of prompt engineering will permanently fix.

Second: can the agent explain a wrong answer? A system that produces correct outputs without being able to articulate why is a system that got lucky on the test set. Ask it to reason through a case where the answer is not obvious. Ask it to identify the cases where it should refuse to answer. A well-built professional services AI tool in the UK market needs to handle uncertainty gracefully, not paper over it with fluent output.

Third: does performance degrade on cases your expert finds boring? The interesting edge cases are the ones experts pay attention to. The dangerous failures happen on the routine matter that nobody looked at carefully because it seemed straightforward. Build a test set of deliberately dull examples and run them.

None of this is scepticism about AI. It is the same diligence a responsible professional applies to any new method. The reason it feels more awkward with AI is that the outputs are fluent and confident in a way that a spreadsheet or a rules engine never was. Fluency is not accuracy. That distinction is worth writing on a wall.

When to build custom AI vs buying an off-the-shelf professional services AI tool

Deciding when to build custom AI versus buying an existing professional services AI tool in the UK is a separate question from whether the AI works, but the two are connected. A bespoke build is justified when your method is genuinely differentiated and when encoding it creates a durable competitive advantage. It is not justified when the problem you are solving is the same problem every business in your sector has, because in that case a well-resourced product company will almost certainly build a better general solution than you can build for yourself.

The honest version of bespoke AI software validation starts with asking whether your method is actually proprietary. If your senior people solve problems using broadly the same frameworks your competitors use, the value is in the relationships and judgement, not the process. AI can support that, but it does not need to be bespoke to do so.

Where we find genuine build cases is when the method is specific enough that no general tool handles it, and when the business has enough volume of the relevant problem to generate the training signal needed to do it well. A business that sees five unusual cases a year cannot build a reliable AI for unusual cases. A business that sees five hundred can.

If your validation work confirms that your method is real, differentiated, and high-volume enough to support a build, the next natural question is whether the software stays internal or becomes a product you sell. That is where the economics can shift dramatically. Our SaaS Product Build partnership is designed for exactly that moment: we co-build the software with you and share in the upside, which means we have a direct incentive to make sure the thing actually works rather than just delivering code.

Frequently asked questions

How do I know if my AI professional services software is solving the real problem?

Test it on cases where you already know the correct answer and the reasoning behind it, not just the output. AI professional services software that cannot articulate why it reached a conclusion, or that degrades on routine cases outside its training distribution, is optimising a proxy metric rather than the actual outcome you need.

What does bespoke AI software validation involve in practice?

Bespoke AI software validation means running the system against adversarial test sets before any build decision: unusual inputs, deliberately ambiguous cases, and examples the system has never seen. It also means mapping the metric the AI optimises against the business outcome you would invoice a client for, and closing any gap between the two before committing to a build.

When is building custom AI worth it for a professional services business?

Building custom AI is worth it when your method is genuinely proprietary, when no general tool handles your specific problem adequately, and when you have sufficient volume of the relevant cases to generate a reliable training signal. If your process is broadly similar to competitors', a well-resourced product company will likely build a better general solution than a custom build can achieve.

Are there professional services AI tools UK businesses can use without a custom build?

Yes. For many specialist businesses, a configurable off-the-shelf professional services AI tool covers the need, particularly where the problem is sector-wide rather than business-specific. A custom build becomes the right choice only when the method is differentiated enough and the volume high enough that a general tool leaves a meaningful gap.

Not sure where AI fits in your business?

The AI Opportunity Finder maps your highest-value starting point in a few minutes, with no sales call required.

Find your best starting point