Somewhere in your business there is a spreadsheet that should not exist. It sits alongside a system you pay a great deal of money for, it has one person's name unofficially attached to it, and every few months somebody suggests replacing it, looks at what else is on the market, and concludes that the alternatives are worse.

That spreadsheet is not evidence of a badly run business. It is the most detailed piece of market research you own, and almost nobody treats it that way. The friction that survives inside expert businesses - the export-to-Excel step, the second system that exists only to feed the first, the report template three people each maintain slightly differently - is the outline of a product nobody has built, described in precise detail, by the only people qualified to describe it.

What it costs you is rarely dramatic, which is why it lasts for years. Nobody raises it at a board meeting, because there is no incident to raise: your senior people spend a slice of every week on work nobody would have designed on purpose, and the capacity you would need in order to grow is already committed to the workaround.

In this blog, Elliott Prince, Managing Director at Ferrous Labs, explains why generic software fails specialist businesses for structural reasons rather than accidental ones, how to read the friction in your own workflows as evidence, and what to do with it once you can see it.

Five signs, in the order they should worry you

These are arranged deliberately, from the sort of thing you would mention in passing to the sort that should change your plans for the year, and each is a stronger signal than the one before it.

Work done by hand that a computer should be doing. The weakest of the five, because it is the easiest to spot and the likeliest to have a dull explanation. It only counts if it takes meaningful time, is done much the same way by most businesses in your field, and is currently answered with a spreadsheet, a paper form, or a general-purpose tool bent into an unfamiliar shape.

Information sitting in systems that refuse to speak to each other. Stronger, because the cost compounds: every decision made on partial information is a small tax on judgement, and your people are paid for their judgement. Be careful with it in isolation, though: scattered data tends to produce integration projects that are expensive to build and hard to sell to anybody outside your own business, so treat it as supporting evidence rather than as the case itself.

Compliance held together by somebody's diligence. If your regulatory position depends on one experienced person remembering to check something, you have a risk with a name attached rather than a process. It outranks the first two because its cost behaves differently: manual work bleeds you steadily, while a compliance gap costs nothing at all until the week it costs a great deal.

Your workflow pushed through somebody else's idea of your workflow. This is one of the two I pay most attention to. Listen to how your team talks about their tools: when people describe working around the software rather than with it, and inducting a new starter involves explaining which fields to ignore and which to fill in with something not quite true, the software is modelling a business that is not yours. No amount of administrative skill will fix that, because the wrong assumptions sit in the data model, and you cannot configure your way out of a data model.

Enterprise pricing for a fraction of the features. The one I would act on fastest. You are quoted thousands a month for a platform whose relevant portion you could describe on one side of paper, and the rest of that quote exists for somebody else's industry. It is the strongest signal because it says two things at once: the need is real enough that a well-capitalised vendor has built something adjacent to it, and the economics of serving you properly are poor enough that nobody has built it exactly.

So why has nobody built the thing?

The obvious objection to all of that is that markets are not usually this inefficient. If the gap is so clear to you, why has a well-funded software company not walked straight into it?

Part of the answer is arithmetic: a vendor serving two hundred thousand businesses across every sector cannot justify a roadmap slot for the four thousand companies in yours, however loudly those four thousand complain. Part of it is that the logic you want encoded is not written down anywhere they could find it, living instead in the heads of people who have done the work for twenty years.

But the part that matters most, and the reason this has shifted recently rather than always having been true, is training data. A general-purpose model learns from what is abundant and public, and your domain's documents are neither: they sit behind non-disclosure agreements, in client folders, in vocabulary that means one thing in your field and something else entirely everywhere else. No amount of clever prompting will show a model something it was never shown.

A benchmark published in April 2026 by Pedro Barbosa de Carvalho Neto puts a number on it. Testing language models on sorting Brazilian appellate court decisions into five legal areas - about as ordinary as classification tasks get - a small open model fine-tuned on the domain, updating just 0.3% of its parameters, reached 87.6% accuracy: twenty-two points ahead of Claude 3.5 Haiku and twenty-eight ahead of GPT-4o mini.

The headline gap is striking, but it is not the figure that should stay with you. Broken down by category, GPT-4o mini scored an F1 of 0.00 on administrative law, meaning it essentially never got that class right, while the fine-tuned model managed 0.91. The commercial models had not done slightly worse; on one part of the domain they had given up and defaulted to the category they had seen most of. One benchmark in Portuguese law proves nothing about yours, but the mechanism travels: a model sounds equally confident everywhere and is reliable only where it has seen enough of the real thing.

What happened when we benchmarked ourselves against Amazon

We ran into this on a project for a property-tech business, where thousands of property legal documents went through manual review, each needing a trained professional to read it and produce a structured output. The documents were the hard part: dense legal language, nested sections, awkward layouts and tables that do not behave like tables, with generic entity recognition missing too much of the property law vocabulary to be much use.

We could have said exactly that and started building, which would have suited us nicely. Instead we treated build-versus-buy as a real question, built our own domain models for entity recognition and summarisation, and benchmarked them honestly against AWS and Azure. Had the cloud providers won, that would have been the recommendation and a cheaper engagement for the client. On the domain-specific tasks, ours matched or exceeded them - a narrow advantage, and narrow is precisely the point. Each document now takes under a minute, against the hours of manual review it replaced.

A tool for you, or a product for your market?

Once you accept that the gap is structural, there are two quite different responses available, and they carry different burdens of proof. Building the tool for yourself is the lower-risk path, and the five signs are sufficient evidence on their own, because you are the customer: the return lands in your own margin and capacity, you own the asset outright, and you need persuade nobody.

Building a product for your market is a different business altogether. Now you need to know that other businesses feel the same friction as sharply as you do, that they would pay to be rid of it rather than tolerate it, and that something about it compounds - switching costs, or a genuine network effect where the product grows more useful to each company as more of them join. Be sceptical about that second one. Colleagues recommending your tool to other colleagues is word of mouth, and word of mouth is how software sells in expert fields, but it is not a network effect and it is not a moat, because it works just as well for whoever arrives second.

Start with the workflow that irritates you most

You will probably agree with all of this and then not act on it, because the workaround is survivable and this quarter has targets in it. That is an entirely reasonable position, and it is also how five years go by.

The argument against waiting is narrower than the usual warnings about disruption. What changed is the cost: adapting a model to a specialist domain used to be a research project with a research budget, and the Brazilian benchmark above came out of fine-tuning on a consumer GPU, updating a fraction of a percent of the model. That window is open to everyone in your field at once, including the business three counties away sitting beside the same spreadsheet.

So start small. Take the one workflow that irritates your best people most and spend two days writing down exactly what happens: every input, every judgement call, every exception, every moment somebody leaves the system to get something done. You will finish holding either a specification or a very clear reason why you do not have one yet. Three questions worth putting to your team while you are at it:

  • Which of our tools do people describe working around rather than working with?
  • What are we paying enterprise money for, and what fraction of it do we genuinely use?
  • If a competitor removed our worst workflow tomorrow, how long would it take us to notice?

If the answers make uncomfortable reading, you are in ordinary company, and that is rather the opportunity. The businesses that end up owning the software in your industry will not be the ones with the best technology or the biggest budgets. They will be the ones who looked at the spreadsheet that should not exist and recognised it for what it was.

Recognised the spreadsheet?

The two paths need different things from you. If you want the tool for your own business, we build custom AI software around the workflow you already have. If the opportunity is bigger than your own business, we co-build the product with you.

Custom AI software Productise your consultancy