Vertical AI
What is vertical AI?
Vertical AI is an AI product built for one sector or one job instead of for everyone. A contract review tool for law firms, a note-taking system for doctors, a progress tracker for a building site: each one assumes a lot about who is using it and what they are trying to get finished.
The word vertical comes from an older split in software. Horizontal software works the same for a bakery and a bank: a spreadsheet, a mailbox, a general chat assistant. Vertical software only makes sense inside one trade, like a practice management system for a notary. Vertical AI applies that split to AI products.
It is a spending category now, not only a marketing word. Gartner's July 2026 forecast puts worldwide end-user spending on AI models and platforms at 64 billion dollars for 2026, up from 39 billion in 2025, with domain-specific language models as the fastest growing slice at 210 percent growth in one year. That tells you where money is moving, not whether the product in front of you is worth buying.
What actually makes a vertical product different
The honest ranking surprises people, because the model sits near the bottom of it.
The workflow
The product knows the shape of the job: that a due diligence review runs per document set and ends in a memo, or that a consultation ends in a note somebody has to sign. A general assistant knows none of that and waits for you to say it, every time.The integrations with that sector's systems
Harvey, a legal product, pulls context from iManage, NetDocuments, LexisNexis and SharePoint and runs inside Word and Outlook. Abridge writes its clinical note drafts into Epic. Buildots compares your BIM model and your schedule against 360 degree site captures, drone flights and laser scans. That integration work is slow and unglamorous, and it is most of what you are buying.The vocabulary and the document formats
Every sector has words that mean something precise and file formats nobody outside it has heard of. A product that reads a Belgian VAT listing or a BIM export without being told what it is looking at saves you a prompt and a lot of correcting.The evaluation set built from real cases
Buyers skip this and it may be the most defensible part. Abridge describes an evaluation set of over a thousand rubrics, built by physicians who each reviewed a real de-identified encounter. Harvey built a legal benchmark out of billed time entries, scoring how much of a lawyer-quality work product a model finishes and how many of its correct statements carry an accurate source, because multiple choice tests do not capture the work lawyers actually bill for.The compliance work
Processing agreements, retention settings, audit trails, sector certifications, and in Europe the paperwork the AI Act asks of a deployer. A vertical vendor has been through all of it with the previous twenty customers.A model actually trained on domain data
This is the rare one. Bloomberg did it properly in March 2023 with a 50 billion parameter model trained on roughly 363 billion tokens of financial text alongside 345 billion tokens of general text. Almost nobody has repeated that at that scale, because general models kept improving faster than such a training bill could be earned back.
Ask which of the six you are actually buying. Prompting plus retrieval on top of somebody else's model can still be an excellent product. It should just not be priced as if a model had been trained for you.
Where vertical AI is furthest along, and why
Law, healthcare, accounting, insurance and construction lead, and they share three traits. They run on documents, so reading is the job rather than a side effect of it. They share a vocabulary, so a product can be built around words that mean the same thing at every customer. And their errors are expensive: a missed clause, a wrong dosage or a beam in the wrong place costs far more than the software does. That last gap is what pays for the extra checks, the source citations and the human sign-off. Where a sector fails all three tests, a general tool with a decent prompt does the job.
Buying vertical versus building on a general model
The case for buying
The vendor has already met your edge cases. The odd invoice layout, the contract in French with a Dutch annex, the client who still sends scans: someone else hit those first and the product absorbed them. You are paying for the last mile, which is expensive to build and boring to maintain. The venture firm CRV made the same argument in investor language in May 2026, naming customer data you cannot get elsewhere, workflow integration, regulatory complexity and the labelled edge cases every deployment produces. None of the four is model quality.
The case against
Price per seat bites hardest when only three people in the firm use the thing daily, and lock-in follows: once your review workflow, your clause library and your history live inside one vendor's product, leaving means rebuilding all of it.
The bigger objection is that general models keep catching up. In November 2023 a team at Microsoft showed that GPT-4 with a general purpose prompting recipe, one that used no medical expertise at all, beat the leading fine-tuned medical model on medical exam question sets, and the same recipe transferred to exams in law, accounting and nursing. Domain tuning is not useless. But a capability gap you pay a premium for today can close on someone else's release schedule.
A study published in December 2023 compared unsupervised fine-tuning against retrieval as two ways of getting new domain knowledge into a model. Retrieval consistently won, both for facts the model had seen and for new ones, and the models struggled to absorb new facts through fine-tuning at all. If a vendor says the domain knowledge lives in their weights, that is the claim to press on.
The Belgian reading
The sum works out differently here than in the United States. A Flemish accounting or construction product that knows the local rules can be worth more than a better model that does not. Filing formats, the Belgian chart of accounts, social secretariat data, the paperwork around registered contractors: none of it is hard reasoning, all of it has to be exactly right, and a general model will cheerfully produce a plausible wrong answer. Silverfin, founded in Ghent in 2013 and part of Visma since 2023, is the standard local example. Language coverage weighs heavier here too: a product that handles Dutch, French and English inside one file clears a bar no US-market product was asked to clear. Ask for a demo on your own documents, in your own languages, before anything else.
A worked example: one supplier contract, two routes
Take a construction SME that wants every incoming supplier contract checked for three things: the payment terms, the retention clause and the price revision formula. Forty contracts a year, and getting one wrong costs real money.
Route one, a vertical contract review product. You upload the contracts and it returns the three clauses with the source paragraph beside each, flagging anything unusual against what it has seen elsewhere. The clause taxonomy, the extraction, the citation and the test set behind all of it came from the vendor. Your work is the account setup, the folder connection and reviewing output.
Route two, a general assistant with your own documents. You write the prompt that defines the three clauses in your own words, connect the folder, and require that every answer quotes its source paragraph. Then comes the part people forget: you collect twenty past contracts where you already know the right answer, and you re-run them every time you touch the prompt or the model version moves underneath you. That test set is now yours to maintain.
On the dimension of who does the last-mile work, that is the entire comparison. Route two is cheaper and more flexible right up to the day a model update lands and your test set is the only thing between you and a quiet regression. At forty contracts a year it usually wins. At four hundred, with three reviewers and a partner signing off, route one starts to win.
What to ask a vertical vendor before you sign
Which model is underneath, and what happens when it is deprecated?
Most vertical products sit on a general model from one of the big labs, which is fine. What matters is whether they test before they move. Abridge states its own discipline plainly: when a new foundation model becomes available, they do not swap it in and hope, they evaluate first.
Is anything actually trained on domain data, or is this prompting plus retrieval?
Both answers are acceptable. Only one is usually true, and knowing which tells you how fast a general model could catch up.
Whose data trained it, and can ours be used?
If customer documents improve the shared model, get that into the agreement rather than the sales deck.
What is in the evaluation set?
How many cases, who wrote them, whether they are real work or invented examples, and what the score measures. A vendor without a serious evaluation set changes the subject to a percentage.
What happens to our data, and what leaves with us?
Retention period, region, sub-processors, and what you get in an export when you leave. Clause libraries and review history are the parts that hurt to rebuild.
What to watch out for with vertical AI
A chat box on a sector website is not vertical AI. If the only sector-specific thing is the logo and the example questions, you are looking at a general assistant with a landing page in front of it. Vertical means the workflow, the integrations and the evaluation are sector-specific, not the marketing.
Accuracy percentages without a definition are not evidence. A number only means something once you know what it was measured on, who wrote the test cases and what counted as correct. A vendor quoting its own evaluation set is legitimate. A figure you cannot interrogate is not.
A thin vertical layer can be absorbed. When a product's only advantage is a good prompt over a general model, the lab that makes the model can ship the same feature. Ask what is still worth paying for if that happens, and if the honest answer is nothing, treat the contract as short term.