Vertical AI
Insurance underwriting AI, why it keeps missing the real cases
Underwriting AI handles the easy cases. The hard ones hinge on a senior underwriter's judgment. Here is what changes when you capture that judgment first.
Key takeaways
- Underwriting automation handles the volume tier well; the hard cases live above the codified rules, in senior underwriter judgment that was never written down.
- Roughly 80% of how work actually gets done inside a company is undocumented, and underwriting is one of the most extreme examples.
- MIT NANDA reported 95% of enterprise AI investments show no measurable business impact, and Gartner projects more than 40% of agentic AI projects will be cancelled by 2027.
- A large multi-brand insurance group cut renewal handling from six minutes to one, roughly 833 labor hours saved per 10,000 renewals per month, because the senior analyst's logic was captured first.
If you run underwriting, you already know the shape of the problem.
The straightforward risks, the small commercial policies, the standard auto renewals, the clean applications with clean data, move fast. Automation handles them. The tools work. The demo looks great.
Then a real case lands on the desk. A mid-market manufacturer with a three-year loss run that reads worse than it actually is, because two of the claims were a subrogation win waiting to happen. A specialty risk where the account executive knows the broker has misclassified the operations. A renewal where the loss ratio spiked because of one shock claim that will never recur.
None of these belong in the automated pile. They belong in front of a senior underwriter. And that senior underwriter, the one everyone in the shop routes the hard files to, is exactly the person your current insurance underwriting AI cannot replicate.
This post is about why that happens, and what has to change before an AI system can actually help on those cases instead of just working around them.
What underwriting AI actually does today
Most systems marketed as insurance underwriting AI do a few specific things well.
They pull data from third-party sources. They score risks against actuarial models. They flag applications that fall outside pre-set rules. They accelerate the intake and triage layer, the stack of small, clean policies that used to sit in an underwriter's queue for two days waiting for a five-minute decision.
That is real value. The bottom of the pyramid, the volume tier, is where automation belongs, and where it pays back quickly.
The problem is that the market keeps positioning this same layer as if it could handle the top of the pyramid too. It cannot. And the reason it cannot is not a model problem. It is a knowledge problem.
The hard cases live above the rules
The rules that automation reads from are the codified version of underwriting: the appetite guide, the pricing matrix, the referral triggers, the declinations list. That is the official process.
The real process is what a senior underwriter does when a case does not fit the official process cleanly. And in complex commercial lines, most cases do not fit cleanly.
Here is what the senior underwriter actually does when a hard file lands:
- Reads the loss run and immediately spots the two claims that are outliers, not trends.
- Recognizes the broker from prior submissions and adjusts trust for known submission quality.
- Knows which industry codes are misclassified in the submission form because the underlying operation is different from what the code says.
- Pulls up a state-specific reg they have running in their head, not in the manual, because a similar risk was flagged by the reinsurer last quarter.
- Decides to write it, but with a higher retention, a specific exclusion, and a scheduled endorsement that a first-year underwriter would never think to add.
None of that is in the underwriting manual. None of it is in the appetite guide. It lives in the underwriter. And research suggests that roughly 80% of how work actually gets done inside a company is undocumented [Tallyfy]. Underwriting is one of the most extreme examples of that number.
When a rule-based or retrieval-based underwriting AI hits this file, it does one of three things: it kicks it to a human (fine, but that is what happens today with no AI), it retrieves the official rule and misses the exception (worse than a human), or it produces a plausible-sounding recommendation with no basis in the real logic (much worse than a human). This is the same reason retrieval alone is not enough for enterprise agents. The answer the underwriter needs is not in the document store.
This is why enterprise AI keeps missing the real cases. Not because the model is weak. Because the substrate the model reads from does not contain the tacit judgment that carries the decision, the operating knowledge that a senior underwriter runs on but never wrote down.
Why this is a substrate problem, not a model problem
MIT NANDA reported that roughly 95% of enterprise AI investments show no measurable business impact [MIT NANDA]. RAND has put the enterprise AI project failure rate at around 80% [RAND]. Gartner projects that more than 40% of agentic AI projects will be cancelled by 2027 [Gartner].
Those numbers are consistent across sectors, but underwriting is where they show up most cleanly, because underwriting is a discipline where the codified logic and the real logic diverge sharply and the divergence is exactly what determines whether a book is profitable.
The pattern is always the same. A carrier buys or builds an underwriting AI. It reads from the appetite guide, the pricing model, some third-party data, sometimes a document store. It performs well on the demo cases and on the clean bottom-of-book files. Then it hits the accounts that actually drive combined ratio, the complex commercial renewals, the specialty lines, the accounts a senior underwriter would rework. And the AI cannot see the real considerations, because those considerations were never written down.
Better prompts do not fix this. A larger context window does not fix this. A more recent model does not fix this. The AI is reading from the wrong substrate. It needs to read from what your senior underwriters actually know, not from what your appetite guide happens to say.
That gap is what expertise mining closes.
What changes when you capture the real underwriting logic
Expertise mining is a structured extraction process. It sits with a senior underwriter, the one everyone routes the hard files to, and pulls out the real decision logic. Not a transcript. Not a nicer wiki page. The reasons, the exceptions, the trusted data sources, the escalation triggers, the "I always check this field first because the other one is stale" heuristics.
The output is a machine-readable expert profile, the same underwriter's judgment, in a form an AI system can actually use.
Once you have that captured expertise, the underwriting AI stops being a rule engine and starts behaving like a governed junior underwriter working under the senior. It reads the submission the way the senior would. It flags the same considerations the senior would flag. And when the case genuinely requires human judgment, it escalates cleanly, with a summary of what it saw, why it stopped, and what the senior needs to decide.
That is what governed agent operations look like inside an insurance shop. Not autonomy. Not a black box. A specialist agent that has captured the senior underwriter's operating knowledge and stays inside the boundary where a human is still in control of any decision that matters.
The 6-to-1 renewal, a concrete example
A large multi-brand insurance group in our portfolio ran a renewal workflow that took roughly six minutes per file, a competent underwriter reviewing loss history, running the risk analysis, pulling exposure data across systems, checking the appetite guide, and cutting the renewal terms.
We captured the senior risk analyst's real logic first: how they read a loss run, when they overrode the default retention, which systems they trusted and which they double-checked, and where they escalated. Then we ran a governed agent on top of that captured expertise, wired into the same systems the analyst used.
The workflow dropped from six minutes to one, roughly 833 labor hours saved per 10,000 renewals per month. The risk team went from eight people to three. Not because the agent replaced them, but because the three now do the work that actually needed their judgment. The rest was the same work the agent was now equipped to do, because it had the same operating logic in its head.
The order mattered. If we had started with the automation layer and skipped the expertise mining step, the agent would have been reading from the appetite guide and the loss run alone. It would have handled the clean renewals and missed everything the senior analyst caught on the hard ones. The gains would have been smaller and the mistakes would have been expensive.
This is the pattern across insurance. Underwriting automation without captured expertise handles the easy cases. Underwriting automation on top of captured expertise handles most of the book, with humans focused on what genuinely needs them.
What this means for a chief underwriting officer
If you are running an underwriting operation and looking at insurance underwriting AI, the questions to ask are not the ones the market is asking.
Not "how accurate is the model." The model is fine. The substrate is the problem.
Not "how much can we automate." The wrong things automated aggressively will hurt your loss ratio faster than the right things automated cautiously.
Ask instead: whose judgment are we losing when a senior underwriter retires, and have we captured it? Which parts of the book depend on tacit decisions that were never written into the manual? If we deployed an AI system tomorrow, what substrate would it be reading from, the codified rules, or the real operating logic of our best underwriters? These are the same diagnostics that surface key person risk in operations before the departure forces the answer, and they are what let a carrier scale a senior underwriter's judgment without losing it.
The carriers who answer those questions honestly are the ones who will get real value out of underwriting AI. The rest will keep running pilots that work on the demo cases and stall out on the accounts that actually move the number.
More on the operational side of this, including how underwriting expertise gets captured before senior staff retire, is in our companion post on knowledge transfer in insurance operations, and in the broader argument for why the knowledge leg of enterprise AI has to be built before the action leg.
If you want to see what expertise mining looks like inside an underwriting workflow specifically, that is what our insurance page walks through.
Mine your first process.


