Enterprise Operations
Key person risk in operations, how to diagnose it before the departure
Key person risk is easy to feel and hard to measure. Here is a concrete diagnostic, and what a lightweight expertise mining engagement looks like as mitigation.
Key takeaways
- One question surfaces the risk in any workflow: if one person disappeared for a month, which cases would break.
- The exposure list falls into four categories: exception cases, judgment calls, trusted-source overrides, and relationship-mediated work.
- Roughly 80% of business processes are undocumented and about 42% of essential expertise lives only in employees' heads, exactly where the list lives.
- Cross-training, documentation, succession planning, and AI pilots all read from a substrate that lacks the operating logic, so the list moves but does not shrink.
- A lightweight expertise mining engagement covers one workflow in about a month, with one to two hours of the expert's time in the first structured sit-down.
Most operations leaders can name their key person risk in about ten seconds.
It is the underwriter whose desk everything hard ends up on. The senior planner who overrides the system when it is wrong. The claims manager who quietly closes out the exceptions nobody else can figure out. The one person you actually worry about losing, not for morale reasons, but because you know what would happen to the work the day after they went.
You know the name. What you probably do not have is a way to size the risk, defend it in a board meeting, or do something about it that is not a two-week shadowing scramble the moment they give notice.
This post is the diagnostic. And it is the mitigation frame that comes after it.
What key person risk actually is in operations
In finance, key person risk usually shows up in insurance contracts and business continuity plans, one named executive, one big policy, a line item everyone signs off on.
Operational key person risk is not that. It is smaller, quieter, and vastly more common. It is the risk that a specific workflow only clears because one specific person is on the desk that day. Not a founder or a CFO. A senior operator. A department expert. A team lead who has been in the seat for eleven years and knows every exception in the book.
The reason it does not get on the risk register is that nothing is broken. The work is getting done. The person is still there. The system, on paper, works. From a distance, it looks like healthy operations.
But the exposure is real. In the U.S. private sector, median employee tenure is 3.5 years [Bureau of Labor Statistics]. In manufacturing, Deloitte and the Manufacturing Institute project that 2.1 million U.S. roles could go unfilled by 2030 as the experienced workforce retires, a retirement cliff nobody planned for. The people your operation quietly depends on are not staying forever, and their replacements are not walking in ready.
A single point of failure in the business is only theoretical until it is not. And when it stops being theoretical, the response window is measured in days, not quarters.
The one-question diagnostic
There is a very simple test that gets you most of the way to a real answer. It works in an operations review, a QBR, or a fifteen-minute conversation with a department head.
Take one workflow. Pick any workflow that matters, renewals, claims, planning, chartering, procurement approvals, incident triage, whatever the operation actually turns on. Then ask one question:
If one person disappeared for a month, which cases in this workflow would break?
Not "would it slow down." Not "would we manage." Which specific cases would either stop moving, get done wrong, or need to be escalated to someone above pay grade because the person who normally handles them is not there.
If the answer is "none," the workflow is genuinely resilient. Congratulate the team. Move on.
If the answer is a non-empty list, even a short one, you have key person risk in that workflow. That is the whole test. You do not need a scoring matrix. You need to be honest about the list.
Repeat it across every workflow that matters. What you end up with is not a heat map. It is a list of workflows with named exposure, and inside each, a named set of cases that only clear because one person is on the desk.
That list is your expert dependency audit. It is a lot more useful than the risk register you already have, because it is grounded in actual work.
What ends up on the list
In our engagements the list tends to fall into a few consistent categories. If you run the diagnostic yourself, expect to see these:
Exception cases. The cases where the official process does not quite fit. Non-standard renewals. Claims that sit between two policy tiers. A supplier situation the system has no code for. In writing, these look like edge cases. In the actual month, they are a meaningful share of the volume and almost all of the risk.
Judgment calls. Cases where the data is fine but a decision needs to be made: approve, escalate, decline, restructure, delay. The person who normally makes the call has a decision logic they have never explained out loud. Their replacement can read the same file and land somewhere different.
Trusted-source overrides. The expert knows which field in the system is stale, which report to ignore, which colleague to call before signing. That knowledge is not in a SOP. It is in their head. When they are out, the substitute takes the data at face value and something goes wrong two weeks later that nobody traces back to the substitution.
Relationship-mediated work. The vendor who only responds when it is that person calling. The regulator who has a working shorthand with the same operator across three cycles. The customer who expects a specific voice on the account. These are not soft skills. They are functional dependencies.
Roughly 80% of business processes are undocumented [Tallyfy]. In practice, that undocumented layer is exactly where the list lives. Panopto's research puts it another way: about 42% of essential expertise lives only in employees' heads [Panopto].
You cannot mitigate what you cannot see. The diagnostic makes the list visible.
Why the standard mitigations do not work
Once the list exists, the reflex is to reach for the tools everyone already has.
Cross-training. In theory, the number-two person on the desk absorbs the expert's work through exposure. In practice, they absorb the visible layer, the routine cases, and the exceptions still route to the expert. Nothing changes about the list.
Documentation. Write it all down. Update the SOPs. Fill in the wiki. This helps at the margins. It does not close the list, because the knowledge on the list is tacit. The expert cannot fully articulate what they do until a real case is in front of them, and by then the documentation project has already ended.
Succession planning. A named successor is identified, groomed, and put on notice. But succession planning for roles nobody has documented transfers the title, not the operating logic. The successor becomes the new single point of failure. The list moves. It does not shrink.
AI pilots. A retrieval or agent tool gets pointed at the workflow. It handles the easy cases well. It fails on the same cases that are on the list, because those cases were never written down, so no retrieval system can reach them. McKinsey has reported that 95% of enterprise AI investments show no measurable business impact. This is one of the reasons why.
The pattern is the same across all four. They all read from a substrate that does not contain the operating logic that makes the list what it is.
What a lightweight expertise mining engagement looks like
The alternative is direct. If the risk is the tacit judgment of one person, the mitigation is to capture that tribal knowledge before it walks out the door, structure it, and make it usable outside that person's head. That is expertise mining.
In practice, an early engagement is small. It runs on one workflow, not the whole operation. It follows a pattern we have repeated across insurance, manufacturing, maritime, and consumer supplements.
Step 1. Sit with the expert. Not a shadowing exercise. A structured conversation, working from real recent cases. The goal is to pull out the decision logic, the exceptions, the escalation triggers, the trusted sources, and the specific reasoning behind why one case gets approved and a nearly identical one does not.
Step 2. Turn it into a machine-readable expert profile. The output is not a transcript. It is not a tidier wiki page. It is a structured artifact, an expert profile, that captures the operating knowledge in a form an AI system can act on and a person can review.
Step 3. Connect it to the systems the work actually lives in. The profile has to know which system holds which field, whether the data is fresh, and what an expert is allowed to do with it: read it, draft with it, or stop and ask a human. This is the Company Expertise layer.
Step 4. Run governed agent operations on top of it. The expert profile does not replace the expert. It handles the volume of cases that were the expert's real workload, drafts the ones that need a decision, and escalates the ones that need a human. Every action that matters asks for approval first.
The timeline for one workflow is about a month, and the expert is involved for one to two hours in that first sit-down, plus incremental review after. Not a two-year knowledge management program. One workflow at a time, in a rhythm the operation can absorb.
We have seen what this changes in production. At a large multi-brand insurance group, the risk team on one line went from eight people to three, and the renewal workflow dropped from six minutes to one. The exception cases stopped routing to a single desk. At a maritime chartering firm, an eight-person team now runs with one person on the desk, because the expertise that used to sit in eight heads now sits in a profile that runs next to the remaining operator.
The list from the diagnostic is what you work down. Workflow by workflow. Expert by expert.
How to run this in your own operation this quarter
You do not need us to run the diagnostic. Do this:
- Pick the five workflows that matter most to the P&L or to compliance.
- In each, ask the department head the question. Which cases would break if one person disappeared for a month?
- Write the list. Names of cases, name of the person the risk sits with.
- Rank the workflows by list length, weighted by the value or exposure of the cases on it.
- Start at the top of the list. Not the biggest workflow, the one with the tightest key person concentration.
That is the expert bottleneck you actually have to solve. Everything else can wait.
And when you decide to do something about it, do not start with tooling. Start with the expert. Their judgment is the asset. Capture it while it is still in the building.
Mine the expertise. Then operate differently.




