Enterprise Operations

Expert dependency risk, how to measure it and what to do about it

Expert dependency risk is an operational risk, not an HR concern. Here is how to measure it across your critical workflows, and what actually mitigates it.

Ataberk Taçar 9 min read
Cover art for measuring expert dependency risk in operations

Key takeaways

  • Expert dependency risk is the exposure carried when the operating logic behind a critical workflow lives in one or two people and cannot be executed at the same quality by anyone else.
  • With median U.S. private-sector tenure at 3.5 years, your most experienced operator is on average three and a half years from leaving.
  • A five-step audit, workflows, case-types, named people, criticality and substitutability scores, red-cell count, puts a defensible number on the exposure.
  • Hiring adds capacity not expertise, documentation captures the official process not the real one, and cross-training misses the exceptions that did not happen that month.
  • Once operating logic is captured, a renewal cycle went from six minutes to one and the risk team shrank from eight people to three without a drop in quality.

Every operating leader I meet already knows the name. When I ask, "if one person disappeared tomorrow, which workflow blows up?" they answer inside three seconds. Sometimes with a laugh. Sometimes not.

That instant recall is the problem. It means the risk is known, it is named, and it is unmanaged.

Expert dependency is not an HR softness. It is a hard operational risk with a measurable exposure, every bit as concrete as a supplier concentration risk or an uncovered position on the balance sheet. Most companies just do not measure it. This post is about how to actually put a number on it, and what to do once you have.

Why this is a risk officer's problem, not HR's

The framing has drifted for years. Because the risk is embodied in a person, the topic gets handed to HR. HR responds with retention programs, succession planning, and onboarding budgets. All of which are useful. None of which touch the underlying exposure.

The exposure is operational. If one senior underwriter walks, a renewal process breaks. If one planner retires, a plant runs blind on which supplier to trust for the next six months. If one senior consultant leaves the practice, a book of complex cases loses its judgment layer overnight. Those are P&L events. They belong on the same risk register as vendor concentration, cyber exposure, and business continuity. This is the operational version of what we have written about elsewhere as key person risk in operations, a diagnosis you want to make before the departure, not after.

The number that keeps this problem alive is boring and public. In the U.S. private sector, median employee tenure is 3.5 years. [Bureau of Labor Statistics, 2024] Your most experienced operator is, on the industry-average timeline, three and a half years from leaving. Not because they are unhappy. Because that is what tenure looks like now.

Manufacturing is worse. Deloitte and the Manufacturing Institute estimate the sector will have 2.1 million unfilled jobs by 2030, driven substantially by retirement of experienced workers. That is not a talent shortage in the abstract. That is a knowledge shortage, dressed as a talent shortage.

The exposure exists whether or not you measure it. The only question is whether you see it before it hits or after.

What "expert dependency risk" actually means

A working definition, so we do not talk past each other.

Expert dependency risk is the exposure your organization carries when the operating logic behind a critical workflow lives in one or two people, and cannot be executed at the same quality by anyone else in the building.

Three parts matter. It has to be a critical workflow, one where degradation shows up in revenue, cost, or compliance. The operating logic has to live in a person, not in a document, not in a system, not in a policy. And the loss has to be non-substitutable in the short term, you cannot re-hire the judgment inside a quarter.

If all three are true, that workflow is a single point of failure operations exposure. The organization is one resignation, one retirement, one health event away from a real gap.

This is different from "we would prefer not to lose this person." Every company would prefer not to lose most of its people. Expert dependency risk is narrower and sharper, it is the subset of that anxiety that has an operational P&L attached.

The lightweight audit, how to measure it

You do not need a consulting engagement to see this. You need one meeting per critical function and a sheet of paper. Here is the audit we run with operating leaders before we ever talk about tooling.

Step 1: List your critical workflows. Not everything. The ten to twenty workflows that, if they stopped or degraded meaningfully, would show up in the next board pack. Underwriting renewals. Claims triage on complex cases. Production planning for the top three SKUs. Chartering match-making. Multi-jurisdiction tax review. Whatever the equivalent is in your operation.

Step 2: For each workflow, list the case-types. Break the workflow into the categories of case that flow through it. Renewals split into standard renewals, mid-market renewals, high-risk renewals, and disputed renewals. Claims split into straightforward, borderline, and contested. Planning splits into stable-demand SKUs and volatile-demand SKUs. This is the granularity that reveals the risk, not the workflow name, but the case-types inside it.

Step 3: For each case-type, name the person. Ask one question: "if we lost one person tomorrow, which case-types would degrade meaningfully, and by how much?" Force a name. If the answer is "it depends" or "the team," push. In our experience, the honest answer is almost always a specific person for the hardest 20% of case-types.

Step 4: Score the exposure. For each named person, score two things on a simple 1-5 scale. First, criticality: how much of the workflow's economic value passes through their judgment. Second, substitutability: how long it would take, in months, to bring someone else to the same quality. A score of (5 criticality, 5 non-substitutability) is a red cell. A workflow with more than one red cell is a workflow you are actively under-managing.

Step 5: Add up the red cells across the enterprise. That is your expert dependency risk exposure. Not perfect. But real, measurable, and defensible in a boardroom.

The reason this is worth doing is that most organizations have never counted. Research suggests that around 42% of unique expertise sits with a small share of employees. [Panopto] Roughly 80% of operational processes are undocumented. [Tallyfy] Which means, on average, the red cells are hiding in plain sight, nobody has done the arithmetic.

For a deeper walk through what actually leaves with the person, we wrote about it here: The expert everyone calls, and what happens when they leave.

Why the usual mitigations fail

Once the audit is on the table, the reflex is always the same. Hire more. Document more. Cross-train more.

Each of these fails for a different reason.

Hire more. More headcount does not create expertise. It creates capacity. The knowledge dependency is on judgment, exception handling, and the operating logic behind decisions, none of which arrive with a new hire. A new senior person will take twelve to twenty-four months to become useful on the hardest case-types. During that ramp, the risk is unchanged.

Document more. This is the most common instinct and the most expensive miss. Documentation captures the official process, what the company says should happen. It does not capture the real process, what actually happens when a real case hits the desk. The gap between the two is where the expert lives, and no SOP has ever closed it. If documentation could solve this, the roughly 80% undocumented number would not have held for a decade.

Cross-train more. Shadowing is useful, but it is limited by two things. Time, the expert is usually the busiest person in the building. And tacit knowledge, the expert often does not know what they know until the situation surfaces. You can shadow for six months and still miss the reasoning behind the exceptions, because the exceptions did not happen that month.

The right frame is not "reduce dependency by copying the person." It is "capture the operating logic in a form that does not depend on the person being present." The same insight is why the expert bottleneck problem cannot be solved by adding more people around the expert, you have to lift the judgment off the person and put it into the workflow.

The actual mitigation, capture the operating logic

The mitigation is extraction, not more copies of the person.

You sit with the expert. You do it structured. You do not ask "what do you do?" You ask "why this case and not that one?" You ask what breaks when the official process is followed literally. You ask which system field they trust and which one they ignore. You ask what makes them stop and escalate. You separate personal habit from operating logic. You validate what one expert says against what another does on the same case. And you package the result as structured, machine-readable operating knowledge that anyone, a new hire, a senior colleague, an AI agent, can act on with the same judgment.

This is what we call expertise mining. It is a different substrate from documentation. Documentation is written by the person, once, in their voice. Expertise mining is extracted from the person, structured, validated, and kept current as the work changes. Different input. Different output. Different result. If you want the underlying definition of the substrate itself, we unpacked it in what is operating knowledge, and why your documents do not have it.

The reason it matters for risk management is straightforward. Once the operating logic is captured, the exposure changes character. The workflow no longer depends on the person being at their desk. Onboarding a replacement no longer takes twelve months, because the replacement inherits the captured expertise instead of rebuilding it from scratch. And if you are running any kind of enterprise AI or agentic system on top of the workflow, the captured expertise is the input that makes those systems handle real cases instead of only the easy ones. MIT NANDA found that 95% of enterprise AI investments show no measurable business impact. A substantial share of that failure traces back to the same substrate problem. The agents were built without the operating logic they needed.

For a practical walk-through of the capture process itself, we wrote a guide: How to capture tribal knowledge before it walks out the door.

What success looks like on the register

The point of measuring expert dependency risk is not to reduce it to zero. That is unrealistic. The point is to get it into the same conversation as every other operational risk, with a number, a trend, and a mitigation plan.

Success looks like this. The red-cell count for each function is reviewed quarterly. Every red cell has a named owner and a specific capture plan: which expert, which case-types, on what timeline. The exposure trend is visible over time, the same way vendor concentration or cyber posture is trended. When a resignation happens, the impact analysis is a lookup, not a scramble. This is also how succession planning for roles nobody has documented stops being a wishful exercise and starts producing real handovers.

On a live operation, the numbers move fast when the mitigation is applied. At a large multi-brand insurance group we work with, once the renewal expert's operating logic was captured and put into the workflow, the renewal cycle went from six minutes to one minute per case, and the risk team shrank from eight people to three without a drop in quality. The exposure did not vanish, a workflow always benefits from an experienced human somewhere in the loop, but it stopped being a single point of failure. The operating logic belonged to the company, not to a person.

That is the shape of the fix. Not fewer experts. Fewer red cells. Fewer workflows where the operating logic can walk out the door.

If your organization has never counted red cells across its critical workflows, that is the first move. The audit takes days, not months. And it is the only honest way to see the exposure you are actually carrying.

Mine the expertise. Then operate differently.

Share this post

Continue reading

Cover art for scaling expert judgment past the bottleneck
Enterprise Operations

The expert bottleneck, how to scale judgment without losing it

Every operations leader eventually meets the same ceiling. One person's judgment holds the operation together, and the throughput of the entire function is capped at what that person can process in a day. You cannot hire more of them. But you can scale the judgment itself.

Cover art for diagnosing key person risk in operations
Enterprise Operations

Key person risk in operations, how to diagnose it before the departure

Most operations leaders can name their key person risk in about ten seconds. What they do not have is a way to size the risk, defend it in a board meeting, or do something about it before the notice period. This post is the diagnostic, and the mitigation frame that comes after it.

Cover art for succession planning in undocumented roles
Enterprise Operations

Succession planning for roles nobody has documented

Most succession plans are org-chart exercises. A name in a box, a successor in the next box, a review date on the calendar. When the person in the box leaves, the successor inherits the title and none of the twenty years of judgment that made the role work.