You already know who it is.
There is one person on your team whose two-week vacation makes you slightly nervous. Not because they slack — the opposite. They are the reliable one. And somewhere in the last two years, without anyone deciding it, a part of your operation moved inside their head and stayed there.
That is a single point of failure: a person, credential, relationship, or piece of knowledge that only one member of your team holds, where losing them stops work rather than slowing it. Assume it is not just one: the one you can name first is rarely the whole list, and the rest stay invisible until the day they are not. (For the term itself, see the bus factor, defined.)
You reduce single points of failure by naming each one, ranking it by how often the task runs and how badly it hurts when it stops, and matching a different countermeasure to each type — cross-training is only part of the answer, and only for three of the eight.
The good news is that these come in recognizable shapes. Once you can name the type, the fix stops being a vague ambition to “document more” and becomes something you can actually put on a calendar. Below are the eight shapes this tends to take, the tell that gives each one away, and the specific countermeasure it needs.
The Short Answer: How to Reduce Single Points of Failure on a Team
The method is naming them, ranking them by frequency and damage, then matching the countermeasure to the type — not trying to make everyone able to do everything. In practice that is four moves:
- Map the tasks, not the people. List the work that has to happen weekly, monthly, and once a year. Beside each one, write the name of everyone who could do it unaided today. Any row with one name is a candidate.
- Rank each candidate by how often it runs and how badly it hurts when it stops. A rare, annoying task is not the same risk as a constant, business-stopping one, and it does not deserve the same investment.
- Match the countermeasure to the type. Documentation fixes knowledge. It does not fix a license, a password, or a relationship. Get this wrong and a coverage project produces a folder of unread wiki pages and no actual resilience — a documented process is no help when the constraint is a license or an admin account.
- Rehearse it once, in the quiet season. A backup who has never done the task is not a backup. A written procedure nobody has followed is a guess.
The reason this is worth an afternoon: turnover is not a rare event. According to the BLS Employee Tenure Summary for January 2024 (opens in new tab), median tenure for wage and salary workers was 3.9 years, and 22 percent of them had been with their employer a year or less. You are not planning for a freak occurrence. You are planning for Tuesday.
And concentration is the default state, not the exception. When Avelino and colleagues estimated “truck factors” across 133 popular open-source projects on GitHub (opens in new tab) in 2016, they found that 65 percent had a truck factor of two or fewer — meaning the loss of one or two developers would leave the project incapacitated. And the study filtered out its smallest-team candidates before measuring — the bottom quartile by number of developers was discarded — so these are not five-person shops. They concentrated anyway. Your five-person shop is not doing better by accident.
If you want to put a number on your own exposure before you start fixing anything, that is a separate exercise — this post is about the shapes, not the scoring. For the scoring, see what the bus factor is and how to measure it on a team.
Match the Fix to the Risk, Not to the Person
Before the eight types, the triage. Two questions decide how much a given single point of failure is worth spending on:
- How often does the task run? Rare and seasonal, or constant and weekly?
- How bad is it when it stops? Merely annoying, or the business stops?
Those two questions produce four responses, and only one of them is expensive:

The same four responses in text:
| Damage by how often | Rare / seasonal | Constant / weekly |
|---|---|---|
| Business stops | Write the runbook. Nobody will remember how this went last time. Document it while it is fresh: the steps, the logins, and who to call. | Duplicate the person — the highest risk of the four. Training is not enough here. You need a second trained — or certified — body, with the renewals staggered. |
| Merely annoying | Write down the number. Not everything deserves a backup. Record who to call, then move on to a risk that actually costs you money. | Cross-train a backup. One named backup who does the task for real, on a rotation — not someone who watched it get done once. |
A runbook, here, means the steps in order that someone who has never done the task could follow without asking anyone.
The quadrant that quietly eats teams is write down the number: rare, and only annoying. It attracts effort because it is easy to document, and it returns almost nothing. Skip it. Write the phone number on the wall and spend the afternoon on duplicate the person — constant, and the business stops — instead.
Hold on to this alongside the types below: the type tells you which fix; the quadrant tells you how much of it to buy, and when.
The 8 Types of Single Point of Failure (and the Fix Each One Needs)
Each one below gives the tell that exposes it and the countermeasure it actually needs — and those countermeasures are not interchangeable. A summary table closes the run.
1. The Undocumented Process Owner
They know the fifteen steps of the month-end close, the reorder thresholds, the exact wording that gets the supplier to expedite. None of it is written anywhere, because writing it down has never been the urgent thing on any given Tuesday.
The tell: questions get answered in real time, instantly, and always by the same person. Nobody consults a document because there is no document to consult.
The fix: have the wrong person write it. Asking the owner to document their own process rarely works — it is boring, it is slow, and they can already do the task faster than they can describe it. Instead, have the intended backup do the task while the owner watches, and have the backup write the steps. The gaps in their draft are exactly the tacit knowledge you were trying to capture.
2. The Relationship Holder
The supplier, the inspector, the difficult key account, the one contractor who answers on a Sunday. The relationship is real and it is with a person, not with your company.
The tell: when someone else emails, the reply comes slower, or does not come. Quotes get worse. The favor that used to be free now has a price.
The fix: a two-name rule on every external relationship that matters. A named second contact gets copied on threads and joins the calls before anyone needs them to — introduced as a colleague, not as a replacement. Relationships cannot be transferred at the moment of departure; they can only be widened slowly beforehand.
3. The Keyholder
They are the admin on the account, the recovery email, the phone that receives the two-factor code, the person with the only physical key to the storage unit.
The tell: imagine locking yourself out of something on a Saturday. If your first thought is “I’d have to call them,” that is a keyholder.
The fix: this one is not a training problem at all — it is an access register. Every system, who has admin, and what happens if that person is unreachable. Move accounts off personal logins and onto role-based ones. That register is what an Onboarding & Offboarding Equipment & Access Tracker is built to hold — a joiner-mover-leaver log of who was given what and when it changed, deliberately a record of access rather than a store of secrets — so the answer on the Saturday is a row in a file rather than a phone call.
4. The Judgment Call
They are not doing a task so much as making a decision: which jobs to take, what to discount, when a batch is not good enough to ship, when to escalate.
The tell: work queues at their desk. Other people are perfectly capable of doing the thing but will not commit to deciding it, so everything waits.
The fix: write down the rule, not the decision. Most judgment is a threshold in disguise. “Discount up to 10 percent without asking; over 10 percent, ask me.” “Reject the batch if more than two in twenty fail the check.” Once the rule is explicit, much of the queue clears without them — and what is left is the genuinely hard call, which is what you actually wanted their attention on.
5. The Certified One
The only person whose license, ticket, food-handler certificate, insurance-required qualification, or regulator-recognized credential makes the work legal.
The tell: one name appears on the compliance line of every job. Their renewal date is the single most load-bearing entry in your calendar and nobody knows what it is.
The fix: you cannot document your way out of this one, and cross-training does not help — the credential is the constraint. You need a second certified person with a staggered renewal date, which is a budget decision made months in advance, not a training decision made in a crisis. Start by knowing what you actually hold and when it lapses; the mechanics of that are covered in how to track employee training and certifications before one quietly expires.
6. The Fixer of the One Weird Machine
The aging piece of equipment, the legacy tool, the spreadsheet with the formula nobody wants to touch. It works, provided one specific human is present to coax it.
The tell: the phrase “just ask them, they know the trick.” A trick is not a process. A trick is a defect with a person taped over it.
The fix: fix the machine, not the person. This is the one type where the correct answer is often to spend money rather than to plan — replace the tool, rebuild the file, buy the service contract. If you genuinely cannot, then it belongs in the write-the-runbook quadrant — rare, but the business stops — so: a real runbook, with photographs, written while the machine is behaving.
7. The Peak-Only Specialist
For fifty weeks of the year, coverage looks fine. During the rush, the fair, the audit, the seasonal spike, one person becomes irreplaceable and everyone quietly knows it.
The tell: your coverage looks healthy in March and terrifying in November. Nobody notices because nobody audits coverage during the peak — they are too busy.
The fix: rehearse in the quiet season. Run the peak task once in the slowest month, with the backup driving and the specialist sitting on their hands. It will be slower and slightly painful, and that discomfort is the entire point — it is the cost of the rehearsal, paid at a time when it does not cost you a customer. A rehearsal is also the only honest way to score depth rather than exposure, which is what the four-step ILUO scale commonly used on Lean skills matrices — in training, limited, unsupervised, operator (able to train others) — is for; a Training & ILUO Skills Matrix is built around that scale.
8. The Willing Martyr
They absorb everything. They stay late, they never hand off, they are genuinely excellent, and they have an alarming amount of unused paid time off. Everyone likes them. They are the most dangerous entry on this list because the risk is disguised as a virtue.
The tell: a growing balance of unused PTO and a shrinking number of tasks anyone else has touched in a year.
The fix: make delegation the measured deliverable. “Train the assistant manager to run the Friday reconciliation by the end of the quarter” is a goal you can review. “Delegate more” is not. Then make them take the vacation, and do not let them take the laptop. A single point of failure who never leaves is not covered — they are just untested.
The eight types at a glance
| Type | The tell | The fix | Cross-training helps? |
|---|---|---|---|
| Undocumented Process Owner | Answers come from a person, never a document | Backup does the task and writes the steps | Yes |
| Relationship Holder | Replies slow down when anyone else emails | Two-name rule, introduced early | Partly |
| Keyholder | “I’d have to call them” on a Saturday | Access register, role-based logins | No |
| Judgment Call | Work queues at one desk | Write the threshold, not the decision | No |
| Certified One | One name on every compliance line | Second credential, staggered renewal | No |
| Fixer of the One Weird Machine | “They know the trick” | Replace the tool, or write a real runbook | Partly |
| Peak-Only Specialist | Coverage looks fine in the off-season | Rehearse the peak task in the quiet month | Yes |
| Willing Martyr | Growing balance of unused PTO, shrinking handoffs | Delegation as a measured deliverable | Yes |
Partly means cross-training closes some of the gap but leaves the real constraint — a relationship, or a machine — in place.
Why Cross-Training Alone Does Not Fix Most of These
Cross-training is the reflex answer to key-person risk, and it is part of the answer for exactly three of the eight types above. For the other five it is either insufficient or beside the point.
Look at the Cross-training helps? column of the summary table. Three types — the Keyholder, the Certified One, and the Judgment Call — are not knowledge problems at all:
- A credential is not knowledge. You can teach someone every step of the work and they still cannot legally sign it off. The fix is a second qualification, purchased and scheduled.
- An access right is not knowledge. Knowing the password is useless if the recovery code goes to a phone in someone else’s pocket.
- Authority is not knowledge. People who understand a decision perfectly will still refuse to make it if nobody has told them they are allowed to.
This matters because the failure mode of a well-intentioned coverage project is easy to picture: you run a training push, everyone shadows someone for a morning, the matrix turns green, and the actual risk is untouched — because the thing that would have stopped you was a renewal date and an admin account.
There is also a legitimate second answer that this post has skirted: sometimes the right response to concentrated capability is not to spread it thinner but to hire someone whose job it is. That trade-off — breadth across the team versus a dedicated specialist — is a real argument with a real dividing line, and it is worked through in cross-train everyone or hire a specialist.
A coverage grid is what keeps the honest version of this visible: which tasks have one name against them, which have two, and which of the twos are real rather than aspirational. A Cross-Training & Coverage Planner is built around exactly that view. Reach for the planner when you need the grid; reach for the ILUO matrix when the grid you already have is the thing lying to you, because a grid that counts “has been shown once” as coverage will always look greener than the team is.
Turn the List Into Typed, Dated Risks
Building the coverage map itself — tasks down the side, people across the top, one mark per person who can run the task unaided — is a well-worn exercise, and it is already worked through step by step in what the bus factor is and how to measure it on a team. If you would rather run the wider version that also scores skill depth and training need, that one is how to identify skills gaps on your team.
Do whichever fits, then come back with the part those exercises hand you and stop at: the list of tasks with exactly one name against them. This is where that list becomes a plan, in two steps:
- Label each one-name row with one of the eight types. This is the step that is easiest to skip, and it is the one that makes the output actionable — a row labeled Certified One and a row labeled Undocumented Process Owner need entirely different money, on entirely different timelines.
- Place each labeled row in the two-by-two. How often across, damage down. Anything landing in duplicate the person — constant, and the business stops — is this quarter’s work. Everything else gets a date or gets accepted on purpose.
The output is not a document. It is a short list of named, typed risks with an owner and a date against each — which is the only form of this exercise that survives contact with a busy quarter.
For the ones that turn out to be genuinely structural — the person whose departure would take a service line with them — a Succession-Planning & Key-Person-Risk Workbook carries the longer-horizon version: who is the named successor, how ready are they today, and what closes the gap.
The takeaway: you cannot eliminate single points of failure on a small team, and chasing zero will burn goodwill you need. What you can do is make sure every one of them is named, typed, and deliberately chosen — so that the only surprises left are the ones you decided to live with.
Common Questions About Single Points of Failure on a Team
Isn’t some key-person risk unavoidable on a small team?
Yes, and pretending otherwise is how coverage plans get abandoned. On a team of four, one person will always be the only one who does something. The goal is not zero single points of failure — it is zero unknown single points of failure, and zero where the task runs constantly and the business stops without it. Everything else you accept on purpose, with the phone number written down.
How do I raise this without making my best employee feel replaceable?
Lead with the vacation. “I want you to be able to take two weeks off without your phone going” is the same project as “I want to survive you quitting,” and it is both truer and easier to hear. For the person who is the single point of failure, being the only one who can do a thing is a tax they are already paying, in interrupted evenings and unused PTO.
How often should we re-run the coverage map?
Twice a year, plus every time the team changes shape. A hire, a departure, a new tool, a new service line, or a new compliance requirement all redraw the map. The twice-a-year cadence exists because concentration re-forms quietly: the person who is best at a task keeps being handed that task, and before long they own it again.