Executive summary
Operational resilience has become a documented discipline in most large organisations and a tested one in very few. There are plans, owners, recovery objectives and an annual exercise, and yet real incidents still produce the same discovery every time: a dependency nobody had written down, a recovery step that assumed a person who has left, a workaround that requires the system that is down.
The gap is not effort. It is that the testing is designed to pass. This paper sets out an approach that produces uncomfortable results on purpose: name the services that matter, set tolerances that can genuinely be breached, map what those services actually depend on rather than what the architecture diagram claims, then run scenarios hard enough to break something and treat the break as the return on the exercise.
Start with services, not systems
Resilience programmes that begin with an application inventory end up defending assets in order of how well documented they are. Begin instead with the handful of things the organisation must keep doing for customers, patients, regulators or the market. Everything else is a dependency of one of those, and its priority follows from that rather than from its owner's seniority.
Most organisations have fewer of these than they expect, usually between five and fifteen. If the list runs past twenty, the definition has slipped from services the outside world depends on to activities the organisation performs, and the exercise loses its edge.
Impact tolerances that bite
A tolerance is the point at which disruption stops being an inconvenience and starts causing real harm. It has to be a number, with a unit and a deadline, and it has to be set by people willing to be held to it.
- Expressed in outcome terms, such as orders unfulfilled or patients unable to be seen, rather than in system uptime
- Set by the executive who owns the service, not by the technology function that will be judged against it
- Tight enough that a plausible scenario can breach it, otherwise the test cannot fail and cannot teach
- Reviewed when the operation changes, because tolerances set for a business you no longer run are decoration
A tolerance nobody could ever breach is not a standard, it is a comfort. The test of a good one is whether the person who set it looks slightly uneasy about it.
Map what the service really depends on
Every important service rests on a chain of people, systems, data, facilities, suppliers and third parties. The map has to be the real one, including the undocumented parts: the spreadsheet that reconciles two systems, the single engineer who knows the failover, the supplier who is subcontracting without telling you.
- People, and specifically the individuals whose absence stops the service
- Systems and the data they depend on, including the integrations nobody owns
- Facilities, utilities and the physical estate
- Suppliers and their suppliers, which is where concentration risk quietly enters the picture
- Third-party platforms, where the tolerance is set by someone else's recovery, not yours
Where a dependency map exists as a static document, assume it is wrong in at least one material way. The version that stays true is the one wired to the systems that would reveal a change, which is the argument we make in detail in autotwins.
Designing a scenario worth running
Severe but plausible is a narrow band, and both edges are failure modes. Too mild and it passes without effort. Too extreme and the room dismisses it, because arguing about whether the scenario is realistic is a very effective way to avoid discussing whether you would survive it.
- Anchor in something that has happened somewhere, to your industry or an adjacent one
- Combine two ordinary failures rather than inventing one spectacular event, since that is how real incidents are built
- Remove the thing the plan quietly assumes, such as the recovery site, the key person, or the ability to communicate
- Run it long enough to reach the second wave, because most plans hold for four hours and fail on day three
- Include the outside world: customers calling, regulators asking, the press, and staff who want to know if they should come in
Three levels of testing
Not everything needs a live exercise, and pretending otherwise is why testing gets deferred. Match the method to what you are trying to learn.
- Tabletop. Cheap, frequent, and good at exposing decision rights, communication and the assumptions in a plan. Weak at proving anything technical works
- Simulation against a model. Run the scenario against a live model of the operation to quantify what actually breaches and when. This is where most of the analytical value now sits, without touching production
- Live exercise. Expensive and disruptive, so reserve it for the failure modes where human response under real pressure is the thing being tested, or where a technical recovery has never actually been performed
The result is the value
The output of a good test is a list of things that did not work, which is uncomfortable to present and far more valuable than a clean report. Protect that by separating the exercise from performance assessment. If a failed test reflects badly on the team that ran it, you will get clean tests forever and learn nothing.
- Record what breached, at what point, and why, without softening the language
- Assign each finding an owner and a date, and track it like any other engineering defect
- Re-test the specific failure, since remediation that is never re-tested is an assumption
- Report to the board in outcome terms: which services, which tolerances, what is not yet fixed
- Feed the findings back into the tolerances, because sometimes the honest answer is that the tolerance was wrong
Common pitfalls
- Designing the scenario so the plan survives it
- Setting tolerances in system uptime, which nobody outside the organisation experiences
- Testing recovery of a system while skipping the people and communications around it
- Stopping the exercise at the four hour mark, before the interesting failures start
- Treating third parties as out of scope when your tolerance depends on their recovery
- Letting findings become a report rather than a tracked backlog with dates
How WAJD Group helps
We help name the services that matter, set tolerances that can be defended, build dependency maps that stay current rather than decaying, and design and run scenarios that are allowed to fail. Where a live model of the operation exists, we run the analytical scenarios against it so the disruption stays in the model. Then we run the remediation loop as a managed service, with board-ready evidence at the end of it. See the connected thinking in engineering manufacturing resilience and our resilience case study.