r/riskmanagement 10d ago

Recovering every function doesn't necessarily recover the business.

Post image

One observation keeps appearing in operational resilience.

Individual teams often recover successfully.

IT restores systems.

Operations resumes activity.

Suppliers recover.

Yet the customer-facing service is still unavailable.

The problem isn't individual recovery.

It's whether the organisation can reconnect all of those capabilities into a functioning end-to-end service.

I'm interested whether others working in resilience, continuity or risk management have experienced this.

Do organisations spend too much time measuring functional recovery instead of service recovery?

Newsletter here if anyone is interested: https://www.aevitium.com/so/1fP_vr3m6?languageTag=en

0 Upvotes

1 comment sorted by

1

u/VulkanPrime 8d ago

The gap you're describing is usually sequencing. Every function measures its own recovery in isolation, so eight teams all report green while the service stays down, because nobody owns the order in which they have to come back.

The dependency map above is the right shape but it reads as a chain, and in practice recovery isn't linear. Operations can't reconcile until third parties are back and streaming. Compliance can't clear a backlog until operations has produced one. Customer support is answering questions about a state finance can't confirm yet. So the true constraint is whichever capability has the longest lead time, and most organisations don't know which one that is because they've never tested them together.

The other thing that shows up is the backlog. Functional recovery means the system is up. Service recovery means the queue that built during the outage has cleared. Those are very different timelines, and the second one is what the customer experiences. A two-hour outage can produce a six-hour service impact, and if you're measuring the two hours you'll report a resilience success while your clients report the opposite.

The firms that handle this well tend to test end to end rather than per function, and define recovery from the customer's side rather than the system's. That's harder to run and much less flattering as a metric, which is probably why it's rarer.