The decisions that determine recovery are rarely made during recovery

Increasing complexities, outages hampering organizations' data center operations: Survey

An outage is rarely the moment an organisation fails. More often, it is the moment the decisions made months or years earlier become impossible to ignore.

The conversation organisations have after an outage is rarely about the outage itself. They talk about the decisions that preceded it. ‘Why was recovery expected to work a certain way?’Why did responsibilities become unclear?’ ‘Why did restoring business operations take longer than anticipated?’ ‘Why were some assumptions never validated until systems were already unavailable?’

By the time these questions are being asked, they are usually too late to influence the outcome. The infrastructure has already been built. Recovery priorities have already been established. Operational ownership has already been defined. The outage simply exposes the quality of those decisions.

That is why I have increasingly come to see disaster recovery differently. It is often treated as an operational capability that begins when systems fail. In practice, recovery is largely determined long before the first recovery plan is ever activated. The infrastructure choices organisations make, the governance they establish, and the operational discipline they build over time often have a greater influence on recovery outcomes than the actions taken during the outage itself.

Perhaps the more important question leaders should be asking is not “How quickly can we recover when something fails?” but “When were the decisions that will determine our recovery actually made?”

Decisions that shape recovery rarely look like recovery decisions

The more useful question we must ask is not how organizations recover during an outage; it should be about the decisions that determine recovery.

Recovery decisions rarely announce themselves as recovery decisions. They appear as infrastructure trade-offs, cost versus redundancy, simplicity versus resilience, centralization versus geographic distribution. They appear as ownership discussions that are postponed because there are more immediate priorities. By the time those decisions are revisited, they are no longer design choices. They have become recovery constraints.

Individually, none of these feels like disaster recovery decisions. They are practical business and infrastructure decisions, often made for perfectly valid reasons. Collectively, however, they determine whether an organisation can recover with confidence long before the first recovery plan is ever activated.

In my experience, recovery has rarely failed because organisations lacked technology. It fails because the technology was expected to compensate for decisions it was never designed to overcome. Infrastructure has always reflected the priorities of the organisations that build it. Recovery is no different.

Increasingly, I have come to see disaster recovery less as an operational capability and more as an outcome of infrastructure design. Recovery cannot simply be added once systems are in production.

This report by Cockroach Labs on outages speaks to where the industry is struggling. Only 20% of organisations describe themselves as fully prepared for outages. To me, that does not point to a shortage of disaster recovery technologies. Most organisations already have backup platforms, replication capabilities, and documented recovery procedures. It points to something more fundamental: recovery is still too often treated as a capability that can be added later, instead of a design objective that shapes infrastructure decisions from the very beginning.

The implications of that gap rarely become visible during normal operations. They become visible when recovery moves from planning to execution.

Recovery is proven under pressure, not on paper

The consequences of treating recovery as something that can be addressed later often remain invisible until organisations are forced to recover under real operational pressure.

The CrowdStrike outage in July 2024 reinforced this in a way few incidents have. A single faulty software update affected an estimated 8.5 million Windows devices worldwide, disrupting operations across industries. Yet organisations experiencing the same incident did not experience the same recovery. Some restored critical operations within hours. Others spent days bringing essential systems back online. (Source: Microsoft, July 2024)

I don’t believe CrowdStrike changed how organisations recover. It exposed how differently organisations had prepared to recover.

Outages don’t just interrupt systems. They audit every infrastructure decision that came before them. The triggering event was common. The recovery wasn’t.

The difference wasn’t determined on the day systems became unavailable. It reflected decisions that had been made much earlier. How well did teams understand application dependencies? Had recovery ownership been clearly established? Were critical business services prioritised before recovery plans were written? Had recovery assumptions ever been validated under real operating conditions?

These are not questions organisations should be answering during an outage. They should already have been answered while infrastructure was being designed, governance was being established, and operational responsibilities were being defined.

That is why I no longer see disaster recovery as something organisations activate during an outage. I see it as a capability they either build into their operating model from the beginning or spend an outage wishing they had. The difference is rarely visible during normal operations. An outage simply brings it into full view.

The leadership conversations that matter the most happen before an outage

If recovery is shaped long before an outage, then the conversations leadership teams need to have also begin much earlier. In fact, they often begin long before the facts surrounding an incident are fully understood. As the recent Bank of Baroda cyber data leak continues to be investigated, organizations across industries don’t need to wait for the final findings to ask themselves a different set of questions. Would our governance withstand similar scrutiny? Are operational responsibilities clearly defined? Have assumptions around access, ownership, and recovery ever been challenged under real conditions?

If I were sitting in a board review today, these are the questions I would want leadership teams to ask long before the next outage occurs:

  • Which recovery decisions would still need to be made during an outage instead of beforehand?
  • Which assumptions about our recovery strategy have never been validated under real operating conditions?
  • Do we know which business services must be restored first, and have those priorities been agreed across technology and business teams?
  • Is ownership of recovery clearly defined, or will critical decisions depend on individuals coming together for the first time during a crisis?
  • Are we evaluating disaster recovery only on technology capabilities, or also on the operational maturity required to execute recovery with confidence?

Confidence doesn’t come from knowing a recovery plan exists. It comes from knowing the organisation has already removed the uncertainty around the decisions that matter most before recovery is ever required.

These questions don’t eliminate outages. They change how organisations experience them.

The next phase of infrastructure leadership will not be defined by who prevents every failure. It will be defined by who designs systems, governance, and operational discipline well enough that recovery becomes predictable rather than hopeful.

Because in the end, an outage rarely changes an organisation. It simply reveals the quality of the decisions it had already made.

Authored by Vinay Chhabra, Co-Founder & MD, AceCloud

Share on