When an outage becomes a threat
RPO
How old may the last data state be?
RTO
By when must the process run again? Shorter than the MTPD.
MTPD
From when does the outage threaten the company? The business decides.

Schematic illustration, not to scale.

“We have two mirrored data centres. Nothing can happen to us.” I hear this surprisingly often in conversations about business continuity. It describes a real achievement by IT and still answers the wrong question.

High availability is not an emergency plan

Redundant data centres protect well against a broken power supply or a failed server. Against a logical attack they help little. If malware encrypts the primary system, the mirroring writes the damage to the second one within moments. In the end, two data centres stand still.

Business continuity begins where high availability ends: with the question of how the company keeps working when IT is not there, and how it restarts in an orderly way.

Three figures management should know

BSI Standard 200-4 and ISO 22301 work with a handful of indicators. They sound technical but are business decisions.

MTPD, the maximum tolerable period of disruption. From when does an outage threaten the company’s existence? Through contractual penalties, lost customers, spoiled raw materials, damage to plants or regulatory consequences. Only the business knows this limit, not IT.

RTO, the recovery time objective. By when must a process be running again? It has to be shorter than the tolerable period of disruption. Set without that limit, it is wishful thinking.

RPO, the recovery point objective. How old may the last data state be that you fall back on? In batch production, even a short loss can mean that an entire batch cannot be released.

If IT estimates these values, the plans are technically sound and miss the business. If they are set in a business impact analysis together with the process owners, they become decisions management can stand behind.

The plan has to work where it happens

“When production stops, a 200-page manual does not help. All that counts is whether the team can still fill a container without a computer.”

An emergency manual stored as a file on a network drive is often unreachable in an emergency. Precisely when it is needed. In production there is a further point: the shift supervisor has to act within minutes, not after studying a lengthy document.

What has proven itself, therefore, are short printed instructions where they are needed: how to stop which plant safely, how to switch to manual operation, how to keep up documentation without systems, whom to inform. The detailed plan remains important, but the first hour is decided at the plant.

An unrehearsed plan is an assumption

Only an exercise shows whether a plan works. A moderated half-day exercise with management, production and IT regularly brings to light what no document contains: who actually decides? Who can be reached when the usual channels fail? Which dependency between production, IT and ERP was nobody aware of?

What follows

  1. Set tolerable outage times. For the critical processes, together with those responsible, documented as a management decision.
  2. Map dependencies. Between production, IT, ERP and service providers. That is where the surprises are.
  3. Plan and rehearse the restart. Short instructions for the first hour, a plan for an orderly restart, an exercise that tests it.

BSI Standard 200-4 allows a staged entry. You do not have to start with a complete management system to be able to act. What matters is that in an emergency someone reaches for a plan instead of the phone.

How is it in your case?

In 30 minutes you will know where you stand.

Book an initial call