Back to Knowledge
BlogMVC & Resilience
Jun 24, 20269 min readYahya Jarraya

The Third Layer Nobody Built

Protection and Recovery are mature. Operational Continuity is not. That is where the damage concentrates.

Executive Summary

Enterprises have built two of the three layers of operational resilience. Protection: everything Cyber Security deploys to keep threats out. Recovery: everything IT rebuilds after an incident. Both are mature. Both are well-funded. Both are necessary.

The third layer, Operational Continuity, is the problem.

When protection fails and a major incident hits, IT does the right thing: it shuts everything down, isolates the network, and starts rebuilding. That process takes more than 4 weeks on average. During those 4 weeks, the business waits. Because nobody prepared it to operate without IT. The BIA identified the critical processes. The BCP documented what should happen. But both are static artifacts: no live business data, no operational governance, no audit trail, no activation mechanism. Built to satisfy a compliance checklist, not to function as an operational instrument during a crisis.

That wait is not a technical failure. It is a strategic one. Paying settlements, honoring deliveries, running payroll, managing inventory, delivering product, registering orders, securing physical assets: none of these stop being obligations because the ERP is offline. And the cost is not the attack. It is the operational silence that fills those 4 weeks.

Most large enterprises have a BCP. They have done a BIA. They have RTOs for individual applications. What they do not have is an executable process that works when systems are down, mapped at the business level, tested by the teams, available in minutes.

They had plans. They lacked execution. They did not have a Minimum Vital Company.

That is the third layer. Not a plan. An execution capability. And it is the one almost nobody has truly built.

1. The Optimization Trap

The more operations are digitized, the more companies lose the ability to function without them. Digitization is mandatory. It is very good. It is a must-have in our digital, automated, and AI era. It works perfectly fine when the machine is online.

Olivier Hamant, biologist and researcher at INRAE, made this exact case in his 2023 essay Antidote to the Cult of Performance (Gallimard): the relentless pursuit of optimization comes at the cost of robustness. Systems built to perform under normal conditions break badly when conditions stop being normal.

This is not an abstract observation. It is the operating reality of every large enterprise today. Every process has been digitized, streamlined, integrated. The same integrations that make operations fast in normal times make them completely unavailable when the underlying systems are down.

2. Three Layers of Resilience

Enterprise resilience has three layers.

Protection is everything the security team has built: firewalls, EDR (Endpoint Detection and Response), detection capabilities, anti-virus, zero trust architecture. Its job is to keep threats out. The owner is Cyber Security.

Recovery is the IT team's plan to rebuild: the DRP (Disaster Recovery Plan), the backup architecture, the infrastructure rebuild timeline. Its job is to restore systems after an incident, in a specific sequence that accounts for dependencies between applications, data flows, and integrations. Not everything comes back at once. Systems are restored in order, each one a prerequisite for the next. The owner is IT.

Continuity is the ability to keep the business running while IT rebuilds. This is what BCP (Business Continuity Plan) and BIA (Business Impact Analysis) are supposed to address. The owner is Business and Operations.

The three layers of enterprise resilience: Protection, Continuity, Recovery

The first two layers are mature. Most large enterprises have invested heavily in Protection and Recovery over the past decade. The third one has not received the same treatment, and that is precisely where the damage concentrates.

3. The Dependency Problem Nobody Talks About

Here is the insight that rarely gets articulated clearly.

When people think about recovering a critical process, they think about recovering the applications that run it. IT gives a Recovery Time Objective (RTO) for each system. Each system has its own timeline. That feels manageable.

The problem is that business processes do not run on one system. They run on chains of systems, with data flowing between them.

Take a concrete example: executing a supplier payment. That single operation requires the supplier portal, the ERP managed by Procurement, and the TMS managed by Treasury, all operational simultaneously, with the right data flowing between them, validated by the right people in the right sequence.

The RTO for that process is not the RTO of the fastest system in the chain. It is the RTO of the last system to come back online, plus the time to rebuild and validate all interfaces and data flows between systems, plus the functional validation that the entire chain is coherent and operationally sound, not just technically restored. A system that IT marks as “back online” and a system that Finance can actually use to execute a payment are not always the same thing.

This logic applies far beyond treasury. Inventory management requires warehouse systems, ERP, and supplier portals to all agree on the same data. Delivery execution requires logistics platforms, order management systems, and customer portals in sync. Production quality requires MES, ERP, and compliance systems to be interconnected. Asset and personnel security requires access control systems, HR databases, and physical security platforms to communicate.

This is the dependency problem. And it often means that when IT says “we will have the systems back online,” they are giving individual system RTOs. Nobody is giving the business an RTO for the process itself, end to end, including interface rebuilds, data flow restoration, and functional validation across all dependencies.

Most large enterprises have never mapped this. They have RTOs for applications. A list of critical applications. They do not have RTOs for business processes. That distinction is where weeks of operational freeze hide.

4. What Actually Happens During an Incident

When a ransomware attack is confirmed, IT does not immediately start restoring systems. The first priority is containment.

The network goes down intentionally. Every system is taken offline to stop propagation. The origin must be identified before anything can be restarted, because restarting an infected system restores the attack, not the operations. Forensics take time. Isolation takes time. Clean infrastructure rebuild from verified backups takes time.

None of this is a failure on IT's part. It is exactly the right response.

But during all of it, the isolation, the forensics, the clean rebuild, the business is in the dark. No ERP. No TMS. No supplier portal. No email system, often. No shared drives. In some cases, no phone directory.

The teams that know how to pay a supplier have never paid one without an ERP. The treasury team that manages daily cash positioning has never done it without their TMS. The logistics team has never processed a delivery without their WMS. They have the knowledge. They do not have the tools, the data, or a step-by-step process designed to work in exactly this situation.

So they wait.

And every day they wait, the meter is running.

5. The Stakes Have Changed

Fifteen years ago, that was acceptable. An IT incident lasted a few days. Systems were simpler, less integrated, faster to restore.

Today, the math is different.

A ransomware event takes more than 4 weeks on average, with systems intentionally kept offline to stop propagation. More than 4 weeks without operations. Not a manageable disruption. Not a blip on the quarterly report.

The documented numbers speak for themselves.

Marks and Spencer: approximately £300M impact on operating income, representing roughly 30% of FY25 operating income, across approximately 7 weeks of disruption.

Jaguar Land Rover: approximately £196M in exceptional direct costs, approximately 8% of FY25 operating income, with 4 to 5 weeks of operational interruption, and the UK government guaranteeing up to £1.5B in commercial borrowings.

These are not worst-case scenarios from a risk model. These are documented recent facts, directly attributable to the inability to execute critical business operations during a crisis.

Add to those figures: market cap damage, customer churn, contract penalties, regulatory exposure, and supplier relationships that do not fully recover. The operational freeze is the most expensive part of a cyber event, more expensive, in many cases, than the ransom itself.

6. Plans Do Not Execute Themselves

Most large enterprises have a BCP. They have done a BIA. They have an emergency contact list and a crisis communication protocol.

What they do not have is a step-by-step executable process that works when systems are down.

The continuity plan exists on paper. The ability to execute it during a crisis is a different problem entirely.

A plan that lives in a SharePoint no one can access during an outage is not a continuity capability. A contact list that exists in an Outlook no one can open is not a crisis tool. A procedure that requires the ERP to be online is not a backup procedure. It is the same procedure, relabeled.

True operational continuity means the critical processes are pre-designed, pre-tested, and executable by the right people with the right data, independently of the primary IT environment. Not as a theory. As a proven operational capability.

That is the definition of a Minimum Vital Company: the minimal set of critical processes and data that an organization must keep executable, to sustain its vital activities during a major disruption. Not a plan. An execution posture.

7. What Good Looks Like

Organizations that have genuinely closed this gap share three characteristics.

First, they have mapped their critical processes at the business level, not the application level. They know which processes must keep running, who owns each one, what data each one requires, and what the end-to-end RTO looks like across all dependencies: system restoration, interface rebuild, data flow validation, and functional sign-off.

Second, they have built executable procedures that work without primary IT systems. Step-by-step. Role-assigned. With the data pre-loaded and kept current. Designed to be activated in minutes, not assembled under pressure. Covering the full spectrum of vital activities: treasury and payments, supply chain and inventory, production and quality, HR and payroll, logistics and delivery.

Third, they have tested them. Not in theory. In practice. The teams have run simulations. They know what to do. The muscle memory is there before the crisis, not built during it.

That is the full picture of operational resilience: protect what you can for as long as you can, recover as fast as the architecture allows, and keep the business running throughout the entire window in between. Three layers. All three built. All three tested.

Not the plan. The capability.

Astran is the operational resilience platform that makes this possible. AlwaysReady® is the independent execution layer that keeps vital activities running when primary systems go down. Not a plan. Not a backup. An execution capability.