Resilience Requires the Ability to Take Action
Design, fund, practice
Resilience does not mean preventing every outage. What matters is which business capabilities must be maintained during a disruption and which courses of action organizations prepare for, fund, and realistically test to ensure this.
Resilience Is Not a Contingency Plan
Many companies have backups, emergency manuals, and technical safeguards. Nevertheless, one crucial question often remains unanswered: What must still function during a major disruption to ensure that its impact on the organization and its stakeholders remains manageable?
Resilience does not mean preventing every failure. It means being able to maintain critical business capabilities even under stress—possibly in a limited capacity, through alternative processes, or by combining existing resources in new ways. In this context, “business capability” does not refer to a single system or an isolated process, but rather to the organization’s ability to deliver value relevant to its stakeholders.
This article brings together perspectives from cybersecurity, digital operational resilience, and all-hazards resilience of critical infrastructure. While their regulatory scopes differ, the key question remains the same: What essential services must continue to be provided even in the event of a disruption?
Resilience Starts with Setting Priorities
A completely fail-safe organization is neither realistic nor economically viable. The key factors are which business capabilities must be maintained, what minimum level of performance is temporarily sufficient, and how long an interruption is tolerable. In addition to financial consequences, obligations to customers, employees, partners, government agencies, or society can also determine where this limit lies.
The Implementing Regulation (EU) 2024/2690 takes up this concept for the NIS2 entities it expressly covers. Its annex requires a business impact analysis; Recital 13 specifies, among other things, the maximum tolerable downtime, as well as targets for recovery time, recovery point, and service delivery, to support a comprehensive assessment.
Resilience doesn't mean keeping everything running. It means keeping the right things running for long enough.
Transparency Takes Precedence over Redundancy
Anyone looking to improve resilience quickly thinks of fallback systems, backups, or additional service providers. But before an alternative can be planned, it must be clear what critical business operations actually depend on: people, processes, information, authorizations, technology, service providers, and supply chains.
The crucial question, therefore, is not just, “Which systems and service providers do we use?” but rather, “What business capabilities do we lose if one of these components fails?”
DORA makes this principle regulatorily binding for the financial sector. Financial institutions must, among other things, assess third-party and concentration risks, as well as the substitutability of ICT service providers for critical functions; exit strategies must be in place for such services.
From this perspective, three architectural principles can be derived: ownership, manageable complexity, and compensability.
Ownership Fosters Decision-Making Ability
Crises rarely provide complete information or clear courses of action. Therefore, before a disruption occurs, it must be clarified who is authorized to shut down a critical service, activate emergency operations, prioritize competing business functions, or temporarily accept risks.
NIS2 requires Member States to ensure that the governing bodies of the covered entities approve cyber risk management measures and oversee their implementation. For the institutions it covers, the Implementing Regulation (EU) 2024/2690 provides clearly defined roles, responsibilities, and powers.
Ownership means responsibility backed by actual decision-making authority. Resilience does not require making the perfect decision in every situation. It requires that the right person be able to make a timely decision.
Manageable Complexity
Complexity is inevitable in modern organizations. The key question is whether it remains manageable even under pressure. During normal operations, experiential knowledge, informal coordination, and manual workarounds mask structural weaknesses. A major disruption reveals whether dependencies are understood, responsibilities are clear, and alternative processes are actually feasible.
An architecture that can only be managed under normal conditions is not resilient.
Mastery creates the conditions necessary to use existing skills in a targeted, different way during a crisis.
Compensatory Capacity Instead of a Perfect Replacement
A nearly identical 1:1 replacement for every failed system, service provider, or resource is often expensive or unrealistic. The key is to maintain critical business operations at the required minimum level.
When considering architecture, three basic patterns can be distinguished:
- Substitution: An alternative module performs the necessary function directly.
- Reconfiguration: Remaining capabilities are recombined. For example, if a payment platform goes down, an alternative bank access method, prepared payment lists, manual checks, and a different approval process can work together to maintain the critical part of the process.
- Controlled degradation: The performance level is intentionally reduced. For example, only payroll payments and time-sensitive supplier invoices could be processed initially, while less urgent payments would be put on hold.
Emergency operations are not a zone free of oversight. Simplifications must remain within predefined legal, safety-related, and ethical boundaries.
The CER Directive requires critical infrastructure operators to implement, among other things, measures to maintain operations, establish alternative supply chains, and conduct exercises. The (non-binding) guidelines published in the 2026 Guidelines from the European Commission also list possible measures such as prioritizing critical functions and sharing resources, equipment, or personnel.
It Takes Practice to Turn a Plan Into a Skill
An emergency plan initially describes only an assumption about how people, processes, and technologies will function under stress. Whether this assumption holds up will only become clear during testing.
The ENISA study *NIS Investments 2025* is based on data from 1,080 experts representing organizations from the EU and from highly critical NIS sectors. Thirty percent of the organizations surveyed had not conducted a cybersecurity assessment in the previous twelve months; 49 percent cited business continuity as a key challenge.
At the ECB’s cyber resilience stress test in 2024, 109 banks were confronted with a severe but plausible cyber incident; 28 institutions underwent an in-depth review, including an IT recovery test. The ECB confirmed the existing response and recovery structures but also identified areas for improvement.
An existing plan is documentation. A tried-and-true plan is a skill.
Insights Must Lead to Change
The ENISA *Cybersecurity Exercise Methodology* distinguishes between *Lessons Identified* and *Lessons Learned*: Insights only become actual lessons learned when their implementation leads to change. The methodology, therefore, combines evaluation with action planning, follow-up, and retesting.
In a speech on June 3, 2026, the ECB reported that nearly three-quarters of the findings from its cyber resilience stress test had now been addressed. A key question here is: Where do we have the same structural weakness, even though nothing has happened there yet? This transforms reactive troubleshooting into systemic prevention.
Resilience Is an Investment Decision
Resilience requires time, personnel, and a budget. However, the size of the budget alone says little about an organization's resilience. What matters is which critical business function an investment protects or helps maintain in the event of a failure.
Not every business function requires the same level of protection. For one, redundant operation may be appropriate; for another, a preplanned reconfiguration or controlled minimal operation may suffice.
Investments in resilience, therefore, do not simply finance “greater security” in the abstract. They create concrete options for action in situations where normal operations are no longer available.
Resilience Becomes Evident Under Stress
Three architectural principles encapsulate the essence:
- Ownership fosters decision-making ability
- Manageable complexity enables control
- The ability to compensate creates options for action.
In practice, five questions remain:
- Which few mental faculties must continue to function in the event of a severe impairment?
- What is the minimum performance level that must be maintained, and for how long?
- Who is authorized to make the necessary decisions in the face of uncertainty?
- How can missing building blocks be compensated for through substitution, reconfiguration, or controlled degradation?
- When were these assumptions last realistically tested in collaboration with academic departments, service providers, and partners?
Anyone who cannot clearly answer these questions must first gain a clear understanding of their own operational reality.
Resilience is not demonstrated by what an organization has planned, but by what actually works under pressure.
Written by
Kai Korla works at Arvato Systems at the intersection of cybersecurity, governance, and architecture. He focuses in particular on how security risks can be not only documented but also actively managed through clear accountability and informed decision-making.