A holistic approach to managing resilience

Our project experience shows that an integrated approach is required, one that brings together organisational, technical and operational measures.

Cyber Resilience Management

Bridging the gap between theory and practice

Policies, guidelines and audit evidence are important, but they do not in themselves guarantee resilience. What matters is whether the status of critical digital services consistently aligns with the target state, and how organisations respond to any deviations. Our practical guide shows how organisations can bridge this common gap between theory and practice. 

When ‘compliant’ does not automatically mean ‘resilient’ 

Many organisations invest a great deal of effort in meeting the requirements of standards and regulatory frameworks – from ISO 27001 and IT-Grundschutz to sector-specific guidelines. This results in policies, guidelines, emergency response manuals and, often, successfully completed audits. Nevertheless, practical experience shows that, in a real-world emergency, it is not the documentation that matters, but whether the environment is actually set up and operated as planned, and how the organisation responds to deviations – whether caused by deliberate changes or unforeseen events. With the targeted implementation of NIS2, DORA or KRITIS, the focus is increasingly shifting from mere documentation towards demonstrable effectiveness in day-to-day operations. 

Cyber resilience management comes into play precisely here: it combines conceptual work (What needs to be achieved?) with technical and operational implementation (How can the environment be maintained in this state in the long term, and how can critical digital services be operated in a controlled manner or restored in a controlled way even in the event of cyber-attacks or disruptions?). 

Why concepts and reality diverge 

 Philipp Kleinmanns, SVP Cyber Resilience Management at Materna, explains: 

“A gap can quickly emerge between concept and operation. This is particularly true when configurations and changes are predominantly implemented manually. Without automation and continuous target-actual comparison, there is no guarantee that the infrastructure will remain permanently in the state that was conceptually planned.” 

  • Manual implementation: The more settings are set manually, the more likely inconsistencies are – particularly under time pressure. 

  • Configuration drift: Even once the target state has been achieved, the environment changes continuously due to patches, new systems and new dependencies, and gradually drifts away from the design. 

  • Complexity due to additional layers: Virtualisation, container platforms, cloud services and APIs increase the number of components and, consequently, the points at which deviations can occur. 

  • Lack of end-to-end visibility: Traditional infrastructure monitoring is insufficient if it cannot also address ‘compliance issues’ (e.g. permissions, baselines, hardening status). 

  • Lack of control logic: Resilience is documented, but is not managed as an operational state requiring continuous control, with clear responsibilities and key performance indicators. 

The cause of this gap is rarely a lack of technology, but rather a lack of accountability for the target state in operations. As long as no one is explicitly accountable for ensuring that systems remain in the defined state at all times, drift occurs systematically and independently of the technology used. AI is becoming increasingly relevant, particularly when dealing with drift and anomalies. Not as an additional function, but as an integral part of operational control: it expands the attack surface, but at the same time enables deviations and patterns to be detected at an early stage during ongoing operations, prioritised and responses automated. It thus becomes a key lever not only for defining the target state, but also for continuously maintaining it.

Resilience by Design: Not just ‘patching things up’ after the event 

Philipp Kleinmanns describes ‘Resilience by Design’ as consistent forward thinking: 

“Resilience and security should be taken into account right from the design stage of platforms and infrastructure, rather than trying, under pressure and after the event, to make an environment ‘somehow’ more resilient through piecemeal measures.” 

At its core, this is not primarily about technology, but about safeguarding business continuity and the ability to deliver: limiting outages, managing recovery, and ensuring decision-making capability. Limiting the scope of outages is also a question of architecture: clear decoupling, a limited blast radius and defined dependencies are crucial to remaining capable of acting in an emergency. 

A recurring pattern: resilience is only ‘called upon’ when something goes wrong. It is more effective to factor it in from the outset – just as with secure software development. This applies to infrastructure and platform decisions (e.g. high availability, sensible decoupling, clear dependencies) just as much as it does to fundamental security principles and deliberate architectural and investment decisions. Designing systems in such a way that partial failures do not immediately bring the entire service to a standstill saves valuable time in the event of an incident. 

Roles, decision-making processes, storage locations: every minute counts during an incident 

Technical measures are only effective during an incident if the organisation is also prepared. Anyone who only starts asking, during a security incident or disruption, who makes the decisions, who communicates and where the relevant documents are kept, wastes time and increases the damage. That is why clear responsibilities, reporting channels, escalation levels and a functioning emergency management system (including business continuity management) are essential components of cyber resilience and must be practised regularly. The ‘offline’ approach is also important: what remains accessible if central systems or standard communication channels fail? 

From a one-off project to a cycle: maintaining cyber resilience on an ongoing basis 

Cyber resilience is not a state that you achieve once and then ‘tick off’ as done. New digital services, new threats and changes to the infrastructure constantly give rise to new risks. Successful organisations therefore establish a recurring process that permanently links strategy and operations. The aim is to establish cyber resilience as a controllable operational state: 

  1. Identify and prioritise critical services and associated risks (impact, dependencies, tolerances). 

  2. Define the target state (architecture, baselines, roles, emergency procedures) – ideally as actionable standards. 

  3. Implement automatically – so that the target state becomes reproducible. 

  4. Monitor continuously – status (compliance) and events (security and observability). 

  5. Improve and adapt – rectify deviations, update concepts, automatically incorporate new systems (discovery). 

What matters here is not the number of controls met, but whether organisations can actively manage and improve their resilience using metrics such as control coverage, drift, and detection and recovery times. 

A mini-checklist to get you started: Make drift visible (actual vs. target), define a few clear baselines, automate recurring changes, close monitoring gaps (compliance and events), and test emergency roles and communication channels in drills – not just ‘for the audit’, but for a real-life emergency. 

Our approach to cyber resilience management

  • A risk-based approach: work together to prioritise the scenarios that pose the greatest threat to business operations. 

  • Target vision and roadmap: translating the target architecture, baselines and operational requirements (including roles and the emergency response plan) into actionable steps. 

  • Implement effectively from a technical perspective: resilience by design, hardening and robust identity and access management as the foundation. 

  • Sustainable operation: Integrate compliance-oriented monitoring, security monitoring and observability so that deviations and incidents are identified at an early stage. 

  • Continuous improvement: regular reviews and exercises, updating documentation and controls in response to new services and risks. 

  • Make gaps measurable: Ensure transparency regarding discrepancies between the target state and actual operations (e.g. baseline coverage, drift, recoverability of critical services). 

Your contact person

Philipp Kleinmanns, SVP Cyber Resilience Management at Materna.