Pillars of Risk Management Operational Resilience

0
0
Hire Risk Expert
You agree to our Terms and Conditions of Use, PDPA & Privacy Policy and Cookies Policy

Pillars of Risk Management Operational Resilience

OPERATIONAL RESILIENCE: KEEPING WHAT MATTERS WORKING WHEN THINGS GO WRONG

OPERATIONAL RESILIENCE AS ONE OF THE PILLARS OF RISK MANAGEMENT

On the main Pillars of Risk Management page, I explained why my team and I kept Operational Resilience alongside Technology Risk, Enterprise Risk Management, Cyber Risk, Business Continuity, Governance & Compliance, Third-Party Risk, Risk Assessment and the Risk Register even as we added newer areas such as A.I. & Risk, Human Behaviour, Future of Work, Trust, The Future Human and Signals.

Operational Resilience is one of those established foundations because no individual, small business or organisation can prevent every disruption. Technology can fail. A supplier can stop delivering. Someone important can suddenly become unavailable. A cyberattack can interrupt access to information. An extreme weather event can affect premises or transport. An unexpected event such as COVID-19 can change how people work, buy and communicate almost overnight.

This is where Operational Resilience asks a different question from many other areas of Risk Management. Instead of concentrating only on “How do we stop this from happening?”, it also asks “If it happens anyway, can we still deliver what really matters?”

That question links naturally to the other Pillars already developed on this website. Technology Risk helps us understand what technology we depend upon. Cyber Risk considers what happens when digital systems or information are deliberately compromised. Enterprise Risk looks at what those disruptions mean for the wider business. Operational Resilience then asks whether the important service or activity can continue, recover or adapt despite the disruption.

The newer areas of this website make the same question increasingly relevant. What happens if a business becomes dependent upon A.I.? What if employees can no longer perform an important task without a particular digital platform? What if several external providers ultimately depend upon the same underlying service? What if an emerging geopolitical or climate event disrupts something the business never realised was critical?

These may sound like new problems, but many lead back to the established Operational Resilience question: what must continue working, what does it depend upon, and what happens when one of those dependencies fails?

WHAT IS OPERATIONAL RESILIENCE?

Operational Resilience is the ability of a business to continue delivering its important products, services or activities when disruption occurs, and to recover or adapt when normal operations are affected.

For a small business, this might mean continuing to take bookings when the online booking system is unavailable, serving customers when one employee is unexpectedly absent, processing payments when the usual payment method fails or finding another way to obtain an important product when a supplier cannot deliver.

Operational Resilience is therefore not simply about having backup technology. A business depends upon people, processes, technology, information, premises, suppliers and other external services. The failure of any one of these can affect the ability to deliver what customers actually need.

The simplest way to understand Operational Resilience is therefore to begin with the outcome rather than the disruption.

Ask:

What does this business absolutely need to keep doing?

Then ask:

What do we depend upon to do it?

Only after that do we start considering what could disrupt those dependencies and what alternatives might be available.

HOW DID OPERATIONAL RESILIENCE COME ABOUT?

The idea behind Operational Resilience is not new. Individuals and businesses have always found ways to continue when normal arrangements were disrupted.

A shopkeeper might have used a manual receipt book when the electronic system failed. A small company might have kept spare keys, an alternative supplier or another person who knew how to perform an important task. A family business might have stored copies of important documents somewhere other than the main workplace.

People did these things because experience taught them that something eventually goes wrong.

Over time, professional Risk Management developed more formal disciplines around continuity, disaster recovery, incident management and resilience. As businesses became more dependent upon interconnected technology, external providers and complex processes, it became increasingly clear that recovering one system or maintaining one contingency plan was not always enough.

A customer does not care whether the technology team successfully restored one server if the service the customer needs is still unavailable. A business owner does not benefit from knowing that every individual supplier passed a risk assessment if several of those suppliers fail together because they depend upon the same underlying provider.

Operational Resilience therefore developed towards a broader question: can the important service still be delivered when disruption occurs across the chain of things required to provide it?

That is why contemporary Operational Resilience places so much emphasis on dependencies and end-to-end services rather than looking at each component in isolation.

WHAT OPERATIONAL RESILIENCE QUESTIONS DID INDIVIDUALS AND SMALL BUSINESSES USED TO ASK?

Individuals and small-business owners have asked Operational Resilience questions for years without calling them Operational Resilience.

A shop owner might ask, “What do I do if the cash register stops working?”

A small clinic might ask, “What happens if the receptionist cannot come to work?”

A tradesperson might ask, “What happens if my van breaks down tomorrow?”

A restaurant owner might ask, “What happens if my usual supplier cannot deliver?”

A consultant running a one-person business might ask, “If my laptop dies today, can I still work tomorrow?”

These are all versions of the same resilience question: what do I depend upon, and do I have another way of continuing if that dependency fails?

The answers were often simple. Keep a spare device. Train another employee. Maintain an alternative supplier. Keep copies of important information. Know how to perform a critical activity manually. Have emergency contact details somewhere accessible.

Those measures may look informal compared with professional resilience frameworks, but the underlying principle is exactly the same.

WHAT OPERATIONAL RESILIENCE TOOLS CAN A SMALL BUSINESS USE?

A small business can begin building Operational Resilience by identifying the few activities that matter most, understanding what those activities depend upon and deciding what alternatives are available when something fails.

The first useful tool is a simple critical-activity or important-service assessment. Instead of listing everything the business does, identify the activities whose disruption would have the greatest impact on customers, income or the ability to operate.

A small wellness business, for example, might identify appointment booking, payment collection, access to customer records and the actual delivery of the service as especially important. A small online retailer may depend upon order processing, payment, inventory information and delivery.

The next tool is dependency mapping. For each important activity, consider the people, technology, information, premises and suppliers needed to make it work.

A booking process may depend upon a cloud platform, internet access, customer information and someone who knows how to operate the system. A delivery service may depend upon stock, a courier, customer addresses and access to the order-management system.

The next question is “What could we do instead?” This introduces fallback arrangements, alternative suppliers, backup systems, manual processes, cross-training and recovery plans.

A small business does not need a complicated resilience programme to answer these questions. It needs to know what is genuinely important and avoid discovering its most critical dependency only after that dependency has failed.

HOW IS OPERATIONAL RESILIENCE DIFFERENT FROM BUSINESS CONTINUITY?

Operational Resilience and Business Continuity are closely related, but they are not exactly the same. Business Continuity focuses more specifically on how a business prepares to continue or recover operations during disruption, while Operational Resilience takes a broader end-to-end view of whether important services can withstand, respond to and recover from disruption.

A Business Continuity Plan may tell employees what to do when the office becomes unavailable or an important system fails. It may contain contact details, recovery priorities, alternative working arrangements and procedures.

Operational Resilience asks a broader question: even if each individual continuity plan works, can the service that matters to the customer actually continue?

For example, a business may have a technology recovery plan, an alternative workplace and a supplier contingency plan. But if all three arrangements depend upon the same internet service or the same key employee, there may still be a hidden vulnerability.

The two disciplines therefore complement one another. Business Continuity provides important practical recovery tools. Operational Resilience helps the business look across those tools and dependencies from the perspective of the service that must continue.

The detailed treatment of Business Impact Analysis, recovery objectives and Business Continuity Plans belongs under the Business Continuity sub-index. This page focuses on the wider resilience question.

HOW DOES OPERATIONAL RESILIENCE LINK TO TECHNOLOGY RISK?

Technology is now one of the most important dependencies behind many business services, so Technology Risk and Operational Resilience increasingly overlap.

Technology Risk asks whether the systems, applications and platforms a business relies upon are reliable, recoverable and appropriately managed. Operational Resilience asks what happens to the business service if that technology becomes unavailable anyway.

Consider a small business whose entire booking, customer records and payment process sits on one online platform. The dependency itself is a Technology Risk. The Operational Resilience question is whether the business can still serve customers when that platform becomes unavailable.

This might require a manual appointment list, another way to collect payment, accessible backup information or an alternative provider.

Technology Risk therefore helps identify and manage the technology dependency. Operational Resilience focuses on whether the important activity survives the loss of that dependency.

HOW DOES CYBER RISK CONNECT TO OPERATIONAL RESILIENCE?

A cyberattack is one possible cause of operational disruption.

Ransomware may prevent a business from accessing records. An account takeover may interrupt an online service. A cyberattack against a supplier can indirectly disrupt a business that was never attacked itself.

Cyber Risk focuses on preventing, detecting and responding to malicious digital activity. Operational Resilience asks what happens when the cyber controls do not prevent the disruption and the business must continue operating anyway.

This creates an important progression across the Pillars. Cyber Risk asks how the attack is prevented and contained. Technology Risk looks at the affected systems and recovery. Business Continuity provides the plans and alternative arrangements. Operational Resilience asks whether the important service is actually still being delivered.

A single event can therefore move across several Pillars without the Pillars becoming duplicates of one another.

HOW DOES OPERATIONAL RESILIENCE BECOME AN ENTERPRISE RISK ISSUE?

Operational Resilience becomes an Enterprise Risk issue when disruption threatens the wider objectives, financial health, customers or reputation of the business.

A payment outage lasting ten minutes may be inconvenient. The same outage lasting several days may affect revenue and customer confidence. If the business is already experiencing cash-flow pressure, the consequences may become much more serious.

A supplier disruption can similarly move from an operational issue into an Enterprise Risk if the business cannot fulfil orders, loses major customers or has no affordable alternative supplier.

Enterprise Risk therefore helps us see what a resilience failure means for the business as a whole. Operational Resilience concentrates more specifically on the ability to continue delivering what matters through the disruption.

HOW HAS OPERATIONAL RESILIENCE CHANGED?

Operational Resilience has changed because the things businesses depend upon have changed.

One major shift has been digital dependency. A small business may now rely upon cloud software, online payments, email, digital identity services and third-party platforms for activities that once had local or manual alternatives. A disruption to one online service can therefore affect several business activities at once.

Another change is the growth of outsourcing and platform-based business models. Small companies can now operate with very few internal resources because payroll, payments, accounting, marketing, communications and technology can all be obtained from external providers. This creates flexibility, but it also means that much of the business may depend upon organisations it does not control.

Remote and hybrid work have changed dependencies too. Employees may no longer sit together in one workplace. Collaboration tools, home internet connections, cloud services and remote access have become part of the operating environment.

Demographic change creates another form of resilience risk. An ageing workforce can increase the importance of succession and knowledge transfer. Small businesses may rely heavily on one experienced employee who understands a process that has never been documented. At the other end of the workforce, newer employees may rely more heavily on digital tools and A.I., creating questions about whether important skills remain available when the technology is removed.

Consumer behaviour has changed as well. Customers increasingly expect immediate online access, rapid responses and digital transactions. A disruption that might once have been tolerated for a day may now be noticed within minutes.

Supply chains have become more interconnected. Geopolitical developments, transportation disruption, climate events and shortages can affect even small businesses through suppliers several steps removed from them.

Operational Resilience has therefore moved beyond preparing for a single obvious disaster. It increasingly involves understanding chains of dependency and combinations of disruption.

HOW DID COVID-19 CHANGE OPERATIONAL RESILIENCE?

COVID-19 was a particularly powerful Operational Resilience lesson because it showed that disruption can affect people, premises, suppliers, technology, customer behaviour and working arrangements at the same time.

A traditional disruption scenario might have assumed that the office was unavailable but employees were otherwise able to work. COVID-19 created much wider questions. Could people work from home? Did they have suitable technology? Could suppliers continue operating? Were customers still buying in the same way? What happened when several employees became unavailable together?

Many businesses discovered that the ability to work remotely depended upon cloud services and home connectivity that had previously been considered convenient rather than critical.

COVID also demonstrated the importance of adaptability. Some businesses did not simply activate an existing continuity plan. They changed the way they delivered their product or service, introduced new digital channels and redesigned processes in response to circumstances that were continuing to evolve.

This is an important distinction. Resilience is not always about returning immediately to the old way of working. Sometimes resilience involves adapting to a new way of operating.

That insight remains relevant well beyond the pandemic.

HOW HAVE CHANGING WORK BEHAVIOURS AFFECTED OPERATIONAL RESILIENCE?

Hybrid work, remote work, flexible employment, freelancing and greater use of external specialists have changed where operational capability sits.

A business may no longer have all the people it depends upon physically together. It may use a bookkeeper working remotely, an external technology provider, freelance marketing support and cloud applications supplied from different countries.

This creates flexibility, but it also changes the resilience map.

A business now needs to consider not only whether an employee is available but whether that person has access to the systems and information required to work. It may need to understand whether an external specialist is the only person who knows how an important process operates.

Knowledge therefore becomes an important resilience dependency.

This is why documentation, cross-training, succession planning and knowledge continuity are becoming more significant, particularly for small businesses where one person's absence can have a disproportionately large effect.

This connects Operational Resilience naturally with Future of Work → Workforce Resilience, Future Skills and Ageing Workforce.

HOW HAS THIRD-PARTY DEPENDENCY CHANGED OPERATIONAL RESILIENCE?

Third-party dependency has changed Operational Resilience because many important services now rely upon organisations outside the business.

A small company might use an external payment processor, cloud accounting platform, booking service, courier, website host and A.I. provider. Each supplier makes the business easier to operate, but each can also become part of the chain required to deliver a service.

The first step is to understand which third parties are genuinely important. The next is to understand what happens when they fail.

The toolkit therefore extends beyond simply maintaining a supplier list. Businesses may consider alternative suppliers, data portability, contractual arrangements, service monitoring, exit plans and concentration.

Concentration becomes particularly interesting when apparently separate suppliers rely upon the same underlying service. Several applications may use the same cloud provider. Different A.I. products may depend upon the same foundation model.

The business can therefore have several vendors but still possess only one real dependency underneath them.

The detailed assessment of vendors belongs under Third-Party Risk. Operational Resilience looks at whether the business can continue delivering its important service if that external dependency disappears.

HOW IS THE OPERATIONAL RESILIENCE TOOLKIT CHANGING?

The Operational Resilience toolkit is evolving from mainly maintaining recovery plans towards understanding important services, end-to-end dependencies, concentration and the ability to withstand more complex combinations of disruption.

Established tools remain important. These include Business Impact Analysis, Business Continuity Plans, recovery objectives, dependency mapping, scenario testing, alternative arrangements, incident management and exercises.

The difference is increasingly in how those tools are used.

A traditional dependency list might record which technology and suppliers support a business activity. Contemporary end-to-end service mapping attempts to show how people, processes, technology, information, facilities and external providers connect across the whole service.

Scenario testing is also developing beyond single-event scenarios. Instead of considering only “What happens if the office is unavailable?”, businesses can examine combinations such as a technology outage occurring while key people are absent, or a supplier failure occurring during a period of already high demand.

Concentration analysis is becoming more useful because different services can share the same hidden dependency.

The toolkit is therefore changing in response to a world that is more digital, distributed, outsourced, interconnected and fast moving.

WHAT ARE RISK PRACTITIONERS BEGINNING TO USE OR ADAPT NOW?

Risk practitioners are increasingly adapting Operational Resilience tools towards end-to-end service mapping, more realistic scenario testing, continuous monitoring of critical dependencies and greater attention to concentration and substitutability.

One important shift is from asking “Which processes are critical?” towards asking “Which services or outcomes really matter, and what has to work for us to keep delivering them?”

That shift changes the map. Instead of looking only within one department or process, the practitioner follows the service across people, technology, information and suppliers.

Another development is the use of severe-but-plausible scenarios. The objective is not to predict the exact event that will occur. It is to ask whether the business can withstand a sufficiently difficult but credible disruption.

Scenario testing is also becoming more interconnected. A technology failure may be combined with staff absence, supplier problems or high customer demand rather than tested in isolation.

Practitioners are also looking more closely at single points of failure, concentration, substitutability and recovery capability. Knowing that an alternative supplier exists is useful; knowing whether that supplier can actually take over quickly enough is more useful.

In more technology-intensive environments, some organisations are also experimenting with continuous resilience testing and forms of chaos engineering, deliberately testing how systems behave when selected components fail. These approaches will not be necessary for most small businesses, but the idea behind them is useful: resilience should be tested rather than merely assumed.

HOW IS A.I. CHANGING OPERATIONAL RESILIENCE?

A.I. changes Operational Resilience when it becomes sufficiently embedded in an important activity that the business begins to depend upon it for continued delivery.

Suppose a small company initially uses an A.I. assistant only to draft marketing material. If the A.I. becomes unavailable, the inconvenience may be minor.

Now suppose the business gradually uses A.I. to handle customer enquiries, analyse orders, prepare documents, schedule appointments and support important decisions. The dependency has changed.

The Operational Resilience questions now include whether the A.I. service is always available, whether another provider could be substituted, what happens if the underlying model changes, whether a human can perform the task manually and how long the business can function without the tool.

This is where traditional resilience tools can be adapted. Dependency mapping can include A.I. services and foundation-model providers. Scenario testing can include loss of the A.I. service. Fallback planning can determine whether a manual process or alternative provider is available.

The deeper A.I.-specific risks belong under A.I. & Risk. Operational Resilience asks the more specific question: if we become dependent upon A.I., can the important service continue when that dependency fails?

WHAT HAPPENS WHEN PEOPLE LOSE THE ABILITY TO WORK WITHOUT A.I.?

A.I. introduces another form of resilience dependency that is not purely technological.

If people increasingly rely upon A.I. to analyse information, write documents, solve problems or make recommendations, they may become more productive. Over time, however, the organisation needs to consider whether critical human knowledge is being lost.

The A.I. service may itself be highly resilient, but resilience can still weaken if people can no longer perform the work when it is unavailable or when its output is unsuitable.

This expands the toolkit beyond technical recovery towards knowledge continuity, cross-training, documentation, manual fallback procedures and deliberate retention of critical human capability.

It also creates a direct bridge to Future of Work and The Future Human. Operational Resilience asks whether the organisation can continue. Future of Work asks what happens to roles and skills as technology changes the work. The Future Human takes the question further by examining what increasing technological dependence might mean for human capability itself.

The risks overlap, but each page asks a different question.

HOW DO CLIMATE AND GEOPOLITICAL RISKS AFFECT OPERATIONAL RESILIENCE?

Climate and geopolitical developments can become Operational Resilience issues when they disrupt the people, suppliers, infrastructure, transport, energy or technology required to deliver an important service.

A small business does not need to become a geopolitical analyst to recognise that a supplier may be affected by trade restrictions, conflict, transportation disruption or energy shortages.

Likewise, climate-related events can affect physical premises, transport, utilities, supply chains or employees even when the business itself is not located in the directly affected area.

The resilience question remains practical: which important activities depend upon something that could become unavailable because of an external event, and what alternatives exist?

These risks connect Operational Resilience with Signals → Geopolitical Risk and Climate Risk. Signals helps the reader notice what may be changing. Operational Resilience asks whether the business could continue if that change produces a disruption.

This is another example of how newer risks connect back to established Pillars.

HOW CAN SIGNALS AND HORIZON SCANNING IMPROVE OPERATIONAL RESILIENCE?

Operational Resilience traditionally focuses heavily on preparing for disruption. Signals and horizon scanning can add an earlier layer by helping businesses notice changes that may affect important dependencies before disruption occurs.

A business might notice increasing supply delays, persistent problems at a key technology provider, changes in regulation, extreme weather patterns or emerging geopolitical tension affecting a supplier location.

Not every development requires immediate action. Some may simply belong on a watch list. Others may justify scenario testing or the identification of an alternative supplier before the situation becomes urgent.

This creates a natural progression between Signals and Operational Resilience:

Notice what is changing → consider what it could disrupt → understand the dependency → test the consequences → strengthen alternatives where worthwhile.

The objective is not perfect prediction. It is to reduce the number of occasions on which the business discovers an important dependency only after it has already failed.

WHAT OPERATIONAL RESILIENCE QUESTIONS ARE PEOPLE ASKING NOW?

People rarely search for the phrase “Operational Resilience framework” when they first encounter a resilience problem. They ask the question created by the situation in front of them.

A small-business owner may ask, “Can my business operate if I am suddenly unavailable?” That may involve Operational Resilience, Business Continuity, key-person dependency and Enterprise Risk.

Someone may ask, “What happens if the cloud system my entire business uses goes down?” That links Operational Resilience with Technology Risk, Third-Party Risk and Business Continuity.

Another may ask, “What happens if my staff become too dependent on A.I.?” This begins as an Operational Resilience question about dependency and fallback capability, but it also leads into Future of Work and The Future Human.

A business using overseas suppliers may ask, “How do I prepare for geopolitical disruption when I cannot predict what will happen?” Signals can help identify changing conditions, while Operational Resilience helps the business examine alternative suppliers, inventories and dependencies.

Someone else might ask, “How can a small business prepare for another unexpected event like COVID?” The answer is not to attempt to predict the next pandemic exactly. It is to understand which activities matter most, what they depend upon, where the single points of failure are and how the business could adapt when normal assumptions suddenly stop being true.

These sound like different questions, but underneath them sits the same Operational Resilience concern:

If something important changes or fails, can I still deliver what matters?

WHICH NEWER RISKS MAKE THE EVOLUTION OF OPERATIONAL RESILIENCE USEFUL?

The evolution of Operational Resilience is particularly relevant to A.I. & Risk, Future of Work, The Future Human and Signals, while continuing to connect strongly with Technology Risk, Cyber Risk, Enterprise Risk, Business Continuity and Third-Party Risk.

A.I. creates new operational dependencies where intelligent systems become embedded in important activities. Future of Work introduces workforce resilience questions around skills, knowledge, remote work and changing employment structures. The Future Human becomes relevant when technological dependence begins affecting people's own ability to perform important work.

Signals adds the forward-looking layer. Geopolitical Risk, Climate Risk, Technology Trends and Emerging Risk can all identify developments that may eventually disrupt important services.

The established resilience tools therefore do not disappear. Business Impact Analysis, dependency mapping, recovery planning, scenario testing and continuity arrangements remain valuable. What changes is the range of dependencies and scenarios that those tools need to consider.

End-to-end service mapping, concentration analysis, more complex scenario testing, knowledge continuity and A.I. dependency assessment are extensions of the same objective: understand what the business needs to keep working and build enough adaptability that one unexpected disruption does not bring everything to a halt.

OPERATIONAL RESILIENCE IS ABOUT CONTINUING WHAT MATTERS

Operational Resilience is sometimes made to sound far more technical than the underlying idea really is. At its simplest, it begins by asking what a business must continue doing for its customers, employees or owners, what that activity depends upon and what could be done if one of those dependencies disappeared.

The answers have become more complex because businesses themselves have become more complex. A small company can now depend upon cloud services, payment platforms, remote employees, external specialists, overseas suppliers and A.I. systems despite having only a handful of employees.

This means the resilience toolkit has had to evolve. Traditional continuity plans and recovery arrangements increasingly sit alongside end-to-end dependency mapping, service mapping, concentration analysis, severe-but-plausible scenario testing, knowledge continuity and assessments of external technology and A.I. dependency.

Yet the underlying question remains familiar.

A shopkeeper keeping another way to accept payment, a small-business owner training a second employee to perform an important task and a modern company testing what happens when its cloud platform or A.I. provider becomes unavailable are separated by very different technologies and circumstances.

They are nevertheless responding to the same human concern:

Something I depend upon may fail. How do I make sure that what really matters can still continue?

That is why Operational Resilience remains one of the Pillars of Risk Management.