Operational Resilience
Operational Resilience is the ability of an organisation to prevent, adapt to and recover from disruptions.

Operational resilience is about keeping key services running when things go wrong. Not avoiding every fault, which no one can promise, but taking the hit and carrying on. In payments that means shoppers can still pay, payouts still leave, and the books still balance while a part of the stack is unwell. It covers tech, people, suppliers and process together. It treats disruption as a normal event to be planned for, rather than an accident to be explained afterwards.
Rule makers pushed the term into wide use. The Basel Committee's Principles for operational resilience, published in March 2021, frame it as the power to withstand events that could cause serious faults or wide disruption. In the UK, the PRA's supervisory statement SS1/21 asks firms to name their important business services and set impact tolerances for them. The EU has its own regime for digital operational resilience, and other markets have their own again. Which rules bind a given firm depends on its licences and where it trades, so this is a place to take local advice.
Which Services Matter Most
The first move is to name what matters. Not systems, but services a customer would notice losing. Taking a card payment. Paying out to a seller. Answering a dispute. Each one crosses several systems and several suppliers, which is the point of framing it that way. A firm that maps only its systems tends to guard the wrong things. One that maps services can see that a single quiet supplier sits under three of its most important flows.
Setting A Limit On Harm
An impact tolerance is a limit on harm. It answers a blunt question: how long can this service be down, or degraded, before the damage becomes too much? Setting it in plain terms, such as hours of outage or number of failed payments, makes the number testable. It also forces a choice rather than a wish. Once a tolerance exists, the work is clear. Build, test and rehearse until the service can stay inside it. Or accept the gap and say so.
Mapping And Third Parties
Payment stacks are chains of other people's services. A gateway, an acquirer, a fraud tool, a card scheme, a cloud region, a message queue. Resilience work means drawing that chain out fully, including the parts nobody owns in house. Overlap is the risk to look for. Two suppliers that appear separate may sit in the same data centre, or lean on the same upstream service. Contracts matter here as much as design. A supplier with no duty to report an incident is a blind spot.
Spare Routes In The Payment Path
The practical answer is more than one route. More than one acquirer per market, more than one region for the gateway, and a way to shift traffic without a code release. Dynamic routing is what turns a spare route into a working one. A second route that needs a manual switch is often too slow to help. Watching uptime figures per route, rather than one blended number, is what makes a problem visible early. This piece on how orchestration supports uptime and redundancy covers the pattern.
Failure Modes Worth Rehearsing
Some faults are loud. A provider goes dark, and traffic stops. Those are easier to handle than the quiet ones. A gateway timeout that leaves payments in an unknown state. A slow issuer that turns a healthy checkout into a queue. A settlement delay that leaves finance unable to close the day. A data breach at a supplier, where the response is legal as much as technical. Rehearsing the quiet ones is where most of the learning sits.
Test It, Do Not Just Write It
A plan that has not been run is a hope. Severe but plausible scenario testing is the part regulators ask about, and it is also the part that finds the real gaps. Cut a region and watch what happens. Fail an acquirer during a busy hour in a test setup. Try the manual process the runbook describes and time it honestly. Firms that do this tend to find that the technical failover works and the human steps around it do not, which is useful to know before a real event.
People And Process
Resilience is not just a tech matter. Someone has to notice, decide and tell people. On-call rotas, clear ownership, and a decision right to switch routes without a committee all shape how long an incident lasts. Customer messaging matters too: a shopper who sees a clear message tends to retry, while one who sees nothing tends to leave. Round the clock cover is part of that picture, which is why 24/7 support is treated as a resilience control rather than a comfort.
Where To Start
Name three key services and map them end to end, suppliers included. Set one impact tolerance per service and write it in hours and volumes. Test the worst plausible case, not the convenient one. Keep a second route live rather than dormant, and route a small share of traffic through it so you know it works. Review compliance duties per market, since the rules differ. And treat every incident as a source of evidence, not blame. This piece on why redundancy in payment gateways matters is a good next read.
Frequently Asked Questions
It depends on where a business operates and what licences it holds. The Basel Committee published principles in 2021, the UK regulators set expectations around important business services and impact tolerances, and the EU has its own regime for digital operational resilience. Requirements differ by jurisdiction, so local advice is usually needed.
A stated limit on harm: how long a service can be unavailable or degraded before the damage becomes unacceptable. Expressing it in plain terms, such as hours of outage or numbers of failed payments, makes it testable. It also turns a general aspiration into a target that can be built and rehearsed against.
Disaster recovery focuses on restoring systems after a failure. Operational resilience starts from the service a customer relies on and asks whether it keeps running at all, including through supplier problems, process gaps and staffing issues. Technical failover is part of it, but so are contracts, communication and decision rights.
Severe but plausible scenarios rather than convenient ones. Losing a region, losing an acquirer during a busy hour, a gateway timeout that leaves payments in an unknown state, or a settlement delay that stops the day closing. Teams often find the technical failover works while the human steps around it take far longer than expected.
It generally does, provided the second route is live rather than dormant and traffic can be shifted without a code release. Routing a small share of volume through it keeps it exercised. It is also worth checking for hidden overlap, since two providers can share a data centre or an upstream dependency.

Still Have Questions?
Let’s Find the Right Solution for You
Stay Connected with Us!
Follow us on social media to stay up to date with the latest news, updates, and exclusive insights!


