AFQY | A Few Quiet Yarns
← All newsSecurity

When one vendor sneezes: two years of big outages and the new resilience conversation

21 July 2026· AFQY News

When one vendor sneezes: two years of big outages and the new resilience conversation

Two years ago this week, a single faulty file did what no cyber attack ever had. On 19 July 2024, CrowdStrike distributed a bad configuration update for its Falcon security software, and Microsoft estimated 8.5 million Windows devices crashed. That was less than one percent of the world’s Windows machines, and it was still enough to be called the largest outage in the history of information technology.

New Zealand felt it immediately. More than half of Retail NZ’s members were affected, Woolworths closed around a dozen stores when checkouts refused to restart, and shoppers stood in aisles unable to pay with money they knew was in their accounts. Offshore, the bill was eye-watering: analysis from insurance specialist Parametrix put direct losses to Fortune 500 companies at US$5.4 billion, with only 10 to 20 percent of that expected to be covered by insurance.

If CrowdStrike felt like a one-off, 2025 corrected the record. In October, a DNS failure in AWS’s US-EAST-1 region knocked out its DynamoDB database service and cascaded across dozens of other services for around nine hours, taking household names like Snapchat and Ring with it. Weeks later, a database permissions change at Cloudflare doubled the size of a bot-management configuration file, crashing traffic-routing software across its global network in what the company called its worst outage since 2019. Neither was an attack. Small internal changes, systemic consequences.

The uncomfortable maths of concentration

The common thread is concentration. A handful of providers now carry an extraordinary share of the world’s digital load: CrowdStrike alone holds an estimated quarter of the endpoint detection and response market. And as one analyst observed after the AWS incident, many organisations were hit indirectly because their software supply chain relies on AWS even when they don’t realise it. You can be a casualty of concentration risk without ever signing the contract.

Regulators have noticed. The EU’s Digital Operational Resilience Act has applied since January 2025, and it goes further than box-ticking: critical ICT third-party providers to the financial sector now face direct oversight from European supervisory authorities. That is a formal acknowledgement that vendor risk has become systemic risk.

New Zealand has no DORA, but the direction of travel is similar. The Reserve Bank’s latest Financial Stability Report names cyber resilience and business continuity planning among its key supervisory priorities, and it detailed a first-quarter outage of ESAS, the high-value settlement system, that disrupted around $4.5 billion of transactions before services were restored within about three hours. Then there is geography. The government is working through ten initiatives to protect the undersea cables that connect us to the world, and ran its first simulated cable-break exercise in March. Losing one international cable is manageable because traffic reroutes across spare capacity. Losing more than one means overseas sites simply stop loading.

From uptime to resilience

Something has shifted in how boards talk about all this. The question is no longer “what was our uptime last quarter?” but “how fast do we recover, and what keeps trading while we do?” One analyst summed up the new mood well: cloud resilience is no longer a technical metric, it is an enterprise capability. Nor is the answer simply to buy more of everything. Gartner has cautioned that pursuing multicloud resilience can cost more than it saves, adding complexity without removing systemic risk. The real work is quieter: knowing your dependencies, deciding which failures you can genuinely tolerate, and rehearsing for the ones you can’t.

Worth remembering: on the morning CrowdStrike broke the world, cash and eftpos kept New Zealand trading. Resilience is rarely glamorous, and sometimes it is the oldest rail in the stack. The leaders who come through the next outage well won’t be the ones who never went down. They’ll be the ones who knew exactly what to do when they did.