Additional 9,600 Flights Cancelled Since July 20 After 46,000 Grounded on July 19 Because of CrowdStrike Update
Lead StoryBusiness
The bizarre story of CrowdStrike’s software update continues to leave computers, businesses, and consumers stuck, more than three days after the disastrous mistake propagated through the internet.
Data gathered by Microsoft Corporation and CrowdStrike estimate that originally over 8.5 million computers were crippled by bugs in a single software update rolled out the end of the last week by CrowdStrike.
CrowdStrike, based in Austin, Texas, is a cybersecurity company which provides a protective buffer layer which isolates software on individual computers and servers from multiple categories of computer threats. It has advanced technologies which can detect hacking attempts as they start and then block them and provides similar capabilities for the many companies which use cloud-based services such as Microsoft’s Azure, Google Cloud, and Amazon Web Services make available to enterprises around the globe.
CrowdStrike drew in $2.24 billion in revenues in 2023 because of the wide popularity of its products. Those revenues were also up by 54.4% over the previous year’s results. Up until now, that sharp upward growth curve looked like it might be maintained. After the debacle over its latest software update, which forced Windows computers and servers to boot into uninterruptable and continuous reboot cycles without ever working, and which CrowdStrike says was its fault, CrowdStrike might not survive the onslaught of lawsuits which are coming their way. It could be forced into Chapter 11 bankruptcy. Even worse, it could forever lose the trust of its customer base, forcing it to give itself up for purchase on a fire sale basis to someone else.
The software glitch, part of the company’s latest update to its Falcon sensor cybersecurity software, affected almost every kind of industry imaginable. From stockbrokers to banks, from pharmacies to hospitals, from order management to shipping fulfillment services, streaming entertainment services, television broadcasting companies, and even social media services, computers went down “en masse” within hours of each other beginning on Friday, July 19. And so did the businesses those computers supported.
As an example of how diverse the problems were, in the United Kingdom the National Health Service suffered outages which made setting appointments for doctors, arranging specific medical treatments, and the distribution of prescription medicine close to impossible for a while.
Elsewhere in Europe, German hospitals reported cancelling elective surgeries throughout their networks, in order to allow for slower manual scheduling to be prioritized for emergency care.
In the U.S. similar problems occurred. In Boston, Brigham and Women's Hospital was forced to cancel all non-emergency surgeries. Seattle Children’s Hospital shuttered its outpatient services area starting as soon as the computers went down too. So also did New York City’s Memorial Sloan Kettering Cancer Center, which put on hold any medical care requiring anesthesia for some time.
In other regions, entire city services were put on hold. In Portland, Oregon, the situation was so critical that Mayor Ted Wheeler was forced to issue an Emergency Declaration, to put the region on notice that many services might not be available for an indefinite period. There were even widespread reports of 9-1-1 emergency calling services outages, as well as the automatic scheduling means of responses to support the emergency calls which did get through.
In the shipping services industries, FedEx said it “activated contingency plans to mitigate impacts” from the CrowdStrike updates but warned customers to expect some shipping delays for packages scheduled for delivery on July 19.
Starbucks was also affected by the outage. In some regions it suspended pre-ordering services using its mobile app as a result.
For the public, perhaps the most visible sign of the impacts of CrowdStrike’s failed update was for those awaiting takeoff of flights across the globe. On July 19 alone, the air travel tracking application FlightAware reported over 46,000 flights were canceled. The reasons why varied, some related to the process of issuing tickets, creating boarding passes, and managing luggage delivery. In other cases, such as Delta Air Lines, the glitch disabled logistics resource management tools which allocate flight crew assignments throughout its international network. That is something which is sufficiently automated over such a wide region that it is impossible to manage workaround by hand, which is something small regional airlines were able to do more easily.
On that first day of the computer failures, Delta was the hardest hit of all U.S. flight carriers, with 1,200 of its regularly scheduled flights forced to remain grounded. 649 of United Airlines flights were also canceled, as were 408 of American Airlines.
Since the problem was only discovered on Friday and most computers over the weekend could be fixed only on a one-by-one manual intervention basis, most air carriers continued to cancel additional – though not all – flights. On July 21, for example, Delta was forced to delay an additional 1,600 flights and cancel another 1,300.
Most airlines have attempted to pacify the sizeable angry group of passengers who were stuck because of the software mess with hotel accommodations and meal vouchers in the short term, plus offers to rebook as soon as possible.
As Secretary of Transportation Pete Buttigieg said in a post on social media this Sunday, those airlines also have a mandatory responsibility to provide immediate cash payouts to those affected by the crisis.
“Delta must provide prompt refunds to consumers who choose not to take rebooking, free rebooking for those who do, and timely reimbursements for food and hotel stays to consumers affected by these delays and cancellations, as well as adequate customer service assistance,” he wrote.
His comments came after it appeared Delta’s initial automated responses to customer requests for help made no mention of their obligation to provide full refunds nearly immediately.
The situation does continue to heal somewhat, especially as individual companies – aided by Microsoft and CrowdStrike – deploy semi-automated tools to remove the botched code and path it with a working fix. On July 22, the total number of flights cancelled as a direct result of this glitch fell to just 1,500.
An analysis of a possible root cause for what seems an unimaginable breach of engineering protocols over such a critical software update, it now appears the latest software patches the company rolled out somehow skipped a verification process known as “sandboxing” as one part of why so much went wrong so fast. That refers to the concept of setting up a fully operational but isolated test system to verify code functionality and flaws in a “virtual sandbox”. As to why that might have happened, experts with more direct knowledge of CrowdStrike’s business processes say the company had a history of uploading an increasing number of updates all the time. With cyberattacks themselves growing more serious and frequent in recent years, not only has that made products like CrowdStrike’s a critical part of the software portfolio companies subscribe to regularly, but it may also have pushed CrowdStrike perhaps to an internal breaking point in how it manages so many updates so quickly.
CrowdStrike CEO George Kurtz quickly took responsibility for the software rollout failures, and emphasized that despite the appearance, what brought done a large fraction of the world’s computer networks on July 19 was not from outside hostile actors,
“This is not a security incident or cyberattack. The issue has been identified, isolated and a fix has been deployed,” Kurtz wrote on July 19 in a post on the social media platform X.
While CrowdStrike has owned up to the failure, downloading and using new semi-automated tools the company is supplying in partnership with Microsoft to remove and replace the bad software looks like it could take weeks to complete. Further flight delays and disruptions of a wide range of services will probably continue throughout that time, but on a far smaller scale than from July 19 to 22.
Coming next will be the lawsuits over what happened. But for most companies the more important long-term issue is how to avoid this sort of technological insanity ever happening again.
The House of Representatives has already summoned CEO George Kurtz to a special hearing to explain precisely how this mistake happened. The hearing has not yet been scheduled.