<img height="1" width="1" style="display:none;" alt="" src="https://px.ads.linkedin.com/collect/?pid=10643465&amp;fmt=gif">

How Snowflake's poor release governance caused a global outage

Snowflake spent months building regional redundancy, then undermined it with a single faulty update. Here's why vendor dependency failure is inevitable, and what it really takes to be resilient.
Ben Espach
Last updated:

Most companies invest heavily in resilience: multi-cloud architecture, geographic redundancy, reputable vendors with strong SLAs. All of it is designed to survive the same class of failure: dependency risk.

On December 16, 2025, Snowflake's customers had all of that in place, and none of it helped. You see, the failure came from inside Snowflake's own release process, a failure class the standard resilience playbook isn't built to catch.

TL;DR

Snowflake's schema update took 10 regions offline for 13 hours

Snowflake's cloud data platform powers just over 20% of the global data warehouse market, making it the single most-used data warehouse technology worldwide. Its regions share a central rulebook that governs how every region reads and processes data. An update rewrote that rulebook in a way older regions couldn't understand, bringing 10 out of 23 to a stop at once.

Four service categories went down at once: query execution, Snowpipe file ingestion (Snowflake's automated data loading service), Snowpipe Streaming (its real-time data pipeline), and data clustering. The only workaround was a manual failover, but only customers who had already configured their failover before the outage could do so.

Snowflake serves more than 13 300 customers. Those in the ten affected regions were locked out of their data for 13 hours.

Just seven weeks earlier, Snowflake had publicly praised "Snowgrid", a cross-cloud replication and failover system designed to keep customer workloads running even if an entire cloud provider went down. When AWS experienced a regional outage in October 2025, Snowgrid worked exactly as promised — over 300 critical customer workloads failed over automatically without interruption. Snowflake called it a non-event for prepared customers. It was a genuine engineering achievement, built on the right principles.

It wasn't enough. When their internal update shipped in December, every Snowflake region shared the same broken rulebook. There was no healthy region to route to. Snowgrid is built to survive external infrastructure failures. This failure came from inside Snowflake's code — a completely different failure class, and one the tool was never designed to catch.

For most customers, losing a full working day wasn't even the full cost. Operators of streaming event pipelines were told their data from the outage window was permanently unrecoverable. Snowplow confirmed: "We do not have the ability to retroactively replay good events that were missed during the incident window."

The outage resolved. But customer data was still gone for good — and no SLA credit can change that.

» Uncover how much risk your systems carry with this quick assessment.

This is what downstream vendor dependency risk looks like in practice

In 2025, OWASP elevated Software Supply Chain Failures to the A03 position, the third most critical risk category in web application security. The definition covers breakdowns in the process of building, distributing, or updating software, and specifically flags failure to test updated library compatibility as a key vulnerability. The Snowflake incident fits the definition like a glove. If your business depends on vendors to operate, you fall inside this risk class.

It's not just Snowflake dropping the ball either. The CrowdStrike event of July 2024 shared the exact structural fingerprint: a trusted vendor's own update process caused a major outage. Validation was bypassed. Geographic redundancy was irrelevant. That single faulty update produced $5.4 billion in direct losses across Fortune 500 companies, an average of $43.6 million per affected company. Healthcare absorbed $1.9 billion; banking lost $1.4 billion.

» Take a look at the damage that CrowdStrike event caused for businesses around the world.

CrowdStrike is a cybersecurity company. Snowflake is ranked among the most future-ready businesses on the planet by Fortune's Future 50 growth index. Neither status protected their customers from updates gone wrong — meaning if you're basing your business continuity on the idea that giants like these have their act together, you might need to reevaluate your resilience plan.

And it's not just high-profile names — the pattern shows up at the infrastructure level too. Faulty software updates produced 515 million lost user hours across EU telecom infrastructure in 2024; more than cable cuts, hardware failures, and power events combined.

Every vendor makes this kind of error eventually. So, focusing on selecting reliable providers to run your business on isn't enough; you need to ensure your architecture is designed to absorb the moments when even reliable vendors fail. That moment, for Snowflake's customers, was a normal December morning.

» Find out how you can ensure your software uptime.

How to protect your software against downstream vendor dependency risk

You can't vet a vendor thoroughly enough to guarantee their next release won't cause an outage. What you can control is what happens on your side when it does. Eight in ten operators believe better processes could have prevented their most recent major outage — your exposure isn't invisible; you can see it coming. But knowing the risk exists and having the infrastructure to survive it are two different things. Here is what that infrastructure actually looks like:

  • Failure: When your cloud vendor goes down, you need your complete stack available somewhere independent: source code, databases, configurations, credentials. That's exactly the issue the outage identified — customers who hadn't set up independent failovers in advance had nowhere to turn. Our SaaS Escrow service maintains a continuously updated copy of your provider's entire cloud environment, stored outside their infrastructure, so recovery doesn't depend on their availability.

  • Attacks: If your assets live exclusively inside a vendor's environment, a breach of that environment is a breach of your assets. In a separate Snowflake security incident, credential-based attackers accessed data belonging to approximately 165 Snowflake customers, including AT&T and Ticketmaster. With Software Backup, your critical assets are stored in an immutable vault, encrypted, tamper-proof, and isolated from events inside the vendor's platform.

  • Non-compliance: Regulators don't distinguish between your failure and your vendor's. For financial entities under the Digital Operational Resilience Act (DORA), an outage affecting critical data operations is reportable regardless of where it originated — and that obligation requires tested continuity provisions in place before an incident. Our Verification (Certified) service runs a full build and deployment test of your deposited code and configurations, and issues a Software Resilience Certificate that satisfies DORA, NIS2, and ISO 27001 auditors.

  • Broken: When a vendor's update breaks critical operations, automated systems keep firing requests at a platform that can't respond — compounding the problem on your side while you wait. If that vendor is unable or unwilling to fix what broke, you need access to the underlying code to resolve it yourself. Our Software Escrow service guarantees that access, ensuring the source code, documentation, and build instructions are available to you if your vendor's release process ever leaves you stranded.

If this happened to your business

It's 03:00 UTC on a Tuesday. Your automated pipelines start failing silently. No alerts point to your own infrastructure. By 04:30, your on-call engineer confirms it: Snowflake is down across your region. The incident started 90 minutes ago. You have no failover configured. You are waiting.

By 06:00, your data team is fielding questions from every direction. Customer-facing analytics are dark. Snowflake revises its ETA before midday, then again. Each revision adds an hour. Your European teams arrive at 08:00 to the same broken environment.

By early afternoon, the outage resolves. But thirteen hours of event data is gone — your pipeline type wasn't eligible for re-ingestion. Your compliance team is now determining whether this triggers a DORA reporting obligation. In the span of a day, you would've lost revenue, lost data, and faced contract breaches.

Your vendor's release cycle is now part of your risk surface

Snowflake ranked #1 on Fortune's Future 50 in 2025, for the second time in three years, and reported $1.16 billion in quarterly product revenue. Those metrics measure capability and market confidence. They also measure the scale of how many organizations carry Snowflake's release cycle as part of their own risk surface. That day, the surface cracked for all of them at once.

Even if you choose a reputable vendor, invest in resilience, and follow the standard resilience strategy, it just doesn't cover modern dependency risk. The only way to distance your business from your provider's instability is guaranteed access to their source code through Software Escrow — so you can run their services independently if you ever need to.

Explore how Codekeeper's software escrow and verification solutions ensure independent recovery regardless of what your vendor's next release does. 

» Book a call with our experts today to find out more

Share this article
Share on facebook Share on linkedin Share on twitter Share on email
blog_book_a_demo_cta_3x
Have questions about protecting your software?
Our escrow experts are standing by to help.
Book a free demo