Cloud Infrastructure

The us-east-1 Outage Was a Dependency Lesson, Not a DNS Lesson

Late on 19 October a latent race condition in DynamoDB’s automated DNS management emptied the records for the service endpoint in us-east-1. Applications could no longer resolve the database. DynamoDB itself was largely recovered within a few hours; the cascade took roughly fifteen.

Why it spread so far

The interesting part is not the DNS bug. It is what depended on it. EC2’s droplet workflow manager stores state in DynamoDB, so its queries timed out, leases expired, and instances were marked unavailable. Network state propagation added hours to recovery. Internal AWS tooling was affected too, which slowed the response.

Snapchat, Fortnite, Roblox, Ring, banking apps, and a government tax site were all visibly down. Downdetector logged reports across more than a thousand services.

The uncomfortable question for multi-region architectures

Plenty of affected companies had deployed across regions and still failed. The reason is control plane placement: global functions like IAM updates and certain cross-region features have critical infrastructure in us-east-1. A healthy region could serve traffic but could not always perform control operations.

  • Know which of your dependencies are regional in name only. Global tables, identity, and certificate operations often route somewhere specific.
  • Test failover without the control plane. If your runbook requires an API call that is also down, it is not a runbook.
  • Design for degraded, not just down. Read-only mode beats an error page and is usually achievable.
  • Retry storms amplify. Exponential backoff with jitter is unglamorous and it is what keeps recovery from stalling.
Resilience is not measured by how many regions you deploy to. It is measured by what still works when the one you did not think about goes away.

What we changed

On client platforms this pushed two things up the priority list: an explicit inventory of single-region dependencies including the ones inherited from managed services, and a degraded-mode path for the handful of user journeys that actually matter. Neither is expensive. Both are much harder to add during an incident.

Sources & further reading

← All articles

Ready to Take Your Business to the Next Level?

Partner with us to create impactful digital solutions that fuel your business growth and deliver tangible results.

Start Your Project With Us