• DevOps
  • AWS
  • Reliability

AWS US-EAST-1 Recovery: EC2 Launch Failures Resolved and Services Restored

What happened in the October 2025 AWS us-east-1 outage, why EC2 launches failed during recovery, and how to design workloads that ride out a regional event.

By ZeroTwo SolutionsUpdated 3 min read

The AWS disruption in the Northern Virginia (us-east-1) Region on October 19 and 20, 2025 was one of the most widely felt cloud incidents in years. Even after the initial DNS problem was fixed, teams could not launch new EC2 instances for hours, and AWS throttled launches until it could bring capacity back to pre-event levels. AWS has since published a detailed post-event summary. Here is what happened, why EC2 recovery took so long and what engineering teams should take from it.

What happened?

According to AWS, the event ran from 11:48 PM PDT on October 19 to 2:20 PM PDT on October 20, with three distinct periods of customer impact:

  1. DynamoDB API errors (11:48 PM to 2:40 AM). A latent race condition in DynamoDB’s automated DNS management system left the regional endpoint, dynamodb.us-east-1.amazonaws.com, with an empty DNS record. Anything that needed to resolve that endpoint, including other AWS services, started failing.
  2. EC2 instance launch failures (2:25 AM to 10:36 AM). New instance launches failed, and some instances launched after recovery began had network connectivity problems until 1:50 PM.
  3. Network Load Balancer errors (5:30 AM to 2:09 PM). Some load balancers saw increased connection errors caused by health check failures in the NLB fleet.

Why did EC2 launches keep failing after DNS was fixed?

The DNS record was restored early in the morning, but EC2 launches did not recover with it. AWS explained the cascade:

  • A hidden dependency. The EC2 subsystem that manages the physical servers instances run on keeps “leases” on those servers and stores its state in DynamoDB. While DynamoDB was unreachable, those leases expired, so the servers could not be used for new launches.
  • Congestive collapse during recovery. When DynamoDB came back, the subsystem tried to re-establish leases across the entire fleet at once. The backlog was so large that work timed out and was retried faster than it could complete. Engineers had to throttle incoming work and restart hosts to drain it.
  • A second backlog in networking. Once launches succeeded again, the system that propagates network configuration to new instances was behind, so some new instances started without working network connectivity.
  • Load balancer side effects. NLB health checks were evaluating new instances whose network state had not yet propagated. They failed those checks, and capacity was removed from service.

To stabilize the region, AWS temporarily throttled operations such as EC2 instance launches and Lambda’s polling of SQS queues, then gradually restored those throttles to pre-event levels as the backlogs cleared.

Who was impacted?

Any workload in us-east-1 that needed new capacity during the event was affected: Auto Scaling groups trying to scale out, Spot and On-Demand launches, container platforms adding nodes, and batch or data pipelines that start instances on demand. Services that depend on DynamoDB or other affected AWS APIs in the region saw errors even if their own instances stayed healthy. Already-running instances generally kept running, which is why teams with enough pre-provisioned capacity rode it out far better than teams relying on just-in-time scaling.

What AWS is changing

In its summary, AWS described several fixes, including disabling the affected DNS automation worldwide until the race condition is corrected, adding safeguards to how Network Load Balancer removes capacity during health check failures, adding recovery testing for the EC2 lease subsystem, and improving throttling so backlogs drain without collapsing.

What teams should do now

  • Do not depend on launching capacity during an incident. Keep enough warm capacity, or use reserved capacity, so a launch outage degrades you gracefully instead of taking you down.
  • Design for multi-AZ, and decide deliberately about multi-region. Multi-AZ protects against a single data center problem. This event was regional, so critical services need a tested plan for running from another region, even if only in a reduced mode.
  • Map your hidden dependencies. Know which managed services, regional endpoints and third-party APIs your application needs at startup, during scaling and during failover.
  • Make retries polite. Use exponential backoff with jitter and circuit breakers, so your own clients do not amplify a provider’s recovery problem.
  • Alarm on the symptoms you care about. Set CloudWatch alarms on failed launches, Auto Scaling activity errors and unhealthy target counts, and subscribe to the AWS Health Dashboard for your accounts.
  • Assess and rehearse. Tools like AWS Resilience Hub can assess your architecture, but nothing replaces a practiced incident response playbook and regular failover drills.

Key takeaway

No cloud provider is immune to outages, and recovery can take longer than the initial fault. The teams that recover fastest are those with spare capacity, multi-region plans for critical paths, sensible retry behavior and well-rehearsed incident response. Use this event as a prompt to review your own resilience strategy.

More articles

All articles
Start a project

Have an AI or FinTech product to build?

Tell us what you want to ship. You will get a technical roadmap and a fixed-scope proposal within 48 hours.

  • Reply within 1 business day
  • NDA on request
  • You own 100% of the IP