When one cloud region took a chunk of crypto offline

The 2025 AWS outage began on 20 October, when a DNS fault in Amazon’s DynamoDB service in one AWS region took down a large part of the internet for about three hours. Snapchat, Hulu, Venmo, and a long list of others.
It also took down a large part of crypto, which is supposed to be the industry that cannot be switched off.
What went down
Coinbase’s trading platform went offline. So did Base, the layer 2 that Coinbase itself operates. Robinhood had problems. Infura had problems. The outage ran roughly from 06:48 to 09:40 UTC, and Coinbase was down for over three hours.
Base did not simply stop. It degraded in the way distributed systems do: latency climbed, transaction submissions backed up, and the queue took most of the working day to clear even after AWS recovered.
None of the underlying blockchains had a problem. Ethereum was fine. Base’s own sequencer was reachable in principle. What broke was the infrastructure layer sitting between users and those chains, and that layer turned out to live in one place.
The number that explains it
Reporting after the incident put a figure on the concentration: one analysis estimated around 37% of Ethereum execution-layer nodes were hosted on AWS, and that the share went past 42% once layer 2s were included. Other analyses use different denominators, so treat those percentages as reported rather than audited. Within AWS, a large share of blockchain instances sat in US-East-1, the same region that failed. The stronger evidence is where the major providers say they run.
The picture is a decentralised ledger reached almost entirely through a centralised funnel.
Who stayed up, and why
This is the part that makes the incident useful rather than just depressing.
Kraken and OKX were among the services reported as staying online. Not because they were lucky, and not because they had predicted a DynamoDB DNS fault. The architectural explanation is that teams running across multiple regions and multiple clouds had somewhere else to send traffic when one path died.
That is the whole difference. Same day, same outage, same dependency on the same chains. One group had a second path and one group did not.
The uncomfortable read for the rest of us
Most teams are in the second group, and for understandable reasons. Multi-region is expensive. Multi-cloud is harder. A second RPC provider means a second bill and a second integration, and the first one has been fine for two years.
The counter-argument is simply arithmetic. Three hours of downtime on a payment product is not three hours of lost revenue. It is every transaction in that window failing, every customer who tried and could not seeing that you did not work, and a support queue that outlives the outage by days.
And the failure keeps arriving from a direction nobody watched. In 2020 it was a client version number. In 2025 it was a DNS record in a database service that most affected teams had never configured.
You cannot predict the next one. You can make sure it is not fatal.
What redundancy actually means here
It does not mean running your own nodes, unless you want to. For most teams it means something much smaller: have more than one way to reach each chain, and make the switch automatic.
Automatic is the word doing the work. A second provider you have to manually swap to during an incident is not redundancy, it is a to-do item you will be doing at 3am while your customers are already gone.
The next post is about what automatic failover actually involves, and why the obvious version of it makes outages worse rather than better.
Sources (checked 22 September 2026):
- AWS official update on the US-EAST-1 DNS incident
- CoinDesk analysis of the AWS outage and crypto infrastructure
- Metrika post-mortem on the outage effect on blockchain networks
Aurpay is a non-custodial crypto payment gateway. Our RPC routing layer is open source at the RPC Gateway repository.

