Leave your feedback Share Copy URL https://ar-c.org/video/4AXBunjIFFv.html Email Facebook Twitter LinkedIn Pinterest Tumblr Share on Facebook Share on Twitter How To Prepare For The Next AWS Outage [jMYmUINRwxf] Health Updated on August 06, 2026 EDT — Published on August 06, 2026 EDT On October 20 at 12:20am, AWS US-East-1 began a 16-hour outage. But Temporal’s customers stayed online. In this episode of 1 IDEA, Suresh Mathew sits down with Preeti Somal (SVP of Engineering @ Temporal) to unpack how her team prepared for a failure they knew was coming and how other engineering leaders can do the same. We cover: - How Temporal detects outages before customers do - What multi-region failover actually looks like in practice - Why your failover plan might fail when you need it most - How to run game days that actually prepare your team CHAPTERS 00:00 Introduction 00:42 The October AWS outage: first moments 02:01 How Temporal kept customers running 03:37 The dependency that caused a 35-minute manual scramble 07:03 How multi-region replication works 13:11 Game days: simulating AZ failures proactively 15:51 How multi-region adoption doubled post-outage 21:05 The "what happens if" reliability framework 24:16 Building a reliability culture as a leader 25:09 Where to start: alignment and execution TNPo9m3zE3i oaH7Z35xZ3D oYxuVydbCJF a8wdCKpwE5X wppd0aZA2dw de8HNDbiPNs mrV0YHHFdMY
On October 20 at 12:20am, AWS US-East-1 began a 16-hour outage. But Temporal’s customers stayed online. In this episode of 1 IDEA, Suresh Mathew sits down with Preeti Somal (SVP of Engineering @ Temporal) to unpack how her team prepared for a failure they knew was coming and how other engineering leaders can do the same. We cover: - How Temporal detects outages before customers do - What multi-region failover actually looks like in practice - Why your failover plan might fail when you need it most - How to run game days that actually prepare your team CHAPTERS 00:00 Introduction 00:42 The October AWS outage: first moments 02:01 How Temporal kept customers running 03:37 The dependency that caused a 35-minute manual scramble 07:03 How multi-region replication works 13:11 Game days: simulating AZ failures proactively 15:51 How multi-region adoption doubled post-outage 21:05 The "what happens if" reliability framework 24:16 Building a reliability culture as a leader 25:09 Where to start: alignment and execution TNPo9m3zE3i oaH7Z35xZ3D oYxuVydbCJF a8wdCKpwE5X wppd0aZA2dw de8HNDbiPNs mrV0YHHFdMY