Leave your feedback Share Copy URL https://ar-c.org/video/IctyzX4cpp5.html Email Facebook Twitter LinkedIn Pinterest Tumblr Share on Facebook Share on Twitter How To Prepare For The Next AWS Outage [JDZxRPSiBCZ] Health Updated on August 07, 2026 EDT — Published on August 07, 2026 EDT On October 20 at 12:20am, AWS US-East-1 began a 16-hour outage. But Temporal’s customers stayed online. In this episode of 1 IDEA, Suresh Mathew sits down with Preeti Somal (SVP of Engineering @ Temporal) to unpack how her team prepared for a failure they knew was coming and how other engineering leaders can do the same. We cover: - How Temporal detects outages before customers do - What multi-region failover actually looks like in practice - Why your failover plan might fail when you need it most - How to run game days that actually prepare your team CHAPTERS 00:00 Introduction 00:42 The October AWS outage: first moments 02:01 How Temporal kept customers running 03:37 The dependency that caused a 35-minute manual scramble 07:03 How multi-region replication works 13:11 Game days: simulating AZ failures proactively 15:51 How multi-region adoption doubled post-outage 21:05 The "what happens if" reliability framework 24:16 Building a reliability culture as a leader 25:09 Where to start: alignment and execution 2wFaI70tf0G LFv4Qn8fzQN 5CQ6bKqHyu3 68O7V7a14nu GI7yUQkFMYc zzk5fTdBPIw JqHKfpxORRz nrJKU35oWFg
On October 20 at 12:20am, AWS US-East-1 began a 16-hour outage. But Temporal’s customers stayed online. In this episode of 1 IDEA, Suresh Mathew sits down with Preeti Somal (SVP of Engineering @ Temporal) to unpack how her team prepared for a failure they knew was coming and how other engineering leaders can do the same. We cover: - How Temporal detects outages before customers do - What multi-region failover actually looks like in practice - Why your failover plan might fail when you need it most - How to run game days that actually prepare your team CHAPTERS 00:00 Introduction 00:42 The October AWS outage: first moments 02:01 How Temporal kept customers running 03:37 The dependency that caused a 35-minute manual scramble 07:03 How multi-region replication works 13:11 Game days: simulating AZ failures proactively 15:51 How multi-region adoption doubled post-outage 21:05 The "what happens if" reliability framework 24:16 Building a reliability culture as a leader 25:09 Where to start: alignment and execution 2wFaI70tf0G LFv4Qn8fzQN 5CQ6bKqHyu3 68O7V7a14nu GI7yUQkFMYc zzk5fTdBPIw JqHKfpxORRz nrJKU35oWFg