Services

Open source

EN

KUBICAST #189 - DevOps and the AWS outage

Has a dead zone turned your life into a mess too?

mansplainer

João Brito

In this episode we break down the recent major AWS outage from a DevOps/SRE perspective: from the initial hypothesis (DNS/DynamoDB and cross-regional dependencies) to the practical implications for resilient architecture. We talked with Lucas Azevedo about how to diagnose incidents that seem like "just another hitch" and turn into a global outage. We cover versioning wars in DNS automations, control planes that still depend on us-east-1, and why "it's always DNS... even when it isn't."

We also explore multi-region resilience strategies and the real trade-offs between cost, complexity, and RTO/RPO. We discuss how to map blast radius, prioritize actionable runbooks, and create resilience exercises that go beyond chaos monkey. We bring cases of cascading failures, impacts on managed services (KMS, EKS, IAM, Support), and live observability practices to shorten MTTR when the provider is down.



Finally, we wrap up with key takeaways for product and platform teams: feature flags for integration fallback, alternative routes for control planes, circuit breakers in clients, and playbooks for communication with stakeholders. Two topics that deserve special attention in this talk: multi-region resilience in practice and how to prepare your organization for "almost impossible" incidents.

Important Links:

  • Lucas Azevedo - https://www.linkedin.com/in/lazevedo-devops/

  • DevOps Community on Discord - https://discord.com/invite/k6wPagw4tV

  • João Brito - https://www.linkedin.com/in/juniorjbn/



    🎧 Also listen to Kubicast on Spotify, and share it with the whole team working on the new DR plan!

Newsletter Getup.

Atualizações sobre Kubernetes e Software Supply Chain Security todos os meses.

Operating Kubernetes in production for more than 13 years. With Quor, this experience extends to software supply chain security as well.