Hero guide: Reliability and resilience
This guide highlights sessions and learning opportunities at re:Invent 2024 for those who are on a mission to find out how to improve the resilience, reliability, and availability of their applications using AWS offerings.

Sathyajith Bhat
AWS Container Hero
Reliability is a strong focus for AWS, as it is considered one of the six pillars of the AWS Well-Architected Framework. As a service owner and maintainer, you already know that ensuring the reliability of your service can be a challenging endeavor.
This guide focuses on talks, workshops, and labs that help attendees build, operate, and maintain reliable systems. Whether you’re an experienced reliability engineer looking to enhance your knowledge or a keen engineer looking to embark on your reliability engineering journey, these guides should set you on the right path.
Reliability engineering sessions often make use of advanced features of many AWS offerings and are therefore categorized at the 300 and 400 levels. You will find that many of my session suggestions align with these levels. It might be advantageous to explore the 200-level sessions on those services or do some light reading before attending the sessions.
Breakout session
KUB301 | How to build scalable platforms with Amazon EKS
A wide range of companies, from the most innovative startups to the world’s leading enterprises, are running their internal platforms on Amazon EKS, helping them to accelerate developer velocity and increase the pace of innovation. In this session, learn about best practices that AWS has developed over years of helping thousands of customers build and scale their internal platforms on Amazon EKS.
STG301 | Data protection and resilience with AWS storage
Data continues to be one of the world’s most valuable resources. Data protection and resilience are critical to business success, helping to keep workflows running and allow business recovery from data loss events such as unplanned outages, natural disasters, and ransomware. Join this session to dive deep on how AWS storage offers organizations defense-in-depth data protection and resilience for application data across recovery point and time objectives, helping to mitigate risks with immutable solutions, restore testing, policy-based access controls, encryption, and auditing and reporting.
Builders' session
COP401-R | Connecting the dots: Get hands-on experience with distributed tracing [REPEAT]
Distributed tracing is a crucial observability technique for understanding complex, microservices-based applications on AWS. In this builders’ session, learn how to use Amazon CloudWatch to instrument your applications and collect tracing data, allowing you to identify performance bottlenecks, debug issues, and optimize your systems. Leave this session with the knowledge and skills to unlock the full potential of distributed tracing and enhance the observability of your AWS-based applications, driving better decision-making for your business. You must bring your laptop to participate.
SVS305-R | Handle scaling with AWS Lambda [REPEAT]
In this interactive builders’ session on AWS Lambda scaling, learn how serverless workloads can handle massive spikes. Explore best practices to generate high throughput using load testing tools, and analyze scaling behavior with Amazon CloudWatch metrics. Discover real-world use cases, covering potential pitfalls like cold starts, concurrency limits, and upstream/downstream bottlenecks. Leave with a tested approach to verify application scaling, even at high throughput levels. Gain confidence that your applications can truly scale on demand by pushing serverless to its limits. You must bring your laptop to participate.
Chalk talk
DAT311-R | Design secure and resilient relational database architectures on AWS [REPEAT]
Overwhelmed with all the security best practices and trying to understand what makes sense for you to implement to protect your data? Join this chalk talk to learn how to deploy a secure and resilient architecture using all the security features that effectively integrate with your databases running on Amazon Aurora and Amazon RDS. Strengthen your enterprise’s security posture by discovering best practices and avoiding common pitfalls. Leave this talk equipped with all the available security features, knowing how to set them up and configure meaningful alerts to safeguard your data.
DAT303-R | Making your Amazon Aurora cluster more resilient [REPEAT]
Amazon Aurora is a cloud-native relational database with unparalleled high performance and availability at global scale. Hundreds of thousands of organizations trust Aurora for their workloads today. In this chalk talk, learn about features such as Aurora Global Database and Amazon RDS Proxy, and discover best practices to achieve high availability for your Aurora clusters. The talk includes a whiteboarding session on high availability and disaster recovery solutions to help you maximize your application’s resilience.
SVS317 | Understanding serverless resilience: Patterns and best practices
In this chalk talk, learn considerations for building resilient serverless applications by understanding business requirements. Starting with a simple serverless architecture, explore how nonfunctional requirements drive continuous enhancement and design of architectural patterns and improvements. Gain insights into operational excellence, resilience patterns, common business requirements, and best practices. Leave with practical knowledge for translating business needs into robust, resilient serverless architectures that separate business logic from infrastructure concerns.
NET302-R | Ask me anything about networking [REPEAT]
This chalk talk covers the breadth of AWS networking services. Are you facing challenges figuring out how you can optimize Elastic Load Balancing (ELB), where Amazon VPC Lattice fits in to your architecture, how you can migrate to AWS PrivateLink, how you should implement AWS Verified Access, whether you should migrate from AWS Transit Gateway to AWS Cloud WAN, or any other networking topic? Bring your questions to explore in detail through interactive conversations and whiteboarding.
HYB301-R | Building highly available and fault-tolerant edge applications [REPEAT]
Organizations need to be able to deploy highly available and fault-tolerant applications at the edge. These workloads are often some of the most demanding in terms of durability and availability requirements. Deploying workloads on AWS hybrid cloud and edge computing services like AWS Local Zones and AWS Outposts requires planning for networking, compute, storage, and other failure modes. In this chalk talk, dive deep into each of these services, review use cases, and examine reference architectures. Learn design considerations, deployment patterns, and best practices for high availability and disaster recovery.
Workshop
ARC401-R | Advanced cross-Region DR patterns on AWS [REPEAT]
Join this hands-on workshop to explore a resilient, cloud-native architecture that surpasses the stringent availability and recovery regulations for financial markets utility providers. This highly available design uses Amazon ECS, Java Spring Boot, Amazon MQ, Amazon Kinesis, Amazon DynamoDB, Amazon Aurora, Amazon Application Recovery Controller, and AWS Systems Manager. Delve into key design decisions, including cross-Region data stores, messaging, and cross-Region state replication. Also learn about exactly-once transaction processing, reliable global traffic shifting, and best practices for fail-safe disaster recovery orchestration. You must bring your laptop to participate.
ARC303-R | Building, operating, and testing resilient Multi-AZ applications [REPEAT]
In this workshop, get hands-on experience building, operating, and testing a resilient Multi-AZ application. Use Amazon CloudWatch dashboards, insights rules, and composite alarms to observe the health of your application. Then, inject randomized faults using AWS Fault Injection Service to simulate a variety of Single-AZ impairments. Also, learn how to use AWS CodeDeploy to perform zonal deployments and experience deployment failures. Finally, use Amazon Application Recovery Controller zonal shift to recover from these failures and protect your customer experience. You must bring your laptop to participate.
ARC304 | Disaster recovery on AWS
In this workshop, get hands-on experience implementing different AWS disaster recovery strategies while building a Unishop application. Each module throughout the workshop highlights the AWS strategies and services that have features that make it easy to implement a disaster recovery plan. You must bring your laptop to participate.
STG403-R | Safeguard and audit data protection with AWS Backup [REPEAT]
Backup is critical for organizations to meet business operation and regulatory compliance needs. It also helps protect against disasters, accidental deletions, and ransomware. In this workshop, explore different scenarios using AWS Backup. Find out how you can define policies to centrally protect and restore AWS services; choose among Amazon EC2 instances, Amazon EBS volumes, Amazon S3 objects, Amazon Aurora databases, and Amazon EFS file systems; configure AWS Backup Vault Lock to enable write-once, read-many (WORM) backup data; and deliver simple, auditor-ready compliance monitoring and reporting with AWS Backup Audit Manager. You must bring your laptop to participate.
COP305 | Using observability for effective incident response
Effective incident management is crucial for business continuity. This hands-on workshop simulates an incident, and you discover how to collect, analyze, and correlate data from various sources to gain a holistic understanding of your system’s behavior. Explore techniques for setting up effective alerting and automated workflows to proactively identify and respond to incidents using Amazon CloudWatch and AWS Systems Manager. You must bring your laptop to participate.
Closing comments
Whether this is your first re:Invent or you’re a multiyear veteran, I hope you find this guide helpful in navigating the session catalog. Remember to make the most out of networking options available. Whether via PeerTalk, “hallway track,” the community booths, or the hero lounges, don’t hesitate to meet and talk to people.