Beyond the Checkbox: What Breaks When You Actually Stress-Test Cloud Incident Response
M Harvey (AWS)
fwd:cloudsec North America 2026 · Day 1
Overview
In the realm of cloud security, the true test of an organization's incident response (IR) capabilities often comes not from theoretical discussions but from real-world chaos. Matthew Harvey, a founding member of AWS's customer incident response team, presented a compelling talk at fwd:cloudsec on the critical disconnect between how organizations think their IR processes work and how they actually perform under pressure. His presentation, "Beyond the Checkbox: What Breaks When You Actually Stress-Test Cloud Incident Response," delves into the shortcomings of traditional tabletop exercises and introduces a more rigorous, Dungeons & Dragons-inspired approach to simulation.

Key moments
- 0:40 Introduction to AWS SERS and its purpose
- 2:00 Explaining SERS: A Dungeons & Dragons approach
- 2:50 Anatomy of a SERS scenario: overview, expectations, criteria
- 5:00 Observation: Customers often stack the odds
- 6:00 Observation: Disconnects in response speed and risk
- 7:00 Observation: "It's not our job" mentality
- 8:00 Solution: Stress-testing and executing playbooks
Beyond the Checkbox: What Breaks When You Actually Stress-Test Cloud Incident Response
Speakers: M Harvey, AWS
Conference: fwd:cloudsec
YouTube: https://www.youtube.com/watch?v=L3NlGD_lV9I
Overview
In the realm of cloud security, the true test of an organization's incident response (IR) capabilities often comes not from theoretical discussions but from real-world chaos. Matthew Harvey, a founding member of AWS's customer incident response team, presented a compelling talk at fwd:cloudsec on the critical disconnect between how organizations think their IR processes work and how they actually perform under pressure. His presentation, "Beyond the Checkbox: What Breaks When You Actually Stress-Test Cloud Incident Response," delves into the shortcomings of traditional tabletop exercises and introduces a more rigorous, Dungeons & Dragons-inspired approach to simulation.
Harvey's insights are born from over 50 Security Incident Response Simulations (SERS) conducted with diverse AWS customers. The core message is clear: many organizations inadvertently "stack the odds" in their favor during simulations, leading to a false sense of security. This talk is a vital resource for any security professional, CISO, or incident responder seeking to move beyond superficial compliance and genuinely fortify their cloud incident response posture against the inevitable. It provides actionable methodologies to uncover hidden flaws and improve organizational resilience before a high-stakes security event demands perfection.
Background
▶ Watch: Introduction to AWS SERS and its purpose (0:40)
Matthew Harvey's extensive experience at AWS, spanning over 12 years, positions him uniquely to discuss the intricacies of cloud incident response. As a founding member of the AWS customer incident response team, his primary mission is to guide customers through security events occurring on their side of the shared responsibility model. However, a crucial secondary mission of his team is to proactively help customers identify and address their security vulnerabilities before a real incident necessitates their intervention. This preventative approach is primarily facilitated through tabletop exercises, which AWS has refined into its proprietary Security Incident Response Simulation (SERS) methodology.
Traditional tabletop exercises, while common, often fall short. They typically involve a group of executives or technical IR staff discussing a hypothetical event, moving "pieces across the board" to visualize a resolution. Harvey highlights that these generic simulations frequently suffer from a fundamental flaw: participants often subconsciously (or consciously) manipulate the scenario to ensure a favorable outcome. This can manifest as an overestimation of response speed, a simplification of complex technical challenges, or a reluctance to acknowledge inter-team dependencies. The problem, therefore, isn't the concept of simulation itself, but the lack of realistic stress-testing that genuinely exposes the friction points and single points of failure within an organization's incident response framework. Harvey's work aims to inject a level of realism and ambiguity into these simulations, mirroring the unpredictable nature of actual security incidents.
Key Findings
▶ Watch: Anatomy of a SERS scenario: overview, expectations, criteria (2:50)
Based on his experience running over 50 SERS exercises, Matt Harvey identified four critical observations about how organizations typically approach incident response simulations, alongside several methodologies to counter these tendencies:
Observations on Flawed Tabletop Approaches:
- Stacking the Odds: Customers frequently manipulate scenarios to their advantage. This includes:
- Modifying Scope of Breach: Suggesting a smaller system or narrower permissions than initially presented in the scenario.
- Faster Containment: Assuming immediate and effortless containment, akin to "Ghostbusters" effortlessly trapping a ghost, rather than acknowledging the complex, multi-step process involved in real-world containment.
- Reframing High-Risk as Low-Risk: Dismissing a threat actor's continued presence in a system as benign ("they can't do anything") even after initial permissions are removed, ignoring the inherent risk of persistent access.
- "It's Not Our Job": Teams deflecting responsibility for specific actions (e.g., "that's Identity and Access Management's problem") rather than collaboratively addressing the full incident lifecycle.
- Overestimating Data Retrieval Timelines: A staggering 90% of organizations in these tabletops drastically overestimate how quickly they can perform a log dive. While participants might assume a 30-minute log retrieval, the reality for a multi-account, large-scale event (e.g., spanning 1,000 EC2 instances) could be weeks. This discrepancy is critical because an active threat actor continues to operate during this prolonged data analysis period. The root causes include:
- Syntax Challenges: Lack of familiarity with specific log query syntax.
- Scope Challenges: Difficulty coordinating and querying logs across numerous accounts or services.
- Delays: The sheer volume of data and the manual effort required. Management often fails to grasp the complexity and time required for thorough log analysis.
- Single Points of Failure (SPOFs): These can be technical or non-technical systems where a breakdown in logic can cripple a larger response. Examples include:
- Human SPOFs: A critical individual (e.g., "Jim") holding an essential One-Time Password (OTP) or unique knowledge, but being unavailable (on vacation, not picking up the phone).
- Technical SPOFs: Storing crucial security logs on the very system that has been compromised, rendering them inaccessible or untrustworthy during an incident.
- Lack of Metrics: A general absence of data to identify and track SPOFs.
- Process Decay: The natural degradation of processes due to organizational turnover, lack of updates, and neglect.
- Outdated Playbooks: Playbooks are not updated when personnel change, leading to calls to former employees who no longer have a stake in the organization's security.
- Knowledge Silos: Critical knowledge resides with individuals who eventually leave, without proper documentation or transfer.
Methodologies for Improvement:
To address these findings, Harvey proposes several methodologies to make tabletop exercises more effective:
- Actually Run Through Playbooks: Instead of simply stating "we have a playbook for that," facilitators should demand that participants pull up and execute the playbook in real-time. This includes making actual "war room" calls or paging on-call personnel to test their responsiveness and the validity of contact information.
- Embrace the "Insider Threat" Mindset: When scenarios seem far-fetched, introducing the concept of an insider threat immediately makes them plausible and increases participant engagement and seriousness.
- Prioritize "Water Cooler Conversations": The informal, spontaneous problem-solving that happens during in-person breaks is invaluable. Remote exercises often lose this critical element, hindering organic collaboration and discovery.
- Implement a Stakeholder Safety Pledge: To combat the fear of reprisal (e.g., during performance review periods), a CISO or senior leader should proactively pledge that identifying flaws during the exercise is a positive action and will not lead to negative consequences. This fosters an environment of psychological safety, encouraging open and honest discussion.
- Design Scenarios in the "Goldilocks Zone": Scenarios should be neither too complex nor too simple. They must contain enough ambiguity to force teams to interact, communicate, and innovate, rather than follow a rigid, pre-determined path. They should be grounded in credible, data-driven injects (e.g., an open S3 bucket with a ransom note) and avoid "landing helicopters on roofs" — overly dramatic or unrealistic elements. Each scenario should have a clear inject, potential solutions, and measurable outcomes.
- Optimize Timing and Pacing:
- Technical Teams: 4-hour sessions, broken into 1-1.5 hour segments, are ideal for deep dives.
- C-level/Stakeholders: 90-minute sessions focusing on communication, legal, PR, and customer interaction are more effective, using a single, evolving inject rather than multiple mini-scenarios.
Technical Deep Dive
▶ Watch: Observation: Customers often stack the odds (5:00)
The AWS Security Incident Response Simulation (SERS) is designed as a sophisticated, Dungeons & Dragons-style adventure for incident response. The facilitator acts as the Game Master, guiding participants through a scenario and occasionally introducing randomized elements or challenges using a d20 die, which, while initially met with strange looks, helps to inject realism and unpredictability.
Each SERS scenario is structured around three key components:
- Overview: A concise description of the security event (e.g., "security company Xactor contacts via email about a public AWS resource containing 'secret sauce' like PHI or credit card information").
- Expectations (Guard Rails): Limits and boundaries for the scenario to prevent discussions from going "off the rails" and keep the focus tight.
- Success Criteria: Defines how the customer and facilitator will know if the desired outcomes of the simulation have been met.
- Outcomes: Emphasizes the critical need for metrics to derive tangible improvements from the discussion.
A major technical challenge highlighted is data retrieval timelines, particularly concerning log management. In a real-world multi-account event involving potentially thousands of EC2 instances, coordinating and querying logs becomes incredibly complex. Harvey stresses that "you're not going to have a playbook that has every potential syntax for every potential log dive." The challenges include:
- Syntax Complexity: Different log sources, services, and query languages require specialized knowledge.
- Cross-Account Scope: Aggregating and correlating logs across numerous AWS accounts is a significant architectural and operational hurdle.
- Coordination: Identifying who initiates the log queries, who tracks their progress, and where the results are centrally stored and analyzed. Without robust Security Information and Event Management (SIEM) or centralized logging solutions, this becomes a critical bottleneck.
The concept of Single Points of Failure (SPOFs) extends beyond human elements to technical systems. A critical example cited is storing security logs on the very system that a threat actor has compromised. If the attacker gains control of the logging infrastructure, the ability to investigate their actions is severely hampered or entirely lost. This underscores the importance of resilient, out-of-band logging mechanisms, often involving separate accounts, services, or even external solutions for log aggregation and retention.
Crucially, the SERS methodology deliberately injects ambiguity into scenarios. For instance, an "open S3 bucket with a ransom note" is a straightforward premise, but the subsequent actions and inter-team coordination required to address it are fraught with unknowns. This ambiguity forces teams to interact, communicate, and make decisions under uncertainty, mimicking the real-world pressures of an incident far more effectively than a step-by-step checklist. The focus is not on a "correct" answer but on observing how teams collaborate, identify gaps, and adapt. The constant emphasis on metrics throughout the talk serves as the technical underpinning for continuous improvement, allowing organizations to quantify "how many single points of failure do you have in your playbooks?" and whether "that number [is] trending up or down."
Demo / Proof of Concept
▶ Watch: Observation: "It's not our job" mentality (7:00)
While Matthew Harvey's talk did not feature a live technical demonstration of a specific tool or a traditional proof-of-concept exploit, the entire methodology of the Security Incident Response Simulation (SERS) itself serves as a practical, real-world "proof of concept" for improving incident response capabilities. The "demo" described is the structure and execution of these Dungeons & Dragons-style tabletop exercises.
Harvey walked the audience through an example scenario directly from an AWS SERS deck. This scenario involved "security company Xactor contacting the email address to boast that their new security product has found a public AWS resource that contains what appears to be some sort of secret sauce" (e.g., PHI, credit card information). This credible inject sets the stage. The facilitator (Game Master) then guides the participants, introducing elements of unpredictability, sometimes even by having them roll a d20 to determine certain outcomes, reflecting the chaotic nature of real incidents.
The "proof" of this concept lies in its ability to expose critical operational gaps that traditional, less rigorous tabletops often miss. By forcing participants to:
- Actively execute playbooks: Not just discuss them, but simulate paging on-call, pulling up actual documents, and following procedures.
- Grapple with ambiguity: Scenarios are designed to lack clear-cut solutions, pushing teams to collaborate and innovate.
- Confront realistic timelines: The simulation highlights the actual, often extended, time required for tasks like multi-account log retrieval.
- Identify human and technical SPOFs: The exercise naturally surfaces individuals or systems that represent single points of failure.
The "outcomes" and subsequent metrics derived from these simulations are the ultimate proof points. As Harvey notes, customers who undergo these rigorous exercises often return six months later with detailed action plans, having addressed knowledge gaps, process decay, and identified SPOFs, demonstrating the tangible impact of this stress-testing approach. The SERS itself is the demonstration of a more effective way to prepare for, and respond to, cloud security incidents.
Defensive Implications
▶ Watch: Solution: Stress-testing and executing playbooks (8:00)
The insights from Matthew Harvey's talk provide critical guidance for defenders looking to strengthen their cloud incident response posture beyond mere compliance. The core defensive implications revolve around moving from theoretical understanding to practical, stress-tested readiness.
- Embrace Realistic Stress-Testing: Organizations must move beyond checkbox exercises. Design and execute simulations that deliberately introduce ambiguity, challenge assumptions, and prevent participants from "stacking the odds." Incorporate scenarios that are grounded in real-world threats and data, including the possibility of insider threats to enhance plausibility and engagement.
- Validate Playbooks Through Execution: It's insufficient to simply have playbooks. Defenders must regularly execute them in a simulated environment. This means pulling up the actual documents, attempting to contact on-call personnel (even if only a simulated page), and following every step. This process will inevitably expose outdated contact information, missing steps, and procedural gaps that would be catastrophic during a real incident. Establish a regular cadence for playbook review and update to combat process decay.
- Fortify Log Management and Observability: The significant overestimation of data retrieval timelines is a glaring defensive weakness. Organizations must invest in robust, centralized, and cross-account log aggregation and analysis solutions (e.g., a well-configured SIEM or data lake). Implement strong tagging strategies and consistent logging across all cloud resources (e.g., AWS CloudTrail, VPC Flow Logs, S3 access logs). Critically, incident responders need practical experience with query syntax and a clear understanding of the realistic time and resources required for complex log dives, especially across multiple accounts. Ensure logs are stored securely and independently from the systems they monitor to avoid them becoming a single point of failure.
- Identify and Mitigate Single Points of Failure (SPOFs): Conduct thorough assessments to identify both technical and human SPOFs. Technically, this means ensuring critical systems (like log storage) are resilient and isolated. Operationally, it means eliminating reliance on single individuals for critical knowledge or access (e.g., ensuring multiple people have access to critical OTPs or documentation). Implement multi-factor authentication, robust access management, and knowledge transfer programs.
- Cultivate Psychological Safety: Implement a Stakeholder Safety Pledge at the outset of any incident response exercise. Senior leadership must clearly communicate that identifying flaws and gaps is a positive contribution, not a cause for punishment. This fosters an environment where teams feel safe to openly discuss weaknesses, which is crucial for genuine improvement.
- Focus on Metrics and Continuous Improvement: Treat incident response as an engineering discipline. Define clear metrics for incident response capabilities, such as time to detection, time to containment, time to resolution, number of SPOFs identified, and playbook effectiveness. Track these metrics over time to measure progress, identify areas needing improvement, and demonstrate the value of security investments. Regularly generate reports detailing findings from simulations to ensure identified gaps are addressed.
By adopting these defensive strategies, organizations can move "beyond the checkbox" and build a truly resilient cloud incident response capability, prepared for the unpredictable realities of modern cyber threats.
Key Takeaways
- Human Telemetry is Critical: Effective communication and collaboration among security professionals and stakeholders are irreplaceable. No tool can automate the nuanced human interaction required to navigate complex security incidents.
- Psychological Safety is Paramount: Implement a "Stakeholder Safety Pledge" to create an environment where teams feel secure in openly identifying security gaps and weaknesses without fear of reprisal, leading to more honest and productive simulations.
- Concrete Data and Metrics Drive Improvement: Go beyond qualitative assessments. Track measurable outcomes from simulations, such as the number of single points of failure, playbook execution times, and log retrieval efficiency, to drive continuous improvement in IR capabilities.
- Stress-Test, Don't Just Tabletop: Design and execute incident response simulations that deliberately introduce ambiguity, challenge assumptions, and prevent participants from "stacking the odds" to uncover true operational and technical weaknesses.
- Validate Playbooks in Practice: Don't assume playbooks work. Actively simulate their execution, including paging on-call personnel and attempting real-time data retrieval, to expose outdated procedures, contact information, and operational gaps.
- Proactively Address Process Decay and SPOFs: Establish a regular cadence for reviewing and updating playbooks, documenting knowledge transfer, and identifying and mitigating both technical and human single points of failure before an incident occurs.
About the Speaker(s)
Matthew Harvey is a highly experienced security professional with over 12 years at AWS. He is one of the founding members of the AWS customer incident response team, whose primary mission is to assist customers in navigating security events that occur within their side of the AWS shared responsibility model. Beyond reactive incident support, Harvey's team also focuses on proactive measures, helping customers identify and understand potential flaws in their security posture before they lead to a full-blown incident. He is instrumental in developing and conducting AWS's Dungeons & Dragons-inspired Security Incident Response Simulations (SERS), which are designed to stress-test and improve organizational incident response capabilities.
Reviews
Dr. Zero (Offensive Security Researcher) — SOLID
Competent practitioner talk from someone who's clearly done the work — 50+ SERS engagements gives Harvey real pattern-recognition on where IR programs rot. The observations are valid and the methodology is sensible, but none of this is surprising to anyone who's run tabletops seriously. Fills a slot well for an audience that hasn't already internalized these lessons.
Heather Calloway (CISO) — SOLID
Harvey's talk is operationally grounded and the findings are real — 90% of organizations can't accurately estimate log retrieval timelines, playbooks rot, and single points of failure hide in plain sight. The problem is the talk stops at diagnosis: it tells security teams what breaks, but doesn't give executives or governance structures a clear accountability model for fixing it.