One SOC, The Whole SOC, and Nothing But The SOC, So Help Me

Carson Zimmerman

BSidesSF 2025 — Here Be Dragons · Day 1 · Main

Overview

Carson Zimmerman, architect of Microsoft's Security Operations Center and author of MITRE's Eleven Strategies of a World-Class Cybersecurity Operations Center, argues that SOCs fail not from lack of tooling or talent, but from structural mistakes that fragment the functions that must work together. His prescription: keep the SOC atomic, integrate engineering and threat intelligence directly into the operation, build genuine capacity for non-incident work, and make truth-telling — not just threat detection — a first-class SOC function. ---

Watch on YouTube

Visual summary for One SOC, The Whole SOC, and Nothing But The SOC, So Help Me by Carson Zimmerman
Visual summary for One SOC, The Whole SOC, and Nothing But The SOC, So Help Me by Carson Zimmerman

Key moments

  1. 2:00 Universal SOC truth: every SOC feels like major struggle regardless of resources
  2. 4:00 SOC atomicity: six functions must stay together or the SOC fails every time
  3. 12:00 Mistake 1: engineering outside SOC creates wall-throwing and cowboy operators
  4. 15:59 Detection is engineering: v-team model enables hour-turnaround detection writing
  5. 17:59 Mistake 2: hunting team elitism and exempting them from ops rotation is toxic
  6. 21:59 Mistake 4: reserve 20-50% SOC capacity for non-incident improvement work
  7. 23:59 Post-incident show-and-tell: free training by walking through every query run

One SOC, The Whole SOC, and Nothing But The SOC, So Help Me

Speaker: Carson Zimmerman

Conference: BSidesSF 2025 — April 26-27, 2025, San Francisco

YouTube: https://www.youtube.com/watch?v=0WtATprKq60

Reading time: ~8 minutes

TL;DR

Carson Zimmerman, architect of Microsoft's Security Operations Center and author of MITRE's Eleven Strategies of a World-Class Cybersecurity Operations Center, argues that SOCs fail not from lack of tooling or talent, but from structural mistakes that fragment the functions that must work together. His prescription: keep the SOC atomic, integrate engineering and threat intelligence directly into the operation, build genuine capacity for non-incident work, and make truth-telling — not just threat detection — a first-class SOC function.

Introduction

Every SOC feels like it is in crisis. Carson Zimmerman has seen this across organizations of every size and maturity level, across more than two decades of SOC work — as an analyst, engineer, architect, and, most recently, as the architect of Microsoft's own security operations center. The constant struggle is not the exception; it is the baseline condition. What distinguishes SOCs that improve from those that stagnate is not better technology or larger headcount, but whether they have structured themselves correctly in the first place.

At BSidesSF 2025, Zimmerman distilled this experience into a deceptively simple thesis: the SOC is atomic. Certain functions belong inside it, and every time those functions are separated out, the SOC degrades — without exception. His talk walks through what belongs in the SOC, what doesn't, and four specific structural mistakes that organizations make over and over again, along with the patterns that fix them. The framework draws directly from his book, Eleven Strategies of a World-Class Cybersecurity Operations Center, available free from MITRE at mitre.org/11strategies.

The Atomic SOC: What Belongs Inside

▶ Watch: Defining the atomic SOC and its core functions (04:00)

The SOC's job is to turn the OODA loop — Observe, Orient, Decide, Act — faster and more accurately than the adversary. Every function that supports this loop belongs in the SOC. Every function that dilutes it or pulls resources toward something else introduces risk.

Zimmerman identifies the non-negotiable atomic components:

  • Triage, analysis, and response — the people looking at alerts, investigating, coordinating incidents, performing containment and eviction
  • Hunt and detection creation — proactively searching for adversary presence and building the detections that drive the triage function
  • Threat intelligence — understanding the adversary, its TTPs, and how they apply to the specific enterprise
  • SOC engineering — the tooling, telemetry, and automation that the operations functions depend on
  • SOC leadership and communications — situational awareness, training, and the ability to communicate what the SOC knows

Some functions are contextually appropriate additions — firewall management in some organizations, vulnerability scanning, PSIRT capabilities, composite asset inventory. Others look tempting but are traps: pen testing, compliance management, and especially anti-fraud and anti-abuse work. The last one is particularly dangerous because it is "SOC-like" in many ways, but when incidents are quiet, analysts will gravitate toward the interesting abuse data and lose focus on their core mission.

The rule Zimmerman invokes is the Cheesecake Factory test: just because everything is on the menu doesn't mean you should order everything. A SOC that tries to do too many things will do none of them well.

Mistake One: Engineering Outside the SOC

▶ Watch: The engineering anti-patterns and how to fix them (12:01)

The first and perhaps most common structural mistake is treating SOC engineering as a separate organizational entity. Zimmerman is blunt: "All SOCs have engineering in them. The only difference is whether you acknowledge it and fund it or not." When engineering is organizationally distant, requirements get thrown over a wall, tools get forced onto analysts who didn't ask for them, and the engineers — however smart and well-meaning — develop their own picture of what the SOC needs that diverges from reality.

The fix is proximity, but proximity alone is insufficient. Engineers embedded in the SOC face the same incident treadmill as analysts, which means strategic work never happens unless it is explicitly protected. Zimmerman recommends a detection v-team model: a structured rhythm of business that brings engineers and analysts together regularly, with dedicated time carved out from incident response for detection development, telemetry improvement, and tool work. The engineers should talk to analysts every day. "I literally wrote the book on this," Zimmerman notes, "and when I was in an engineering role I talked to my investigations lead at least once a week because there was stuff he knew about the mission that I didn't."

Mistake Two: Excluding Hunt and Threat Intel

▶ Watch: Hunt and threat intel integration (16:02)

The second mistake is treating the hunting and threat intelligence functions as elevated, separate disciplines that are exempt from the operational work of the SOC. Zimmerman has heard the attitude directly: "I've gone to the hunting team. I'm better than you. I don't have to do shift work anymore." His response is categorical: everyone in the SOC should be in an ops rotation of some kind. Hunters, threat intel analysts, detection engineers — all of them have operational roles to fill, and exempting them from the work creates hierarchy where there should be unity.

The skills used in hunting — working with large datasets to find adversary presence, orienting the SOC toward the right threats — are not fundamentally different from the skills used in triage and investigation. The difference is mostly in how they are applied. When hunt and threat intel are siloed away from operations, tools and processes stagnate, knowledge doesn't flow in either direction, and both sides wonder what the other is actually doing.

The goal Zimmerman sets is that everyone feels like they are in the fight. Career progression should include opportunities for analysts to move toward detection engineering, hunting, or management — and hunters and threat intel staff should be able to move back toward operations. No one should feel above the work.

Mistake Three: Excluding Truth-Telling

▶ Watch: Truth-telling as a SOC function (20:02)

The third mistake is subtle but consequential: failing to organize the SOC around its natural truth-telling capability. Every significant hunt, every major incident, every thorough post-incident review produces information about the state of the enterprise — misconfigured systems, missing telemetry, asset inventory gaps, process failures — that is separate from the adversary finding that prompted the investigation. This information "falls off the back of the truck" in the course of normal SOC work.

Zimmerman argues that the SOC should embrace this truth-telling role structurally, not just incidentally. This means formal post-incident review processes that capture what went wrong in the enterprise, not just what the adversary did. It means feeding vulnerability scan data, asset inventory gaps, and process failure observations into an accountable improvement track. It means the SOC has a mechanism for turning its incidental findings into action, rather than accumulating them in informal notes that nobody reads.

The SOC is uniquely positioned to be the enterprise's most honest voice about its own security posture, precisely because the SOC sees the adversary exploiting real weaknesses in real time. Failing to formalize that capability is leaving a major organizational asset unused.

Mistake Four: No Capacity for Improvement

▶ Watch: Building non-incident capacity and avoiding stagnation (22:02)

The fourth and final mistake is the one Zimmerman calls the stagnation trap: building a SOC that only has capacity for incidents. Every organization says it wants to improve, but almost none builds the structure, the rhythm of business, or the accountability mechanisms to make improvement happen. The result is a team that is perpetually reactive, perpetually exhausted, and perpetually losing its best people to organizations that offer something more than an endless incident queue.

His prescription is specific: reserve at least 20% of SOC capacity — ideally 50% or more — for non-incident work. This includes SOP and process improvement, tool and automation development, detection review, hunt and exploration, and post-incident learning sessions. This is not wasted capacity — it is surge capacity. When a major incident strikes, the people who have been doing improvement work are already warmed up and ready to contribute, rather than being pulled from elsewhere in a scramble.

The closed-loop principle matters here: the people who write detections should experience the consequences of those detections. Did the detection fire? Did it generate false positives? Did the analyst know what to do with it? Detection engineers who are isolated from operations will never develop the feedback loops that make their work better. Similarly, post-incident technical share-outs — where analysts walk through the queries they ran, the joins they used, the pivots they made — are nearly free team-building and knowledge-transfer opportunities that most SOCs leave on the table.

Notable Quotes

"All SOCs have engineering in them. The only difference is whether you acknowledge it and fund it, or not."

— Carson Zimmerman [[▶ 12:01]](https://www.youtube.com/watch?v=0WtATprKq60&t=721s)

"If you're not doing improvement work, you're stagnating, and if you're stagnating, you're failing, and people are leaving. Build at least twenty percent — ideally fifty percent or more — of capacity for non-incident work."

— Carson Zimmerman [[▶ 22:02]](https://www.youtube.com/watch?v=0WtATprKq60&t=1322s)

"When I was talking to executives about security, I wasn't talking about security anymore. I was talking about their mission, their business, what matters to them. Reframe the argument around that, and instantly executives pay attention."

— Carson Zimmerman [[▶ 28:04]](https://www.youtube.com/watch?v=0WtATprKq60&t=1684s)

Key Takeaways

  • The SOC is atomic. Triage/analysis/response, detection engineering, hunt, threat intelligence, and SOC engineering belong together. Separating any of them causes consistent, predictable failure — without exception.
  • Engineering belongs in the SOC, not adjacent to it. Engineers who don't talk to analysts daily will build the wrong things. Proximity enables the feedback loops that make detection work effective.
  • No one is above ops rotation. Hunters and threat intel analysts who are exempt from operational work create hierarchy and stagnation. Everyone should have roles to fill in the operational rhythm.
  • Truth-telling is a SOC function, not a byproduct. Post-incident reviews, asset inventory gaps, and process failures are things the SOC discovers constantly. Making those findings actionable requires organizational structure, not just goodwill.
  • Improvement requires protected capacity. SOCs that allocate 100% of resources to incident response have no capacity to get better. Reserve explicit time for non-incident work — it is also surge capacity when the big incident eventually arrives.

Reviews

Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT

Zimmerman wrote the book on SOC design — literally, MITRE's Eleven Strategies — and he delivers the distilled version here with the authority of someone who architected Microsoft's SOC and watched the same structural failures repeat across every organization he's touched. The four-mistake framework is concrete, the prescriptions are specific, and the 20-50% non-incident capacity target is the kind of uncomfortable truth most SOC leaders already know and need to hear again.

Heather Calloway (CISO) — STRONG ACCEPT

Zimmerman has built Microsoft's SOC and written the MITRE book on world-class security operations, and his diagnosis of how SOCs fail — by externalizing engineering, exempting hunt and threat intel from operations rotation, failing to formalize truth-telling, and building zero capacity for non-incident work — is grounded in two decades of direct experience. The 20-50% capacity-for-improvement recommendation is the one most organizations will resist and most need.

→ Top-rated talks at BSidesSF 2025 — Here Be Dragons

All talks from BSidesSF 2025 — Here Be Dragons