Making WAF Mainstream: From Static Defenses to Living, Learning Protection
Roy Weisfeld (Co-founder and CTO · Huskys), Surya Pentakota (Enterprise to Edge Security Lead · Bytedance/TikTok)
BSidesSF 2026 · Day 2 · AMC Theatre 10
Overview
This talk, "Making WAF Mainstream: From Static Defenses to Living, Learning Protection," delivered by Roy Weisfeld and Surya Pentakota, addresses the pervasive frustrations associated with traditional Web Application Firewalls (WAFs) and proposes a revolutionary approach to transform them into intelligent, adaptive edge security layers. The speakers argue that conventional WAFs, with their static rule sets and fragmented visibility, are inadequate for today's complex, multi-cloud, and global application environments, leading to significant operational overhead, revenue loss, and security gaps.
Key moments
- 0:00 Welcome and speaker introductions
- 2:20 What will be covered in this talk
- 3:00 WAF frustrates everyone: The core problem
- 4:20 Engineers' pain points with traditional WAFs
- 6:40 Executives' concerns: WAF's business and revenue impact
Making WAF Mainstream: From Static Defenses to Living, Learning Protection
Speakers: Roy Weisfeld, Co-founder and CTO at Huskys; Surya Pentakota, Enterprise to Edge Security Lead at ByteDance/TikTok
Conference: BSides SF
YouTube: https://www.youtube.com/watch?v=lNOkQcH3t7Y
Overview
This talk, "Making WAF Mainstream: From Static Defenses to Living, Learning Protection," delivered by Roy Weisfeld and Surya Pentakota, addresses the pervasive frustrations associated with traditional Web Application Firewalls (WAFs) and proposes a revolutionary approach to transform them into intelligent, adaptive edge security layers. The speakers argue that conventional WAFs, with their static rule sets and fragmented visibility, are inadequate for today's complex, multi-cloud, and global application environments, leading to significant operational overhead, revenue loss, and security gaps.
Roy Weisfeld, drawing on his extensive background as a web hacker and infrastructure expert, and Surya Pentakota, who leads edge security for the massive scale of TikTok, combine their perspectives to highlight the critical need for change. They contend that WAFs, despite being the first line of defense for many organizations, are often neglected or poorly managed due to their inherent complexities and the difficulty in correlating security events with business context. The talk aims to demystify WAF management and introduce a unified framework that leverages Artificial Intelligence (AI) to create a living, learning protection system.
The core message of the presentation is that while WAFs themselves are not going away, their operational model must fundamentally shift. By integrating AI for context enrichment, automated remediation, and predictive analysis, organizations can overcome the limitations of static defenses. This transformation promises not only enhanced security posture but also significant improvements in operational efficiency, incident response times, and overall business resilience, particularly for global platforms like TikTok that operate at immense scale and face a constantly evolving threat landscape.
Background
▶ Watch: Welcome and speaker introductions (0:00)
The current state of WAFs is universally described as "sucking" by the speakers, causing significant frustration for both engineers and executives. From an engineering perspective, the challenges are numerous and deeply ingrained in legacy approaches. Static rule sets are a primary culprit, being difficult to map against dynamic infrastructure and specific business logic. This often leads to false positives, blocking legitimate traffic and resulting in revenue loss. Engineers also struggle with visibility gaps due to the sheer volume of telemetry and millions of logs generated by WAFs, making it nearly impossible to derive meaningful insights or proactively detect threats. The common refrain of "it's always at fault" highlights how WAFs become scapegoats for application issues, further complicating troubleshooting and ownership. Critically, application context is missing from raw WAF logs, making it hard to prioritize attacks based on business impact or understand the revenue flow. Lastly, orchestrating policies across multiple WAF technologies (CDN edge, infrastructure edge, application layer, load balancer) in a defense-in-depth strategy is a monumental challenge, as a "one policy for all" approach is impractical.
For executives, these technical headaches translate directly into critical business risks. Revenue impacts are a major concern, as exemplified by a simple policy change blocking a checkout API, leading to immediate financial loss. The blast radius of attacks is often unknown, making it difficult to assess the full extent of damage. Ownership finger-pointing wastes valuable time and resources, frequently resulting in paging the wrong engineers during critical incidents. The constant threat of blocking live business traffic due to false positives makes executives wary of WAFs, often leading to calls to "just shut down WAF" – a non-starter for security practitioners. Ultimately, the lack of clarity on financial risk and mitigation strategies due to WAF chaos is a major point of contention for CFOs and leadership.
The prevailing architecture exacerbates these issues, presenting a "chaos" of disparate systems. Organizations typically manage traffic flowing through multiple layers: CDN providers (Cloudflare, CloudFront, Akamai), WAFs (rate limiting, DDoS protection), load balancers (VPCs, ALBs, security groups), and finally, the virtual machines hosting web servers. This multi-layered defense, while conceptually sound, translates into fragmented visibility across four or more different dashboards, each with its own context. Assets and configurations are not uniformly mapped, and diverse traffic types—legitimate customers, automation, AI agents, bots—get lost in this complexity. This current state prevents a holistic understanding of "what is blocked, where is it blocked, and on which layer do you even need to look at when you try to investigate it." The fundamental problem is that "you can't really protect what you can't see," and the traditional WAF approach fails to provide this comprehensive visibility.
Key Findings
▶ Watch: What will be covered in this talk (2:20)
The talk presents several key findings and contributions aimed at revolutionizing WAF management:
- WAFs Fundamentally "Suck" Today: The core problem identified is the inherent limitations of traditional WAFs, characterized by static rule sets, lack of visibility into business context, high operational overhead, and a tendency to generate false positives that impact revenue. This leads to a reactive and frustrating security posture.
- Shift from WAF to Unified Edge Security Layer: The speakers advocate for moving beyond the narrow concept of a WAF to a holistic edge security layer. This integrated approach combines CDN, WAF, load balancer, and compute layer configurations and telemetry into a single, understandable view, addressing the current chaos of disparate dashboards and unmapped assets.
- AI is Not a Buzzword, But a Necessity for Scale: A central finding is that Artificial Intelligence (AI) is indispensable for modern WAF management. It enables context enrichment, predictive analysis, and the creation of agentic flows for automated, intelligent remediation, which is crucial for handling the scale and dynamism of today's attack surface.
- AI-Driven Agentic Flow for Automated Remediation: The proposed AI framework facilitates a multi-step, adaptive remediation process. This includes detecting anomalies, enriching context with business logic, generating specific mitigation actions, backtesting against past traffic in log-only mode to minimize false positives, and then safely deploying and validating policies at scale.
- TikTok Case Study Validates the Approach: TikTok's implementation of this unified, AI-driven framework serves as a powerful validation. It demonstrated a drastic reduction in policy deployment time (from over 30 minutes to under 3 minutes), significant reduction in error rates and false positives, enhanced bot detection and rate limiting, and the ability to generate custom, context-rich executive reports.
- Unified Framework Addresses Vendor Lock-in and Migration Challenges: The framework's ability to normalize and correlate configurations across different CDN and WAF vendors provides a solution to the long-standing problem of vendor lock-in. It enables the translation and migration of WAF rules between providers, turning years-long projects into manageable tasks.
- Enhanced Visibility Leads to Actionable Insights: By correlating data from external intelligence, WAF/CDN layers, cloud infrastructure, diverse traffic sources, and security context, the framework offers unparalleled visibility. This allows for identifying configuration drifts, potential bypasses, and assessing the overall security posture with a clear scoring mechanism.
- AI as a "Force Multiplier," Not a Replacement: The speakers emphasize that AI's role is not to replace security engineers or existing WAFs/CDNs, but to act as a force multiplier, enabling teams to manage increasing traffic and attack complexity more efficiently and effectively without necessarily increasing manpower.
Technical Deep Dive
▶ Watch: WAF frustrates everyone: The core problem (3:00)
The proposed solution to the WAF dilemma is a unified framework built upon an intelligent architecture, designed to provide a holistic view and automated response capabilities. This framework is composed of three main components: Data Ingestion, the AI Core, and Outputs and Actions.
The process begins with Data Ingestion, where context is gathered from five critical sources:
- External Intelligence: Information about current attacker tactics, techniques, and procedures (TTPs), digital footprint exposures, and known vulnerabilities.
- WAF and CDN Layer: Configurations, rule sets, routing policies, and logs from various WAFs (e.g., Azure Front Door, AWS WAF, Cloudflare) and CDNs (e.g., Akamai, CloudFront). This captures how assets are currently protected and traffic is routed.
- Cloud Layer: Details about business applications, infrastructure, and the assets being protected within the cloud environment (e.g., AWS, GCP, Azure).
- Traffic Sources: Differentiated traffic patterns from mobile devices, desktops, AI agents, bots, and internal services, each requiring specific handling.
- Security Context: Data from internal security programs like bug bounties, penetration testing reports, and CNAPP tools that enrich existing asset and configuration information.
Once ingested, these disparate data streams undergo normalization to create a consistent and actionable format. This is where the AI Core takes over, operating in three stages:
- Correlation Engine: This engine connects the normalized data across different layers – CDN, WAF, routing, and crucially, business logic. It identifies relationships and dependencies that are invisible in fragmented systems, allowing for a comprehensive understanding of traffic flow and policy enforcement.
- Intelligence Engine: With the correlated context, this engine analyzes various signals to identify anomalies, potential threats, and misconfigurations. It moves beyond simple pattern matching to understand the intent and impact of events.
- Signal Fusion and Verdict Engine: This final AI component aggregates findings from the intelligence engine, applies confidence scoring to different signals (e.g., a managed rule blocking an entire application has lower confidence than a custom rule on a specific endpoint), and performs deduplication and prioritization. This ensures that only high-fidelity, actionable insights are presented, reducing alert fatigue.
A key technical innovation is the AI agentic flow for automated remediation. When an anomaly is detected (e.g., a spike in 4xx responses blocking 3,000 requests), the AI system orchestrates the following:
- Contextual Enrichment: Beyond raw logs, the AI integrates business context (e.g., identifying a blocked endpoint as a "booking endpoint" or related to an "ad campaign") and analyzes specific payload details (e.g., cookies containing two hyphens).
- Remediation Generation: Based on the enriched context, the AI generates a highly specific and fine-tuned remediation rule (e.g., an exclusion for cookies with two hyphens on a particular booking API endpoint).
- Backtesting and Iteration: The proposed rule is rigorously backtested against past traffic in a log-only mode. This crucial step simulates the rule's impact without affecting live production, identifying potential false positives and legitimate traffic that might be blocked. The AI iterates on the rule until a high confidence score is achieved, ensuring minimal impact on legitimate users.
- Deployment and Validation: Once refined, the policy is automatically deployed to the relevant WAF (e.g., Azure Front Door). The system then validates its effectiveness across all deployment zones and logs the decision for auditability, notifying relevant application owners. This entire process is designed to happen at scale, addressing environments with thousands of rules and multiple applications.
This framework also inherently addresses the vendor lock-in problem. By abstracting and normalizing WAF configurations, it creates an intermediate representation that can be translated between different WAF vendors. For instance, an 8-line AWS WAF rule for blocking an admin route could be translated to a 2-line Cloudflare rule. This capability significantly streamlines migrations between providers, allowing automatic translation of many rules and highlighting those requiring manual review, reducing migration projects from years to weeks or months. The goal is to make WAFs not just protective, but also manageable, agile, and adaptable to evolving business and infrastructure needs.
Demo / Proof of Concept
▶ Watch: Engineers' pain points with traditional WAFs (4:20)
The talk presented several real-world examples and a detailed case study to demonstrate the efficacy of their unified, AI-driven framework.
The first example illustrated handling a spike in 4xx responses on a booking endpoint:
- Detection: The system identified a significant increase in 4xx errors on a critical booking API endpoint, which was not immediately correlated with known attack patterns.
- Investigation & Isolation: Using specialized queries (e.g., an Azure Front Door query to filter blocked logs), the engine pinpointed that requests to "bookings" and "tickets" API endpoints were being blocked. Further analysis revealed that these blocked requests contained cookies with two hyphens (
--). This pattern is often associated with SQL injection attempts, where hyphens are used to comment out parts of a query. - Contextual Enrichment: The AI system then enriched this technical pattern with business context. It discovered that the hyphens were originating from a metapixel associated with an ad campaign, which generated campaign IDs that legitimately included two hyphens.
- Automated Remediation: With this complete context, the framework could automatically generate a highly specific remediation: an exclusion rule for cookies containing two hyphens, but only for the specific booking and ticketing API endpoints, thereby preventing legitimate ad campaign traffic from being blocked without compromising overall security. This entire process, from detection to fine-tuned remediation, could be completed in minutes, a drastic improvement over the hours or days it would take with traditional, manual WAF management.
A second example focused on visualizing and securing the ingress to application traffic flow:
- Asset and Configuration Assessment: The framework correlates configurations across different layers—CDN (e.g., Cloudflare), WAF, and cloud infrastructure (e.g., AWS)—to create a comprehensive map of the traffic path.
- Network Flow Analysis: It analyzes the actual publicly accessible paths from the internet ingress to the application.
- Desired Path vs. Actual Path: By comparing the desired, secure path (CDN -> WAF -> Load Balancer -> Application) with the actual flow, the system can quickly identify misconfigurations. For instance, it might reveal that the load balancer is directly accessible from the internet, effectively bypassing all WAF protection. This critical exposure, often hard to spot across fragmented dashboards, is immediately surfaced in the unified view.
- Optimization: The framework can even suggest ways to rebuild or slim down the path by removing unnecessary components or consolidating configurations, an impossible task without this holistic visibility.
The most compelling proof of concept came from the TikTok case study, demonstrating the framework's effectiveness at an unprecedented scale:
- Multi-CDN Architecture: TikTok operates with multiple CDNs globally to ensure resilience and optimal content delivery to over a billion monthly active users across thousands of domains. This complex setup traditionally required opening four separate CDN consoles to block an attack, involving multiple Terraform files and extensive coordination, taking over 30 minutes.
- AI-Driven Automation: With the unified framework, the AI learned the policy structures of multiple CDNs and could automatically deploy a policy change across all of them in under 3 minutes.
- Reduced Error Rates: The system significantly reduced error rates and false positives by first running policies in log-only mode. A local AI agent would study traffic patterns, identify potential false positives, and eliminate them before enforcing policies in production.
- Advanced Bot Mitigation: The framework enhanced TikTok's ability to limit bot activities by learning new bot patterns and scraping techniques from live traffic, adapting to the dynamic nature of bot attacks.
- Executive Reporting: It also enabled the generation of custom, AI-powered reports for executives, providing tailored financial risk assessments for CFOs and attack pattern insights for CISOs, based on each mitigated attack.
A specific attack timeline for a JS challenge for fingerprint cluster attack at TikTok showcased the framework's speed:
- Detection: An attack pattern was identified in less than a minute.
- Mapping: Within 2 minutes, the system mapped the attack to exact foreign assets, identified potentially vulnerable similar hosts, and infrastructure dependencies.
- Analysis: An AI engine completed the analysis within 5 minutes, proposing mitigation actions like rate limiting or targeted blocks. The policy was tailored instantly.
- Log-Only Deployment: Within 3 minutes, the policy was deployed in log-only mode to study traffic, ensure no false positives, and prevent business impact.
- Enforcement: After 3 minutes of thorough traffic analysis in log-only mode, the policy was enforced across multiple CDNs through a strict change management process.
- Total Resolution Time: The entire process, from detection to full policy enforcement for a complex attack, took only 12 minutes—a remarkable feat in the industry.
These examples underscore that the framework is not merely conceptual but delivers tangible, measurable improvements in security posture, operational efficiency, and business resilience.
Defensive Implications
▶ Watch: Executives' concerns: WAF's business and revenue impact (6:40)
The insights from this talk provide critical guidance for defenders looking to modernize their edge security posture and overcome the limitations of traditional WAFs.
- Adopt a Holistic Edge Security View: Organizations must move beyond thinking of WAFs as isolated components. Instead, a unified edge security layer that integrates CDNs, WAFs, load balancers, and compute infrastructure is essential. This holistic perspective enables defenders to see the entire traffic flow and identify bypasses or misconfigurations that are invisible in fragmented systems.
- Embrace AI as a Force Multiplier: AI is no longer optional; it's a necessity for scale and efficacy. Defenders should leverage AI for:
- Contextual Enrichment: Integrating business logic, application context, and external threat intelligence with raw WAF logs to prioritize threats based on actual business impact.
- Predictive Analysis: Identifying emerging attack patterns and potential vulnerabilities before they are exploited.
- Automated Remediation: Developing agentic flows that can generate, backtest, and deploy precise mitigation rules rapidly.
- Operational Efficiency: Automating repetitive tasks, reducing manual intervention, and allowing security teams to focus on strategic initiatives rather than log analysis.
- Prioritize Visibility and Inventory: A fundamental step is to gain complete visibility into all assets and configurations across the entire edge. This includes mapping upstream to downstream components, understanding network flows, and continuously assessing what is publicly accessible versus what is intended. "You can't protect what you can't see."
- Implement Safe Automation with Log-Only Mode and Backtesting: To counter the increasing speed of attacks (CVEs now exploited in less than -2.6 days, meaning before they are even public), automation of WAF rule deployment is critical. However, this must be done safely. Defenders should always deploy new rules in log-only mode first, using AI to backtest against past traffic and analyze potential false positives before moving to block mode. This prevents business disruption while enabling rapid response.
- Conduct Continuous WAF Assessment: WAF configurations and rules should not be static or assessed only once or twice a year. They are business-critical assets and require continuous assessment. Defenders should regularly review existing rules, especially those on "skip" or "bypass," to ensure their necessity and effectiveness. This also involves understanding the top blocked routes and whether the traffic breakdown makes sense.
- Address Vendor Lock-in Proactively: Organizations should recognize the challenges of vendor lock-in when dealing with multiple WAF/CDN providers, particularly during mergers, acquisitions, or infrastructure migrations. The unified framework's ability to translate rules between vendors offers a strategic advantage, enabling smoother transitions and preventing security gaps during migration periods.
- Focus on Business Value: CISOs and security leaders should shift their focus to measuring the effectiveness of their edge security in terms of business value. This includes quantifying the reduction in business losses, improving team efficiency without necessarily increasing headcount (through AI), and understanding the financial risk mitigation provided by security controls.
- Educate and Empower Engineers: A critical defensive implication is the need for engineers to deeply understand how WAFs work in the context of their specific business logic and infrastructure. Building internal tooling and expertise around WAF management, even for a single CDN, is a valuable starting point.
By adopting these defensive implications, organizations can transform their WAFs from frustrating, static defenses into dynamic, intelligent, and business-aware edge security layers capable of adapting to the modern threat landscape.
Key Takeaways
- WAFs must evolve from static defenses to living, learning edge security layers. Traditional WAFs are failing due to static rules, lack of context, and operational complexity, necessitating a shift to a unified, intelligent approach.
- Artificial Intelligence (AI) is indispensable for scaling and enhancing WAF effectiveness. AI enables contextual enrichment, predictive analysis, and automated, agentic remediation flows, drastically reducing response times and human intervention.
- A unified framework integrating diverse data sources provides comprehensive visibility and actionable insights. By correlating external intelligence, CDN/WAF configurations, cloud infrastructure, traffic patterns, and security context, organizations can identify and address security gaps across the entire edge.
- Automated, backtested policy deployment in log-only mode is crucial for rapid, safe remediation. This approach, demonstrated by TikTok's 12-minute incident response, minimizes false positives and business impact while countering fast-evolving threats.
- Continuous assessment and automation of WAF rules are non-negotiable. Regularly reviewing and automating rule management, rather than relying on manual, infrequent assessments, is critical to maintaining a robust security posture against increasing attack velocity.
- The new approach tackles vendor lock-in, enabling seamless WAF migrations. By abstracting and translating WAF configurations, the unified framework allows organizations to move between different CDN/WAF providers without extensive manual rewrites, simplifying complex migration projects.
About the Speaker(s)
Roy Weisfeld is the co-founder and CTO at Huskys, an early-stage startup focused on building a new digital ecosystem, specifically an edge security management platform. His extensive background includes many years as a web hacker, network DevOps, and infrastructure expert. Roy is passionate about taking broken platforms and systems and finding ways to fix them rather than always replacing them, a philosophy central to his approach to WAFs.
Surya Pentakota serves as the Enterprise to Edge Security Lead at ByteDance, the parent company of TikTok. In this role, he oversees security engineering, automated controls validation, and independent security testing for one of the world's largest short-form media platforms. Surya brings deep expertise in managing security at a global scale, particularly concerning WAFs, which are crucial for TikTok's edge-centric operations. He has been instrumental in solving the complex challenges associated with WAFs at TikTok, working to introduce a new standard for edge security.
Reviews
Dr. Zero (Offensive Security Researcher) — WEAK
A vendor pitch wearing a case study costume. Roy is selling Huskys, his WAF management startup, and TikTok's Surya provides enterprise credibility cover. The problem statement is real and well-articulated, but everything after that is product marketing dressed as research.
Heather Calloway (CISO) — WEAK
Technically competent practitioners presenting a real operational problem — WAF fragmentation at scale — but the talk never escapes the product pitch frame. The TikTok case study is genuinely interesting, but the solution architecture is Huskys' platform, which means the 'framework' on offer is a vendor demo, not transferable guidance.