Hacker Rock and Roll: Visualizing the 20 Year Evolution of ShmooCon Research
Greg Conti, Danielle Scalera
ShmooCon XX (Final) · Day 2 · Bring It On
Overview
In a poignant tribute to the "last ShmooCon," Greg Conti and Danielle Scalera presented a comprehensive and analytical retrospective on two decades of hacker research showcased at the influential conference. Their talk, "Hacker Rock and Roll: Visualizing the 20 Year Evolution of ShmooCon Research," embarked on an ambitious journey to collect, categorize, and visualize every main stage talk ever given at ShmooCon. The project aimed to not only celebrate the conference's profound impact but also to illuminate the dynamic evolution of the hacker community's innovation, offering unprecedented insights into emerging trends, enduring challenges, and the unique cultural fabric of ShmooCon itself.

Key moments
- 3:20 Welcome and talk's mission to celebrate ShmooCon
- 3:50 Uncovering 20 years of ShmooCon research, visualization inspiration
- 5:45 Speakers Greg Conti and Danielle Scalera introduce themselves
- 7:20 Why analyze ShmooCon's history? Data sharing goals
- 8:50 How technological evolution and events shape conferences
- 9:40 Introducing a timeline of 20 years of tech advances
Hacker Rock and Roll: Visualizing the 20 Year Evolution of ShmooCon Research
Speakers: Greg Conti, Professor; Danielle Scalera, Master's Student
Conference: ShmooCon
YouTube: https://www.youtube.com/watch?v=yIutY_X2FcU
Overview
In a poignant tribute to the "last ShmooCon," Greg Conti and Danielle Scalera presented a comprehensive and analytical retrospective on two decades of hacker research showcased at the influential conference. Their talk, "Hacker Rock and Roll: Visualizing the 20 Year Evolution of ShmooCon Research," embarked on an ambitious journey to collect, categorize, and visualize every main stage talk ever given at ShmooCon. The project aimed to not only celebrate the conference's profound impact but also to illuminate the dynamic evolution of the hacker community's innovation, offering unprecedented insights into emerging trends, enduring challenges, and the unique cultural fabric of ShmooCon itself.
Conti and Scalera meticulously gathered data from 801 talks spanning from 2005 to 2025, confronting the challenge of disparate data sources and the absence of a unified, analyzable format. Their work culminates in a robust taxonomy and a suite of compelling visualizations, including a "Hacker Rock and Roll" art piece inspired by Edward Tufte's iconic data art. This research serves as a vital resource for understanding the historical trajectory of cybersecurity research, providing a framework for future analysis, and underscoring the critical importance of archiving community-driven knowledge. For anyone interested in the history of information security, the dynamics of hacker innovation, or the power of data visualization, this talk offers a rich tapestry of insights.
Background
▶ Watch: Welcome and talk's mission to celebrate ShmooCon (3:20)
The genesis of this ambitious project stemmed from a simple yet profound question: what research and talks were presented at ShmooCon over its two-decade run? The speakers' initial assumption that this data would be readily available in a single, easily analyzable format proved incorrect. They discovered the information was scattered across approximately six different online repositories, necessitating a painstaking process of triangulation and data consolidation. This challenge highlighted a broader issue within the hacker community: the lack of centralized, standardized archiving for its invaluable research output.
The inspiration for visualizing this data came from an art piece by Rebi Garo, featured in one of Edward Tufte's information visualization books. This piece depicted the evolution of rock music genres and market share over time, prompting Conti and Scalera to envision a similar "Hacker Rock and Roll" visualization for ShmooCon. They sought to illustrate which security topics, akin to musical genres, rose and fell in popularity, and how the overall landscape of research evolved.
To contextualize the talks, Danielle Scalera presented a timeline of key technological advances and global events that likely influenced ShmooCon's research focus. This included the creation of Facebook (2004), the first ShmooCon and YouTube (2005), the release of the iPhone (2007) which subsequently led to talks on smartphone security around 2010, the invention of Bitcoin (2008), the founding of US Cyber Command (2010), the disclosure of Operation Aurora (2010), and the significant impact of Snowmageddon on ShmooCon 6. More recently, the COVID-19 pandemic (2020-2021) and the public release of ChatGPT (2022) were identified as major inflection points, demonstrating how current events and technological shifts directly shape the hacker community's research agenda. This dynamic interplay between external factors and research topics underscored the necessity of a structured approach to understanding ShmooCon's contributions.
Key Findings
▶ Watch: Speakers Greg Conti and Danielle Scalera introduce themselves (5:45)
The core achievement of this project was the successful collection and categorization of 801 main stage talks from ShmooCon's 20-year history, including the final 2025 conference. This massive dataset, which they have made publicly available, provides an unprecedented resource for the cybersecurity community.
A primary finding was the immense diversity and breadth of research presented at ShmooCon. The taxonomy they developed ultimately comprised 45 high-level categories, ranging from deep technical topics like "Exploitation," "Reverse Engineering," and "Cryptography" to broader discussions on "Law and Policy," "Community," "Education and Training," "Intelligence," and even "Philosophy" and "Art." This diversity challenges the common perception that security conferences are solely focused on offensive techniques, revealing ShmooCon's unique blend of technical depth, community engagement, and societal reflection.
The project also highlighted the continuous evolution and growth of the research landscape. The speakers initially expected the taxonomy to converge after about 10 years, but it continued to expand in "interesting ways," indicating a relentless pace of innovation and the emergence of entirely new areas of focus within the hacker community. This perpetual expansion demonstrates that the field of cybersecurity is far from static, constantly adapting to new technologies and threats.
Specific insights into the categories included:
- "Exploitation" was consistently the largest category, which the speakers broadly defined as "the hacking of things." This category alone encompassed everything from AI and Cloud to Firmware, Hardware, and even humans, underscoring the pervasive nature of offensive security research.
- "Community" emerged as a strong category, with 46 talks over 20 years, reflecting ShmooCon's deliberate efforts to foster engagement, including "Heidi's Own the Con Talks" and discussions on culture and charity.
- "Defense" and "Detection" were also significant, challenging the notion of an exclusive focus on offense. "Detection," for instance, featured 32 talks, covering topics from the data science of detection to identifying malicious USBs.
- "Law and Policy" stood out with 43 talks, indicating ShmooCon's unique engagement with legal experts and its role as a forum for discussing the implications of technology within a legal framework.
- "Intelligence" was another major category, likely reflecting the background of many attendees and speakers from the intelligence community.
The project also uncovered interesting details about talk titles, from the exceedingly long (e.g., "Hack the Earth Hemisphere: How We Legally Broadcasted Hacker Content to All of North America and Beyond Using End of Life Geosynchronous Satellite" at 152 characters) to the whimsical ("Can a drunk person authenticate using brain waves #notalcoholicsjustresearchers") and the cryptically intriguing ("Defending Against Targeted Attacks Using Duct Tape Popsicle Sticks and Legos"). These titles, while often entertaining, underscored the challenge of accurately categorizing talks solely based on their names, necessitating a deeper dive into abstracts and video content.
Ultimately, the key findings underscore ShmooCon's role as a vibrant crucible of innovation, a platform for diverse research, and a unique reflection of the hacker community's evolving interests and societal impact.
Technical Deep Dive
▶ Watch: Why analyze ShmooCon's history? Data sharing goals (7:20)
The technical undertaking of this project involved several critical steps: data collection, data cleaning, taxonomy development, talk labeling, and visualization.
Data Collection and Cleaning:
The initial hurdle was gathering all talk metadata. The team relied heavily on a triangulation approach, leveraging several key online archives:
- infocon database.org
- archive.org
- citationthink.com
These sources, while invaluable, each had gaps, requiring cross-referencing to compile a complete list of 801 talks. The project focused exclusively on main stage talks, intentionally excluding "fire talks" due to inconsistent availability across years.
Data cleaning was a significant effort. Talk titles often contained special characters, inconsistent formatting, or were intentionally vague or humorous. The team had to standardize titles and parse them effectively. They also noted the future potential to expand the dataset to include fire talks and speaker information, which was beyond the scope of this initial effort due to the project's already considerable size.
Taxonomy Development and Labeling:
A central component was the creation of a "straw man taxonomy" – a hierarchical categorization of all talks. This was a bottom-up, or emergent approach, where the categories and subcategories were derived directly from the content of the talks themselves, rather than imposing a predefined structure. This iterative process ensured 100% coverage of all ShmooCon talks.
To inform their taxonomy, Conti and Scalera consulted existing frameworks and structures from various domains, recognizing that no single one perfectly fit the hacker research landscape:
- Black Hat briefing tracks
- InfoSec certification categories (e.g., (ISC)² CISSP, ACM)
- Military frameworks (intelligence, electronic warfare)
- Industry frameworks (e.g., MITRE ATT&CK, DISARM framework for influence operations)
- The European Commission Cyber Security Taxonomy was identified as the most comprehensive and useful external reference.
A critical design decision was the balance between depth and breadth in the taxonomy. They opted for a balanced approach, leaning towards precision to capture the nuanced nature of hacker research. For instance, if a category like "reverse engineering" grew too large, it was promoted to a high-level category, with sub-branches for "Hardware Reverse Engineering," "Software Reverse Engineering," and "Systems Reverse Engineering." This resulted in the 45 high-level categories mentioned earlier.
Labeling each of the 801 talks was a meticulous process guided by specific heuristics:
- Review the title: While often misleading, it was the first step.
- Read the abstract: This was the minimum requirement for applying a label.
- Review the video or other materials: If the abstract was insufficient, they would watch the talk or consult supplementary content.
- Read news stories: In rare cases where videos were unavailable, external news coverage was consulted.
- 51% Rule: A talk was labeled based on the primary content (51% or more). For talks covering multiple areas (e.g., Dan Kaminsky's talks), a more general, higher-level category was applied if a single dominant theme couldn't be identified.
- Intent mattered: Understanding whether a talk was defense-oriented, exploitation-focused, or simply an explainer of new technology was crucial for correct categorization.
- Expanding the taxonomy: If a talk presented a novel idea not covered by the existing taxonomy, the taxonomy itself was expanded to accommodate it, demonstrating its organic growth.
Visualization using Graphviz and Python:
To visualize the complex, hierarchical taxonomy, the team employed Graphviz, an open-source graph visualization software. Graphviz uses the DOT language (Graph Description Language) to define graphs, nodes, and edges. A key challenge with DOT was its strictness regarding node names, requiring unique, non-alphanumeric identifiers, which often meant concatenating terms (e.g., "DeceptionCaseStudy") to ensure uniqueness and prevent errors.
To automate the generation of Graphviz code from their spreadsheet of labeled talks, they developed a Python parser. This custom script, approximately 350 lines long, takes a tab-separated value (TSV) file as input (their master spreadsheet) and outputs valid DOT language code. This automated approach was essential given the scale of 801 talks and the intricate hierarchical structure of the taxonomy. The Python script also generated statistics about the talk distribution.
The visualizations produced included:
- A comprehensive, multi-layered graph of the entire taxonomy, demonstrating the depth and breadth of the research.
- Smaller, focused graphs for individual high-level categories (e.g., "Community," "Education and Training," "Detection," "Law and Policy," "Exploitation," "Cryptography"), illustrating their internal sub-structures.
- Time-series visualizations, like the evolution of "Deception" talks year-by-year, showing how topics emerge, deepen, and diversify over time.
This technical deep dive reveals the rigorous methodology applied to transform raw, dispersed data into a structured, analyzable, and visually compelling representation of ShmooCon's intellectual output.
Demo / Proof of Concept
▶ Watch: How technological evolution and events shape conferences (8:50)
The "demo" of this talk was the presentation of the various visualizations created from their analyzed data, serving as a powerful proof of concept for their methodology. These visualizations were not merely static charts but dynamic representations designed to convey the evolution and interconnectedness of ShmooCon research.
The primary demonstration was the complete taxonomy graph. While too large to be fully legible on a single slide, it visually represented the hierarchical structure of the 45 high-level categories and their numerous subcategories, showcasing the depth (up to six levels deep in some areas) and breadth of the research landscape. This graph, generated using Graphviz, vividly illustrated how diverse topics like "Exploitation," "Law and Policy," "Defense," and "Community" were organized. The speakers also presented close-ups of specific categories, such as "Community" (46 talks), "Education and Training" (20 talks), "Detection" (32 talks), "Law and Policy" (43 talks), "Exploitation," and "Cryptography," allowing the audience to appreciate the granular detail within each domain. For instance, the "Exploitation" subcategories included "Hardware," "Biotech," "Cameras," "Keyboards," "Jukeboxes," and "Pinball Machines," highlighting the sheer creativity and unexpected targets of hacker research.
A particularly engaging demonstration was the bar chart race visualization. This animated graph depicted the cumulative number of talks in the largest categories over the 20 years of ShmooCon. As the years scrolled by, the bars representing categories like "Exploitation," "Community," "Defense," "Law and Policy," and "Intelligence" grew and shifted in rank, offering a dynamic perspective on the fluctuating prominence of different research areas. This visualization effectively connected the research trends to the historical timeline of technological and societal events discussed earlier, allowing viewers to infer potential causal relationships.
The culmination of their visualization effort was the "Hacker Rock and Roll" art piece. This unique visualization mapped the 20 years of ShmooCon on the horizontal axis and the percentage of talks dedicated to each subject on the vertical axis. Inspired by J.R.R. Tolkien's art and Edward Tufte's data visualization principles, the piece metaphorically depicted the "sea" of early research evolving into "grasslands," "mountains," and "snow-capped peaks" as topics matured and diversified. The high-resolution version of this art piece was made available online, designed to be printed on canvas, transforming complex data into a tangible work of art.
The entire presentation served as a proof of concept for their methodology: that by systematically collecting, categorizing, and applying visualization tools like Graphviz and custom Python parsers, it is possible to extract meaningful insights and create compelling representations of vast, unstructured data sets from the hacker community. The public release of their data, code, and visualizations further empowers others to build upon their work, refining the taxonomy or applying similar methodologies to other conferences.
Defensive Implications
▶ Watch: Introducing a timeline of 20 years of tech advances (9:40)
While the talk primarily focused on meta-analysis and visualization of research trends, its findings hold significant defensive implications for individuals, organizations, and the broader cybersecurity community. Understanding the evolution of hacker research, as illuminated by Conti and Scalera's work, can inform defensive strategies in several key ways:
- Anticipating Emerging Threats: The "Exploitation" category, consistently the largest, reveals the constantly expanding attack surface. Defenders can infer from the long-term trends and the diversity of subcategories (e.g., AI, Cloud, Firmware, Biotech implants, IoT devices like Jukeboxes and Pinball Machines) where attackers are focusing their efforts. This historical perspective allows for a more proactive stance, helping defenders anticipate future attack vectors rather than merely reacting to current exploits. The year-over-year growth in certain exploitation areas, or the emergence of new subcategories, signals areas requiring increased defensive attention and resource allocation.
- Informing Defense and Detection Strategies: The significant presence of "Defense" and "Detection" as major categories demonstrates that the community is actively developing countermeasures. By examining the subcategories within "Detection" (e.g., data science of detection, hardware manipulation detection), defenders can identify cutting-edge techniques and tools being discussed. This can guide the adoption of new defensive technologies, the refinement of existing security controls, and the development of more effective incident response and threat hunting capabilities. Understanding what the community considers innovative in defense helps practitioners stay ahead.
- Understanding the Legal and Policy Landscape: The robust "Law and Policy" category at ShmooCon highlights the increasing interplay between technology, law, and ethics. For defenders, this means understanding the legal ramifications of their actions, the evolving regulatory environment, and how new laws (e.g., data privacy, cybercrime legislation) impact their operations. Awareness of these discussions can help organizations ensure compliance, mitigate legal risks, and contribute to policy-making efforts.
- Community Engagement and Knowledge Sharing: The strong "Community" and "Education and Training" categories underscore the value of collective intelligence. Defenders can leverage the insights from this historical analysis to identify critical areas for skill development, training initiatives, and fostering internal knowledge sharing. Engaging with the broader hacker community, as ShmooCon exemplifies, provides invaluable insights into practical attack and defense methodologies that might not yet be published in academic or industry reports.
- Strategic Resource Allocation: The bar chart race and other time-series visualizations can inform strategic resource allocation. If certain attack techniques or technologies are consistently gaining prominence, organizations can prioritize investments in security solutions, training, and personnel for those specific areas. Conversely, understanding when certain topics "fall off" can help deprioritize less relevant or already-solved problems.
- The Value of Depth: The talk emphasized the value of "depth" in taxonomy and analysis. For defenders, this translates to the importance of deep, specialized knowledge in specific domains. A superficial understanding of security threats is often insufficient; true resilience comes from expertise in specific protocols, systems, or attack surfaces, mirroring the detailed subcategories found in the ShmooCon taxonomy.
In essence, Conti and Scalera's work provides a macroscopic view of the cybersecurity battleground's evolution. By studying this historical data, defenders gain a unique vantage point to understand the past, interpret the present, and better prepare for the future of information security.
Key Takeaways
- The immense value of archiving hacker community research: The project underscored the critical need for systematic documentation and preservation of talks, abstracts, and videos from conferences like ShmooCon, serving as a historical record of innovation.
- ShmooCon's unique breadth and depth of research: Over 801 talks across 20 years revealed 45 high-level categories, encompassing everything from chip-level exploitation to discussions on philosophy, art, law, and community, challenging the perception of a purely offense-focused conference.
- The dynamic and evolving nature of cybersecurity topics: The taxonomy continuously expanded over two decades, demonstrating the relentless pace of innovation and the emergence of new research areas, rather than a convergence of topics.
- ShmooCon's balanced focus: The conference exhibits a strong emphasis on "Community," "Defense," "Detection," and "Law and Policy," alongside "Exploitation," reflecting a holistic approach to information security beyond purely offensive techniques.
- Power of data visualization for complex datasets: Utilizing tools like Graphviz and custom Python parsers enabled the transformation of disparate talk data into compelling and insightful visualizations, such as the bar chart race and the "Hacker Rock and Roll" art piece.
- A call for continued community effort: The project provides a foundational dataset, code, and taxonomy, encouraging future community research to refine, expand, and apply these methodologies to other conferences or to automate the labeling process, potentially with AI in the future.
About the Speaker(s)
Greg Conti is an experienced figure in the cybersecurity and academic communities. He formerly taught at West Point and has worked with prominent government organizations such as the NSA and US Cyber Command. With a long history in the hacker community, he has been attending ShmooCon since ShmooCon 3 and is a well-known trainer for both Black Hat and Defcon. Currently, he works at Cydan, a cybersecurity research company he co-founded, which focuses on training, research, and consulting.
Danielle Scalera is an emerging talent in the field of cybersecurity. She is currently a first-year Master's student at the New Jersey Institute of Technology (NJIT), having earned her Bachelor of Science degree from Mars College (now Mars University). Danielle works as a part-time security researcher with Cydan, where she assisted with Black Hat training in 2022. This talk marked her debut as a speaker at ShmooCon.
Reviews
Dr. Zero (Offensive Security Researcher) — MUST SEE
This talk, while not a zero-day or live exploit, is a foundational piece of work for understanding the history and evolution of hacker research. The speakers undertook a monumental task, meticulously collecting, cleaning, and categorizing 20 years of ShmooCon talks. The resulting custom taxonomy and visualizations offer an unparalleled look into the community's intellectual journey, demonstrating genuine skill, persistence, and a novel approach to historical self-reflection. It's a deep dive into the very fabric of our shared technical heritage.
Heather Calloway (CISO) — STRONG ACCEPT
This retrospective on ShmooCon's 20-year research evolution provides a critical, data-driven lens for security leaders to understand the shifting landscape of threats and defenses. By meticulously categorizing over 800 talks, Conti and Scalera offer a unique historical perspective on where hacker innovation has focused, highlighting the enduring diversity and continuous evolution of cybersecurity challenges. While not a tactical guide, this work is invaluable for informing strategic risk discussions, justifying resource allocation, and anticipating future attack vectors at the institutional level. It underscores the importance of community knowledge archiving and offers a practical…