
TALK: Break it, Catch it, Harden it: AI Agents Running a Full-Cycle Detection Pipeline
Benoit Denis, CSE, Ayesha Bashir, CSE
ABSTRACT
The CCCS has built a set of LLM-driven agents that streamline the entire loop from detection engineering to triage, enabling analysts to do more. Our pipeline uses red-team and blue-team agents in an automated purple loop: initial detections are authored, evasion variants are generated and detonated in an isolated lab by the red-team agents, telemetry is verified using open-source monitoring and logging tools by the blue agents, and the analytic is hardened. This process repeats until we get to a behavioral-tier detection.
A dedicated triage agent autonomously cross-references live endpoint telemetry, analytic detection documentation, enriched process trees, and environmental baselines while conducting parallel threat-hunting before an analyst ever engages the alert. The agent then generates structured feedback that informs detection logic tuning and analytic threshold adjustments, closing the loop between triage outcomes and detection engineering.
This talk covers architecture decisions, failure modes we encountered, and what actually moved the needle for our analysts versus what merely looked impressive in a demo.
Description
Detection engineering sits at the intersection of prevention and response, building the analytics that decide whether an alert fires or an attack goes unnoticed. Our team is primarily developers with minimal IT infrastructure experience, which means standing up realistic test environments to validate detections against live attack techniques has historically been slow, manual, and painful. Meanwhile, our triage team is drowning in alert volume, spending time on repetitive classification work before ever reaching the alerts that matter.
These two problems, slow detection validation and unsustainable triage volume, are why we built an agentic workflow.
We will start with a brief overview of detection engineering and the challenges of building preventive analytics without dedicated lab infrastructure. From there, we introduce our purple-team lab: an isolated environment where AI agents automate the red/blue cycle. They author detections, generate evasion variants, detonate proof-of-concepts, verify telemetry, and harden analytics until they anchor on behavioral patterns rather than brittle signatures. The same infrastructure doubles as an investigative tool during real compromises, scouting across all telemetry tables to rapidly build a broad picture of activity.
On the triage side, a dedicated agent cross-references live telemetry, detection documentation, process trees, and environmental baselines while conducting parallel threat-hunting and producing structured reports before an analyst ever engages the alert. Each verdict feeds directly back into the detection engineering pipeline: rule enhancements for true positives, tuning adjustments for false positives, all gated on analyst approval while closing the loop in both directions.
This talk walks through the engineering journey that got us here: the architecture decisions, the iterations that failed, and the ones that stuck. We will compare results before and after, be transparent about where the model struggles and still requires human judgment, and show the actual analytics and triage reports the agents produce. We will close with a detailed look at both agents' core components and what we plan to add next.