Greetings & Acknowledgements
ℹ️ Heads Up: I write specifically keeping editing to a minimum. There are likely spelling, grammatical mistakes. I prefer to keep my thoughts as close to raw as possible, and sometimes that means going off on tangents. Thank you for reading.
Hi folks,
Before we begin, I want to take a moment to extend a special thanks to my SANS Research Advisor, Lenny Zeltser, as well as my student advisor Betty Deeb w/ SANS. Second, I want to extend another massive thank you to the 5 participants who assisted me in collecting the data and providing their insights for our human baseline scores.
TL;DR – What was the research about?
I have been working in infosec, specifically DFIR for about 16 years. Over that time, I have witnessed and have been directly impacted by long hours, stressful engagements, over-zealous clients, and challenging leadership. All of which – led to burnout. This led me to wanting to explore solutions for how we can mitigate the threat of burnout for incident responders and general security operators.
It dawned on me one day while I was knee-deep in case work at a previous employer. I had 8 cases open, 2 of which were declared, active, ongoing security incidents that I was ICing (Incident Commanding). The other 6 cases were bug bounty reports, regulatory work etc. Not as in-depth as a full-fledged security incident, but still required some of my time to manage on a day to day basis.
As I was balancing these cases, my manager informed me that our new CISO was explicitly interested in knowing what we were doing about this When Your Scanner Becomes the Weapon: From Trivy to LiteLLM. Thus, I had to stop everything I was doing, review the blog, check intelligence sources, review internal code bases, check to see if artifacts had executed across the network, etc. While important, it raised the question for me:
If I am struggling with managing this case load, at a super-advanced, AI research company, how are others managing? Do solutions exist? What does the data say?
It just so happened this was the exact inspiration I needed to complete my thesis and begin my work for my master’s degree, I had finally found a topic I enjoyed, cared about, and was directly impacted by. I wanted to see what we can do to improve work-life for security operators utilizing the latest technology.
Research Concerns
As I began my literature review, it became very, very, very, clear to me that there was simply not enough actual, hard data on burnout for cyber security workers. One article did stand out to me however, and I encourage anyone reading this to spend the time to read the paper:
A Survey-Based Quantitative Analysis of Stress Factors and Their Impacts Among Cybersecurity Professionals – (Sunil Arora, John D. Hastings, 2024)
Excerpt: This study investigates the prevalence and underlying causes of work-related stress and burnout among cybersecurity professionals using a quantitative survey approach guided by the Job Demands-Resources model. Analysis of responses from 50 cybersecurity practitioners reveals an alarming reality: 44% report experiencing severe work-related stress and burnout, while an additional 28% are uncertain about their condition.
50 practitioners, very impressive, but anyone who is reading this, likely works in the industry and knows that burnout is prevalent. I was shocked to uncover that there really wasn’t a lot of data out there from researchers or industry veterans given that it seems to be unceremoniously accepted that working in InfoSec means it comes with a cost to mental health. Thus, I believe that AI tools can help operators to manage their cognitive load, it did for me.
⚠️If you’re reading this, and have been burnt out in the past, think about recording your experiences or expanding research like mine, or conducting your own!
AI As Assistance Not Decision Making
Attackers, adversaries, APTs, etc. They are using AI tools. Bug bounty, vuln researchers, they are using AI tools as well. Defenders, responders, operators, need to think about how we can also leverage these technologies as cognitive aids. What do I mean by that?
Simple. AI performs triage, collection and presentation of data to aid the operator, but the operator maintains judicious control over the case itself.
AI, LLMs, should never, ever, ever, be in the ‘decision-making’ chair for an active ongoing security incident. There is too much liability and risk trusting AI tools to proactively manage and make decisions on engagements that could very well impact the longevity of an organization.
Overall
I started using AI tools (with minimal configuration) to ingest, analyze, and provide cited sources for complex vulnerability disclosures. The goal being to quickly review new threat data, so that I could then prescribe actions based on the risks that the new vulnerability introduced, while simultaneously ensuring my case-load did not falter.
In my opinion, responders should consider leveraging modern AI tools to assist specifically in triage, and tuning data pipelines. Use the tools, specifically to drive automation, but to act as a triage agent to analyze, and relay findings to humans who are managing multiple cases and need assistance to manage the workload.
The deluge of data, cases, vulnerabilities, incidents, will not cease, we should consider how we can use these emerging technologies to improve our overall response posture from organization to organization.
Summary of findings:
In my paper, I hypothesize that the use of LLMs can act as a cognitive aid for incident responders who are currently active working in security operations and managing/juggling multiple cases. These are the high-level findings:
-
Experiment: Across 27 experiments — three models, three prompt levels, three reports — 16 cleared the 80% Human-Aligned Mitigation Score. HAMS measures a model’s five recommendations against a gold standard built from 75 recommendations submitted by five tenured practitioners, clustered into eight control themes, with the top five per report treated as consensus. Hitting 4 of those 5 is a pass. So free-tier models can provide initial triage; they just don’t do it reliably, and require specific prompting conditions.
-
Initial Findings: The prompt ’level’ moved the numbers more than the specific model did. Level 1 (zero-shot) failed on the narrative report across all three models. Level 2 forced each recommendation to cite the specific telemetry or IoC backing it, and Level 3 put the model in a senior SOC analyst persona — those two constraints recovered Report 3 for Gemini and ChatGPT. Gemini 3.5 Flash at Level 3 was the only configuration that passed all three reports. Claude Sonnet 4.5 tied Gemini at 6 of 9 overall and was the steadiest across prompt levels — its only score below 3/5 was the zero-shot run on Report 3, which every model failed.
-
Where AI Models Help: telemetry-heavy DFIR post-mortems. On the Bissa Scanner report, 8 of 9 model/prompt combos hit 4/5 gold-standard themes — collapsing a 15 minute median human triage into under a minute.
-
Where Models Failed: Strategic narratives with no explicit indicators caused every single model to have challenges in their triage. Every model failed “Beyond the Battlefield” at zero-shot, ChatGPT never passed the LiteLLM report at any level, and its single worst score in the whole matrix (1/5) came at Level 3. Persona prompting cannot manufacture telemetry the document doesn’t contain.
- 📄 Full paper: SANS Technology Institute (published 20 Jul 2026)
- 📊 Raw data: github.com/0xpsilocyber/LLM-IR-Stress-Research (CC BY 4.0)
Music that inspired this post :)
![]() |
![]() |
![]() |
|---|---|---|
| Storm | Make Water | When We Were Young |
| Night Tapes | Pearly Drops | Architects |


