Beyond the Hype: Evaluating LLM Integration and Practical Limitations in Security Operation Centers
Elnaz Rabieinejad, Ali Dehghantanha, Fattane Zarrinkalam, and Sarina Dastgerdy
Twenty SOC practitioners, not a vendor deck. LLMs as drafts and leads for logs and first-pass triage. Nobody in the room trusted them for the high-stakes call.
Open PDF↗- Title
- Beyond the Hype: Evaluating LLM Integration and Practical Limitations in Security Operation Centers
- Authors
- Elnaz Rabieinejad, Ali Dehghantanha, Fattane Zarrinkalam, and Sarina Dastgerdy
- Year
- 2026
- Publisher
- arXiv
- Tags
- llm, soc, interviews, hallucination, operations
I keep this one because it interviews the people who would have to live with the output. Not a leaderboard. Not a product pitch. Twenty SOC practitioners across 13 organizations, University of Guelph Cyber Science Lab and College of Engineering, submitted 17 August 2026 as arXiv:2608.17154 (cs.CR). Authors: Elnaz Rabieinejad, Ali Dehghantanha, Fattane Zarrinkalam, and Sarina Dastgerdy. DOI: 10.48550/arXiv.2608.17154. License: CC BY-NC-ND 4.0. The PDF stays on arXiv; I am not hosting a copy.
The method is the point. Semi-structured interviews, June-December 2024, codebook thematic analysis, saturation by interview 15, N=20 to be sure. Front-line analysts, managers, detection engineers, a few researchers who still touch a queue. Mostly Ontario, one California. Self-report, not SIEM telemetry, and they say so. They also stay out of the adversarial-prompt circus: this is about the model being wrong on a Tuesday, not about someone jailbreaking it.
Four places the model actually enters the workflow: log summarization, first-pass triage steps, detection-rule drafts, CTI synthesis. Every one of those has a human gate after the text shows up. The gate is cheap when you can falsify the answer against a raw log line or a linter. It is expensive when the answer is a story.
They refuse to treat "hallucination" as one blob. Seven failure modes, from the transcripts:
- FM1 fabricated facts - IOCs, TTPs, incident details that were never in the evidence.
- FM2 invalid detection logic - invented operators that look copy-pasteable until the SIEM rejects them. Or worse, does not.
- FM3 misleading investigative direction - a plausible next step that steers you off the telemetry.
- FM4 wrong interpretation - FP/FN in the label or the writeup.
- FM5 opaque claims - an answer with no "why," so you rebuild the reasoning by hand.
- FM6 guessing under novelty - zero-day, thin intel, the model fills the gap anyway.
- FM7 verification tax - you saved an hour drafting and spent half of it checking every line.
The quote I will steal: an LLM is an assistant, never the final authority.
The maturity rubric is the part I will actually reuse. L0 is individual intuition ("if it feels off"). L1 is an undocumented never-trust-it culture. L2 is prompt libraries and evidence-only prompting. L3 is schema checks, dual control, audit logs. Their sample lived mostly in L0/L1. That matches every SOC I have watched try this: a chat pane in a side window, a senior who still reads the logs, and no ticket field for "the model said so."
If a vendor tells you the model will close the ticket, send them this paper. Then keep a human on the gate.