Reading Room/beyond-the-hype-llm-soc
PDF2026arXiv

Beyond the Hype: Evaluating LLM Integration and Practical Limitations in Security Operation Centers

Elnaz Rabieinejad, Ali Dehghantanha, Fattane Zarrinkalam, and Sarina Dastgerdy

Twenty SOC practitioners, not a vendor deck. LLMs as drafts and leads for logs and first-pass triage. Nobody in the room trusted them for the high-stakes call.

Open PDF↗
Title
Beyond the Hype: Evaluating LLM Integration and Practical Limitations in Security Operation Centers
Authors
Elnaz Rabieinejad, Ali Dehghantanha, Fattane Zarrinkalam, and Sarina Dastgerdy
Year
2026
Publisher
arXiv
Tags
llm, soc, interviews, hallucination, operations

I keep this one because it interviews the people who would have to live with the output. Not a leaderboard. Not a product pitch. Twenty SOC practitioners across 13 organizations, University of Guelph Cyber Science Lab and College of Engineering, submitted 17 August 2026 as arXiv:2608.17154 (cs.CR). Authors: Elnaz Rabieinejad, Ali Dehghantanha, Fattane Zarrinkalam, and Sarina Dastgerdy. DOI: 10.48550/arXiv.2608.17154. License: CC BY-NC-ND 4.0. The PDF stays on arXiv; I am not hosting a copy.

The method is the point. Semi-structured interviews, June-December 2024, codebook thematic analysis, saturation by interview 15, N=20 to be sure. Front-line analysts, managers, detection engineers, a few researchers who still touch a queue. Mostly Ontario, one California. Self-report, not SIEM telemetry, and they say so. They also stay out of the adversarial-prompt circus: this is about the model being wrong on a Tuesday, not about someone jailbreaking it.

Four places the model actually enters the workflow: log summarization, first-pass triage steps, detection-rule drafts, CTI synthesis. Every one of those has a human gate after the text shows up. The gate is cheap when you can falsify the answer against a raw log line or a linter. It is expensive when the answer is a story.

They refuse to treat "hallucination" as one blob. Seven failure modes, from the transcripts:

  • FM1 fabricated facts - IOCs, TTPs, incident details that were never in the evidence.
  • FM2 invalid detection logic - invented operators that look copy-pasteable until the SIEM rejects them. Or worse, does not.
  • FM3 misleading investigative direction - a plausible next step that steers you off the telemetry.
  • FM4 wrong interpretation - FP/FN in the label or the writeup.
  • FM5 opaque claims - an answer with no "why," so you rebuild the reasoning by hand.
  • FM6 guessing under novelty - zero-day, thin intel, the model fills the gap anyway.
  • FM7 verification tax - you saved an hour drafting and spent half of it checking every line.

The quote I will steal: an LLM is an assistant, never the final authority.

The maturity rubric is the part I will actually reuse. L0 is individual intuition ("if it feels off"). L1 is an undocumented never-trust-it culture. L2 is prompt libraries and evidence-only prompting. L3 is schema checks, dual control, audit logs. Their sample lived mostly in L0/L1. That matches every SOC I have watched try this: a chat pane in a side window, a senior who still reads the logs, and no ticket field for "the model said so."

If a vendor tells you the model will close the ticket, send them this paper. Then keep a human on the gate.

Beyond the Hype: Evaluating LLM Integration and Practical Limitations in Security Operation Centers · Abraxas Labs