socbench· LLM triage benchmark
Measurement report · v1.0 · May 2026
socbench

Can an LLM doan L1 analyst's job?

We measured eight large language models on 600 realistic SIEM alert scenarios spanning twenty categories. Each model made a decision on every scenario; the decisions were compared against ground truth labeled by human analysts. Accuracy, latency and cost must be read together — any one alone misleads.

Most benchmarks report one accuracy number. socbench reports what that accuracy costs — in money, in seconds, and in the alerts a model waves through.

Every model saw the same 600 alerts drawn from live SIEM data across twenty categories, and returned a verdict on each: true positive, false positive, or needs confirmation. Ground truth was labelled by working analysts, not by another model.

600 scenarios · 20 categories · 8 models · ~20,000 evaluations

leaderboard

all 8 models · sorted by strict accuracy
  1. 01
    Gemini 3.1 Pro
    Google
    92%strict
    92%
  2. 02
    Grok 4.3
    xAI
    89%strict
    89%
  3. 03
    GLM 5.1
    Zhipu AI
    86%strict
    86%
  4. 04
    Kimi K2.6
    Moonshot AI
    83%strict
    83%
  5. 05
    GPT 5.5
    OpenAI
    79%strict
    79%
  6. 06
    Qwen 3.6 35B A3BMoEself-host
    Alibaba
    79%strict
    79%
  7. 07
    Claude Opus 4.7thinking max
    Anthropic
    78%strict
    78%
  8. 08
    DeepSeek V4 Pro
    DeepSeek
    69%strict
    69%
bar = strict accuracy, exact-match scoringsee full ↓

Strict scoring demands an exact match. Merging “true positive” with “needs confirmation” — operationally the same call — lifts every model, some far more than others.

Rank here is accuracy alone. Five of these eight are beaten on cost by a model that is both cheaper and more accurate; six are beaten the same way on latency.

01 / Leaderboard

Same exam, two readings

Strict scoring requires an exact match: TP=TP, FP=FP, CCN=CCN. Lenient scoring merges TP and CCN — operationally both point the same way: "don't close it, act on it." The distance between them measures the model's hesitation between a real threat and a case that still needs confirmation; a long line means an indecisive model.

Strict accuracy · eight-model distribution600 scenarios · ~20,000 evaluations
StrictLenient (TP↔CCN merged)

How to read this.These scores are specific to this scenario set; they are not a measure of general model capability. They measure one thing only: whether the model makes the same call an L1 analyst would, matching the human label. You can sort by clicking any column header.

ModelProviderStrictLenientC / P / WLatencyTotal$ / correct

Slow thinking doesn't pay off. It backfires.

The two longest-reasoning configurations — Claude Opus 4.7 (thinking · max) and Kimi K2.6, at up to 104 seconds per alert — land in the bottom half of the ranking. The fastest model, Grok 4.3, is also the second most accurate. On this task, extra thinking time does not buy accuracy.
02 / Cost

The price of accuracy spans a 32× range

Top-left is the corner you want: cheap and accurate. Bottom-right: expensive and weak. The most expensive model costs 32× as much as the cheapest and is only 8 points more accurate — hence the logarithmic horizontal axis.

03 / Latency

Queues don't wait

In a SOC queue, 100 seconds per alert rules out real-time triage. This chart asks whether there's a trade-off between speed and accuracy. The point cloud is shapeless: no trade-off.

04 / Dataset

A distribution with no easy wins

Scenarios are weighted to mirror the composition of real SOC queues. One in five is an ambiguous case that needs customer confirmation; one in eight is an outright false positive — so a model that calls everything "malicious" cannot score well.

VerdictWhat it meansCountShare
TPGenuine malicious activity; critical cases trigger automatic escalation to L2383
CCNAuthorization or scope must be confirmed — ask the customer122
FPFalse positive; the alert can be closed68
IDNot enough information to reach a verdict27
Twenty categories
05 / Method

How this was measured

Envelope

Each scenario was built as an envelope close to what an analyst would see on screen: raw log fields, asset context, and the title of the rule that fired. The model was not asked for free text — it had to pick one of four options.

Ground truth

Labels were assigned by human analysts and were not shown to the models at any stage. Critical TPs were flagged as requiring automatic L2 escalation.

Repetition

Each scenario was run across roughly 30 variants; the reported scores are the aggregate of those runs. A single run is not evidence — only repetition separates model-to-model differences from noise.

Cost

Calculated from provider list prices as of the measurement date. Qwen 3.6 ran on our own hardware; the figure shown is the equivalent API price, not our electricity cost.

SOC Benchmark v1.0 — May 10, 2026

600 scenarios · 8 models · 20 categories · ~20,000 evaluations.

This is an independent measurement study; no model provider sponsored it. Scores are specific to this scenario set and should not be read as a measure of general model capability.