Alessandro AleddaInsider Risk Advisory

Contact me

ArticlesInsider Risk KPIs: What the Numbers Reward

The key performance indicators of an insider risk management program are usually the figures at hand rather than the ones that inform a decision, and several of them, read alone, reward the conduct the program exists to prevent.

6 OCTOBER 2026  |  9 MIN READ

At some point an insider risk management program starts producing numbers, and sooner or later it has to report them to its sponsors as evidence of its effectiveness. Whether the program began with its mandate or, as often happens, with a set of policies put into production, by the end of the first quarter someone has built a slide with the usual numbers: total alerts, total false positives, cases opened and closed. These figures are available, defensible, and exact to the unit, which is why they end up in front of the sponsor. Taken alone, they say little about risk, and several of them reward the conduct the program exists to prevent.

I have run the function and I have built those slides. The problem is seldom bad faith. Each of the metrics a traditional security operations center (SOC) relies on, carried over to an insider threat function, brings with it a distortion that is known, visible in practice, and usually ignored, because the metric is easy to produce and the distortion is awkward to explain. This article lists some of the distortions, states the pairing that corrects each one and the comparison each figure is judged against, and sets out what should reach the people who decide on the program.

The usual figures#

The number of alerts is the first figure a program has and usually the last it should show. An alert count measures the configuration of the detection platform more than the risk in the population. A rise can mean better coverage or worse tuning; a fall can mean that a threshold was raised, that a policy was retired, or that a data source stopped feeding, and the count alone will usually say nothing about which. Read as success, it rewards the analyst who leaves noisy rules in place; read as failure, it rewards the one who quietly narrows them.

Insider detection adds a reason of its own. Many insider risk engines include adaptive components, whose view of what is normal for the actions and conditions they monitor is first established and then updated, so the count moves while the engine learns, particularly early in deployment. A count driven by deterministic conditions is steadier, but it remains a function of how the policies were designed and where the thresholds were set, not a direct measure of the risk in the population. Early in a program, then, the alert count describes the engine’s learning more than the population it watches.

The number of confirmed incidents is worse, because it looks like an outcome. Read as success, it rewards escalation: a case that could have been closed as a benign positive at triage is escalated one step further, where the metric records it. Read as failure, it rewards concealment, and a function whose confirmed incidents are compared with a target soon learns which outcome code keeps the quarter quiet.

Time to close rewards premature closure. A triage analyst measured on it learns to close at the first reading, and the cases that return are counted as new cases, which keeps the average down and the exposure where it was. The false positive rate, taken alone, rewards narrow detection: the fastest way to lower it is to detect less, and a false positive rate of zero can be evidence of precision or of a rule tuned until it fires on nothing; the rate alone cannot tell which. The absence of incidents can read as safety what may instead be a lack of visibility, and the figure alone cannot distinguish the two.

Metrics per person or per small unit are a different kind of risk. They are accurate and often requested, and used to rank employees they turn operational evidence into an evaluation of individual conduct, which may exceed the purpose and the governance basis under which the evidence was collected. A ranking of employees by alert count is a list of the people whose roles touch the most data, read as a list of suspects. A ranking of small units is the same list, aggregated one step.

The pairings#

The correction is usually a second metric rather than a better one, read beside the first and chosen because it exposes the distortion of the first. The false positive rate of a rule is read with the rule’s unique signal catches, the incidents that this rule alone detected: a rule with a low false positive rate and no unique catches is either a rule that corroborates others or a rule that has been tuned into silence, and the pair, with the share of the rule’s alerts that arrive together with another rule’s, tells which. Time to close is read with the rate of reversals, the cases reopened or reclassified after closure; an improving time to close beside a rising reversal rate is a triage queue being cleared, which is a different thing from a queue being worked.

The alert count is read with the changes in coverage and configuration over the same period, which say why the volume moved, and with the yield of automated closure against the manual closures that give the same reason, which says what the volume costs. Confirmed incidents are read with reversals and with the themes of benign positive closures, because the benign positives are where the awareness work is and the reversals are where the escalation pressure shows. The times to first-line and second-line feedback, the figures that measure the other functions, are read with the reversals that follow the feedback and with the cases pending by function, since a fast answer of low content improves the first number and worsens the second. The absence of incidents is read with coverage: the scenarios validated, the vectors verified as visible, and the gaps declared beside them.

MetricWhat it rewards, read aloneRead with
Number of alertsNoisy rules kept, or quiet narrowingChanges in coverage and configuration; yield of automated closure against manual closures with the same reason
Confirmed incidentsEscalation as success, concealment as failureReversals; themes of benign positive closures
Time to closePremature closure at triageRate of reversals
False positive rate per ruleNarrow detectionUnique signal catches of the same rule; the share of its alerts that coincide with another rule’s
Time to feedback from other functionsFast answers of low contentReversals after feedback; cases pending by function
Absence of incidentsBlind spots read as safetyScenarios validated; vectors verified as visible; declared gaps
Figures per person or per small unitIndividual conduct turned into a rankingSet aside

The baseline#

A metric also needs something to be compared with, and here the common practice is to look outside: a peer, a sector average, a level on a maturity scale. My position is that the program is compared with its own earlier state. Before any change, a baseline is taken from the program’s own records, with the source and the period of each figure, and every later change is judged against the baseline taken before it. A tuning change is judged on the false positive rate and the unique catches of the rule before and after; a staffing decision on the times and the pending cases before and after; a new data source on the vectors visible before and after.

A baseline is comparable only if the exposure behind each figure is comparable too. A count without its denominator can move because the monitored population, the coverage, or the observable activity changed while the behavior did not. Where that exposure can change materially, it is recorded with the metric and carried across the comparison: the monitored population, the observable activity, or another denominator that reflects the opportunity for the measured event to occur.

The comparison also holds only if each figure keeps its definition and its source across the change. If alerts are grouped into cases after the baseline is taken, the alert count falls and the program has not changed; if a figure is rebuilt each quarter from whatever the platform exports that week, the comparison measures the export. Where the definition of a figure is not written down, the quarterly review becomes an argument about what the figure counts, and a figure with no recorded source is estimated again each time it is asked for.

I do not use maturity models, and I do not accept cross-organization maturity scores as measures of program performance. Two programs with the same alert count have different populations, different data sources, different thresholds, different legal regimes, and different definitions of what an alert is, and a benchmark that compares the two compares the definitions. A maturity assessment describes the capabilities a program has and how consistently it runs them; it does not measure how the program performs against its own population, which is the question the sponsor is paying to have answered. The program’s own baseline is the comparison that bears that weight.

The sponsor’s figures#

The hard problem is selection. The function usually has dozens of figures, and the executive sponsor and the governance body should see a few, chosen so that what they read says what the program does and nothing it does not. The test for each candidate is the decision it informs. The governance body decides on the roadmap, on lawfulness, and on the residual risk it accepts; it reads coverage, effectiveness, the state of the data protection record, and the progress of the roadmap. Executive management decides on resources, on priorities, and on the cooperation it will require from the business; it reads behavioral trends, exposure by aggregate and population, the capacity gap, and the decisions requested. A figure that informs none of those decisions stays on the operational dashboard, where the queue, the service levels, and the rules under tuning belong, and where the decisions are the function’s own.

Figures shown to the business are aggregated to a level at which no person is reasonably identifiable in the reporting context, and they state the behavior and the exposure of the unit, so that each unit sees its own risk. No metric ranks persons, or units small enough to identify a person. This is partly a matter of lawfulness and partly a matter of what the figures are for: a business unit that sees its own rate of self-forwarding beside the program’s trend will usually act on it, and a business unit that sees itself ranked will usually argue with the ranking.

Each figure travels to the sponsor with its reading attached. “Confirmed incidents rose from four to nine” is a number; “confirmed incidents rose from four to nine, of which five came from a scenario validated this quarter, with reversals unchanged” is a finding. The reading is the part the sponsor cannot produce alone, and it is the part that keeps the number from being read in the way that rewards the wrong conduct. Someone has to write it: each figure has an owner, who keeps its definition, answers for its reading, and says when it has stopped meaning what it meant. A figure nobody owns keeps its name on the slide while what it counts drifts with every change to the rules and the data sources.

The rule#

The rule is short, and each of its parts answers one of the failures described above. Every metric has a definition, a source, an owner, an exposure base where one is required, and a decision it informs. A metric with no definition is argued about; one with no source is estimated; one with no owner drifts; one with no exposure base cannot distinguish behavior from scope; one with no decision is removed, however exact it is and however long it has been on the slide. Each metric is read with at least one other that exposes its distortion, and each is read against the baseline taken before the change it is meant to judge.

How many of the figures on your dashboard would change a decision if they moved?