Skip to content
SZ-MCP
Get Support

Alert quality report

The alert quality report shows, rule by rule, whether your alerts are worth getting: how long acknowledgement and recovery took, how many problems cleared by themselves within minutes, and how many told someone for nothing. It lists the noisiest rules first, each with a suggestion. Open it on Monitor › Reports (the Alert quality card, then Run), or ask Claude, which uses alerts.quality. Every member of the organisation can run it. It reads only the alert engine’s own history: there are no controller calls and no AI model.

The last 30 days by default. On the card, Period offers Last 7 days, Last 30 days, Last 90 days or Dates… (a From and To). A period can be at most 400 days. Rule narrows it to one rule.

The report covers the HARD problem episodes that began in the period. SOFT problems that never became HARD are not episodes. Flapping starts are kept for 30 days, so in a longer period only the last 30 days of flapping are counted.

ColumnMeaning
ProblemsHARD episodes that began in the period
OpenEpisodes not yet recovered
Never ackedThe share of episodes nobody acknowledged
To ack (median)Median time to acknowledge, over the acknowledged ones (the mean is in the CSV)
To recover (median)Median time to recover, over the closed ones
Short-livedEpisodes that cleared within 10 minutes with nobody acknowledging them
FlappingHow many times the rule started flapping
NoiseEpisodes that sent a PROBLEM notification, were never acknowledged, and closed within 30 minutes, as a share of all episodes
Most problems onThe device with the most episodes, linked to its post-mortem
SuggestionWhat to change (below)

An episode held back by downtime or by an upstream outage told nobody, so it is never counted as noise. Claude can change the two thresholds: shortMinutes (1–240, default 10) and actionableMinutes (1–1440, default 30).

A summary line above the table totals the problems, acknowledgements, mean times, overall noise and the number of noisy rules. Rules with no problems in the period are listed underneath as “No problems: …”. A deleted rule shows as (deleted) and a disabled one as (off).

The verdicts are checked in this order, and the first one that matches wins:

VerdictWhen
TOO FEWFewer than 3 episodes, too few to judge
NOISYNoise is 50% or more
IGNORED80% or more of episodes were never acknowledged, and the rule did notify someone
SLOW ACKHalf the acknowledgements took 60 minutes or more
OKNone of the above

Each suggestion is a fixed rule of thumb, not AI-written:

  • Noisy threshold or nodata rule with many blips (30% or more short-lived). It suggests raising maxCheckAttempts by 2 (at most 10), so a problem must last longer before anyone is told. It includes a ready backtest over the last 7 days, which Claude can run before you change anything.
  • Noisy event rule. It suggests narrowing the match (codes, severities), or lowering the state to WARNING.
  • Other noisy rules. It suggests tightening the threshold or the window.
  • Ignored. It suggests lowering the state, turning notification off, or routing the rule to an escalation level whose people will act on it.
  • Slow ack. It suggests checking who is told (contacts, notification periods) and adding an escalation level.
  • One device behind half the episodes. It suggests looking at that device, or putting it in downtime while it is fixed, before changing the rule.
  • Flapping twice or more. It suggests a smoothed value (avg_over_time over 10–15 minutes), or warn and crit further apart.

Yes: after a run, CSV downloads one row per rule and Print / PDF prints the report. The CSV adds the mean times, the instance counts and the suggestions. The same per-rule table is alert-quality.csv in the audit evidence pack.

The quality, digest and post-mortem reports and MOP documents share a limit of 20 reports a minute per organisation. Past it, the reply is 429 “at most 20 reports a minute”.

MessageCause
at most 400 daysThe period is longer than 400 days
the period must end after it starts, and start before nowFrom is after To, or in the future
Refused by the call’s argument checkClaude gave shortMinutes outside 1–240 or actionableMinutes outside 1–1440 (the dashboard card doesn’t set them)
rule too longA rule name over 63 characters