Alert quality report
The alert quality report shows, rule by rule, whether your alerts are worth
getting: how long acknowledgement and recovery took, how many problems
cleared by themselves within minutes, and how many told someone for nothing. It
lists the noisiest rules first, each with a suggestion. Open it on Monitor ›
Reports (the Alert quality card, then Run), or ask Claude, which uses
alerts.quality. Every member of the organisation can run it. It reads only the
alert engine’s own history: there are no controller calls and no AI model.
What period does it cover?
Section titled “What period does it cover?”The last 30 days by default. On the card, Period offers Last 7 days, Last 30 days, Last 90 days or Dates… (a From and To). A period can be at most 400 days. Rule narrows it to one rule.
The report covers the HARD problem episodes that began in the period. SOFT problems that never became HARD are not episodes. Flapping starts are kept for 30 days, so in a longer period only the last 30 days of flapping are counted.
What does it measure for each rule?
Section titled “What does it measure for each rule?”| Column | Meaning |
|---|---|
| Problems | HARD episodes that began in the period |
| Open | Episodes not yet recovered |
| Never acked | The share of episodes nobody acknowledged |
| To ack (median) | Median time to acknowledge, over the acknowledged ones (the mean is in the CSV) |
| To recover (median) | Median time to recover, over the closed ones |
| Short-lived | Episodes that cleared within 10 minutes with nobody acknowledging them |
| Flapping | How many times the rule started flapping |
| Noise | Episodes that sent a PROBLEM notification, were never acknowledged, and closed within 30 minutes, as a share of all episodes |
| Most problems on | The device with the most episodes, linked to its post-mortem |
| Suggestion | What to change (below) |
An episode held back by downtime or by an upstream outage told nobody, so it is
never counted as noise. Claude can change the two thresholds:
shortMinutes (1–240, default 10) and actionableMinutes (1–1440, default 30).
A summary line above the table totals the problems, acknowledgements, mean times, overall noise and the number of noisy rules. Rules with no problems in the period are listed underneath as “No problems: …”. A deleted rule shows as (deleted) and a disabled one as (off).
What do the verdicts mean?
Section titled “What do the verdicts mean?”The verdicts are checked in this order, and the first one that matches wins:
| Verdict | When |
|---|---|
| TOO FEW | Fewer than 3 episodes, too few to judge |
| NOISY | Noise is 50% or more |
| IGNORED | 80% or more of episodes were never acknowledged, and the rule did notify someone |
| SLOW ACK | Half the acknowledgements took 60 minutes or more |
| OK | None of the above |
What does it suggest?
Section titled “What does it suggest?”Each suggestion is a fixed rule of thumb, not AI-written:
- Noisy threshold or nodata rule with many blips (30% or more short-lived).
It suggests raising
maxCheckAttemptsby 2 (at most 10), so a problem must last longer before anyone is told. It includes a ready backtest over the last 7 days, which Claude can run before you change anything. - Noisy event rule. It suggests narrowing the match (codes, severities), or lowering the state to WARNING.
- Other noisy rules. It suggests tightening the threshold or the window.
- Ignored. It suggests lowering the state, turning notification off, or routing the rule to an escalation level whose people will act on it.
- Slow ack. It suggests checking who is told (contacts, notification periods) and adding an escalation level.
- One device behind half the episodes. It suggests looking at that device, or putting it in downtime while it is fixed, before changing the rule.
- Flapping twice or more. It suggests a smoothed value (
avg_over_timeover 10–15 minutes), or warn and crit further apart.
Can I get it as a file?
Section titled “Can I get it as a file?”Yes: after a run, CSV downloads one row per rule and Print / PDF prints the
report. The CSV adds the mean times, the instance counts and the suggestions.
The same per-rule table is alert-quality.csv in the
audit evidence pack.
What limits and errors apply?
Section titled “What limits and errors apply?”The quality, digest and post-mortem reports and MOP documents share a limit of
20 reports a minute per organisation. Past it, the reply is 429 “at most 20 reports a
minute”.
| Message | Cause |
|---|---|
at most 400 days | The period is longer than 400 days |
the period must end after it starts, and start before now | From is after To, or in the future |
| Refused by the call’s argument check | Claude gave shortMinutes outside 1–240 or actionableMinutes outside 1–1440 (the dashboard card doesn’t set them) |
rule too long | A rule name over 63 characters |