Skip to content
SZ-MCP
Get Support

Alert post-mortems

A post-mortem writes up one alert episode for a review: how it went from the first sign to recovery, what was around it when it became HARD, and what else happened on the device. It ends with empty review headings for you to fill in. It is assembled from the alert record and the stored history, with no AI model and no controller calls. Every member of the organisation can open one.

From a link next to an alert episode, or by asking Claude. The printable page is Monitor › Reports › Post-mortem, at /reports/postmortem. It is linked from:

WhereLink
The Alerts page, a problem’s drawer (HARD problems only)Post-mortem so far
The Alerts page, History tab: a HARD state changepost-mortem
An availability report’s Incidents tablePost-mortem
The alert quality report’s Most problems on columnThe device name

Claude uses alerts.postmortem({ rule, instance, at? }). It returns the write-up as data and as Markdown, a link to the printable page and a download link. An episode is chosen by at: the episode that time falls in, else the latest one that began before it. Without at, it is the newest episode.

The page header gives the facts, then the sections follow:

  • Header: the rule (with its description), the device, First sign (the first non-OK state, SOFT included), HARD from, Recovered (or “not yet”) with the duration, the worst state, and how many times this rule fired on this instance in the record.
  • What happened: every state change and acknowledgement in the episode, oldest first. Renotifications are left out.
  • Context at the start: the context pack as a notification would have carried it at the HARD start. It shows where the device sits, what it hangs off, what else was wrong, recent events and changes, baselines and earlier occurrences.
  • Runbook: the rule’s runbook field, when it has one.
  • Site notes: notes on the device, what contains it, its location and the organisation.
  • Timeline around it: everything else recorded on the device from 2 hours before the first sign to 30 minutes after recovery, up to 100 entries.
  • Review: empty headings for the cause, why it wasn’t caught sooner, and what you’ll change.

Only episodes that became HARD count. A SOFT blip that recovered by itself has no post-mortem.

Yes: Download .md on the page saves it as Markdown, and Print prints it. The file is named postmortem-<rule>-<state>-on-<device>.md. Write the review into the downloaded file, or ask Claude to draft the review sections from what you tell it. Times are in the scheduled digest’s time zone if one is set, otherwise UTC.

HARD state changes and recoveries are kept 400 days, so that is how far back a post-mortem can reach.

MessageCause
No HARD episode of <rule> on <instance> in the alert record (HARD changes and recoveries are kept 400 days).The instance never became HARD, or the episode is older than 400 days
rule and instance are requiredA link without both
rule or instance too longA hand-edited page link with a rule over 63 or an instance over 300 characters
at is an ISO time or epoch msA hand-edited page link with a bad at. Claude’s alerts.postmortem says `at` is not a time instead
at most 20 reports a minuteThe limit shared with alert quality, the on-demand digest and MOP documents

When the page can’t find an episode, it shows the error and a link to Alert history.