Alert post-mortems
A post-mortem writes up one alert episode for a review: how it went from the first sign to recovery, what was around it when it became HARD, and what else happened on the device. It ends with empty review headings for you to fill in. It is assembled from the alert record and the stored history, with no AI model and no controller calls. Every member of the organisation can open one.
Where do I open one?
Section titled “Where do I open one?”From a link next to an alert episode, or by asking Claude. The printable page
is Monitor › Reports › Post-mortem, at /reports/postmortem. It is linked
from:
| Where | Link |
|---|---|
| The Alerts page, a problem’s drawer (HARD problems only) | Post-mortem so far |
| The Alerts page, History tab: a HARD state change | post-mortem |
| An availability report’s Incidents table | Post-mortem |
| The alert quality report’s Most problems on column | The device name |
Claude uses alerts.postmortem({ rule, instance, at? }). It returns the
write-up as data and as Markdown, a link to the printable page and a download
link. An episode is chosen by at: the episode that time falls in, else the
latest one that began before it. Without at, it is the newest episode.
What does it contain?
Section titled “What does it contain?”The page header gives the facts, then the sections follow:
- Header: the rule (with its description), the device, First sign (the first non-OK state, SOFT included), HARD from, Recovered (or “not yet”) with the duration, the worst state, and how many times this rule fired on this instance in the record.
- What happened: every state change and acknowledgement in the episode, oldest first. Renotifications are left out.
- Context at the start: the context pack as a notification would have carried it at the HARD start. It shows where the device sits, what it hangs off, what else was wrong, recent events and changes, baselines and earlier occurrences.
- Runbook: the rule’s
runbookfield, when it has one. - Site notes: notes on the device, what contains it, its location and the organisation.
- Timeline around it: everything else recorded on the device from 2 hours before the first sign to 30 minutes after recovery, up to 100 entries.
- Review: empty headings for the cause, why it wasn’t caught sooner, and what you’ll change.
Only episodes that became HARD count. A SOFT blip that recovered by itself has no post-mortem.
Can I download it?
Section titled “Can I download it?”Yes: Download .md on the page saves it as Markdown, and Print prints it. The
file is named postmortem-<rule>-<state>-on-<device>.md. Write the review into
the downloaded file, or ask Claude to draft the review sections from what you
tell it. Times are in the scheduled digest’s time zone if one is set, otherwise
UTC.
What limits and errors apply?
Section titled “What limits and errors apply?”HARD state changes and recoveries are kept 400 days, so that is how far back a post-mortem can reach.
| Message | Cause |
|---|---|
No HARD episode of <rule> on <instance> in the alert record (HARD changes and recoveries are kept 400 days). | The instance never became HARD, or the episode is older than 400 days |
rule and instance are required | A link without both |
rule or instance too long | A hand-edited page link with a rule over 63 or an instance over 300 characters |
at is an ISO time or epoch ms | A hand-edited page link with a bad at. Claude’s alerts.postmortem says `at` is not a time instead |
at most 20 reports a minute | The limit shared with alert quality, the on-demand digest and MOP documents |
When the page can’t find an episode, it shows the error and a link to Alert history.