Monitoring the monitoring
An alerting system that has quietly stopped looks exactly like a healthy
network. SZ-MCP judges its own pipeline — are rules running, is data fresh,
are notifications going out — and exposes that three ways: on the dashboard’s
Alerts card under Self-monitoring, at GET /health/org for your own
monitoring system to poll, and as a heartbeat it sends only while things are
working, so silence means trouble.
This judges the monitoring, not the network. A network full of CRITICAL problems is a healthy monitor.
What does the health check look at?
Section titled “What does the health check look at?”Each check is OK, WARNING or CRITICAL, and the overall status is the worst of them.
| Check | CRITICAL when | WARNING when |
|---|---|---|
engine | The alert engine did not answer | |
rules | An enabled rule is more than 5 minutes overdue: the engine’s clock has stopped | No rules are enabled, or some rules fail to evaluate |
data | Data is on but none has arrived yet, or the newest data is stale | Nothing feeds the organisation (polling is off) |
collector | The controller refused the saved login, so polling is paused until the credentials change | The last poll failed for another reason |
delivery | Notifications have waited more than 5 minutes: delivery isn’t running | Rules are enabled but there is no channel, or deliveries failed in the last hour |
inventory | The last inventory sync failed |
Data is stale after 15 minutes, or after 2.5 poll intervals when polling has
slowed down (it backs off when the controller is slow), whichever is longer.
Claude can read the same report with alerts.health().
How do I poll it from Nagios, Zabbix or Uptime Kuma?
Section titled “How do I poll it from Nagios, Zabbix or Uptime Kuma?”There are two endpoints:
| Endpoint | Auth | Answers |
|---|---|---|
GET https://sz-mcp.lanpulse.com/health | None | Whether the service and its database answer: {"status":"ok","service":"sz-mcp"} (200) or 503. Says nothing about any organisation |
GET https://sz-mcp.lanpulse.com/health/org | The organisation’s scrape token | This organisation’s monitoring health |
/health/org takes the same scrape token as the
metrics endpoint, as
Authorization: Bearer szms_… or as the password of HTTP Basic auth (any
username). It returns 200 for OK or WARNING and 503 for CRITICAL, with
this JSON:
{ "status": "WARNING", "at": 1790431320000, "summary": "WARNING - delivery: no notification channel: problems are recorded but nobody is told", "checks": [ { "name": "rules", "state": "OK", "message": "…" }, { "name": "data", "state": "OK", "message": "…" }, { "name": "delivery", "state": "WARNING", "message": "…" } ]}Add ?format=text for just the one-line summary, Nagios plugin style (only
the checks that are not OK are listed, unless everything is OK). It is limited
to 30 checks per minute per organisation; polling once a minute is plenty.
curl -s -H "Authorization: Bearer $SZMS_TOKEN" \ "https://sz-mcp.lanpulse.com/health/org?format=text"What is the dead-man’s switch?
Section titled “What is the dead-man’s switch?”A heartbeat: while the monitoring is not CRITICAL, SZ-MCP POSTs to a URL you choose every few minutes. Point it at a service that alerts when pings stop — Healthchecks.io, Cronitor, Better Stack or an Uptime Kuma push monitor. If the pings stop, whether because something in SZ-MCP broke or the service itself is down, that service tells you.
An admin sets it under Dead-man’s switch on the Alerts card:
| Field | Meaning |
|---|---|
| Heartbeat URL | https on a public host. It usually holds the receiver’s secret, so it is stored encrypted and never shown again (only a hint). Leave blank when editing to keep it |
| Every (minutes) | 1 to 60, default 5. Set the receiver’s grace period to at least twice this |
Pause / Resume, Send now and Remove manage it. The card shows when the last ping went out, or that it failed or was not sent because the status was CRITICAL.
Each ping is a JSON POST with User-Agent: sz-mcp-heartbeat:
{ "source": "sz-mcp", "org": "…", "status": "OK", "at": "2026-09-26T14:05:00.000Z", "summary": "OK - …" }