How alerts work
SZ-MCP watches your SmartZone network with Nagios-style alert rules. Every rule is evaluated in your organisation’s own store (every minute by default) against the metrics, alarms and inventory that polling collects. A problem is confirmed over several reports before anyone is told, and notifications go out within about a minute of that.
You set rules up by asking Claude, which uses the alerts.* functions in code
mode: it previews each rule and can backtest it against stored history before
saving. You can also install the starter pack
in one click, or import rules as YAML.
Where does the data come from?
Section titled “Where does the data come from?”API polling. With polling on (step 3 of the setup checklist, or the API polling switch on the dashboard’s Metrics card), SZ-MCP reads the controller every 5 minutes: metrics, the controller’s outstanding alarms, and rogue access points. Rules read what polling stored; they never call the controller themselves.
Some metrics, such as AP and switch CPU and memory, are only delivered by SmartZone’s Northbound streaming, which the hosted service cannot receive yet. Rules on those metrics have nothing to evaluate and stay quiet. See Metrics and the scrape endpoint.
What states can a problem be in?
Section titled “What states can a problem be in?”Each rule produces one instance per thing it checks: per AP, per port, per alarm code on a device, per rogue BSSID. Each instance has a state:
| State | Meaning |
|---|---|
OK | Fine |
WARNING | Past the rule’s warn threshold, or a rule set to raise WARNING |
CRITICAL | Past the rule’s crit threshold, or a rule set to raise CRITICAL (the default) |
UNKNOWN | The check could not be judged |
Only instances that are, or recently were, a problem are stored, so a healthy network costs almost nothing to watch.
What is the difference between SOFT and HARD?
Section titled “What is the difference between SOFT and HARD?”A new problem starts SOFT and only becomes HARD — and notified — after
maxCheckAttempts consecutive failing results. This is Nagios’ rule, and it
stops a single bad sample from paging anyone.
- Attempts count new reports, not evaluations. A rule that runs every minute
over data that arrives every 5 minutes does not re-count the same sample, so
maxCheckAttempts: 3means three consecutive failing reports. - The default is 3 attempts, and 1 for event and rogue rules (which go HARD at once). The starter pack sets its own values per rule.
- A change between problem states while HARD (WARNING to CRITICAL) is a HARD change and is notified.
- Recovering from a SOFT problem is silent. Recovering from a HARD problem that was notified sends a RECOVERY.
On the Alerts page, SOFT problems are marked (soft) and drawn faded.
What counts as flapping?
Section titled “What counts as flapping?”An instance that keeps changing state is marked flapping and its
notifications pause. SZ-MCP uses Nagios’ method: the weighted share of state
changes over the last 21 results (newer results weigh more). Flapping starts
above 30% and stops below 20%. A FLAPPINGSTART and a
FLAPPINGSTOP notification bracket the period. A rule can turn this off with
flapDetection: false; the starter pack’s ap-flapping and port-flapping
rules do, because they count flaps themselves.
How do I stop being told about something I’m handling?
Section titled “How do I stop being told about something I’m handling?”Two ways, both available to operators and above, on the Alerts page or through Claude:
- Acknowledge a problem. It stops re-notification and sends an ACKNOWLEDGEMENT notification with your comment. Acknowledgements are sticky by default: they survive WARNING ↔ CRITICAL changes and are dropped when the problem recovers. A non-sticky one is dropped on any state change.
- Schedule downtime. Problems it covers are still tracked but not notified. Downtime can cover a rule, a device, everything within a zone, switch group or location, or a combination (all must match). It lasts up to 90 days.
A HARD problem that was held back by downtime, an acknowledgement, flapping or an upstream outage is notified when that ends, if it is still a problem.
What does “unreachable” mean?
Section titled “What does “unreachable” mean?”When something a device depends on is down, the device’s own problem is marked unreachable and not notified. SZ-MCP knows the physical topology from its inventory — an AP’s switch port, the port’s switch. If an upstream entity has a HARD CRITICAL problem in any rule, problems downstream of it are suppressed, so a failed switch produces one page, not one per AP behind it.
Because a switch can report later than its APs, a new problem on something with
an upstream waits 5 minutes (dependencyWait) before it is notified, giving
the upstream failure time to show up. A rule can opt out with
dependencies: false.
Does a problem repeat until someone deals with it?
Section titled “Does a problem repeat until someone deals with it?”Only if the rule says so. By default a HARD problem is notified once. A rule
with renotify (for example 4h) repeats the PROBLEM notification at that
interval until it is acknowledged, recovers, or is put in downtime. The starter
pack’s ap-offline and switch-offline repeat every 4 hours.
The Alerts page
Section titled “The Alerts page”The Alerts page at /alerts (linked from the dashboard’s Alerts card as
Open the alerts page) refreshes every minute. It has five parts:
- The grid — one row per device with a problem, one column per rule that has
one. Each cell shows
CRIT,WARNorUNKN, with markers:✓acknowledged,⏸in downtime,⇡unreachable (something upstream is down),~flapping. Soft or handled problems are faded. With nothing wrong it reads All OK. - The detail pane — click a cell for the output, how long it has been a problem, and any acknowledgement or downtime. Operators and above get a Comment box and Acknowledge (or Remove acknowledgement), Downtime this check 1h, 4h or 24h, and Downtime the device 4h.
- Downtime — active and scheduled downtime, each with Cancel.
- Rules — every rule with its kind, expression, current problems, last run (or its error) and version, plus the Starter checks and Rules as code sections described in Alert rules.
- Recent alert log — state changes, acknowledgements, downtime and rule changes, with the notification each one sent.
The alert history (state changes, acknowledgements, downtime and who changed which rule) is kept for 30 days, and every rule change keeps the previous version.
Who can do what?
Section titled “Who can do what?”Roles are per organisation. Reading is open to everyone; changing things needs at least the role shown.
| Action | Viewer | Operator | Engineer | Admin / owner |
|---|---|---|---|---|
| See problems, rules, history and downtime | ✓ | ✓ | ✓ | ✓ |
| Backtest a rule, export rules as YAML | ✓ | ✓ | ✓ | ✓ |
| Acknowledge, schedule or cancel downtime | ✓ | ✓ | ✓ | |
| Create, change or delete rules; import YAML; install the starter pack | ✓ | ✓ | ||
| Email contacts, webhook, Jira, escalations, heartbeat, test send | ✓ | |||
| Create or revoke the scrape token | ✓ |
Through Claude, a role that is too low gets write_blocked with a message
naming the role that can. Previews (dry runs) of rule changes work for every
role; only apply: true is gated.