Skip to content
SZ-MCP
Get Support

How alerts work

SZ-MCP watches your SmartZone network with Nagios-style alert rules. Every rule is evaluated in your organisation’s own store (every minute by default) against the metrics, alarms and inventory that polling collects. A problem is confirmed over several reports before anyone is told, and notifications go out within about a minute of that.

You set rules up by asking Claude, which uses the alerts.* functions in code mode: it previews each rule and can backtest it against stored history before saving. You can also install the starter pack in one click, or import rules as YAML.

API polling. With polling on (step 3 of the setup checklist, or the API polling switch on the dashboard’s Metrics card), SZ-MCP reads the controller every 5 minutes: metrics, the controller’s outstanding alarms, and rogue access points. Rules read what polling stored; they never call the controller themselves.

Some metrics, such as AP and switch CPU and memory, are only delivered by SmartZone’s Northbound streaming, which the hosted service cannot receive yet. Rules on those metrics have nothing to evaluate and stay quiet. See Metrics and the scrape endpoint.

Each rule produces one instance per thing it checks: per AP, per port, per alarm code on a device, per rogue BSSID. Each instance has a state:

StateMeaning
OKFine
WARNINGPast the rule’s warn threshold, or a rule set to raise WARNING
CRITICALPast the rule’s crit threshold, or a rule set to raise CRITICAL (the default)
UNKNOWNThe check could not be judged

Only instances that are, or recently were, a problem are stored, so a healthy network costs almost nothing to watch.

What is the difference between SOFT and HARD?

Section titled “What is the difference between SOFT and HARD?”

A new problem starts SOFT and only becomes HARD — and notified — after maxCheckAttempts consecutive failing results. This is Nagios’ rule, and it stops a single bad sample from paging anyone.

  • Attempts count new reports, not evaluations. A rule that runs every minute over data that arrives every 5 minutes does not re-count the same sample, so maxCheckAttempts: 3 means three consecutive failing reports.
  • The default is 3 attempts, and 1 for event and rogue rules (which go HARD at once). The starter pack sets its own values per rule.
  • A change between problem states while HARD (WARNING to CRITICAL) is a HARD change and is notified.
  • Recovering from a SOFT problem is silent. Recovering from a HARD problem that was notified sends a RECOVERY.

On the Alerts page, SOFT problems are marked (soft) and drawn faded.

An instance that keeps changing state is marked flapping and its notifications pause. SZ-MCP uses Nagios’ method: the weighted share of state changes over the last 21 results (newer results weigh more). Flapping starts above 30% and stops below 20%. A FLAPPINGSTART and a FLAPPINGSTOP notification bracket the period. A rule can turn this off with flapDetection: false; the starter pack’s ap-flapping and port-flapping rules do, because they count flaps themselves.

How do I stop being told about something I’m handling?

Section titled “How do I stop being told about something I’m handling?”

Two ways, both available to operators and above, on the Alerts page or through Claude:

  • Acknowledge a problem. It stops re-notification and sends an ACKNOWLEDGEMENT notification with your comment. Acknowledgements are sticky by default: they survive WARNING ↔ CRITICAL changes and are dropped when the problem recovers. A non-sticky one is dropped on any state change.
  • Schedule downtime. Problems it covers are still tracked but not notified. Downtime can cover a rule, a device, everything within a zone, switch group or location, or a combination (all must match). It lasts up to 90 days.

A HARD problem that was held back by downtime, an acknowledgement, flapping or an upstream outage is notified when that ends, if it is still a problem.

When something a device depends on is down, the device’s own problem is marked unreachable and not notified. SZ-MCP knows the physical topology from its inventory — an AP’s switch port, the port’s switch. If an upstream entity has a HARD CRITICAL problem in any rule, problems downstream of it are suppressed, so a failed switch produces one page, not one per AP behind it.

Because a switch can report later than its APs, a new problem on something with an upstream waits 5 minutes (dependencyWait) before it is notified, giving the upstream failure time to show up. A rule can opt out with dependencies: false.

Does a problem repeat until someone deals with it?

Section titled “Does a problem repeat until someone deals with it?”

Only if the rule says so. By default a HARD problem is notified once. A rule with renotify (for example 4h) repeats the PROBLEM notification at that interval until it is acknowledged, recovers, or is put in downtime. The starter pack’s ap-offline and switch-offline repeat every 4 hours.

The Alerts page at /alerts (linked from the dashboard’s Alerts card as Open the alerts page) refreshes every minute. It has five parts:

  • The grid — one row per device with a problem, one column per rule that has one. Each cell shows CRIT, WARN or UNKN, with markers: ✓ acknowledged, ⏸ in downtime, ⇡ unreachable (something upstream is down), ~ flapping. Soft or handled problems are faded. With nothing wrong it reads All OK.
  • The detail pane — click a cell for the output, how long it has been a problem, and any acknowledgement or downtime. Operators and above get a Comment box and Acknowledge (or Remove acknowledgement), Downtime this check 1h, 4h or 24h, and Downtime the device 4h.
  • Downtime — active and scheduled downtime, each with Cancel.
  • Rules — every rule with its kind, expression, current problems, last run (or its error) and version, plus the Starter checks and Rules as code sections described in Alert rules.
  • Recent alert log — state changes, acknowledgements, downtime and rule changes, with the notification each one sent.

The alert history (state changes, acknowledgements, downtime and who changed which rule) is kept for 30 days, and every rule change keeps the previous version.

Roles are per organisation. Reading is open to everyone; changing things needs at least the role shown.

ActionViewerOperatorEngineerAdmin / owner
See problems, rules, history and downtime✓✓✓✓
Backtest a rule, export rules as YAML✓✓✓✓
Acknowledge, schedule or cancel downtime✓✓✓
Create, change or delete rules; import YAML; install the starter pack✓✓
Email contacts, webhook, Jira, escalations, heartbeat, test send✓
Create or revoke the scrape token✓

Through Claude, a role that is too low gets write_blocked with a message naming the role that can. Previews (dry runs) of rule changes work for every role; only apply: true is gated.