Skip to content
SZ-MCP
Get Support

Alert rules

An alert rule says what counts as a problem. There are five kinds: threshold, nodata, event, presence and rogue. You normally create rules by describing what you want to Claude, which writes the rule, shows what it would flag right now, and saves it only when asked. An organisation can have up to 200 rules.

Ask Claude, for example “alert me when a radio’s airtime stays above 80% for 15 minutes”. Claude uses these alerts.* functions:

FunctionWhat it does
create_ruleDry run by default: validates the rule and previews what it would flag now. apply: true saves it
update_ruleMerges changes into a rule (a field set to null returns to its default). Dry run unless apply: true. The previous version is kept
delete_ruleDeletes a rule and its current problems; its versions and history stay
backtestReplays a rule over stored history without saving anything (below)
check_nowEvaluates one rule, or every enabled rule, immediately
rules / get_ruleLists rules with their problem counts, last run and last error; versions: true adds a rule’s change history

Saving, changing and deleting need the engineer role or above. Disabling a rule clears its current problems.

KindRaises a problem whenRequired fieldsOne instance per
thresholdA PromQL expression crosses warn and/or critexpr, and crit or warnSeries the expression returns
nodataA reporter has sent nothing for longer than maxAgefamily, maxAgeSilent reporter
eventA controller alarm (or event) matchesmatchDevice and code
presenceA registered device is not seen on the networkselectDevice
rogueAn access point that isn’t yours advertises a watched SSIDssids and/or ssidsFromRogue BSSID

expr is PromQL over the metrics catalogue (Claude reads it with metrics.catalog()). op is one of > (default), >=, <, <=, ==, !=. Give crit, warn or both.

- name: radio-airtime
kind: threshold
expr: avg_over_time(radio_airtime_percent[15m])
warn: 60
crit: 80

Every reporter of family (ap, switch or cluster) that has been silent for longer than maxAge (for example 20m). within limits it to a zone, switch group or location. The family stream watches Northbound streaming itself.

match selects log entries by kinds (alarm, event), codes, severities (for example Critical, Major) and text. An alarm stays a problem until SmartZone clears it. An event goes back to OK after clearAfter (default 30m) without a repeat.

Polling reads the controller’s outstanding alarms every 5 minutes, so alarm rules work without streaming. Controller events (as distinct from alarms) arrive only by Northbound streaming, which the hosted service cannot receive yet.

select is an inventory selector for devices you have registered as ones that must be present (for example { tags: { critical: 'yes' } }). A device that is not seen as a wireless client, a wired client or an LLDP neighbour is a problem. Ask Claude to register devices and tag them; that is inventory.add_device and inventory.tag.

ssids lists SSIDs to watch (compared ignoring case), and/or ssidsFrom selects inventory WLANs whose SSIDs are watched, for example { type: 'wlan', tags: { facility: 'yes' } }. A rogue must have been seen within maxAge (default 30m); within limits it to rogues seen by APs in a zone or location. The problem sits on the AP that saw it. A rogue marked known, ignored or allow-listed on the controller is OK.

Polling reads rogues every 5 minutes. The zone must have rogue detection turned on in SmartZone.

FieldDefaultMeaning
namerequiredLower-case letters, digits, _ . -, up to 63 characters
descriptionnoneShown on the Alerts page and in notifications
enabledtrue
stateCRITICALThe state an event, presence, nodata or rogue problem raises (WARNING or CRITICAL)
maxCheckAttempts3; 1 for event and rogue rulesConsecutive failing reports before the problem is HARD and notified (1–20)
every1mHow often the rule is evaluated
renotifynone (notify once)Repeat an unacknowledged HARD problem this often, e.g. 4h
flapDetectiontruePause notifications while an instance flaps
dependenciestrueSuppress a problem whose upstream (the AP’s switch port or switch) is down
dependencyWait5mHow long a new problem with an upstream waits before notifying
notifytruefalse records problems without notifying
labelsnoneFree-form, e.g. { team: facilities }; carried on notifications and used by escalations

Durations are written like 30s, 5m, 1h, 7d. See How alerts work for what SOFT, HARD, flapping and dependencies mean.

Backtest it. alerts.backtest replays a rule (a saved one by name, or a definition) over stored history and reports how many instances would have had problems, the episodes, the notifications that would have gone out, flapping starts, and the longest problems. The default window is 24 hours (up to 30 days); past 48 hours it runs on hourly rollups. Threshold rules replay over the metrics (default step 5 minutes); event rules replay over the stored log.

Ask: “Backtest a rule that fires when port CRC errors exceed 50 in 15 minutes over the last week, and tell me how noisy it would have been.”

27 rules for SmartZone, each labelled pack: starter. Install them from step 5 of the setup checklist, from Starter checks on the Alerts page (Preview, then Add N checks), or by asking Claude (alerts.starter_pack). Installing adds only the rules you don’t have and never changes one that exists, so rules you have tuned survive a re-install and new checks arrive when you run it again.

AreaRuleFires on
Controllercontroller-cluster-downThe cluster is not In_Service
controller-node-downA controller node is not In_Service
controller-cpu-highNode CPU over 80% (warn) / 95% (crit)
controller-memory-highNode memory over 85% / 95%
controller-disk-highNode disk over 80% / 90%
controller-silentNo controller report for 30 minutes
stream-silentNorthbound streaming silent for 10 minutes
sz-alarm-criticalA Critical SmartZone alarm (CRITICAL until SmartZone clears it)
sz-alarm-majorA Major SmartZone alarm (WARNING)
certificate-expiringA certificate the controller uses expires within 30 days / 7 days
license-expiringA licence expires within 60 days / 14 days
license-capacityA licence pool 90% / 100% used
Wi-Fiap-offlineAn AP is not online (repeats every 4 h)
ap-flappingAn AP went offline and back 4 / 8 times in an hour
ap-cpu-highAP CPU over 80% / 95% for 15 minutes
ap-memory-highAP memory over 85% / 95% for 15 minutes
radio-airtime-highChannel airtime over 70% / 85% for 15 minutes
radio-noise-highNoise floor above −85 / −75 dBm for 15 minutes
clients-below-usualA zone’s clients below 50% / 25% of the same time last week
rogue-facility-ssidA rogue advertising a facility SSID (CRITICAL)
rogue-own-ssidA rogue advertising any of your SSIDs (WARNING)
Switchingswitch-offlineA switch is not online (repeats every 4 h)
switch-cpu-highSwitch CPU over 80% / 95% for 15 minutes
switch-memory-highSwitch memory over 85% / 95% for 15 minutes
switch-poe-budgetPoE drawn over 80% / 95% of the budget
port-crc-errors10 / 100 CRC errors on a port in 15 minutes
port-flappingA port went down and up 4 / 10 times in an hour

Some of these stay quiet on the hosted service by design:

  • ap-cpu-high, ap-memory-high, switch-cpu-high and switch-memory-high use metrics only Northbound streaming delivers, so they have nothing to evaluate while polling is the data source.
  • stream-silent watches the stream, which is not in use.
  • rogue-facility-ssid watches WLANs tagged facility: yes, so it is quiet until you tag one. Ask Claude: “Tag the BMS-Sensors WLAN as a facility SSID.”
  • clients-below-usual skips zones that usually have under 20 clients, and a zone has no baseline until it has a week of data. A holiday reads as a drop: acknowledge it or schedule downtime.

Yes: export them as YAML, edit, and import the file back. The export writes each rule in its shortest form (fields at their default are left out) with a fixed field order and rules sorted by name, so a diff shows only what someone changed.

  • Export: Download the rules as YAML on the Alerts page (any member), or ask Claude (alerts.export_rules).
  • Import: paste the document into Rules as code on the Alerts page and click Preview changes, then Apply (engineer or above). Claude’s alerts.import_rules works the same way, as a dry run until apply: true.

The document looks like this (a bare list of rules also works):

version: 1
rules:
- name: ap-offline
kind: threshold
expr: ap_up
op: "<"
crit: 1

Import creates rules that are new and updates ones that differ. Tick Delete rules the document doesn’t list (Claude: prune: true) to also delete rules missing from the file. Every rule is validated before anything changes: one invalid rule rejects the whole import, with the rule’s position, name and the problem. The document must be version: 1, at most 1 MB, and may only have the top-level keys version and rules.