Alert rules
An alert rule says what counts as a problem. There are five kinds:
threshold, nodata, event, presence and rogue. You normally create
rules by describing what you want to Claude, which writes the rule, shows what
it would flag right now, and saves it only when asked. An organisation can have
up to 200 rules.
How do I create a rule?
Section titled “How do I create a rule?”Ask Claude, for example “alert me when a radio’s airtime stays above 80% for
15 minutes”. Claude uses these alerts.* functions:
| Function | What it does |
|---|---|
create_rule | Dry run by default: validates the rule and previews what it would flag now. apply: true saves it |
update_rule | Merges changes into a rule (a field set to null returns to its default). Dry run unless apply: true. The previous version is kept |
delete_rule | Deletes a rule and its current problems; its versions and history stay |
backtest | Replays a rule over stored history without saving anything (below) |
check_now | Evaluates one rule, or every enabled rule, immediately |
rules / get_rule | Lists rules with their problem counts, last run and last error; versions: true adds a rule’s change history |
Saving, changing and deleting need the engineer role or above. Disabling a rule clears its current problems.
Which rule kind do I need?
Section titled “Which rule kind do I need?”| Kind | Raises a problem when | Required fields | One instance per |
|---|---|---|---|
threshold | A PromQL expression crosses warn and/or crit | expr, and crit or warn | Series the expression returns |
nodata | A reporter has sent nothing for longer than maxAge | family, maxAge | Silent reporter |
event | A controller alarm (or event) matches | match | Device and code |
presence | A registered device is not seen on the network | select | Device |
rogue | An access point that isn’t yours advertises a watched SSID | ssids and/or ssidsFrom | Rogue BSSID |
Threshold
Section titled “Threshold”expr is PromQL over the metrics catalogue (Claude reads it with
metrics.catalog()). op is one of > (default), >=, <, <=, ==,
!=. Give crit, warn or both.
- name: radio-airtime kind: threshold expr: avg_over_time(radio_airtime_percent[15m]) warn: 60 crit: 80Nodata
Section titled “Nodata”Every reporter of family (ap, switch or cluster) that has been silent
for longer than maxAge (for example 20m). within limits it to a zone,
switch group or location. The family stream watches Northbound streaming
itself.
match selects log entries by kinds (alarm, event), codes,
severities (for example Critical, Major) and text. An alarm stays a
problem until SmartZone clears it. An event goes back to OK after
clearAfter (default 30m) without a repeat.
Polling reads the controller’s outstanding alarms every 5 minutes, so alarm rules work without streaming. Controller events (as distinct from alarms) arrive only by Northbound streaming, which the hosted service cannot receive yet.
Presence
Section titled “Presence”select is an inventory selector for devices you have registered as ones that
must be present (for example { tags: { critical: 'yes' } }). A device that is
not seen as a wireless client, a wired client or an LLDP neighbour is a
problem. Ask Claude to register devices and tag them; that is
inventory.add_device and inventory.tag.
ssids lists SSIDs to watch (compared ignoring case), and/or ssidsFrom
selects inventory WLANs whose SSIDs are watched, for example
{ type: 'wlan', tags: { facility: 'yes' } }. A rogue must have been seen
within maxAge (default 30m); within limits it to rogues seen by APs in a
zone or location. The problem sits on the AP that saw it. A rogue marked known,
ignored or allow-listed on the controller is OK.
Polling reads rogues every 5 minutes. The zone must have rogue detection turned on in SmartZone.
What fields does every rule take?
Section titled “What fields does every rule take?”| Field | Default | Meaning |
|---|---|---|
name | required | Lower-case letters, digits, _ . -, up to 63 characters |
description | none | Shown on the Alerts page and in notifications |
enabled | true | |
state | CRITICAL | The state an event, presence, nodata or rogue problem raises (WARNING or CRITICAL) |
maxCheckAttempts | 3; 1 for event and rogue rules | Consecutive failing reports before the problem is HARD and notified (1–20) |
every | 1m | How often the rule is evaluated |
renotify | none (notify once) | Repeat an unacknowledged HARD problem this often, e.g. 4h |
flapDetection | true | Pause notifications while an instance flaps |
dependencies | true | Suppress a problem whose upstream (the AP’s switch port or switch) is down |
dependencyWait | 5m | How long a new problem with an upstream waits before notifying |
notify | true | false records problems without notifying |
labels | none | Free-form, e.g. { team: facilities }; carried on notifications and used by escalations |
Durations are written like 30s, 5m, 1h, 7d. See
How alerts work for what SOFT, HARD, flapping and
dependencies mean.
How do I test a rule before saving it?
Section titled “How do I test a rule before saving it?”Backtest it. alerts.backtest replays a rule (a saved one by name, or a
definition) over stored history and reports how many instances would have had
problems, the episodes, the notifications that would have gone out, flapping
starts, and the longest problems. The default window is 24 hours (up to 30
days); past 48 hours it runs on hourly rollups. Threshold rules replay over the
metrics (default step 5 minutes); event rules replay over the stored log.
Ask: “Backtest a rule that fires when port CRC errors exceed 50 in 15 minutes over the last week, and tell me how noisy it would have been.”
What is in the starter pack?
Section titled “What is in the starter pack?”27 rules for SmartZone, each labelled pack: starter. Install them from
step 5 of the setup checklist, from
Starter checks on the Alerts page (Preview, then Add N checks),
or by asking Claude (alerts.starter_pack). Installing adds only the rules you
don’t have and never changes one that exists, so rules you have tuned survive a
re-install and new checks arrive when you run it again.
| Area | Rule | Fires on |
|---|---|---|
| Controller | controller-cluster-down | The cluster is not In_Service |
controller-node-down | A controller node is not In_Service | |
controller-cpu-high | Node CPU over 80% (warn) / 95% (crit) | |
controller-memory-high | Node memory over 85% / 95% | |
controller-disk-high | Node disk over 80% / 90% | |
controller-silent | No controller report for 30 minutes | |
stream-silent | Northbound streaming silent for 10 minutes | |
sz-alarm-critical | A Critical SmartZone alarm (CRITICAL until SmartZone clears it) | |
sz-alarm-major | A Major SmartZone alarm (WARNING) | |
certificate-expiring | A certificate the controller uses expires within 30 days / 7 days | |
license-expiring | A licence expires within 60 days / 14 days | |
license-capacity | A licence pool 90% / 100% used | |
| Wi-Fi | ap-offline | An AP is not online (repeats every 4 h) |
ap-flapping | An AP went offline and back 4 / 8 times in an hour | |
ap-cpu-high | AP CPU over 80% / 95% for 15 minutes | |
ap-memory-high | AP memory over 85% / 95% for 15 minutes | |
radio-airtime-high | Channel airtime over 70% / 85% for 15 minutes | |
radio-noise-high | Noise floor above −85 / −75 dBm for 15 minutes | |
clients-below-usual | A zone’s clients below 50% / 25% of the same time last week | |
rogue-facility-ssid | A rogue advertising a facility SSID (CRITICAL) | |
rogue-own-ssid | A rogue advertising any of your SSIDs (WARNING) | |
| Switching | switch-offline | A switch is not online (repeats every 4 h) |
switch-cpu-high | Switch CPU over 80% / 95% for 15 minutes | |
switch-memory-high | Switch memory over 85% / 95% for 15 minutes | |
switch-poe-budget | PoE drawn over 80% / 95% of the budget | |
port-crc-errors | 10 / 100 CRC errors on a port in 15 minutes | |
port-flapping | A port went down and up 4 / 10 times in an hour |
Some of these stay quiet on the hosted service by design:
ap-cpu-high,ap-memory-high,switch-cpu-highandswitch-memory-highuse metrics only Northbound streaming delivers, so they have nothing to evaluate while polling is the data source.stream-silentwatches the stream, which is not in use.rogue-facility-ssidwatches WLANs taggedfacility: yes, so it is quiet until you tag one. Ask Claude: “Tag the BMS-Sensors WLAN as a facility SSID.”clients-below-usualskips zones that usually have under 20 clients, and a zone has no baseline until it has a week of data. A holiday reads as a drop: acknowledge it or schedule downtime.
Can I keep rules in git?
Section titled “Can I keep rules in git?”Yes: export them as YAML, edit, and import the file back. The export writes each rule in its shortest form (fields at their default are left out) with a fixed field order and rules sorted by name, so a diff shows only what someone changed.
- Export: Download the rules as YAML on the Alerts page (any member),
or ask Claude (
alerts.export_rules). - Import: paste the document into Rules as code on the Alerts page and
click Preview changes, then Apply (engineer or above). Claude’s
alerts.import_rulesworks the same way, as a dry run untilapply: true.
The document looks like this (a bare list of rules also works):
version: 1rules: - name: ap-offline kind: threshold expr: ap_up op: "<" crit: 1Import creates rules that are new and updates ones that differ. Tick Delete
rules the document doesn’t list (Claude: prune: true) to also delete rules
missing from the file. Every rule is validated before anything changes: one
invalid rule rejects the whole import, with the rule’s position, name and the
problem. The document must be version: 1, at most 1 MB, and may only have the
top-level keys version and rules.