Weekly baselines
A weekly baseline is what a metric usually is at this hour of the week: the median of the same UTC hour over up to the last 8 weeks. It lets Claude and your alert rules ask “is this normal for a Tuesday at 10:00?” instead of comparing against one fixed number. Baselines are kept for 11 metrics, and need 3 weeks of data before they give an answer.
How do I ask whether something is normal?
Section titled “How do I ask whether something is normal?”Ask Claude, for example “Is client count in Hall B normal for this time of
the week?” Claude uses metrics.baseline with an entity id:
- For an AP, switch or controller node, it compares each of that device’s baseline metrics (and an AP’s radios) per series.
- For a zone, AP group, switch group or location, it compares the sum over everything under it.
Gauges compare their latest value; counters compare their rate over the last hour. Each item gets a verdict:
| Verdict | Meaning |
|---|---|
usual | Within 3 spreads of the usual value |
above / below | 3 or more spreads above or below. Where there is no spread (a sum, or a spread of 0), under half or over double the usual value |
learning | Fewer than 3 weeks of this hour of the week so far |
If everything is still learning, the answer says “No baseline yet: it needs 3 weeks of the same hour of the week.”
How do I use baselines in PromQL?
Section titled “How do I use baselines in PromQL?”Three functions take a plain metric selector (no range or offset) of a metric that keeps a baseline:
| Function | Returns |
|---|---|
baseline(x) | The median of the same hour of the week over up to 8 past weeks; a counter’s per-second rate |
baseline_mad(x) | Its spread: the median absolute deviation × 1.4826, read like a standard deviation |
baseline_weeks(x) | How many past weeks are behind it |
baseline and baseline_mad return nothing until 3 weeks exist. On a metric
without a baseline, the query fails with <metric> keeps no baseline; metrics.catalog() marks the ones that do (baseline: true).
sum by (zone) (ap_clients) / sum by (zone) (baseline(ap_clients)) < 0.6(ap_clients - baseline(ap_clients)) / baseline_mad(ap_clients) < -3The first finds zones 40% below normal; the second finds APs more than 3 spreads below normal. Both work in threshold alert rules and backtests: only weeks before the hour being evaluated count, so a backtest never sees the future. Hours of the week are UTC, so after a daylight-saving change a local peak sits an hour off for the weeks either side.
Which metrics keep a baseline?
Section titled “Which metrics keep a baseline?”Eleven. Four of them come only from Northbound streaming, which is not available on the hosted service yet, so their baselines never fill there. The other seven also come from API polling, which must be on.
| Metric | Fills on the hosted service |
|---|---|
node_cpu_percent | Yes |
ap_clients | Yes |
radio_clients | Yes |
radio_airtime_percent | Yes |
radio_tx_bytes | Yes |
radio_rx_bytes | Yes |
switch_poe_used_watts | Yes |
ap_lan_rx_bytes | No: streaming only |
ap_lan_tx_bytes | No: streaming only |
switch_rx_bytes | No: streaming only |
switch_tx_bytes | No: streaming only |
metrics.catalog() marks these with baseline: true.
Where do baselines appear?
Section titled “Where do baselines appear?”In the context of every problem notification: the context pack’s Usual section lists the device’s baseline metrics against their usual value for the hour of the week, unusual ones first (up to 6). Metrics still learning are left out. The Alerts page’s problem drawer shows the first 3 of them as Against usual. See Alert context.