Metrics and the scrape endpoint
SZ-MCP keeps time series per AP, radio, switch, port and controller node
for your organisation. Claude queries them with PromQL (metrics.* in code
mode), alert rules and boards are built on them, and your own tools can scrape
the current values from GET https://sz-mcp.lanpulse.com/metrics in the
Prometheus text format.
Where do the metrics come from?
Section titled “Where do the metrics come from?”API polling. An admin turns it on with the API polling switch on the dashboard’s Metrics card, or in step 3 of the setup checklist. SZ-MCP then reads the controller every 5 minutes; ports and controller statistics are read every 15 minutes. The same poll reads the controller’s outstanding alarms and rogue access points for alert rules. Once an hour it also reads licences, licence pools and the certificates the controller uses, for the expiry and capacity metrics.
If the controller is slow, polling backs off (the card shows slowed to every N min). If the controller refuses the saved login, polling pauses until the credentials are updated rather than retrying against SmartZone’s lockout. Engineers can click Poll now.
A few metrics — AP and switch CPU, memory and temperature, per-port byte counters and some radio counters — are only sent by SmartZone’s Northbound streaming, which the hosted service cannot receive yet. They have no data under polling.
How long are metrics kept?
Section titled “How long are metrics kept?”| Tier | Resolution | Kept |
|---|---|---|
| Raw | Every report | 48 hours |
| Hourly rollups | 1 hour | 8 days |
| Daily rollups | 1 day (UTC) | 400 days |
A query that reaches past a tier runs entirely on the next one. For longer or finer history, scrape the endpoint below into your own store.
How does Claude use them?
Section titled “How does Claude use them?”With the metrics functions in code mode:
| Function | What it does |
|---|---|
metrics.catalog() | The 54 metrics, their units, sources and labels |
metrics.query() | A PromQL query, instant or over a range |
metrics.status() | What has been collected and how fresh it is |
metrics.poll() | Poll now |
metrics.set_polling() | Turn polling on or off (admin) |
Series are labelled with the inventory’s names and containers, so you can ask
for things like “clients per zone over the last day” or filter with
{within="zone:…"}.
How do I get a scrape token?
Section titled “How do I get a scrape token?”On the Metrics card, under Prometheus scrape endpoint, an admin clicks
Create token. The token starts szms_ and is shown once, together with
a ready-made Prometheus job; copy it then. Rotate token replaces it (the old
one stops working at once) and Revoke deletes it. The card shows the
token’s prefix, when it was created and when it was last used.
The token reads metrics, and the organisation’s
/health/org report, and nothing else. A
scrape never calls your controller.
How do I configure Prometheus?
Section titled “How do I configure Prometheus?”Send the token as a bearer token:
scrape_configs: - job_name: smartzone scheme: https metrics_path: /metrics scrape_interval: 60s authorization: credentials: szms_... static_configs: - targets: ['sz-mcp.lanpulse.com']For tools that only do HTTP Basic auth, send the token as the password (the username is ignored).
SmartZone reports every 3 to 5.5 minutes, so scraping more often than once a minute gains nothing. By default no sample timestamps are written, so your scraper stamps each value at scrape time and series stay continuous between reports.
What query parameters does /metrics take?
Section titled “What query parameters does /metrics take?”Each can repeat or take a comma-separated list.
| Parameter | Effect | Default |
|---|---|---|
family | Only ap, switch or cluster metrics | All |
metric | Only these metrics, by catalogue name (with or without the sz_ prefix) | All |
within | Only entities under a zone, switch group, location…, e.g. within=zone:<id> | Everything |
max_age | Leave out series not reported for this many seconds (60–7200) | 1200 |
timestamps | 1 writes the controller’s own sample times | Off |
Metrics that are reported late or seldom (controller CPU bins, licence and
certificate expiry) stay in the output for 2 to 3 hours regardless of
max_age.
What does the output look like?
Section titled “What does the output look like?”Every metric name is prefixed sz_ and has # HELP and # TYPE lines. Value
series carry only identifying and grouping labels; one sz_<type>_info series
per entity (value 1) carries the display names and the rest. Join them in
PromQL:
sz_port_rx_kbps * on (port) group_left (name, switch_name) sz_port_infoWhat limits apply?
Section titled “What limits apply?”| Limit | What happens |
|---|---|
| 60 scrapes per minute per organisation | 429 with Retry-After |
| 20 failed token checks per minute per client IP | 429 |
| Missing or wrong token | 401 |
Unknown family or metric, or max_age out of range | 400, naming the problem |
| Output over 24 MB (roughly 140,000 series) | 413. The scrape is refused rather than cut short, since a partial scrape looks like series disappearing. Scrape per family (?family=ap, ?family=switch, ?family=cluster) or per container (?within=switchgroup:…) |
A network of more than about 200 48-port switches needs to be split this way.