Per-gateway download/upload/ping from the hourly librespeed runs, with each link's 7-day median overlaid so the WANSpeedDegraded alert condition is visible at a glance. Label-driven — new sites appear automatically once scraped.
5.9 KiB
Monitoring
Metrics and logs live on control over the ZeroTier mesh; the Grafana dashboards
are also published publicly through web01 (see Dashboards).
Collection
- node_exporter (
:9100) on every machine — CPU, memory, disk, systemd units. Binds all interfaces; the scrape ports are firewall-scoped to the mesh subnet (modules/monitoring/exporters.nix). - knot-exporter (
:9433) onns1/ns2only — reads Knot's control socket, fed by themod-statsmodule (query/response counters per zone). - blackbox_exporter (
127.0.0.1:9115) oncontrolonly — outside-in DNS probes. For every zone it queries each nameserver's public address (v4 and v6) for SOA (is the zone served?) and DNSKEY (is it still signed?). This is the resolver's-eye view that the Knot stats can't see. Probe definitions are shared between the exporter and the scrape jobs inmodules/monitoring/blackbox-probes.nix.
Storage & scraping
VictoriaMetrics on control, bound to 127.0.0.1:8428, 180-day retention
(modules/monitoring/server.nix). It scrapes control over loopback and
ns1/ns2/mx1/web01 over the mesh.
The scraper dials IPv4-only by default, so mesh (IPv6) targets need
extraOptions = [ "-enableTCP6" ]. Without it, ns1/ns2 are dropped with "no suitable address found". Check live target health oncontrol:curl -s http://127.0.0.1:8428/api/v1/targets | jq '.data.activeTargets[] | {i:.labels.instance, h:.health, e:.lastError}'
Dashboards
Grafana on control (:3000), anonymous access disabled. Reachable directly
over the mesh, and publicly at https://grafana.cnx.network via web01's reverse
proxy (TLS termination — see Overview). The admin password is a
clan var:
clan vars get control grafana-admin/password
Dashboards are provisioned from modules/monitoring/dashboards/ (any JSON file
there is picked up):
- CNX DNS (
dns.json) — firing alerts, per-nameserver SOA serials, zone expiry countdowns, query/response rates, host CPU/memory/disk/load, and the outside-in DNS probes. - CNX Backups (
backups.json) — borgbackup job health, time since the last run, and per-job state. See Backups. - CNX Uptime (
uptime.json) — per-host up/down status, current uptime, availability over the selected window, and up/down history. Label-driven, so every scraped host appears automatically. - CNX WAN Speed (
wan.json) — per-gateway WAN speed from the hourly librespeed runs: latest download/upload/ping, throughput history with each link's 7-day median as reference, latency/jitter, and test status. See Site gateways. - parsedmarc — DMARC aggregate/forensic report viewer. Auto-provisioned by
the
parsedmarcmodule (not fromdashboards/); reads its own Elasticsearch datasource, not VictoriaMetrics. See DMARC reports below.
DMARC reports
The cnx.email DMARC record (rua/ruf) points at the dmarc@cnx.email
mailbox on mx1. parsedmarc on control (modules/monitoring/parsedmarc.nix)
polls that mailbox over IMAPS, parses the XML reports, and stores them in a local
Elasticsearch (127.0.0.1:9200, loopback-only); Grafana renders them via the
parsedmarc dashboard + two Elasticsearch datasources (dmarc-ag/dmarc-fo).
Two Grafana-13 gotchas, both handled in
modules/monitoring/parsedmarc.nix:
- The Elasticsearch data source plugin is no longer bundled. Grafana 13 externalised its core data sources and nixpkgs ships no
grafanaPluginsentry for Elasticsearch, so a bare ES datasource fails with "Plugin not registered" and every panel errors (and Explore silently falls back to the default datasource). We pull the official, Grafana-signedelasticsearchplugin from the catalog viaservices.grafana.declarativePlugins.- Datasource variables now resolve by UID, not name. The bundled dashboard binds its datasource through template variables whose stored value is the datasource name (
dmarc-ag/dmarc-fo); Grafana 11+ matches that against the datasource UID. So we provision the two datasources ourselves with their UID pinned to their name (and leave the module's ownprovision.grafana.datasourceoff to avoid a duplicate, uid-less pair).deleteDatasourcesclears any stale uid-less records left by an earlier deploy so the pinned-UID versions can be re-created.
The IMAP fetch rides the mesh, not the public net: control pins
mx1.cnx.email to mx1's mesh address in /etc/hosts, so TLS still validates
against the public cert while the bytes stay on the overlay. The mailbox
passphrase is the shared mail-dmarc-cred clan var (so both mx1's mailserver and
control's parsedmarc see the same value):
clan vars get mx1 mail-dmarc-cred/passphrase
Logs
VictoriaLogs on control (:9428), 30-day retention
(modules/monitoring/server.nix). All three hosts ship journald to it via
systemd's own services.journald.upload → the /insert/journald endpoint
(modules/monitoring/exporters.nix); no extra agent. control uploads over
loopback so its logs survive a mesh outage, the other hosts push over the mesh,
and 9428 is firewall-scoped to the mesh like everything else.
Same IPv4-only default as the scraper: VictoriaLogs binds
0.0.0.0:9428for a bare:9428, so mesh (IPv6) pushes from ns1/ns2 are refused until you passextraOptions = [ "-enableTCP6" ](binds[::]). Verify the bind oncontrol:ss -tlnp | grep 9428 # want [::]:9428, not 0.0.0.0:9428
Query logs from Grafana via the provisioned VictoriaLogs datasource (Explore
view, LogsQL), or directly in the built-in UI at http://[control]:9428/select/vmui.
Logs are tagged with _HOSTNAME and _SYSTEMD_UNIT, so to follow one service
across hosts:
_SYSTEMD_UNIT:"knot.service"