CARC grafana dashboard collection
  • Jsonnet 75.5%
  • Python 14.1%
  • Shell 10.1%
  • Makefile 0.3%
Find a file
Troy Redfearn e9f9398040 Cluster Status replaces the researcher boards; firewall and top-account panels
Dashboards
- Add CARC / Cluster Status (dashboards/carc/cluster-status.jsonnet), which
  replaces CARC / Researcher Overview and CARC / Where Will My Job Run
  (researcher-overview, queue-placement: deleted). It is the first dashboard
  to use the recording rules by name.
- Rework CARC / My Jobs around one running job: per-node core use, memory
  headroom, and compute vs memory/network/filesystem stalls from LDMS.
- CARC / Public Status is now reached through Grafana's "Share externally"
  link: no template variables, Prometheus uid pinned on every panel and
  target, gated by the `public` tag.
- Add CARC Admin / LDMS Fleet (dashboards/carc-admin/ldms-fleet.jsonnet).
- Admin Overview: add a Firewall (bifrost) row (scrape, uptime, CPU, FreeBSD
  memory, temperature, bridge throughput and packet rate, NIC drops/errors,
  CPU by mode). Replace the Accounts & Fairshare table with top-10 bar gauges
  for CPUs running and jobs pending: the table outer-joined 620 mostly idle
  fairshare rows, joined on account without cluster, and
  slurm_account_fairshare is 0 for every account.

Libraries
- grafonnet.libsonnet: dashboard.new defaults to browser timezone instead
  of grafonnet's utc.
- datasource.libsonnet: pinnedRef/vortexUid for the public board only.
- slurm.libsonnet: shared partition, status, out-of-service and pending
  reason helpers.
- rules.libsonnet: the recording rules are live, so expr(key, useRules=true)
  is usable.
- panels.libsonnet: fill in refIds, since Grafana 13 rejects duplicates.

Rules and alerting
- carc-recording.yml: group_left on the load-per-core join.
- ldms.yml: mark the IB throughput rules untrusted, with the cross-check
  against node_exporter.
- carc-alerts.yml: real vortex datasource uids in place of placeholders.

Scripts
- check-dashboards.sh: borrow promtool from the Prometheus image, and
  enforce the public/non-public datasource split by tag.
- New check-alerts.py (structure, uids, --live evaluation) and
  check-dashboards-live.py (every panel target against vortex).
- New retire-default-dashboards.sh and set-home-dashboard.sh (need root on
  vortex).
- install-rules.sh: drop the rule_files instructions, now done.

Host config and docs
- host-config/ldms-compute/: syspapi block and component-id drop-in for the
  compute image.
- CLAUDE.md, README.md, METRICS.md, ROADMAP.md updated to match, including
  a METRICS.md Firewall section for bifrost's FreeBSD node_exporter.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-29 14:27:02 -06:00
alerting Cluster Status replaces the researcher boards; firewall and top-account panels 2026-09-29 14:27:02 -06:00
dashboards Cluster Status replaces the researcher boards; firewall and top-account panels 2026-09-29 14:27:02 -06:00
host-config Cluster Status replaces the researcher boards; firewall and top-account panels 2026-09-29 14:27:02 -06:00
lib Cluster Status replaces the researcher boards; firewall and top-account panels 2026-09-29 14:27:02 -06:00
rules Cluster Status replaces the researcher boards; firewall and top-account panels 2026-09-29 14:27:02 -06:00
scripts Cluster Status replaces the researcher boards; firewall and top-account panels 2026-09-29 14:27:02 -06:00
.gitignore Add the LDMS Prometheus exporter, written against the probed API 2026-09-04 23:31:36 -06:00
CLAUDE.md Cluster Status replaces the researcher boards; firewall and top-account panels 2026-09-29 14:27:02 -06:00
jsonnetfile.json Split dashboards by folder, bump grafonnet to v13, add make check and shared libs 2026-09-04 14:57:12 -06:00
jsonnetfile.lock.json Split dashboards by folder, bump grafonnet to v13, add make check and shared libs 2026-09-04 14:57:12 -06:00
ldms-api-probe.py Retract the papi_sampler finding: it came from a truncated read 2026-09-04 19:03:40 -06:00
ldms-exporter.py Measure the dir() cache honestly, and slow the collection interval 2026-09-04 23:46:19 -06:00
ldms-facts.sh Retract the papi_sampler finding: it came from a truncated read 2026-09-04 19:03:40 -06:00
Makefile Fix nodeTable column defaults (self-referential local) and rebuild on lib changes 2026-09-04 15:18:37 -06:00
METRICS.md Cluster Status replaces the researcher boards; firewall and top-account panels 2026-09-29 14:27:02 -06:00
README.md Cluster Status replaces the researcher boards; firewall and top-account panels 2026-09-29 14:27:02 -06:00
ROADMAP.md Cluster Status replaces the researcher boards; firewall and top-account panels 2026-09-29 14:27:02 -06:00

CARC Grafana Dashboards

Dashboard-as-code for CARC (Center for Advanced Research Computing, UNM) using Grafonnet.

Environment

These dashboards target the PLG stack running as rootful podman quadlets on vortex.alliance.unm.edu:

  • Grafana 13.0.1 — OIDC login via Keycloak (carc realm), reverse-proxied by nginx at https://vortex.alliance.unm.edu/.
  • Prometheus v3 — http://prometheus:9090 inside the monitoring podman network, 90d / 800GB retention, remote-write enabled.
  • Loki 3 — http://loki:3100, ingesting vortex's own systemd journal via its local Alloy plus, as of 2026-08-21, pushed logs from across the fleet over the mTLS ingest endpoint (see below).
  • snmp_exporter — UPS monitoring (Mitsubishi + Liebert units) via job="snmp_ups" in Prometheus. Deployed by ~/carc-snmp-exporter, a real (small) Ansible playbook run locally on vortex — the one piece of this host that actually is Ansible-managed; everything else described below is hand-administered.

nginx's mTLS ingest (loki-ingest.conf, port 8443) now proxies two paths to hosts presenting a client cert signed by the IPA CA with a trusted CN (easley, easley-sn, hopper, hopper-sn, auth, coldfront, vdtn01, ood, freeipa01, freeipa02):

  • /loki/api/v1/push — log shipping (already live before this change, for the cluster nodes; the trust list grew to cover the identity/portal hosts too).
  • /api/v1/write — added 2026-08-21, a Prometheus remote-write proxy so Alloy running on easley-sn/hopper-sn can forward metrics scraped from NATed compute nodes vortex can't reach directly (Prometheus itself has --web.enable-remote-write-receiver set).

Confirmed live via the Grafana datasource-proxy API (label/series queries against Prometheus and Loki), not assumed from config:

  • job="prometheus.scrape.compute_nodes" in Prometheus — full node_exporter (CPU, memory, disk, network, Infiniband, hwmon/EDAC) for all 63 Easley compute nodes (easley001-easley063, role="compute"). Hopper's compute nodes are not wired up yet — Easley only.
  • New Loki job values: gpfs (Easley's mmfs.log), beegfs (Hopper's beegfs-client.log), slurm (slurmctld on both clusters' service nodes, slurmd on Easley compute nodes), system (journal/messages/secure/audit/ cron fleet-wide, including the identity/portal hosts), warewulf (provisioning, both service nodes), and one job per identity/portal service: freeipa, keycloak, coldfront, mokey, ood.
  • Alloy's own self-monitoring (go_/process_ metrics) on easley-sn shows up in Prometheus as job="integrations/self" — not job="slurm". There is no job="slurm" in Prometheus at all (confirmed live 2026-09-04). Don't confuse it with the Loki job="slurm" logs above, which are real scheduler logs.

Config lives under /etc/monitoring/, data under /var/lib/monitoring/ on that host; dashboards from this repo are provisioned into Grafana from two folders, each its own file provider in default.yml:

  • /var/lib/monitoring/grafana/dashboards/carc → CARC folder (provider carc) — cluster-status, my-jobs, public-status, ups-power. Left on default/inherited folder permissions, so anyone with the Viewer org role can see them with no per-dashboard permission step.
  • /var/lib/monitoring/grafana/dashboards/carc-admin → CARC Admin folder (provider carc-admin) — admin-overview, compute-node, storage, security, monitoring-health. This one has the Viewer row deliberately removed from the folder's permissions, restricting it to Editor+.

Both are separate from the slurm-*.json dashboards, which live directly under /var/lib/monitoring/grafana/dashboards (the default provider, root folder) and aren't authored here.

make install (see Usage below) pushes compiled dashboard JSON onto the two provisioned directories above. Everything else on that host (the folders/permissions themselves, provisioning config, the box in general) is hand-administered directly by whoever has root on vortex — there is no Ansible or other config-management layer for it (the one exception, snmp_exporter, is noted below). Treat it as production config outside this repo's normal output: changes there should be deliberate and backed up, and you should ask before assuming this repo should take over more than dashboard content.

Known Prometheus scrape jobs today (/etc/monitoring/prometheus/prometheus.yml):

job targets notes
snmp_ups 2 UPS units via snmp_exporter role=ups, vendor=Mitsubishi|Liebert
slurm_exporter easley-sn, hopper-sn (:9341) cluster=easley|hopper
node easley login + controller, vortex, bifrost (firewall, by IP) role=login|controller|monitoring|firewall
prometheus self

Plus one job that arrives via remote-write rather than a local scrape target (see the mTLS ingest description above), so it won't appear in prometheus.yml:

job source notes
prometheus.scrape.compute_nodes Alloy on easley-sn, forwarding all 63 Easley compute nodes cluster=easley, role=compute; Hopper not wired up yet

The /var/lib/monitoring/grafana/dashboards/slurm-*.json dashboards already on the box are the upstream slurm-exporter defaults, not authored here.

Host config snapshots

host-config/ldms-compute/ holds the fragments (syspapi block, systemd drop-in) that make the PAPI counters survive a compute-node reboot; the runbook that assembles them from the Warewulf image copy is on disk on easley-sn's operator as runbook-ldms-image-syspapi.sh.

vortex has no configuration management, so host-config/ keeps read-only copies of the pieces of /etc/monitoring/ this repo depends on. Nothing here applies them; make install only writes dashboard JSON. See host-config/README.md, which also records why the three dashboard provider paths must stay disjoint.

Layout

dashboards/carc/              Public dashboards (CARC folder) — 4 files
dashboards/carc-admin/        Admin dashboards (CARC Admin folder) — 6 files
lib/
  datasource.libsonnet        Datasource variable references
  grafonnet.libsonnet         Single import point for grafonnet
  panels.libsonnet            Shared panel builders (barGaugePanel, nodeTable,
                              stateTimelinePanel, etc.)
  thresholds.libsonnet        Shared threshold objects
  node.libsonnet              Node selector helpers (nodeSel, hostSel, etc.)
  rules.libsonnet             Recording rule name mappings
  slurm.libsonnet             Slurm scheduler template variables
rules/                        Prometheus recording rules (carc-recording.yml)
alerting/                     Grafana alert rules (carc-alerts.yml)
scripts/
  install-dashboards.sh       Deploy to vortex (run via `make install`)
  install-rules.sh            Deploy rules to vortex (run via `make install-rules`)
  check-dashboards.sh         Validate compiled JSON (run via `make check`)
  check-dashboards-live.py    Evaluate every panel target of a built dashboard
                              against live Prometheus/Loki via the Grafana proxy
  check-alerts.py             Validate alerting/*.yml; --live evaluates every query on vortex
  verify-list.sh              Regenerate the "Awaiting verification" section of METRICS.md from build/
vendor/                       jsonnet-bundler dependencies (gitignored)
build/                        Compiled dashboard JSON (gitignored)
  carc/
  carc-admin/

Usage

make vendor         # jb install — fetch grafonnet v13.0.0 into vendor/
make                # compile dashboards/ -> build/
make check          # validate JSON, UIDs, titles, datasource refs, PromQL/LogQL
make fmt            # jsonnetfmt -i in place
make fmt-check      # fail if not formatted
make install        # run check first, then copy build/{carc,carc-admin}/*.json
                    # onto vortex at the provisioned paths (needs sudo)
make install-rules  # dry run by default; copy rules/*.yml and alerting/*.yml
                    # with FLAGS=--yes-write-to-etc-monitoring

make install runs make check first, validates the compile with jq and promtool (if available), then calls scripts/install-dashboards.sh to refresh dashboard JSON in the live carc / carc-admin provisioning directories on vortex. It removes stale JSON in each directory before copying new files and sets ownership to the Grafana user (472). Grafana's file provider polls every 30s, so no restart is needed for a content-only update.

make install-rules (scripts/install-rules.sh) defaults to a dry run showing what would be copied and the follow-up steps; pass FLAGS=--yes-write-to-etc-monitoring to actually copy. The script prints the manual steps still required on vortex after install: promtool check config and a SIGHUP so Prometheus re-reads /etc/prometheus/rules/*.yml, and a Grafana restart (or an admin POST to /api/admin/provisioning/alerting/reload), because Grafana reads provisioning/alerting/ only at start-up. alerting/carc-alerts.yml carries vortex's real datasource UIDs; make check runs scripts/check-alerts.py, which rejects any other UID, and scripts/check-alerts.py --live evaluates every alert query through Grafana's datasource proxy and prints the value range against each threshold, so run it before deploying.

Every dashboard exposes a datasource template variable (type datasource, query prometheus) and references it as ${datasource} on every panel target. Dashboards also expose loki_datasource wherever they use Loki panels. This keeps dashboards portable instead of hardcoding a datasource UID.

Variables convention

Dashboards use lib/node.libsonnet helpers for any node_exporter metrics. The node job has two targets (easley login, easley-sn controller), so a bare {job="node"} selector returns two series. Use nodeSel(role) to filter by role class (e.g., nodeSel('compute', ', cluster=~"$cluster"')), which matches both job="node" and job="prometheus.scrape.compute_nodes" with a role label, or hostSel(instance) for a single directly-scraped host. Template variables exposed by default: $cluster (from any node_exporter series), $node (scoped by $cluster), and $host (from Loki job="system").

The reason for this convention: job="node" and job="prometheus.scrape.compute_nodes" exist today as separate scrape paths; ROADMAP.md workstream 2 will normalize everything to job="node" with role/cluster labels. Dashboards using nodeSel() will keep working through that transition.

Dashboards

Each dashboard has a UID starting with carc-, a title starting with CARC / or CARC Admin /, and lives in the CARC or CARC Admin folder respectively. Panels whose metrics are not yet catalogued in METRICS.md carry [verify] in their descriptions; the list of those panels and their expressions lives in METRICS.md's "Awaiting verification" section.

  • carc-cluster-status (uid: carc-cluster-status, folder: CARC, audience: Any logged-in viewer) — The researcher landing page. What is usable right now, a per-partition table of idle nodes/CPUs/GPUs and waiting jobs split by reason, node state over time with outage reasons and slurmctld node events, why jobs are waiting (free-text reasons collapsed to a bounded set), queue throughput and a "hours to clear the queue" proxy, the viewer's own jobs and fairshare with a per-job link into My Jobs, and shared-storage fill. Both clusters; filterable by cluster, partition and user. Set as the Grafana home dashboard (see "Home dashboard and retired dashboards" below). Replaced carc-researcher-overview and carc-queue-placement, retired 2026-09-08.

  • carc-my-jobs (uid: carc-my-jobs, folder: CARC, audience: Any logged-in viewer) — One Easley job, running or recently ended: a colour-coded footprint (CPU busy, memory used, load per core, IPC, InfiniBand, OOM kills), the same per node over time, and a per-node table. node_exporter and LDMS hardware counters joined to the job's nodes via the LDMS jobinfo sampler. Deep-linked (with the run's time range) from the job table on carc-cluster-status, and from carc-admin-ldms-fleet.

  • carc-public-status (uid: carc-public-status, folder: CARC, audience: Unauthenticated, via a "Share externally" link) — The website/lobby board. One row per cluster: data feed live/stale, nodes in and out of service (from slurm_node_status, each node counted once), CPU and GPU in-use gauges, jobs running, jobs waiting, jobs started in the last hour, and "hours to clear the queue" (the same Little's-law proxy as Cluster Status; the one wait-expectation number that needs no per-job data, after ARCHER2's public page); then 24 h of jobs running, jobs waiting and CPU use for both clusters. Aggregate only: no usernames, partitions, node names or per-job detail. 30 s refresh, hidden time picker, no template variables and a pinned Prometheus uid, because externally shared dashboards do not interpolate variables (see "Public status board" below). Reworked 2026-09-09.

  • carc-ups-power (uid: carc-ups-power, folder: CARC, audience: Any logged-in viewer) — Battery/power status for both datacenter UPS units (Mitsubishi and Liebert) — charge, load, voltage, temperature.

  • carc-admin-overview (uid: carc-admin-overview, folder: CARC Admin, audience: Editor+) — Scheduler internals, exporter self-health, per-account fairshare/usage, login/controller node health, Easley compute fleet aggregates, and cross-service logs (GPFS/BeeGFS, slurmctld/slurmd, Warewulf, security).

  • carc-admin-compute-node (uid: carc-admin-compute-node, folder: CARC Admin, audience: Editor+) — Per-node drill-down for the Easley 63-node fleet: CPU/memory/load, network/Infiniband throughput, EDAC/thermal, NFS retransmits, and slurmd/kernel logs filtered by selected node(s).

  • carc-admin-storage (uid: carc-admin-storage, folder: CARC Admin, audience: Editor+) — Shared filesystem capacity and trends (home, projects, scratch), compute mount health, login/controller disk I/O, and GPFS/BeeGFS incident logs.

  • carc-admin-security (uid: carc-admin-security, folder: CARC Admin, audience: Editor+) — SSH/sudo/audit/Kerberos/Keycloak/portal activity fleet-wide over Loki (pure log view, no Prometheus metrics yet). Broken out by host and failure type.

  • carc-admin-ldms-fleet (uid: carc-admin-ldms-fleet, folder: CARC Admin, audience: Editor+) — Fleet-wide LDMS feed health, node occupancy, IPC and cache-miss distribution, InfiniBand counters (throughput panels marked UNTRUSTED), with per-job and per-node drill-down links.

  • carc-admin-monitoring-health (uid: carc-admin-monitoring-health, folder: CARC Admin, audience: Editor+) — Is the monitoring stack itself working: scrape/remote-write/rule-eval health, compute metrics freshness, Loki stream liveness, and Prometheus TSDB state.

Home dashboard and retired dashboards

CARC / Cluster Status is meant to be what a researcher sees on login. Two things outside this repo make that true, both hand-applied on vortex as root (see host-config/README.md for why nothing here applies them):

  1. Retire the stock slurm-exporter dashboards ("01 - Cluster Overview" … "10 - All Metrics Reference", uids slurm-*) that the default file provider serves from /var/lib/monitoring/grafana/dashboards/default/ into the General folder. sudo scripts/retire-default-dashboards.sh --yes tars them to /var/lib/monitoring/grafana/retired-<date>.tar.gz and deletes the JSON; the provider has disableDeletion: false, so Grafana removes the dashboards within 30 s, no restart. The empty default provider is harmless; drop it from the provisioning file at the next planned Grafana restart if you like.
  2. Make Cluster Status the home dashboard for everyone who has not set a personal one: sudo scripts/set-home-dashboard.sh --yes adds [dashboards] default_home_dashboard_path = /var/lib/grafana/dashboards/carc/cluster-status.json (the path as the container sees it) to /etc/monitoring/grafana/grafana.ini after backing it up, then you run sudo systemctl restart grafana. Org, team and user preferences set through the UI still override it, which is the intended behaviour for admins who want a different landing page.

Retired from this repo on 2026-09-08 (removed from Grafana by the next make install, which prunes JSON the repo no longer builds): carc-researcher-overview ("CARC / Cluster Overview") and carc-queue-placement ("CARC / Where Will My Job Run"), both superseded by carc-cluster-status.

Access tiers

Grafana OIDC (grafana.ini) already maps Keycloak groups to roles: monitor-admins → Admin, systems → Editor, everyone else → Viewer. That covers the researcher/admin split for free — admin dashboards just need to land in a folder whose permissions require Editor+, and researcher dashboards in a folder any Viewer can read. Resolved as the CARC / CARC Admin folder split described above.

Dashboards that use ${__user.login} (the $user variable on cluster-status and my-jobs) assume Keycloak's preferred_username claim equals the Slurm user name, so the template variable expands to the logged-in user's cluster account. If this assumption doesn't hold in your Keycloak realm, fall back to a user textbox variable for manual filtering.

Public status board

carc-public-status is the only dashboard meant to be seen without a login. Anonymous access stays off ([auth.anonymous] enabled = false, OAuth auto-login on, so any other vortex URL still bounces to Keycloak); instead the one dashboard is shared through Grafana's externally shared ("public") dashboards feature, which is enabled on vortex (publicDashboardsEnabled true in the frontend settings, Grafana 13.0.1). That gives it a https://vortex.alliance.unm.edu/public-dashboards/<token> URL that works with no session and exposes nothing else.

Creating the share is a one-time UI step by an org Admin, not something make install does (the repo has no write credential to Grafana):

  1. Open CARC / Public Status, click Share (top right), then Share externally.
  2. Leave Enable time range and Display annotations off; the board is fixed at "last 24 hours" and has no annotations.
  3. Copy the external link. That is the URL for the website or lobby screen. Append ?kiosk for a chrome-free full-screen view on a TV.

Equivalent API call, with an Admin token (the read-only ops token cannot do this):

curl -X POST -H "Authorization: Bearer $ADMIN_TOKEN" -H 'Content-Type: application/json' \
  https://vortex.alliance.unm.edu/api/dashboards/uid/carc-public-status/public-dashboards/ \
  -d '{"isEnabled": true, "share": "public", "timeSelectionEnabled": false, "annotationsEnabled": false}'

The share is stored in Grafana's database keyed on the dashboard uid, so it survives every make install (file provisioning updates the dashboard in place; it does not change the uid). Pause or revoke it from Dashboards → Shared dashboards, or with PATCH/DELETE on the same API path.

Constraints that the dashboard is written around (enforced by scripts/check-dashboards.sh for anything tagged public):

  • Template variables are not interpolated in shared views. The share backend re-reads the saved JSON and queries whatever uid each target names, so ${datasource} would be sent literally. public-status therefore has no variables and pins the vortex Prometheus uid (lib/datasource.libsonnet prometheus.pinnedRef, uid PBFA97CFB590B2093) on every panel and target. It is the one exception to the ${datasource}-only rule; do not copy the pattern into logged-in dashboards.
  • Only aggregate data. Anyone with the link sees it. Nothing on it may carry a user, account, job id, node name or partition label.
  • Embedding in a web page (an <iframe> on the CARC site) additionally needs [security] allow_embedding = true in /etc/monitoring/grafana/grafana.ini (currently unset, so Grafana sends a deny frame header) and a Grafana restart. A plain link or a kiosk browser does not need it.

Recording and alert rules

rules/carc-recording.yml holds Prometheus recording rules (instant queries pre-computed on a schedule and stored as new time series, saving runtime queries on dashboards). alerting/carc-alerts.yml holds Grafana file-provisioned alert rules. Both target hand-administered destinations on vortex (/etc/monitoring/prometheus/rules/ and /etc/monitoring/grafana/provisioning/alerting/ respectively) — production config outside this repo's normal output, so make install-rules treats writing there as a deliberate, explicitly-flagged exception (see above), not the default path. The recording rules have been live on vortex since 2026-09-05 (rule_files is set). The alert rules are complete and verified against live data but not yet copied to vortex; see make install-rules above and the header of alerting/carc-alerts.yml for the deploy and confirmation steps.