- Jsonnet 75.5%
- Python 14.1%
- Shell 10.1%
- Makefile 0.3%
Dashboards - Add CARC / Cluster Status (dashboards/carc/cluster-status.jsonnet), which replaces CARC / Researcher Overview and CARC / Where Will My Job Run (researcher-overview, queue-placement: deleted). It is the first dashboard to use the recording rules by name. - Rework CARC / My Jobs around one running job: per-node core use, memory headroom, and compute vs memory/network/filesystem stalls from LDMS. - CARC / Public Status is now reached through Grafana's "Share externally" link: no template variables, Prometheus uid pinned on every panel and target, gated by the `public` tag. - Add CARC Admin / LDMS Fleet (dashboards/carc-admin/ldms-fleet.jsonnet). - Admin Overview: add a Firewall (bifrost) row (scrape, uptime, CPU, FreeBSD memory, temperature, bridge throughput and packet rate, NIC drops/errors, CPU by mode). Replace the Accounts & Fairshare table with top-10 bar gauges for CPUs running and jobs pending: the table outer-joined 620 mostly idle fairshare rows, joined on account without cluster, and slurm_account_fairshare is 0 for every account. Libraries - grafonnet.libsonnet: dashboard.new defaults to browser timezone instead of grafonnet's utc. - datasource.libsonnet: pinnedRef/vortexUid for the public board only. - slurm.libsonnet: shared partition, status, out-of-service and pending reason helpers. - rules.libsonnet: the recording rules are live, so expr(key, useRules=true) is usable. - panels.libsonnet: fill in refIds, since Grafana 13 rejects duplicates. Rules and alerting - carc-recording.yml: group_left on the load-per-core join. - ldms.yml: mark the IB throughput rules untrusted, with the cross-check against node_exporter. - carc-alerts.yml: real vortex datasource uids in place of placeholders. Scripts - check-dashboards.sh: borrow promtool from the Prometheus image, and enforce the public/non-public datasource split by tag. - New check-alerts.py (structure, uids, --live evaluation) and check-dashboards-live.py (every panel target against vortex). - New retire-default-dashboards.sh and set-home-dashboard.sh (need root on vortex). - install-rules.sh: drop the rule_files instructions, now done. Host config and docs - host-config/ldms-compute/: syspapi block and component-id drop-in for the compute image. - CLAUDE.md, README.md, METRICS.md, ROADMAP.md updated to match, including a METRICS.md Firewall section for bifrost's FreeBSD node_exporter. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|---|---|---|
| alerting | ||
| dashboards | ||
| host-config | ||
| lib | ||
| rules | ||
| scripts | ||
| .gitignore | ||
| CLAUDE.md | ||
| jsonnetfile.json | ||
| jsonnetfile.lock.json | ||
| ldms-api-probe.py | ||
| ldms-exporter.py | ||
| ldms-facts.sh | ||
| Makefile | ||
| METRICS.md | ||
| README.md | ||
| ROADMAP.md | ||
CARC Grafana Dashboards
Dashboard-as-code for CARC (Center for Advanced Research Computing, UNM) using Grafonnet.
Environment
These dashboards target the PLG stack running as rootful podman quadlets on
vortex.alliance.unm.edu:
- Grafana 13.0.1 — OIDC login via Keycloak (
carcrealm), reverse-proxied by nginx athttps://vortex.alliance.unm.edu/. - Prometheus v3 —
http://prometheus:9090inside themonitoringpodman network, 90d / 800GB retention, remote-write enabled. - Loki 3 —
http://loki:3100, ingesting vortex's own systemd journal via its local Alloy plus, as of 2026-08-21, pushed logs from across the fleet over the mTLS ingest endpoint (see below). - snmp_exporter — UPS monitoring (Mitsubishi + Liebert units) via
job="snmp_ups"in Prometheus. Deployed by~/carc-snmp-exporter, a real (small) Ansible playbook run locally on vortex — the one piece of this host that actually is Ansible-managed; everything else described below is hand-administered.
nginx's mTLS ingest (loki-ingest.conf, port 8443) now proxies two paths to
hosts presenting a client cert signed by the IPA CA with a trusted CN
(easley, easley-sn, hopper, hopper-sn, auth, coldfront, vdtn01,
ood, freeipa01, freeipa02):
/loki/api/v1/push— log shipping (already live before this change, for the cluster nodes; the trust list grew to cover the identity/portal hosts too)./api/v1/write— added 2026-08-21, a Prometheus remote-write proxy so Alloy running oneasley-sn/hopper-sncan forward metrics scraped from NATed compute nodes vortex can't reach directly (Prometheus itself has--web.enable-remote-write-receiverset).
Confirmed live via the Grafana datasource-proxy API (label/series queries against Prometheus and Loki), not assumed from config:
job="prometheus.scrape.compute_nodes"in Prometheus — full node_exporter (CPU, memory, disk, network, Infiniband, hwmon/EDAC) for all 63 Easley compute nodes (easley001-easley063,role="compute"). Hopper's compute nodes are not wired up yet — Easley only.- New Loki
jobvalues:gpfs(Easley'smmfs.log),beegfs(Hopper'sbeegfs-client.log),slurm(slurmctldon both clusters' service nodes,slurmdon Easley compute nodes),system(journal/messages/secure/audit/ cron fleet-wide, including the identity/portal hosts),warewulf(provisioning, both service nodes), and one job per identity/portal service:freeipa,keycloak,coldfront,mokey,ood. - Alloy's own self-monitoring (go_/process_ metrics) on easley-sn shows up
in Prometheus as
job="integrations/self"— notjob="slurm". There is nojob="slurm"in Prometheus at all (confirmed live 2026-09-04). Don't confuse it with the Lokijob="slurm"logs above, which are real scheduler logs.
Config lives under /etc/monitoring/, data under /var/lib/monitoring/ on
that host; dashboards from this repo are provisioned into Grafana from two
folders, each its own file provider in default.yml:
/var/lib/monitoring/grafana/dashboards/carc→CARCfolder (providercarc) —cluster-status,my-jobs,public-status,ups-power. Left on default/inherited folder permissions, so anyone with the Viewer org role can see them with no per-dashboard permission step./var/lib/monitoring/grafana/dashboards/carc-admin→CARC Adminfolder (providercarc-admin) —admin-overview,compute-node,storage,security,monitoring-health. This one has the Viewer row deliberately removed from the folder's permissions, restricting it to Editor+.
Both are separate from the slurm-*.json dashboards, which live directly
under /var/lib/monitoring/grafana/dashboards (the default provider, root
folder) and aren't authored here.
make install (see Usage below) pushes compiled dashboard JSON onto the
two provisioned directories above. Everything else on that host (the
folders/permissions themselves, provisioning config, the box in general)
is hand-administered directly by whoever has root on vortex — there is no
Ansible or other config-management layer for it (the one exception,
snmp_exporter, is noted below). Treat it as production config outside
this repo's normal output: changes there should be deliberate and backed
up, and you should ask before assuming this repo should take over more
than dashboard content.
Known Prometheus scrape jobs today (/etc/monitoring/prometheus/prometheus.yml):
| job | targets | notes |
|---|---|---|
snmp_ups |
2 UPS units via snmp_exporter | role=ups, vendor=Mitsubishi|Liebert |
slurm_exporter |
easley-sn, hopper-sn (:9341) |
cluster=easley|hopper |
node |
easley login + controller, vortex, bifrost (firewall, by IP) | role=login|controller|monitoring|firewall |
prometheus |
self |
Plus one job that arrives via remote-write rather than a local scrape target
(see the mTLS ingest description above), so it won't appear in
prometheus.yml:
| job | source | notes |
|---|---|---|
prometheus.scrape.compute_nodes |
Alloy on easley-sn, forwarding all 63 Easley compute nodes | cluster=easley, role=compute; Hopper not wired up yet |
The /var/lib/monitoring/grafana/dashboards/slurm-*.json dashboards already
on the box are the upstream slurm-exporter defaults, not authored here.
Host config snapshots
host-config/ldms-compute/ holds the fragments (syspapi block, systemd
drop-in) that make the PAPI counters survive a compute-node reboot; the
runbook that assembles them from the Warewulf image copy is on disk on
easley-sn's operator as runbook-ldms-image-syspapi.sh.
vortex has no configuration management, so host-config/ keeps read-only
copies of the pieces of /etc/monitoring/ this repo depends on. Nothing here
applies them; make install only writes dashboard JSON. See
host-config/README.md, which also records why the
three dashboard provider paths must stay disjoint.
Layout
dashboards/carc/ Public dashboards (CARC folder) — 4 files
dashboards/carc-admin/ Admin dashboards (CARC Admin folder) — 6 files
lib/
datasource.libsonnet Datasource variable references
grafonnet.libsonnet Single import point for grafonnet
panels.libsonnet Shared panel builders (barGaugePanel, nodeTable,
stateTimelinePanel, etc.)
thresholds.libsonnet Shared threshold objects
node.libsonnet Node selector helpers (nodeSel, hostSel, etc.)
rules.libsonnet Recording rule name mappings
slurm.libsonnet Slurm scheduler template variables
rules/ Prometheus recording rules (carc-recording.yml)
alerting/ Grafana alert rules (carc-alerts.yml)
scripts/
install-dashboards.sh Deploy to vortex (run via `make install`)
install-rules.sh Deploy rules to vortex (run via `make install-rules`)
check-dashboards.sh Validate compiled JSON (run via `make check`)
check-dashboards-live.py Evaluate every panel target of a built dashboard
against live Prometheus/Loki via the Grafana proxy
check-alerts.py Validate alerting/*.yml; --live evaluates every query on vortex
verify-list.sh Regenerate the "Awaiting verification" section of METRICS.md from build/
vendor/ jsonnet-bundler dependencies (gitignored)
build/ Compiled dashboard JSON (gitignored)
carc/
carc-admin/
Usage
make vendor # jb install — fetch grafonnet v13.0.0 into vendor/
make # compile dashboards/ -> build/
make check # validate JSON, UIDs, titles, datasource refs, PromQL/LogQL
make fmt # jsonnetfmt -i in place
make fmt-check # fail if not formatted
make install # run check first, then copy build/{carc,carc-admin}/*.json
# onto vortex at the provisioned paths (needs sudo)
make install-rules # dry run by default; copy rules/*.yml and alerting/*.yml
# with FLAGS=--yes-write-to-etc-monitoring
make install runs make check first, validates the compile with jq and
promtool (if available), then calls scripts/install-dashboards.sh to refresh
dashboard JSON in the live carc / carc-admin provisioning directories on
vortex. It removes stale JSON in each directory before copying new files and
sets ownership to the Grafana user (472). Grafana's file provider polls every
30s, so no restart is needed for a content-only update.
make install-rules (scripts/install-rules.sh) defaults to a dry run showing
what would be copied and the follow-up steps; pass FLAGS=--yes-write-to-etc-monitoring
to actually copy. The script prints the manual steps still required on vortex
after install: promtool check config and a SIGHUP so Prometheus re-reads
/etc/prometheus/rules/*.yml, and a Grafana restart (or an admin POST to
/api/admin/provisioning/alerting/reload), because Grafana reads
provisioning/alerting/ only at start-up. alerting/carc-alerts.yml carries
vortex's real datasource UIDs; make check runs scripts/check-alerts.py,
which rejects any other UID, and scripts/check-alerts.py --live evaluates
every alert query through Grafana's datasource proxy and prints the value
range against each threshold, so run it before deploying.
Every dashboard exposes a datasource template variable (type datasource,
query prometheus) and references it as ${datasource} on every panel
target. Dashboards also expose loki_datasource wherever they use Loki
panels. This keeps dashboards portable instead of hardcoding a datasource UID.
Variables convention
Dashboards use lib/node.libsonnet helpers for any node_exporter metrics.
The node job has two targets (easley login, easley-sn controller), so a bare
{job="node"} selector returns two series. Use nodeSel(role) to filter by
role class (e.g., nodeSel('compute', ', cluster=~"$cluster"')), which matches
both job="node" and job="prometheus.scrape.compute_nodes" with a role label,
or hostSel(instance) for a single directly-scraped host. Template variables
exposed by default: $cluster (from any node_exporter series), $node (scoped
by $cluster), and $host (from Loki job="system").
The reason for this convention: job="node" and job="prometheus.scrape.compute_nodes"
exist today as separate scrape paths; ROADMAP.md workstream 2 will normalize
everything to job="node" with role/cluster labels. Dashboards using
nodeSel() will keep working through that transition.
Dashboards
Each dashboard has a UID starting with carc-, a title starting with CARC / or
CARC Admin /, and lives in the CARC or CARC Admin folder respectively.
Panels whose metrics are not yet catalogued in METRICS.md carry [verify] in
their descriptions; the list of those panels and their expressions lives in
METRICS.md's "Awaiting verification" section.
-
carc-cluster-status(uid:carc-cluster-status, folder: CARC, audience: Any logged-in viewer) — The researcher landing page. What is usable right now, a per-partition table of idle nodes/CPUs/GPUs and waiting jobs split by reason, node state over time with outage reasons and slurmctld node events, why jobs are waiting (free-text reasons collapsed to a bounded set), queue throughput and a "hours to clear the queue" proxy, the viewer's own jobs and fairshare with a per-job link into My Jobs, and shared-storage fill. Both clusters; filterable by cluster, partition and user. Set as the Grafana home dashboard (see "Home dashboard and retired dashboards" below). Replacedcarc-researcher-overviewandcarc-queue-placement, retired 2026-09-08. -
carc-my-jobs(uid:carc-my-jobs, folder: CARC, audience: Any logged-in viewer) — One Easley job, running or recently ended: a colour-coded footprint (CPU busy, memory used, load per core, IPC, InfiniBand, OOM kills), the same per node over time, and a per-node table. node_exporter and LDMS hardware counters joined to the job's nodes via the LDMS jobinfo sampler. Deep-linked (with the run's time range) from the job table oncarc-cluster-status, and fromcarc-admin-ldms-fleet. -
carc-public-status(uid:carc-public-status, folder: CARC, audience: Unauthenticated, via a "Share externally" link) — The website/lobby board. One row per cluster: data feed live/stale, nodes in and out of service (fromslurm_node_status, each node counted once), CPU and GPU in-use gauges, jobs running, jobs waiting, jobs started in the last hour, and "hours to clear the queue" (the same Little's-law proxy as Cluster Status; the one wait-expectation number that needs no per-job data, after ARCHER2's public page); then 24 h of jobs running, jobs waiting and CPU use for both clusters. Aggregate only: no usernames, partitions, node names or per-job detail. 30 s refresh, hidden time picker, no template variables and a pinned Prometheus uid, because externally shared dashboards do not interpolate variables (see "Public status board" below). Reworked 2026-09-09. -
carc-ups-power(uid:carc-ups-power, folder: CARC, audience: Any logged-in viewer) — Battery/power status for both datacenter UPS units (Mitsubishi and Liebert) — charge, load, voltage, temperature. -
carc-admin-overview(uid:carc-admin-overview, folder: CARC Admin, audience: Editor+) — Scheduler internals, exporter self-health, per-account fairshare/usage, login/controller node health, Easley compute fleet aggregates, and cross-service logs (GPFS/BeeGFS, slurmctld/slurmd, Warewulf, security). -
carc-admin-compute-node(uid:carc-admin-compute-node, folder: CARC Admin, audience: Editor+) — Per-node drill-down for the Easley 63-node fleet: CPU/memory/load, network/Infiniband throughput, EDAC/thermal, NFS retransmits, and slurmd/kernel logs filtered by selected node(s). -
carc-admin-storage(uid:carc-admin-storage, folder: CARC Admin, audience: Editor+) — Shared filesystem capacity and trends (home, projects, scratch), compute mount health, login/controller disk I/O, and GPFS/BeeGFS incident logs. -
carc-admin-security(uid:carc-admin-security, folder: CARC Admin, audience: Editor+) — SSH/sudo/audit/Kerberos/Keycloak/portal activity fleet-wide over Loki (pure log view, no Prometheus metrics yet). Broken out by host and failure type. -
carc-admin-ldms-fleet(uid:carc-admin-ldms-fleet, folder: CARC Admin, audience: Editor+) — Fleet-wide LDMS feed health, node occupancy, IPC and cache-miss distribution, InfiniBand counters (throughput panels marked UNTRUSTED), with per-job and per-node drill-down links. -
carc-admin-monitoring-health(uid:carc-admin-monitoring-health, folder: CARC Admin, audience: Editor+) — Is the monitoring stack itself working: scrape/remote-write/rule-eval health, compute metrics freshness, Loki stream liveness, and Prometheus TSDB state.
Home dashboard and retired dashboards
CARC / Cluster Status is meant to be what a researcher sees on login. Two
things outside this repo make that true, both hand-applied on vortex as root
(see host-config/README.md for why nothing here applies them):
- Retire the stock slurm-exporter dashboards ("01 - Cluster Overview" …
"10 - All Metrics Reference", uids
slurm-*) that thedefaultfile provider serves from/var/lib/monitoring/grafana/dashboards/default/into the General folder.sudo scripts/retire-default-dashboards.sh --yestars them to/var/lib/monitoring/grafana/retired-<date>.tar.gzand deletes the JSON; the provider hasdisableDeletion: false, so Grafana removes the dashboards within 30 s, no restart. The emptydefaultprovider is harmless; drop it from the provisioning file at the next planned Grafana restart if you like. - Make Cluster Status the home dashboard for everyone who has not set a
personal one:
sudo scripts/set-home-dashboard.sh --yesadds[dashboards] default_home_dashboard_path = /var/lib/grafana/dashboards/carc/cluster-status.json(the path as the container sees it) to/etc/monitoring/grafana/grafana.iniafter backing it up, then you runsudo systemctl restart grafana. Org, team and user preferences set through the UI still override it, which is the intended behaviour for admins who want a different landing page.
Retired from this repo on 2026-09-08 (removed from Grafana by the next
make install, which prunes JSON the repo no longer builds):
carc-researcher-overview ("CARC / Cluster Overview") and
carc-queue-placement ("CARC / Where Will My Job Run"), both superseded by
carc-cluster-status.
Access tiers
Grafana OIDC (grafana.ini) already maps Keycloak groups to roles:
monitor-admins → Admin, systems → Editor, everyone else → Viewer. That
covers the researcher/admin split for free — admin dashboards just need to
land in a folder whose permissions require Editor+, and researcher dashboards
in a folder any Viewer can read. Resolved as the CARC / CARC Admin
folder split described above.
Dashboards that use ${__user.login} (the $user variable on cluster-status and my-jobs)
assume Keycloak's preferred_username claim equals the Slurm user name, so the
template variable expands to the logged-in user's cluster account. If this
assumption doesn't hold in your Keycloak realm, fall back to a user textbox
variable for manual filtering.
Public status board
carc-public-status is the only dashboard meant to be seen without a login.
Anonymous access stays off ([auth.anonymous] enabled = false, OAuth
auto-login on, so any other vortex URL still bounces to Keycloak); instead the
one dashboard is shared through Grafana's externally shared ("public")
dashboards feature, which is enabled on vortex (publicDashboardsEnabled
true in the frontend settings, Grafana 13.0.1). That gives it a
https://vortex.alliance.unm.edu/public-dashboards/<token> URL that works
with no session and exposes nothing else.
Creating the share is a one-time UI step by an org Admin, not something
make install does (the repo has no write credential to Grafana):
- Open CARC / Public Status, click Share (top right), then Share externally.
- Leave Enable time range and Display annotations off; the board is fixed at "last 24 hours" and has no annotations.
- Copy the external link. That is the URL for the website or lobby screen.
Append
?kioskfor a chrome-free full-screen view on a TV.
Equivalent API call, with an Admin token (the read-only ops token cannot do this):
curl -X POST -H "Authorization: Bearer $ADMIN_TOKEN" -H 'Content-Type: application/json' \
https://vortex.alliance.unm.edu/api/dashboards/uid/carc-public-status/public-dashboards/ \
-d '{"isEnabled": true, "share": "public", "timeSelectionEnabled": false, "annotationsEnabled": false}'
The share is stored in Grafana's database keyed on the dashboard uid, so it
survives every make install (file provisioning updates the dashboard in
place; it does not change the uid). Pause or revoke it from Dashboards →
Shared dashboards, or with PATCH/DELETE on the same API path.
Constraints that the dashboard is written around (enforced by
scripts/check-dashboards.sh for anything tagged public):
- Template variables are not interpolated in shared views. The share
backend re-reads the saved JSON and queries whatever uid each target
names, so
${datasource}would be sent literally.public-statustherefore has no variables and pins the vortex Prometheus uid (lib/datasource.libsonnetprometheus.pinnedRef, uidPBFA97CFB590B2093) on every panel and target. It is the one exception to the${datasource}-only rule; do not copy the pattern into logged-in dashboards. - Only aggregate data. Anyone with the link sees it. Nothing on it may carry a user, account, job id, node name or partition label.
- Embedding in a web page (an
<iframe>on the CARC site) additionally needs[security] allow_embedding = truein/etc/monitoring/grafana/grafana.ini(currently unset, so Grafana sends a deny frame header) and a Grafana restart. A plain link or a kiosk browser does not need it.
Recording and alert rules
rules/carc-recording.yml holds Prometheus recording rules (instant queries
pre-computed on a schedule and stored as new time series, saving runtime
queries on dashboards). alerting/carc-alerts.yml holds Grafana file-provisioned
alert rules. Both target hand-administered destinations on vortex
(/etc/monitoring/prometheus/rules/ and
/etc/monitoring/grafana/provisioning/alerting/ respectively) — production
config outside this repo's normal output, so make install-rules treats
writing there as a deliberate, explicitly-flagged exception (see above), not
the default path. The recording rules have been live on vortex since 2026-09-05
(rule_files is set). The alert rules are complete and verified against live
data but not yet copied to vortex; see make install-rules above and the
header of alerting/carc-alerts.yml for the deploy and confirmation steps.