No description
  • Python 93.1%
  • Shell 6.1%
  • Dockerfile 0.8%
Find a file
Troy Redfearn e4864da3c2 keys/README: the write token's honest scope -- change on one group, the proxy is the boundary
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-09 23:55:30 -06:00
agent Poller: escalation tags, human feedback readback, and the weekly report 2026-09-09 23:50:21 -06:00
bin Zammad proxy: second write route, POST /api/v1/tags/add, allowlisted tags only 2026-09-09 23:48:32 -06:00
evals retrieval_eval: a case may accept any of several articles (tuple expected) 2026-09-09 22:48:16 -06:00
gateway A/B result: ab-large loses to carc-tools; retire it. FINDINGS #30 records the day 2026-09-09 18:46:48 -06:00
keys keys/README: the write token's honest scope -- change on one group, the proxy is the boundary 2026-09-09 23:55:30 -06:00
models A/B result: ab-large loses to carc-tools; retire it. FINDINGS #30 records the day 2026-09-09 18:46:48 -06:00
openclaw A/B result: ab-large loses to carc-tools; retire it. FINDINGS #30 records the day 2026-09-09 18:46:48 -06:00
policies Move ops onto ab-moe to fix the open-ended-prompt tool loop (FINDINGS #26) 2026-08-27 01:37:22 -06:00
prompts Make main capable on ab-moe instead of delegating (FINDINGS #28) 2026-08-27 14:07:19 -06:00
rollback Isolate MCP gateway identities and remove unsafe snapshots 2026-08-28 16:37:47 -06:00
systemd Poller: own proxy-URL variable and a loopback-only guard 2026-09-09 17:45:20 -06:00
tests Poller: escalation tags, human feedback readback, and the weekly report 2026-09-09 23:50:21 -06:00
.gitignore Ignore .audit-runs/ (local eval output snapshots) 2026-09-09 12:23:20 -06:00
.gitlab-ci.yml Add bin/carc-install-system: stage the stack under a carc service account as system units 2026-09-09 12:39:48 -06:00
agents.yaml Retire nemotron-lightning; carc-tools is the batch draft model; stage the ab-large A/B 2026-09-09 17:06:53 -06:00
ARCHITECTURE.md Docs: escalation is a tag, feedback is three human tags, go-live steps and Zammad tag setup 2026-09-09 23:46:00 -06:00
FINDINGS.md A/B result: ab-large loses to carc-tools; retire it. FINDINGS #30 records the day 2026-09-09 18:46:48 -06:00
OPERATING-RULES.md Retire nemotron-lightning; carc-tools is the batch draft model; stage the ab-large A/B 2026-09-09 17:06:53 -06:00
README.md Docs: escalation is a tag, feedback is three human tags, go-live steps and Zammad tag setup 2026-09-09 23:46:00 -06:00

carc-agents

The home of CARC's agentic setup: local model serving, the gateway that fronts it, the security boundaries around it, and the agents that run on top. Nothing in this repo depends on a hosted model API.

The two hosts

The setup now runs across both CARC DGX Sparks (GB10, 128 GB unified memory each). The pooled-fabric plan in ARCHITECTURE.md is built out through Phase 3 (FINDINGS #23–24):

Host Address Runs
sparky-2 sparky-2.hpc.unm.edu (loopback + docker bridge) the LiteLLM gateway, the MCP gateway, the Zammad proxies + reverse tunnel, the OpenClaw agent container, the ticket poller. No model backend since 2026-09-09; the GPU is free. Everything in this repo's systemd/system/ units.
sparky-1 10.100.32.2, the dedicated DGX Spark interconnect only — never the LAN, still VLLM_API_KEY-gated the pooled inference backends: carc-tools, carc-fast, carc-embed, ab-moe.

Splitting the backends across the two boxes is what removed the GPU time-share constraint: the triage pipeline (carc-fast + ab-moe + carc-embed, all on sparky-1) and the sandbox (nemotron-lightning, alone on sparky-2) now run at the same time.

This repo's remote lives at git.repo.alliance.unm.edu/CARC/carc-agents. Its companion, CARC-knowledge-store, holds the knowledge base content, the retrieval pipeline, and the MCP gateway.

What we're building An AI assistant over Zammad tickets: answer the simple ones, escalate the complex ones, grounded in the CARC knowledge base and able to consult Grafana for diagnosis.
Where we are Serving stack, security boundaries, and the pooled two-host fabric are running. The interactive agent frontend is a self-hosted OpenClaw container (openclaw/, bin/carc-agent) that answers grounded ticket/KB/Grafana questions end to end; it replaced the NemoClaw sandbox on 2026-09-09. The Zammad trigger doesn't exist yet and will be built on agent/triage.py, not on the agent runtime.
Design rationale ARCHITECTURE.md
Research log FINDINGS.md — 30 numbered findings, chronological. The measurements behind most decisions here.

Layout

models/*.yaml        one vLLM backend spec per model (image, revision, GPU budget, flags, bind addr)
bin/                 launchers; every service has exactly one
bin/carc-openclaw-build     build the carc-openclaw image from both repos' tracked files (no secrets)
bin/carc-openclaw-run       run the agent container + squid egress proxy on an internal docker network
bin/carc-agent              one non-interactive agent turn: bin/carc-agent <agent> "<message>"
bin/carc-install-system     stage everything as system-scope units under the `carc` service account (sudo)
systemd/system/      system units for /etc/systemd/system (the target layout; see "System units")
systemd/             legacy user units — still what runs today until bin/carc-install-system --cutover
gateway/             LiteLLM config: one OpenAI endpoint over all backends
agents.yaml          the agent roster: models, tool allowlists, subagent policy (rendered into openclaw.json)
prompts/*.AGENTS.md  per-agent operating instructions, staged into each agent's workspace at start
openclaw/            Dockerfile, entrypoint (renders openclaw.json), static base config, squid allowlist
policies/*.yaml      legacy OPA egress presets for the retired NemoClaw sandbox (kept for reference)
bin/carc-sandbox-*   legacy NemoClaw sandbox tooling (retired 2026-09-09; see ARCHITECTURE "Known drift")
agent/triage.py      ticket triage pipeline: classify -> retrieve -> draft -> verify citations
evals/               retrieval and draft-quality evals against the real stack
keys/                gitignored secrets — see keys/README.md
rollback/            how to restore the pre-cutover ad-hoc vLLM

Service map

Everything binds loopback only, with two deliberate exceptions on the docker bridge address (172.18.0.1) so sandboxes can reach in, plus the sparky-1 backends on the private interconnect. Nothing is published on the LAN.

On sparky-2 (this repo's units)

Port Bind Service Unit
8000 127.0.0.1 LiteLLM gateway — the only entry point for model calls carc-litellm
8000 172.18.0.1 bridge forwarder to the above, for sandboxes carc-bridge-forward
8010 127.0.0.1 Zammad read-only egress proxy carc-zammad-proxy
8011 127.0.0.1 Zammad-AI shim — thinking-off + model pin in front of the gateway carc-zammad-ai-proxy
— outbound reverse SSH tunnel so Zammad's built-in AI reaches :8011 without a LAN bind carc-zammad-ai-tunnel
8443–8447 172.18.0.1 MCP gateway (HTTPS) — one port and bearer per identity; see below carc-gateway@<identity>

The exploratory user-level launcher assigns fixed endpoints: kb-researcher 8443, ticket-triage 8444, support 8445, ops 8446, and engineering 8447. Each gets its own keys/gateway-http-<identity>.key; this is local beta plumbing, not a prescribed production service layout. Only kb-researcher (8443) and ops (8446) are currently running; the other three have keys and port assignments but their carc-gateway@<identity> units are not started.

On sparky-1 (reached over the interconnect)

Port Service Unit
8001 vLLM carc-tools (Qwen3.8-27B NVFP4, 128k ctx) carc-vllm@carc-tools
8003 vLLM carc-fast (Qwen3-8B NVFP4, classify/route/summarize) carc-vllm@carc-fast
8004 vLLM carc-embed (Qwen3-Embedding-0.6B) carc-vllm@carc-embed
8012 vLLM ab-moe (Qwen3.6-35B-A3B NVFP4, 128k ctx) carc-vllm@ab-moe-nvfp4

Models registered with LiteLLM: carc-tools, carc-fast, carc-embed, ab-moe, zammad-assist (a carc-fast alias for the Zammad-AI shim). Check with:

curl -s -H "Authorization: Bearer $(cat keys/litellm.key)" \
  http://127.0.0.1:8000/v1/models | python3 -m json.tool

GPU budget

gpu_memory_utilization is a fraction of each box's 128 GB. Each model has one job; the roster was collapsed on 2026-09-09 (see "Which model does what").

Host Backend util job
sparky-1 carc-embed 0.05 KB vector search; needed by everything
sparky-1 carc-fast 0.12 Zammad summarize (60 s hard timeout), the pipeline's classify step
sparky-1 ab-moe 0.30 interactive agents (main, ops, kb-researcher, ticket-triage)
sparky-1 carc-tools 0.36 batch drafting in agent/triage.py and the poller
sparky-2 (none) — free since the A/B; see below

nemotron-lightning was retired 2026-09-09. Its NVIDIA-validated 0.65 recipe existed for the NemoClaw sandbox; once that was replaced, the model's only consumer was the unused coder agent.

Which model does what, and why

evals/draft_quality_eval.py --models carc-tools,ab-moe --repeat 2 on 2026-09-09: carc-tools 90% grounded / 100% escalate-correct at 27.6 s per case; ab-moe 70% / 80% at 7.9 s, every miss being the model degenerating under the JSON-schema grammar (~1 call in 4, FINDINGS #28). So the split is by who reads the output: a human reading an interactive answer gets ab-moe's speed and 131k window; the poller, where nobody is waiting and escalation correctness is the whole point, gets carc-tools.

The A/B: can one large model replace both? (run 2026-09-09: no)

With sparky-2 freed, Qwen3.5-122B-A10B NVFP4 (ab-large) ran alone there against carc-tools and ab-moe on evals/draft_quality_eval.py --repeat 2, same tickets, same retrieval, same prompts (FINDINGS #30):

model grounded escalate-correct unsupported citations fallbacks s/case
carc-tools 85% 95% 0 1 (classifier) 25.3
ab-moe 60% 80% 0 4 (grammar degeneration) 7.5
ab-large 90% 90% 4 0 27.5

The decision rule was "beats carc-tools on grounded and escalate-correct". It did not: it lost on escalation, it invented citations twice (paths retrieval never returned -- carc-tools has done that zero times across every eval run), and it was no faster than the 27B dense model. Retired the same evening; the spec is in git history. Model capacity is not what limits ticket answers here -- KB coverage and ab-moe's grammar bug are -- so tensor-parallel across the two Sparks stays off the table, and there is no larger NVFP4 model it would unlock anyway (the next size, 397B-A17B, is ~244 GB, more than both boxes). sparky-2 is free GPU again; the next use of it should be a KB-coverage or classifier improvement, not a bigger model.

Operating it

# what's up (run on whichever host)
systemctl --user list-units 'carc-*'
docker ps --filter label=carc.agents/managed=true

# logs
journalctl --user -u carc-litellm -f
journalctl --user -u carc-vllm@carc-embed -n 200

# restart a backend (engine init is slow: ~2 min, and the unit allows an hour)
systemctl --user restart carc-vllm@carc-fast

Adding a model is three steps: write models/<name>.yaml (set bind: to the interconnect address if it runs on sparky-1), add a model_list entry to gateway/litellm.config.yaml, then systemctl --user restart carc-litellm. Check the GPU budget table first.

Rotating keys is in keys/README.md. Restoring the pre-cutover setup is in rollback/README.md.

Install the daily smoke-test timer once:

cp systemd/carc-smoke-test.* ~/.config/systemd/user/ && \
  systemctl --user daemon-reload && \
  systemctl --user enable --now carc-smoke-test.timer

It fails loudly on purpose — see the comment in systemd/carc-smoke-test.service for how to read the last result.

Running an agent

# one non-interactive turn against the OpenClaw container (replaces `nemoclaw carc-claw agent`)
bin/carc-agent kb-researcher "how do I extend a job time limit?"
bin/carc-agent ops "how many Easley nodes are down? use count(up == 0)"
echo "<ticket text>" | bin/carc-agent ticket-triage -
docker exec carc-openclaw openclaw gateway health

# rebuild the image after editing agents.yaml, prompts/, openclaw/, or the knowledge store,
# then restart the container (the entrypoint re-renders openclaw.json and re-ingests the KB)
bin/carc-openclaw-build && systemctl --user restart carc-openclaw   # or plain systemctl after cutover

# the triage pipeline and the evals need numpy etc., which the system python3
# lacks -- run them with the knowledge store's venv interpreter. retrieval_eval
# also needs the KB_VLLM_* vars from the knowledge store's .env.carc to reach
# the vllm embedder (triage.py defaults them itself; retrieval_eval does not).
PY=~/CARC-knowledge-store/.venv/bin/python
$PY agent/triage.py "my job died with Killed, the node had 200GB free"

# evals
$PY evals/draft_quality_eval.py --repeat 1
(set -a; source ~/CARC-knowledge-store/.env.carc; set +a; $PY evals/retrieval_eval.py)

# fast live-stack verification (no writes to Zammad/Grafana)
bin/carc-smoke-test

The roster (agents.yaml): main, kb-researcher, ticket-triage, and ops, all on ab-moe. (coder and classifier were retired 2026-09-09; classification lives in the pipeline.) openclaw/entrypoint.py renders that file, prompts/*.AGENTS.md, and openclaw/openclaw.base.json into openclaw.json at every container start, so there is no separate "apply" step and nothing is lost on recreate.

The ticket loop

agent/poller.py is the trigger the target design was missing: it polls new HPC Support tickets through the read-only proxy, runs each through agent/triage.py (classify on carc-fast, retrieve, draft on carc-tools, verify citations), and posts the result as an internal note through the proxy's one write route, which forces internal:true server-side. Nothing it does is visible to a customer.

PY=~/CARC-knowledge-store/.venv/bin/python
$PY agent/poller.py --once --dry-run --limit 2 --verbose   # print what would be posted; nothing written
$PY agent/poller.py --once --write                  # post internal notes (or CARC_TRIAGE_WRITE=1)
$PY agent/poller.py --once --since 2026-09-09T00:00:00Z   # override the high-water mark
  • Dry run is the default. The system timer (carc-triage-poller.timer, every 10 minutes) stays in dry-run until CARC_TRIAGE_WRITE=1 is added to /etc/carc-agents/env. Read a week of dry-run notes in the journal first.
  • State (CARC_TRIAGE_STATE, /var/lib/carc/triage/state.sqlite under the unit) holds ticket ids, outcomes (drafted / escalated / error), reasons and cited paths. It never stores ticket text.
  • KB candidate queue (kb-queue.jsonl beside the state DB): one line per escalated ticket retrieval could not ground, with the search query and top paths only. This is where the knowledge store's next articles come from.
  • Every failure escalates. A model error, an unreadable ticket, or a proxy error records an error outcome for that ticket and the batch continues.
  • Tags are the escalation mechanism and the audit trail. After posting a note the poller tags the ticket through the proxy's second write route, POST /api/v1/tags/add, whose allowlist is fixed in the proxy: ai-drafted always, ai-escalated when it escalated, ai-no-kb-coverage when the ticket went to the KB queue. Nothing automated ever removes a tag. Staff close the loop with one of three tags applied in the Zammad UI -- ai-draft-used, ai-draft-edited, ai-draft-rejected -- which the poller reads back into its state on later runs (21-day window). agent/triage_report.py turns the state DB into the weekly tally: drafted/escalated per category, feedback counts, and the articles cited in rejected drafts, which is the correction signal for the knowledge store.
  • Going live is three steps, in this order: (1) provision keys/zammad-write.token (see keys/README.md, including creating the six tags in Zammad once); (2) sudo bin/carc-install-system so the key reaches /opt/carc, then sudo systemctl restart carc-zammad-proxy; (3) add CARC_TRIAGE_WRITE=1 to /etc/carc-agents/env. The next timer run posts notes and tags. Watch the first day's notes by hand.

The OpenClaw container

docker network carc-agents (--internal: no route off the bridge)
  ├─ carc-openclaw   OpenClaw 2026.7.1 + the knowledge store's MCP gateway
  │                  (kbstore, Zammad-MCP, mcp-grafana) as three stdio servers,
  │                  one per identity, audited to the state dir
  └─ carc-egress     squid; the only exit. openclaw/squid.conf allows exactly:
                     LiteLLM on the network gateway, carc-embed 10.100.32.2:8004,
                     support.alliance.unm.edu:443, vortex.alliance.unm.edu:443
host: carc-bridge-forward binds the network's gateway address -> LiteLLM :8000

The container runs as the invoking uid (read-only rootfs, all capabilities dropped), mounts keys/litellm.key, keys/vllm.key, and the knowledge store's gateway/credentials.yaml read-only under /run/secrets, and keeps its state (rendered openclaw.json, sessions, KB index, gateway audit DB, squid config) in CARC_STATE_DIR (default ~/.local/state/carc-openclaw). The image is built from git ls-files of both repos, so gitignored secrets cannot reach it. KB content (kb/) is bind-mounted from the knowledge-store checkout at run time and re-ingested at every start, so a docs re-import or an article correction reaches the agents with carc-install-system + systemctl restart carc-openclaw; only code and policy changes need bin/carc-openclaw-build.

System units

systemd/system/ is the target layout: system-scope units under a dedicated carc service account (docker group), checkouts at /opt/carc/, runtime knobs in /etc/carc-agents/env, no dependency on a login session or linger. bin/carc-install-system stages all of it; run it as root on each Spark:

sudo bin/carc-install-system --dry-run     # show the plan for this host's role
sudo bin/carc-install-system               # create carc, clone to /opt/carc, copy secrets, enable units
sudo bin/carc-install-system --cutover     # stop the user units, move the HF cache, start, smoke-test

After cutover every command in this README drops --user (systemctl status carc-litellm, journalctl -u carc-vllm@ab-moe-nvfp4).

Current state, honestly

Working and verified. All five vLLM backends (four on sparky-1, Nemotron on sparky-2), the LiteLLM gateway, the loopback/bridge/interconnect network boundaries, the read-only Zammad egress proxy, and the MCP gateway itself. mcp-zammad and mcp-grafana are installed and verified end-to-end against production Zammad and Grafana with dedicated service-account tokens. triage.py produces grounded, citation-verified drafts.

Working — the OpenClaw container. Verified 2026-09-09 with one live turn per identity, run concurrently: kb-researcher cited kb/policies/job-time-limit-extension-policy.md; ops queried production Prometheus (count(up == 0), real datasource UID) through the egress proxy; ticket-triage drafted a reply citing the OOM runbook with Hopper's real 2938 MB/CPU default. ~28 s per turn on ab-moe. The in-container gateway audit DB attributes every call per identity. The egress proxy refuses every host outside its allowlist and the agent network has no direct route out (both checked by the smoke test). Maximally open-ended cluster sweeps still don't converge (FINDINGS #26/#28); hand ops bounded questions.

Working — Zammad's built-in AI. Summarize through carc-zammad-ai-proxy + carc-zammad-ai-tunnel, verified in the Zammad 7.0.0 UI (FINDINGS #29). The shim requires its own inbound bearer and caps requests at 512 KiB before injecting the LiteLLM key. Units enabled; the tunnel's SSH key still lives on the Zammad host by hand. Suggest-reply quality on carc-fast not yet judged.

Retired. The NemoClaw/OpenShell sandbox (carc-claw) and its bin/carc-sandbox-provision / carc-sandbox-trust tooling. FINDINGS #12–#28 document why: recreate-survival, baked trust, an upstream sanitizer bug that blocked every backup, and a manifest that couldn't express MCP servers. The OpenClaw runtime itself carries over unchanged; only the wrapper is gone.

Known open. query_prometheus on an empty result set is logged as a JSONDecodeError deny in the gateway audit (the agent recovers; FINDINGS #20 noted the same formatting issue). ab-moe degenerates under a JSON-schema grammar about one call in four, which fails the ticket-triage eval gate (83% escalate-correct at --repeat 3); carc-tools is the working batch fallback for triage.py.

Not built yet. The Zammad trigger — nothing watches for incoming tickets. It will be a systemd timer around agent/triage.py posting internal notes through a forced-internal:true route on carc-zammad-proxy. See ARCHITECTURE.md.

Troubleshooting

Agent says it has no KB (or Zammad/Grafana) tools

For the OpenClaw container: docker logs carc-openclaw shows the rendered roster and the kb ingest result; docker exec -u 13 carc-egress tail /var/log/squid/access.log shows what the MCP backends tried to reach and whether squid allowed it; the audit DB is $CARC_STATE_DIR/carc/gateway_audit.sqlite. An un-prefixed tool name in agents.yaml is still silently stripped (OpenClaw namespaces MCP tools <server>__<tool>, FINDINGS #20).

The steps below are for the retired NemoClaw sandbox and are kept for anyone reading FINDINGS #16–#28.

Check in order:

  1. The sandbox's in-container MCP stack didn't survive a recreate. Run bin/carc-sandbox-provision --apply --verify — it re-applies the venv, the KB index, the stdio server registrations, the custom OPA presets, the AGENTS.md files, and the context-hardening knobs, then runs a grounded end-to-end check per identity. FINDINGS #22–26 for what each piece is.
  2. agents.yaml's tools.allow stripped them. It is an absolute allowlist that also removes MCP-bridged tools; the log says tool policy removed N tool(s) via agents.<id>.tools.allow. The real tool names are namespaced by server (carc-kb-stdio__kb_search), not kb_search (FINDINGS #20).
  3. The sandbox can't verify the gateway certificate — the case below.

Do not trust manual mcporter/curl reproduction from docker exec: OPA cannot attribute an exec'd process, so those requests are denied for unrelated reasons. Only a real agent turn tests the proxy path.

(Legacy sandbox) can't verify the gateway certificate

Confirm:

C=$(docker ps --format '{{.Names}}' | grep carc-claw)
docker exec $C curl -s -o /dev/null -w '%{http_code} %{ssl_verify_result}\n' \
  https://172.18.0.1:8443/mcp/

000 18 means DEPTH_ZERO_SELF_SIGNED_CERT — the cert isn't trusted. 401 with -k added confirms the network path is fine and it is purely trust. Compare fingerprints:

openssl x509 -in keys/gateway-cert.pem -noout -fingerprint -sha256
docker exec $C sh -c 'for f in /usr/local/share/ca-certificates/nemoclaw-corporate-ca-*.crt; do
  openssl x509 -in $f -noout -fingerprint -sha256; done'

If the live cert isn't in that list, the cert was regenerated after the sandbox was last built. There is no way to refresh sandbox trust without a rebuild, so that is the fix:

bin/carc-sandbox-trust              # show what would change
bin/carc-sandbox-trust --rebuild    # apply it

bin/carc-sandbox-trust exists because sandbox trust is baked at image build time, which makes "regenerate the gateway cert" and "rebuild the sandbox" one operation rather than two. It assembles the CA bundle from this host's anchors plus our gateway cert — never just ours, since an explicit NEMOCLAW_CORPORATE_CA_BUNDLE replaces NemoClaw's automatic host scan rather than merging with it, and dropping the FreeIPA ALLIANCE.UNM.EDU root would cost the sandbox its trust for Zammad and Grafana.

It also re-bakes NEMOCLAW_CONTEXT_WINDOW, which is baked at the same moment and wrong in the same way — reading the real max_model_len from the vLLM backend rather than a hand-copied number. See ARCHITECTURE.md.

This recurs every time the gateway cert is regenerated — which keys/README.md tells you to do on bridge-IP change or key rotation. A FreeIPA-issued cert retires the trap entirely (see the comment at the top of bin/carc-gateway-run).

(Legacy sandbox) bridge IP changed

If the openshell-docker docker network is recreated, 172.18.0.1 may move. bin/carc-bridge-forward and bin/carc-gateway-run both resolve it dynamically, but the TLS cert has the old IP in its SAN. Delete keys/gateway-*.pem, restart carc-gateway@*, then rebuild the sandbox.

A backend won't start

No available memory for the cache blocks means the gpu_memory_utilization fraction is too small for weights + activations + CUDA graphs + the whole KV pool, or something else is resident. vLLM checks the fraction against currently free memory, so a co-tenant backend's overhead counts. Check docker ps against the GPU budget table above. FINDINGS #1 and #24 have the sizing math.