- Python 93.1%
- Shell 6.1%
- Dockerfile 0.8%
|
|
||
|---|---|---|
| agent | ||
| bin | ||
| evals | ||
| gateway | ||
| keys | ||
| models | ||
| openclaw | ||
| policies | ||
| prompts | ||
| rollback | ||
| systemd | ||
| tests | ||
| .gitignore | ||
| .gitlab-ci.yml | ||
| agents.yaml | ||
| ARCHITECTURE.md | ||
| FINDINGS.md | ||
| OPERATING-RULES.md | ||
| README.md | ||
carc-agents
The home of CARC's agentic setup: local model serving, the gateway that fronts it, the security boundaries around it, and the agents that run on top. Nothing in this repo depends on a hosted model API.
The two hosts
The setup now runs across both CARC DGX Sparks (GB10, 128 GB unified memory each). The pooled-fabric plan in ARCHITECTURE.md is built out through Phase 3 (FINDINGS #23–24):
| Host | Address | Runs |
|---|---|---|
| sparky-2 | sparky-2.hpc.unm.edu (loopback + docker bridge) |
the LiteLLM gateway, the MCP gateway, the Zammad proxies + reverse tunnel, the OpenClaw agent container, the ticket poller. No model backend since 2026-09-09; the GPU is free. Everything in this repo's systemd/system/ units. |
| sparky-1 | 10.100.32.2, the dedicated DGX Spark interconnect only — never the LAN, still VLLM_API_KEY-gated |
the pooled inference backends: carc-tools, carc-fast, carc-embed, ab-moe. |
Splitting the backends across the two boxes is what removed the GPU
time-share constraint: the triage pipeline (carc-fast + ab-moe +
carc-embed, all on sparky-1) and the sandbox (nemotron-lightning, alone on
sparky-2) now run at the same time.
This repo's remote lives at
git.repo.alliance.unm.edu/CARC/carc-agents.
Its companion,
CARC-knowledge-store,
holds the knowledge base content, the retrieval pipeline, and the MCP gateway.
| What we're building | An AI assistant over Zammad tickets: answer the simple ones, escalate the complex ones, grounded in the CARC knowledge base and able to consult Grafana for diagnosis. |
| Where we are | Serving stack, security boundaries, and the pooled two-host fabric are running. The interactive agent frontend is a self-hosted OpenClaw container (openclaw/, bin/carc-agent) that answers grounded ticket/KB/Grafana questions end to end; it replaced the NemoClaw sandbox on 2026-09-09. The Zammad trigger doesn't exist yet and will be built on agent/triage.py, not on the agent runtime. |
| Design rationale | ARCHITECTURE.md |
| Research log | FINDINGS.md — 30 numbered findings, chronological. The measurements behind most decisions here. |
Layout
models/*.yaml one vLLM backend spec per model (image, revision, GPU budget, flags, bind addr)
bin/ launchers; every service has exactly one
bin/carc-openclaw-build build the carc-openclaw image from both repos' tracked files (no secrets)
bin/carc-openclaw-run run the agent container + squid egress proxy on an internal docker network
bin/carc-agent one non-interactive agent turn: bin/carc-agent <agent> "<message>"
bin/carc-install-system stage everything as system-scope units under the `carc` service account (sudo)
systemd/system/ system units for /etc/systemd/system (the target layout; see "System units")
systemd/ legacy user units — still what runs today until bin/carc-install-system --cutover
gateway/ LiteLLM config: one OpenAI endpoint over all backends
agents.yaml the agent roster: models, tool allowlists, subagent policy (rendered into openclaw.json)
prompts/*.AGENTS.md per-agent operating instructions, staged into each agent's workspace at start
openclaw/ Dockerfile, entrypoint (renders openclaw.json), static base config, squid allowlist
policies/*.yaml legacy OPA egress presets for the retired NemoClaw sandbox (kept for reference)
bin/carc-sandbox-* legacy NemoClaw sandbox tooling (retired 2026-09-09; see ARCHITECTURE "Known drift")
agent/triage.py ticket triage pipeline: classify -> retrieve -> draft -> verify citations
evals/ retrieval and draft-quality evals against the real stack
keys/ gitignored secrets — see keys/README.md
rollback/ how to restore the pre-cutover ad-hoc vLLM
Service map
Everything binds loopback only, with two deliberate exceptions on the
docker bridge address (172.18.0.1) so sandboxes can reach in, plus the
sparky-1 backends on the private interconnect. Nothing is published on the LAN.
On sparky-2 (this repo's units)
| Port | Bind | Service | Unit |
|---|---|---|---|
| 8000 | 127.0.0.1 |
LiteLLM gateway — the only entry point for model calls | carc-litellm |
| 8000 | 172.18.0.1 |
bridge forwarder to the above, for sandboxes | carc-bridge-forward |
| 8010 | 127.0.0.1 |
Zammad read-only egress proxy | carc-zammad-proxy |
| 8011 | 127.0.0.1 |
Zammad-AI shim — thinking-off + model pin in front of the gateway | carc-zammad-ai-proxy |
| — | outbound | reverse SSH tunnel so Zammad's built-in AI reaches :8011 without a LAN bind |
carc-zammad-ai-tunnel |
| 8443–8447 | 172.18.0.1 |
MCP gateway (HTTPS) — one port and bearer per identity; see below | carc-gateway@<identity> |
The exploratory user-level launcher assigns fixed endpoints: kb-researcher
8443, ticket-triage 8444, support 8445, ops 8446, and engineering
8447. Each gets its own keys/gateway-http-<identity>.key; this is local beta
plumbing, not a prescribed production service layout. Only kb-researcher
(8443) and ops (8446) are currently running; the other three have keys and
port assignments but their carc-gateway@<identity> units are not started.
On sparky-1 (reached over the interconnect)
| Port | Service | Unit |
|---|---|---|
| 8001 | vLLM carc-tools (Qwen3.8-27B NVFP4, 128k ctx) |
carc-vllm@carc-tools |
| 8003 | vLLM carc-fast (Qwen3-8B NVFP4, classify/route/summarize) |
carc-vllm@carc-fast |
| 8004 | vLLM carc-embed (Qwen3-Embedding-0.6B) |
carc-vllm@carc-embed |
| 8012 | vLLM ab-moe (Qwen3.6-35B-A3B NVFP4, 128k ctx) |
carc-vllm@ab-moe-nvfp4 |
Models registered with LiteLLM: carc-tools, carc-fast, carc-embed,
ab-moe, zammad-assist (a carc-fast alias for the Zammad-AI shim). Check with:
curl -s -H "Authorization: Bearer $(cat keys/litellm.key)" \
http://127.0.0.1:8000/v1/models | python3 -m json.tool
GPU budget
gpu_memory_utilization is a fraction of each box's 128 GB. Each model has one
job; the roster was collapsed on 2026-09-09 (see "Which model does what").
| Host | Backend | util | job |
|---|---|---|---|
| sparky-1 | carc-embed |
0.05 | KB vector search; needed by everything |
| sparky-1 | carc-fast |
0.12 | Zammad summarize (60 s hard timeout), the pipeline's classify step |
| sparky-1 | ab-moe |
0.30 | interactive agents (main, ops, kb-researcher, ticket-triage) |
| sparky-1 | carc-tools |
0.36 | batch drafting in agent/triage.py and the poller |
| sparky-2 | (none) | — | free since the A/B; see below |
nemotron-lightning was retired 2026-09-09. Its NVIDIA-validated 0.65 recipe
existed for the NemoClaw sandbox; once that was replaced, the model's only
consumer was the unused coder agent.
Which model does what, and why
evals/draft_quality_eval.py --models carc-tools,ab-moe --repeat 2 on
2026-09-09: carc-tools 90% grounded / 100% escalate-correct at 27.6 s per
case; ab-moe 70% / 80% at 7.9 s, every miss being the model degenerating under
the JSON-schema grammar (~1 call in 4, FINDINGS #28). So the split is by
who reads the output: a human reading an interactive answer gets ab-moe's
speed and 131k window; the poller, where nobody is waiting and escalation
correctness is the whole point, gets carc-tools.
The A/B: can one large model replace both? (run 2026-09-09: no)
With sparky-2 freed, Qwen3.5-122B-A10B NVFP4 (ab-large) ran alone there
against carc-tools and ab-moe on evals/draft_quality_eval.py --repeat 2,
same tickets, same retrieval, same prompts (FINDINGS #30):
| model | grounded | escalate-correct | unsupported citations | fallbacks | s/case |
|---|---|---|---|---|---|
| carc-tools | 85% | 95% | 0 | 1 (classifier) | 25.3 |
| ab-moe | 60% | 80% | 0 | 4 (grammar degeneration) | 7.5 |
| ab-large | 90% | 90% | 4 | 0 | 27.5 |
The decision rule was "beats carc-tools on grounded and escalate-correct". It did not: it lost on escalation, it invented citations twice (paths retrieval never returned -- carc-tools has done that zero times across every eval run), and it was no faster than the 27B dense model. Retired the same evening; the spec is in git history. Model capacity is not what limits ticket answers here -- KB coverage and ab-moe's grammar bug are -- so tensor-parallel across the two Sparks stays off the table, and there is no larger NVFP4 model it would unlock anyway (the next size, 397B-A17B, is ~244 GB, more than both boxes). sparky-2 is free GPU again; the next use of it should be a KB-coverage or classifier improvement, not a bigger model.
Operating it
# what's up (run on whichever host)
systemctl --user list-units 'carc-*'
docker ps --filter label=carc.agents/managed=true
# logs
journalctl --user -u carc-litellm -f
journalctl --user -u carc-vllm@carc-embed -n 200
# restart a backend (engine init is slow: ~2 min, and the unit allows an hour)
systemctl --user restart carc-vllm@carc-fast
Adding a model is three steps: write models/<name>.yaml (set bind: to the
interconnect address if it runs on sparky-1), add a model_list entry to
gateway/litellm.config.yaml, then systemctl --user restart carc-litellm.
Check the GPU budget table first.
Rotating keys is in keys/README.md. Restoring the
pre-cutover setup is in rollback/README.md.
Install the daily smoke-test timer once:
cp systemd/carc-smoke-test.* ~/.config/systemd/user/ && \
systemctl --user daemon-reload && \
systemctl --user enable --now carc-smoke-test.timer
It fails loudly on purpose — see the comment in
systemd/carc-smoke-test.service for how
to read the last result.
Running an agent
# one non-interactive turn against the OpenClaw container (replaces `nemoclaw carc-claw agent`)
bin/carc-agent kb-researcher "how do I extend a job time limit?"
bin/carc-agent ops "how many Easley nodes are down? use count(up == 0)"
echo "<ticket text>" | bin/carc-agent ticket-triage -
docker exec carc-openclaw openclaw gateway health
# rebuild the image after editing agents.yaml, prompts/, openclaw/, or the knowledge store,
# then restart the container (the entrypoint re-renders openclaw.json and re-ingests the KB)
bin/carc-openclaw-build && systemctl --user restart carc-openclaw # or plain systemctl after cutover
# the triage pipeline and the evals need numpy etc., which the system python3
# lacks -- run them with the knowledge store's venv interpreter. retrieval_eval
# also needs the KB_VLLM_* vars from the knowledge store's .env.carc to reach
# the vllm embedder (triage.py defaults them itself; retrieval_eval does not).
PY=~/CARC-knowledge-store/.venv/bin/python
$PY agent/triage.py "my job died with Killed, the node had 200GB free"
# evals
$PY evals/draft_quality_eval.py --repeat 1
(set -a; source ~/CARC-knowledge-store/.env.carc; set +a; $PY evals/retrieval_eval.py)
# fast live-stack verification (no writes to Zammad/Grafana)
bin/carc-smoke-test
The roster (agents.yaml): main, kb-researcher, ticket-triage, and ops,
all on ab-moe. (coder and classifier were retired 2026-09-09; classification
lives in the pipeline.)
openclaw/entrypoint.py renders that file, prompts/*.AGENTS.md, and
openclaw/openclaw.base.json into openclaw.json at every container start, so
there is no separate "apply" step and nothing is lost on recreate.
The ticket loop
agent/poller.py is the trigger the target design was missing: it polls new
HPC Support tickets through the read-only proxy, runs each through
agent/triage.py (classify on carc-fast, retrieve, draft on carc-tools,
verify citations), and posts the result as an internal note through the
proxy's one write route, which forces internal:true server-side. Nothing it
does is visible to a customer.
PY=~/CARC-knowledge-store/.venv/bin/python
$PY agent/poller.py --once --dry-run --limit 2 --verbose # print what would be posted; nothing written
$PY agent/poller.py --once --write # post internal notes (or CARC_TRIAGE_WRITE=1)
$PY agent/poller.py --once --since 2026-09-09T00:00:00Z # override the high-water mark
- Dry run is the default. The system timer (
carc-triage-poller.timer, every 10 minutes) stays in dry-run untilCARC_TRIAGE_WRITE=1is added to/etc/carc-agents/env. Read a week of dry-run notes in the journal first. - State (
CARC_TRIAGE_STATE,/var/lib/carc/triage/state.sqliteunder the unit) holds ticket ids, outcomes (drafted/escalated/error), reasons and cited paths. It never stores ticket text. - KB candidate queue (
kb-queue.jsonlbeside the state DB): one line per escalated ticket retrieval could not ground, with the search query and top paths only. This is where the knowledge store's next articles come from. - Every failure escalates. A model error, an unreadable ticket, or a proxy
error records an
erroroutcome for that ticket and the batch continues. - Tags are the escalation mechanism and the audit trail. After posting a
note the poller tags the ticket through the proxy's second write route,
POST /api/v1/tags/add, whose allowlist is fixed in the proxy:ai-draftedalways,ai-escalatedwhen it escalated,ai-no-kb-coveragewhen the ticket went to the KB queue. Nothing automated ever removes a tag. Staff close the loop with one of three tags applied in the Zammad UI --ai-draft-used,ai-draft-edited,ai-draft-rejected-- which the poller reads back into its state on later runs (21-day window).agent/triage_report.pyturns the state DB into the weekly tally: drafted/escalated per category, feedback counts, and the articles cited in rejected drafts, which is the correction signal for the knowledge store. - Going live is three steps, in this order: (1) provision
keys/zammad-write.token(seekeys/README.md, including creating the six tags in Zammad once); (2)sudo bin/carc-install-systemso the key reaches/opt/carc, thensudo systemctl restart carc-zammad-proxy; (3) addCARC_TRIAGE_WRITE=1to/etc/carc-agents/env. The next timer run posts notes and tags. Watch the first day's notes by hand.
The OpenClaw container
docker network carc-agents (--internal: no route off the bridge)
├─ carc-openclaw OpenClaw 2026.7.1 + the knowledge store's MCP gateway
│ (kbstore, Zammad-MCP, mcp-grafana) as three stdio servers,
│ one per identity, audited to the state dir
└─ carc-egress squid; the only exit. openclaw/squid.conf allows exactly:
LiteLLM on the network gateway, carc-embed 10.100.32.2:8004,
support.alliance.unm.edu:443, vortex.alliance.unm.edu:443
host: carc-bridge-forward binds the network's gateway address -> LiteLLM :8000
The container runs as the invoking uid (read-only rootfs, all capabilities
dropped), mounts keys/litellm.key, keys/vllm.key, and the knowledge store's
gateway/credentials.yaml read-only under /run/secrets, and keeps its state
(rendered openclaw.json, sessions, KB index, gateway audit DB, squid config)
in CARC_STATE_DIR (default ~/.local/state/carc-openclaw). The image is built
from git ls-files of both repos, so gitignored secrets cannot reach it. KB
content (kb/) is bind-mounted from the knowledge-store checkout at run
time and re-ingested at every start, so a docs re-import or an article
correction reaches the agents with carc-install-system + systemctl restart carc-openclaw; only code and policy changes need bin/carc-openclaw-build.
System units
systemd/system/ is the target layout: system-scope units under a dedicated
carc service account (docker group), checkouts at /opt/carc/, runtime knobs
in /etc/carc-agents/env, no dependency on a login session or linger.
bin/carc-install-system stages all of it; run it as root on each Spark:
sudo bin/carc-install-system --dry-run # show the plan for this host's role
sudo bin/carc-install-system # create carc, clone to /opt/carc, copy secrets, enable units
sudo bin/carc-install-system --cutover # stop the user units, move the HF cache, start, smoke-test
After cutover every command in this README drops --user
(systemctl status carc-litellm, journalctl -u carc-vllm@ab-moe-nvfp4).
Current state, honestly
Working and verified. All five vLLM backends (four on sparky-1, Nemotron on
sparky-2), the LiteLLM gateway, the loopback/bridge/interconnect network
boundaries, the read-only Zammad egress proxy, and the MCP gateway itself.
mcp-zammad and mcp-grafana are installed and verified end-to-end against
production Zammad and Grafana with dedicated service-account tokens.
triage.py produces grounded, citation-verified drafts.
Working — the OpenClaw container. Verified 2026-09-09 with one live turn per
identity, run concurrently: kb-researcher cited
kb/policies/job-time-limit-extension-policy.md; ops queried production
Prometheus (count(up == 0), real datasource UID) through the egress proxy;
ticket-triage drafted a reply citing the OOM runbook with Hopper's real
2938 MB/CPU default. ~28 s per turn on ab-moe. The in-container gateway
audit DB attributes every call per identity. The egress proxy refuses every
host outside its allowlist and the agent network has no direct route out (both
checked by the smoke test). Maximally open-ended cluster sweeps still don't
converge (FINDINGS #26/#28); hand ops bounded questions.
Working — Zammad's built-in AI. Summarize through
carc-zammad-ai-proxy + carc-zammad-ai-tunnel, verified in the Zammad 7.0.0
UI (FINDINGS #29). The shim requires its own inbound bearer and caps requests
at 512 KiB before injecting the LiteLLM key. Units enabled; the tunnel's SSH
key still lives on the Zammad host by hand. Suggest-reply quality on
carc-fast not yet judged.
Retired. The NemoClaw/OpenShell sandbox (carc-claw) and its
bin/carc-sandbox-provision / carc-sandbox-trust tooling. FINDINGS #12–#28
document why: recreate-survival, baked trust, an upstream sanitizer bug that
blocked every backup, and a manifest that couldn't express MCP servers. The
OpenClaw runtime itself carries over unchanged; only the wrapper is gone.
Known open. query_prometheus on an empty result set is logged as a
JSONDecodeError deny in the gateway audit (the agent recovers; FINDINGS #20
noted the same formatting issue). ab-moe degenerates under a JSON-schema
grammar about one call in four, which fails the ticket-triage eval gate
(83% escalate-correct at --repeat 3); carc-tools is the working batch
fallback for triage.py.
Not built yet. The Zammad trigger — nothing watches for incoming tickets.
It will be a systemd timer around agent/triage.py posting internal notes
through a forced-internal:true route on carc-zammad-proxy. See
ARCHITECTURE.md.
Troubleshooting
Agent says it has no KB (or Zammad/Grafana) tools
For the OpenClaw container: docker logs carc-openclaw shows the rendered
roster and the kb ingest result; docker exec -u 13 carc-egress tail /var/log/squid/access.log shows what the MCP backends tried to reach and
whether squid allowed it; the audit DB is $CARC_STATE_DIR/carc/gateway_audit.sqlite.
An un-prefixed tool name in agents.yaml is still silently stripped
(OpenClaw namespaces MCP tools <server>__<tool>, FINDINGS #20).
The steps below are for the retired NemoClaw sandbox and are kept for anyone reading FINDINGS #16–#28.
Check in order:
- The sandbox's in-container MCP stack didn't survive a recreate. Run
bin/carc-sandbox-provision --apply --verify— it re-applies the venv, the KB index, the stdio server registrations, the custom OPA presets, theAGENTS.mdfiles, and the context-hardening knobs, then runs a grounded end-to-end check per identity. FINDINGS #22–26 for what each piece is. agents.yaml'stools.allowstripped them. It is an absolute allowlist that also removes MCP-bridged tools; the log saystool policy removed N tool(s) via agents.<id>.tools.allow. The real tool names are namespaced by server (carc-kb-stdio__kb_search), notkb_search(FINDINGS #20).- The sandbox can't verify the gateway certificate — the case below.
Do not trust manual mcporter/curl reproduction from docker exec: OPA
cannot attribute an exec'd process, so those requests are denied for unrelated
reasons. Only a real agent turn tests the proxy path.
(Legacy sandbox) can't verify the gateway certificate
Confirm:
C=$(docker ps --format '{{.Names}}' | grep carc-claw)
docker exec $C curl -s -o /dev/null -w '%{http_code} %{ssl_verify_result}\n' \
https://172.18.0.1:8443/mcp/
000 18 means DEPTH_ZERO_SELF_SIGNED_CERT — the cert isn't trusted. 401
with -k added confirms the network path is fine and it is purely trust.
Compare fingerprints:
openssl x509 -in keys/gateway-cert.pem -noout -fingerprint -sha256
docker exec $C sh -c 'for f in /usr/local/share/ca-certificates/nemoclaw-corporate-ca-*.crt; do
openssl x509 -in $f -noout -fingerprint -sha256; done'
If the live cert isn't in that list, the cert was regenerated after the sandbox was last built. There is no way to refresh sandbox trust without a rebuild, so that is the fix:
bin/carc-sandbox-trust # show what would change
bin/carc-sandbox-trust --rebuild # apply it
bin/carc-sandbox-trust exists because sandbox trust is baked at image build
time, which makes "regenerate the gateway cert" and "rebuild the sandbox" one
operation rather than two. It assembles the CA bundle from this host's anchors
plus our gateway cert — never just ours, since an explicit
NEMOCLAW_CORPORATE_CA_BUNDLE replaces NemoClaw's automatic host scan rather
than merging with it, and dropping the FreeIPA ALLIANCE.UNM.EDU root would
cost the sandbox its trust for Zammad and Grafana.
It also re-bakes NEMOCLAW_CONTEXT_WINDOW, which is baked at the same moment
and wrong in the same way — reading the real max_model_len from the vLLM
backend rather than a hand-copied number. See
ARCHITECTURE.md.
This recurs every time the gateway cert is regenerated — which
keys/README.md tells you to do on bridge-IP change or key rotation. A
FreeIPA-issued cert retires the trap entirely (see the comment at the top of
bin/carc-gateway-run).
(Legacy sandbox) bridge IP changed
If the openshell-docker docker network is recreated, 172.18.0.1 may move.
bin/carc-bridge-forward and bin/carc-gateway-run both resolve it
dynamically, but the TLS cert has the old IP in its SAN. Delete
keys/gateway-*.pem, restart carc-gateway@*, then rebuild the sandbox.
A backend won't start
No available memory for the cache blocks means the gpu_memory_utilization
fraction is too small for weights + activations + CUDA graphs + the whole KV
pool, or something else is resident. vLLM checks the fraction against
currently free memory, so a co-tenant backend's overhead counts. Check
docker ps against the GPU budget table above. FINDINGS #1 and #24 have the
sizing math.