
I don’t just use AI tools. I build the plumbing that lets them do real work safely across an MSP’s stack: ticketing, RMM, endpoint security, Microsoft 365 and backups. This page covers the harnesses I run, how they connect, how my self-hosted agents remember, and the rules that keep all of it reversible.
Design rule for every agent I build: read-only by default, and nothing changes without a way to undo it.
The harnesses I use
I don’t bet on one AI product. I run five harnesses side by side, two of them self-hosted, and pick the one that fits the job. They all follow the same rules and leave the same paper trail.
OpenClaw
What it is: a self-hosted agent platform that runs a small team of agents: a platform agent that owns cross-project design and coordination, a lightweight worker for small, reversible fixes, and a research sub-agent for web research and source gathering.
Where it runs: my Linux lab host, openclaw01, reachable only over Tailscale. It uses open-weight models through OpenCode Go, such as MiniMax M2.7 and DeepSeek V4.
What I use it for: building and maintaining my internal operations tooling: designing systems, wiring up automations, debugging what broke and keeping the documentation good enough to pick work back up after an interruption.
Hermes
What it is: the open-source Hermes Agent, running as my operator and orchestrator. It takes a request, does the work with real tools and reports back what actually happened. I added my own memory and embeddings build on top of it.
Where it runs: the same lab host as OpenClaw. Everyday turns run on fast models such as DeepSeek V4 Flash and MiniMax M2.7 through OpenCode Go; harder problems go to GPT 5.5.
What I use it for: delegating coding tasks to other agents, triaging and drafting replies for support tickets, cited technical research, and the scheduled jobs for both agents: backups, health reports and Microsoft Entra cleanup.
Claude: Chat, Cowork and Code
What it is: Anthropic’s assistant in three forms: Chat for quick questions and drafting, Cowork for agentic work on my desktop, and Code for coding in the terminal.
Where it runs: Cowork runs on my Windows PC and reaches the lab host over SSH. Claude is also the assistant my MSP team uses every day.
What I use it for: it built this site and has written most of its Lab Notes. With the MCP servers and skills below, my techs and I use it for ticket triage, Microsoft 365 breach investigations and Google Workspace investigations through its built-in browser.
ChatGPT: Chat, Work and Codex
What it is: OpenAI’s assistant in three forms: Chat for quick questions and drafting, Work for agentic tasks and Codex for coding.
Where it runs: OpenAI’s apps, reaching the lab host over SSH like the other harnesses.
What I use it for: Codex rebuilt this site’s snapshots so unchanged files are shared and storage stays bounded, and it automated the photo counts on my sim rig galleries. ChatGPT Work competes in my bake-offs.
OpenCode
What it is: an open-source coding agent that isn’t tied to one model vendor.
Where it runs: a terminal, reaching the lab host over SSH. I run it with several models: DeepSeek v4.1, Qwen 3.8 Flash and MiMo v2.6 Flash have each entered a bake-off.
What I use it for: testing what cheaper and open models can do on real work, under exactly the same rules as the big commercial harnesses.
How they connect
Five harnesses could easily mean five ways of doing everything. Instead they share a few common layers: shell access ties the local side together, and Composio ties the external side together.
| Layer | What it does |
|---|---|
| Composio | One integration layer for every harness: Google Workspace, Microsoft 365, Freshdesk, GitHub, Asana, Notion and web search. Each agent also gets its own AgentMail inbox through it for agent-to-agent email. |
| Shell access | PowerShell or bash in every harness, so an agent can run the same diagnostics I would, like pulling the Windows event logs from the time something broke. |
| SSH to the lab host | ChatGPT, Claude and OpenCode reach openclaw01 over SSH on Tailscale to administer it, keep it simple to run and run the backups of OpenClaw and Hermes themselves. |
| MCP servers | Scoped, mostly read-only access to the MSP tools my techs use every day (listed below). |
| Obsidian vault | A shared knowledge base both self-hosted agents read and write. The lab host’s documentation lives there as a source of record any harness can read. |
I treat every Composio connection like an admin credential: each one exists for a clear, documented purpose, because an email sent or a ticket created through it is real.
How my self-hosted agents remember
A model doesn’t remember anything between conversations. The memory lives in the agent: before each turn it assembles what the model needs from stores I can inspect, search and back up. OpenClaw and Hermes solve this differently.
OpenClaw: a stack of five layers
- Working memory: the current conversation, instructions and recent tool results. The most detailed layer, but it only lasts as long as the context window.
- Memory files: plain Markdown I can read and edit: a curated long-term file for stable facts, preferences and decisions, plus daily notes for recent continuity.
- QMD search: a local search engine over those files that combines keyword and semantic search with reranking and query expansion.
- clawmem: a durable knowledge store that keeps facts, decisions, preferences, instructions and entities as small individual documents. Deletes are soft, so mistakes are recoverable.
- lossless-claw: when a conversation gets long, it compacts older turns into linked summaries instead of dropping them, and can expand a summary back into the original detail on demand.
The rule of thumb: QMD searches files, clawmem retrieves durable knowledge, and lossless-claw recovers old conversation. Before older context is compacted, the agent can flush important facts to its memory files, which prevents the worst failure in long agent sessions: silently forgetting what was decided earlier.
Hermes: a session store and a fact store
- Session store: every message from every conversation, with a keyword index, a trigram index that tolerates partial words and typos, a vector embedding per session and a pre-computed summary per session. Raw transcripts age out, but the searchable record stays.
- Fact store: thousands of durable facts about me and my environment. Facts come from a background review of recent turns and, at the end of each session, from pattern matching on things like “I prefer” and “we decided” plus an LLM pass over the transcript. Named entities are linked across facts, and each fact carries a trust score that rises or falls with feedback.
- Memory files: short, size-capped notes about me and the environment, mirrored into the fact store.
When Hermes searches its past, it checks the fact store first for instant answers, reuses cached session summaries, then runs keyword, trigram and vector search together. An LLM only summarizes a conversation when no cached summary exists.
Local search models and backups
The memory search runs on small models on the lab host’s own GPU: embeddinggemma-300M for embeddings, Qwen3-1.7B for query expansion and Qwen3-Reranker-0.6B for reranking. Embedding and ranking stay on my own hardware and cost nothing per query.
Both agents’ memory, configuration, skills and custom code are backed up automatically. A backup only counts once it passes its checks: checksums, a completion marker written last, and database integrity checks. Because the vectors travel inside the databases, the memory can move to new hardware without re-embedding anything, and my custom memory code travels as a patch that I’ve verified rebuilds an identical code tree. A test restore matched the live databases row for row.
How I work
This site is my proving ground. Every harness works on it under the same rules, and every result is recorded here.
- One rulebook in Git: every harness reads the same rules before it touches anything. Harnesses without SSH get the same rules through the WordPress MCP.
- Snapshot before each change: a database and file snapshot comes before every write, no exceptions.
- Typed commits: one logical change per commit, labeled with its type and the harness that made it. Commits without a harness name are refused.
- A Lab Note per session: what I asked for, what changed, how, and an honest account of what worked and what didn’t. Read them on the Lab Notes page.
- Rollback always available: any change can be reverted to an earlier commit or snapshot, and every rollback is itself a recorded commit. History is never rewritten.
- Read-only by default: most of my MCP servers are read-only, and my self-hosted agents follow four change-control levels: read-only work needs no approval, reversible work gets a plan first, system-impacting changes need my explicit approval and a rollback, and anything destructive or external, like sending an email, needs my approval every time.
The Harness Scoreboard ranks every harness on changes, Lab Notes, rollback rate and guideline compliance, computed straight from Git. It keeps everyone honest, Claude included: its first run scored Claude at 92% on snapshots because a few back-to-back changes had shared one. Claude still does most of the work here and writes the most Lab Notes, while Codex and the OpenCode models hold strong compliance on smaller samples.
To compare harnesses and models fairly I run harness bake-offs: the same task, rules and starting point, one contestant at a time, with process scored automatically from Git. In Round 1, a Featured Projects page, five contestants (Claude with Opus 5.5 High, ChatGPT with Sol 6.1 High, and OpenCode with DeepSeek v4.1 High, Qwen 3.8 Flash Medium and MiMo v2.6 Flash) each built the page in under ten minutes. Four earned full process points; one lost points for editing another contestant’s page, which is exactly the kind of slip the scoring is built to catch. This page is Round 2.
The Lab Notes also show where people still matter. When a layout on my Experience page didn’t line up, I caught it, not the AI: its screenshot check confirmed the page rendered, but nobody was judging alignment until I looked.
What I’ve built with them
MCP servers
Model Context Protocol servers that give Claude structured, scoped access to the tools our techs use every day. Most are read-only by design.
| Server | What it does |
|---|---|
| Microsoft 365 Management | Tenant administration with preview-then-execute changes |
| M365 Security Investigation | Read-only sign-in, audit and mailbox forensics for account breach investigations |
| N-sight RMM | Read-only device, check, patch and backup data from N-able N-sight |
| WatchGuard EPDR | Endpoint protection status, security events, risk and patch posture |
| Freshdesk | Ticket triage, private notes and knowledge base access |
Agent skills
- Freshdesk ticket triage (the team’s most-used): investigates alerts across M365, endpoints, EDR and RMM, writes a clean private note and closes informational tickets
- M365 breach report: turns an affected account and a rough timeframe into a structured compromise investigation
- Google Workspace: safe editing of Docs, Sheets and Slides
- Morning brief and inbox sweep: executive-assistant style daily briefs and inbox triage
I also wrote a training guide for my techs on using Claude day to day. Its core principle: AI amplifies what you already know, so understand the process first, and verify everything before it reaches a client.
Automation pipelines
- n8n phishing triage: Freshdesk webhook triggers analysis and filing of reported phishing email
- Weekly backup review: collects the week’s backup alert tickets from Freshdesk, generates a report and flags the backup jobs that need a closer look
- AppSheet approval and filing app
- Composio integration giving each tech’s Claude account API access across the tool stack
Self-hosted agent lab
- Hermes and OpenClaw agents on a Linux host reachable only over Tailscale
- Scheduled jobs for Entra cleanup, PIM digest cleanup, backup reviews, backups and memory backfill
- GPU-backed memory search and a vault-based portable skill framework
- Local models with LM Studio, Ollama and CUDA builds of llama.cpp, plus ComfyUI workflows
- This WordPress site, run in Docker and managed by AI agents with snapshot, change, verify and rollback on every write
Vendor assessment engine
A reusable engine I built with AI agents to answer vendor security questionnaires: the kind a bank sends as a supplier-controls assessment. Instead of re-answering hundreds of controls from scratch, it drafts answers from a stored knowledge base and cites the evidence behind each one.
- Knowledge base: atomic facts about a client plus a crosswalk mapping any vendor question to the facts that answer it (300+ controls carried over from prior assessments)
- Semantic matching: local sentence embeddings match a brand-new questionnaire to the known controls and draft an answer
- Evidence retrieval: the client’s policies and provider attestations (SOC 2, ISO 27001) are indexed so every answer can quote its source
- Review queue: anything unsupported or unknown is flagged for a person instead of guessed, and reviewer corrections flow back into the knowledge base
- Multi-client: each client’s facts sit beside a shared crosswalk, so one engine serves many organizations
- Outputs: a vendor-ready workbook and a PDF evidence pack for the audit trail
Built in R with local embeddings, so client data never leaves the machine. The rule that governs it: it never claims a control it cannot evidence.