
I don’t just use AI tools. I run them as an operation: a set of harnesses, two self-hosted agents on my Linux lab host, one rulebook in Git, and a paper trail behind every change. This page is what I run, why I run it that way, and how I keep it safe.
The harnesses I use
Two agents are self-hosted on my lab host. The rest are products I use from their vendors, and I pair the open coding harness with different models depending on the job.
| Harness | What it is and where it runs | What I use it for |
|---|---|---|
| OpenClaw | Self-hosted agent platform on my Linux lab host (openclaw01), reachable only over Tailscale | My always-on technical co-pilot: cross-project coordination, research, diagnostics and small reversible changes. It runs a roster of specialist sub-agents, and I talk to it from my phone. |
| Hermes | Self-hosted agent on the same lab host | Operations and continuity: change-gated system work, mail and ticket triage, scheduled automations, and the memory layer that carries knowledge between sessions. It also acts as operator for OpenClaw: health checks, updates and config changes, always with a plan first. |
| ChatGPT | OpenAI’s assistant, in chat and in the app | Chat for everyday questions and drafting, Work for longer projects that need files and continuity, Codex for coding. |
| Claude | Anthropic’s assistant | Chat for analysis and writing, Cowork for project-style collaboration, Code for software work. It is also the harness my MSP team uses day to day, through the MCP servers I built. |
| OpenCode | An open coding agent I run locally, paired with whichever model fits the task | Coding and ops work on this site, SSH administration of the lab host, and bake-off comparisons across models. DeepSeek, MiniMax, Qwen, MiMo, GLM and Kimi all rotate through it. |
How they connect
Composio is the one gateway. It holds the credentials for my cloud apps: Google Workspace, Microsoft 365, Freshdesk, GitHub, Notion, Asana and more. Every harness reaches the same integrations through it, so agents can act on my real services without ever holding my passwords. Accounts are connected deliberately through their own authorization flows, tools are dry-run before anything irreversible, and outward-facing actions ask first.
Every harness has a shell. PowerShell on Windows, bash on the lab host. That is how they read logs, run diagnostics, operate WordPress through WP-CLI, and do the unglamorous work that turns a chat window into real operations.
Three harnesses reach the lab host over SSH. ChatGPT, Claude and OpenCode all administer openclaw01 over Tailscale: keeping it simple to run, and running the backups of the harnesses themselves. The agents that live on the host get the same treatment as everything else there: verified snapshots, not hope.
How the self-hosted agents remember
Both agents run on openclaw01, and both are built on the same idea: nothing important stays in someone’s head. Memory works in layers, and each layer trades size for depth.
OpenClaw: a stack of five layers
- Working memory: the current conversation, instructions and recent tool results. Fast and detailed, but gone when the session ends unless something writes it down.
- Memory files: a long-term memory file for durable facts, preferences and decisions, plus dated daily notes for recent continuity. Plain text, human-readable, editable by hand, loaded at the start of every session.
- Local search: a search engine over those files with keyword search, semantic search and reranking. It runs on the host itself with local models, so recall never leaves my hardware.
- Durable fact store: small structured collections of facts, decisions, preferences and instructions that survive beyond any single thread.
- Conversation recall: long chats are compacted into linked summaries instead of being dropped. Ask what was decided three hours into a conversation and the summary expands back into the detail on demand.
Hermes: three layers and a trust score
- Standing memory: a small, always-loaded note about who I am and how I like to work. Kept tiny on purpose so it fits in every conversation.
- Fact store: thousands of short, atomic facts, each linked to the people, systems and projects it concerns, and each carrying a trust score. Facts that prove useful rise; stale or contradicted ones sink. Facts are written automatically when a session ends, by pattern matching and a small local model, and whenever the agent saves one deliberately.
- Session archive: every conversation is indexed for keyword and semantic search, with a summary cached when the session ends. Old conversations stay findable in seconds, even after the raw transcripts are pruned.
One search reaches facts and transcripts together, so “we established this weeks ago” has a real answer. Embeddings, query expansion and reranking run on local models on the host, and neither agent keeps secrets in memory: they reference where to look, never the value.
Written down, backed up, proven
- One shared source of record: both agents read the same Obsidian vault, and memory discipline is explicit: write it down, no mental notes, document lessons so they are not repeated twice.
- One change-control protocol: read-only work needs no approval, reversible work proceeds with a plan, system-impacting work needs explicit approval with a rollback ready, and destructive work needs approval every time. When unsure, classify higher.
- Backups that verify themselves: the memory databases travel inside the host’s backup snapshots, and a snapshot only counts when it passes its gates: a completion marker written after the copy succeeds, checksums over the files, and integrity checks that open the databases instead of trusting the file listing.
- Restore drills: a drill is not “the files are there”. It is a restore that passes every check, with the outcomes written down.
How I work
The harnesses change, the rules do not. Everything I let an AI do runs through one rulebook in Git, the same file every harness reads before it touches anything, including this page.
- Snapshot before every change: a restorable copy exists before anything is written. No exceptions.
- Typed commits: one logical change per commit, prefixed with its type, so the history reads like a ledger instead of a scrapbook.
- A Lab Note per session: every working session that changes the site publishes what changed, how, and what failed, closed by an automatic change record. They are all on this site.
- Rollback is always one command away: restore any commit or any exact snapshot, with history never rewritten.
- Bake-offs to compare harnesses and models: Round 1 gave every harness the same Featured Projects page to build. Round 2 is this page. Same task, same rules, same starting point, graded side by side.
- A Harness Scoreboard computed from Git history: changes, Lab Notes, rollback rate and guideline compliance per harness, refreshed automatically after every Lab Note and never edited by hand.
- Read-only by default: access is scoped, most integrations can look but not touch, and nothing changes without a way to undo it.
What I have built with them
The plumbing below is what turns conversation into operations: tools my techs use every day at the MSP, and the lab that keeps them honest.
MCP servers
Model Context Protocol servers that give Claude structured, scoped access to the tools our techs use every day. Most are read-only by design.
| Server | What it does |
|---|---|
| Microsoft 365 Management | Tenant administration with preview-then-execute changes |
| M365 Security Investigation | Read-only sign-in, audit and mailbox forensics for account breach investigations |
| N-sight RMM | Read-only device, check, patch and backup data from N-able N-sight |
| WatchGuard EPDR | Endpoint protection status, security events, risk and patch posture |
| Freshdesk | Ticket triage, private notes and knowledge base access |
Agent skills
- Freshdesk ticket triage (the team’s most-used): investigates alerts across M365, endpoints, EDR and RMM, writes a clean private note and closes informational tickets
- M365 breach report: turns an affected account and a rough timeframe into a structured compromise investigation
- Google Workspace: safe editing of Docs, Sheets and Slides
- Morning brief and inbox sweep: executive-assistant style daily briefs and inbox triage
Automation pipelines
- n8n phishing triage: Freshdesk webhook triggers analysis and filing of reported phishing email
- Weekly backup review: collects the week’s backup alert tickets from Freshdesk, generates a report and flags the backup jobs that need a closer look
- AppSheet approval and filing app
- Composio integration giving each tech’s Claude account API access across the tool stack
The self-hosted agent lab
- Hermes and OpenClaw agents on a Linux host reachable only over Tailscale
- Seven scheduled jobs: Entra cleanup, PIM digest cleanup, backup reviews, backups and memory backfill
- GPU-backed memory search and a vault-based portable skill framework
- Local models with LM Studio, Ollama and CUDA builds of llama.cpp, plus ComfyUI workflows
- An email identity for each self-hosted agent through AgentMail, used for outbound reports and notifications
- This WordPress site, run in Docker and managed by AI agents with snapshot, change, verify and rollback on every write
Vendor assessment engine
A reusable engine I built with AI agents to answer vendor security questionnaires: the kind a bank sends as a supplier-controls assessment. Instead of re-answering hundreds of controls from scratch, it drafts answers from a stored knowledge base and cites the evidence behind each one.
- Knowledge base: atomic facts about a client plus a crosswalk mapping any vendor question to the facts that answer it (300+ controls carried over from prior assessments)
- Semantic matching: local sentence embeddings match a brand-new questionnaire to the known controls and draft an answer
- Evidence retrieval: the client’s policies and provider attestations (SOC 2, ISO 27001) are indexed so every answer can quote its source
- Review queue: anything unsupported or unknown is flagged for a person instead of guessed, and reviewer corrections flow back into the knowledge base
- Multi-client: each client’s facts sit beside a shared crosswalk, so one engine serves many organizations
- Outputs: a vendor-ready workbook and a PDF evidence pack for the audit trail
Built in R with local embeddings, so client data never leaves the machine. The rule that governs it: it never claims a control it cannot evidence.
Design rule for every agent I build: read-only by default, and nothing changes without a way to undo it.
The day-to-day detail of this work lives in the Lab Notes, and the Harness Scoreboard tracks how each harness performs against the rules.
This page was written by one of the harnesses it describes, under the same rules as everything else on the site: snapshot, typed commits, verification and a Lab Note.