AI & Automation by OpenCode with Qwen 3.8 Flash Medium

Glowing blue fiber optic strands

I don’t just use AI tools. I build the plumbing that lets them do real work safely across an MSP’s stack: ticketing, RMM, endpoint security, Microsoft 365 and backups. This page is the map: what I run, how it connects, how my agents remember, and how I keep the whole thing safe enough to hand to a machine.

The harnesses I use

A harness is the software wrapped around a model: what it can touch, what it remembers, where it runs. The harness matters as much as the model, so I run several on purpose and compare them like an engineer, not a fan.

OpenClaw, the lab host’s operator

OpenClaw is a self-hosted agent platform and my main one on the Linux lab host. The host sits on my private network and is reachable over Tailscale only.

It keeps a small team of scoped subagents instead of one do-everything bot: a lightweight worker for implementation and cleanup, a researcher for web searches, a main agent for design-heavy work. Each gets only the tools its job needs, and some run on tight command allowlists with the filesystem locked to their workspace. Everyday models are MiniMax M2.7 and DeepSeek v4 Pro/Flash, with bigger models for complex work. Its skills reach the Composio gateway and my shared Obsidian vault, and it runs seven scheduled jobs: Entra cleanup, PIM digest cleanup, backup reviews, backups and memory backfill.

Hermes, the daily assistant

Hermes is the second self-hosted agent on the same host and the one I talk to every day: a desktop UI plus a gateway to messaging platforms. Everyday models are MiniMax M2.7 and DeepSeek v4 Pro/Flash, with GPT-class models for complex jobs. It shares the Composio connections and the vault with OpenClaw, keeps its own skills and plugins, and runs its own scheduled work: embedding and fact backfills, housekeeping, digests and monitoring. How both agents remember is covered below, because that part is the real engineering.

ChatGPT: Chat, Work and Codex

Chat for quick questions and drafting. Work for longer agentic jobs against documents and research. Codex for coding from the terminal against real repos. It has shell access on the machine it runs on, and it reaches the lab host over SSH for admin tasks.

Claude: Chat, Cowork and Code

The same three shapes: Chat for quick work, Cowork for agentic knowledge work, Code for terminal coding. Claude is also the harness behind the MCP server suite further down this page, and the first harness to build this WordPress site, under the same rulebook every harness now follows.

OpenCode, one harness and many models

OpenCode is a terminal harness, and the interesting part is that one harness drives several models. Round 1 of the bake-off below ran OpenCode with Qwen, DeepSeek and MiMo against the same task, side by side with Claude and ChatGPT. This page is Round 2. The same idea sits under my lab agents: a hosted provider serves their model mix, so a cheap model handles summarising while a strong one handles reasoning.

How they connect

  • Composio is the integration gateway in front of the cloud tools, for every harness: Google Workspace, Microsoft 365, Freshdesk, GitHub, Notion, Asana and web search, each behind its own scoped connection instead of shared credentials everywhere.
  • Agent mailboxes: each self-hosted agent gets its own email inbox through AgentMail for agent-to-agent coordination and workflow notifications, so I can always tell which agent sent what.
  • Shell access in every harness: bash on the lab host, PowerShell on the Windows side. Scripts ship with prerequisites, verification and rollback, and the default posture is read-only.
  • SSH into the lab host: ChatGPT, Claude and OpenCode reach openclaw01 over Tailscale with key-based SSH to administer it, keep the setup simple to run, and run the backups of the harnesses themselves. One host, every harness, same paper trail.
  • One shared vault: the agents read and write the same Obsidian vault, synced between my devices and the host. Baselines, SOPs and lessons learned live there, so knowledge outlives any single agent or chat.

How my self-hosted agents remember

A model forgets everything the moment a conversation ends. What makes an agent useful is the memory system wrapped around it, and mine are built from plain, auditable pieces rather than one opaque database.

OpenClaw: a memory mesh

OpenClaw runs several specialized memory layers side by side, each with a clear job:

  • Live context is what the agent sees right now: instructions, workspace notes, recent turns. Fast, detailed, and gone when the session ends.
  • Plain-text memory files hold the durable stuff: a curated long-term file plus daily logs. Human-readable, so I can audit exactly what the agent was taught to remember.
  • A local search sidecar indexes those files with keyword and semantic retrieval, reranking and query expansion, so recall works over notes the agent isn’t currently reading.
  • A durable knowledge vault keeps cross-session objects: facts, decisions, preferences, entities and events, retrievable by keyword, meaning or both.
  • Lossless compaction handles long conversations: older turns collapse into linked summaries instead of being truncated, and any summary can be expanded back into full detail on demand. Nothing important is silently dropped.

The embeddings, reranking and summary models run on the host’s GPU, so recall stays local. The practical rule: search broad, expand the right piece, then answer.

Hermes: session search plus a fact store

Hermes splits memory into two databases and one plugin layer:

  • A session archive records every conversation and indexes it three ways: exact keywords, trigram matching for typos and partials, and vector embeddings for meaning. A search finds the right old conversation without reading all of them.
  • A fact store holds thousands of small durable facts, linked to named entities, each with a trust score that rises when a fact proves helpful and falls when it doesn’t. Wrong memories fade out instead of fossilising.
  • Extraction at session end: pattern rules and a cheap model mine the closed transcript for preferences, decisions and environment facts, and cache a summary of the session itself. Later searches answer from the index instead of re-reading transcripts.

Raw transcripts eventually rotate out, but the index, summaries and vectors survive: they live inside the databases, which is also why a restored machine can search its whole history immediately, with no re-embedding.

Backups of the memories themselves

The same discipline I run on client data applies to the lab: the agent databases and the vault go into scheduled snapshots, checksums prove the copies, and a completion marker written only after the last stage passes is what separates a real backup from a job that quietly did nothing. Drills open the restored databases, run integrity checks and compare counts before anything is called restorable.

How I work: one rulebook for every harness

This site is the working demo. Every AI harness that touches it follows the same rules, stored in Git where any of them can read it before acting:

  • Snapshot before every change, so a rollback is always one command away, and rolling back beats leaving something broken.
  • One logical change per commit, typed messages, and the commit names the harness that made it.
  • A numbered Lab Note per session: the ask, what changed, how, and honestly what didn’t work.
  • The Harness Scoreboard grades every harness from git history: snapshot discipline, typed commits, rollback rate, staying in scope. Computed, not self-reported.
  • Bake-offs put harnesses and models on the same task with the same rules and compare what they build and how carefully they build it.
  • Read-only by default. Most servers and skills can only look; anything that writes shows a preview first and keeps a way to undo.

MCP servers

Model Context Protocol servers that give Claude structured, scoped access to the tools our techs use every day. Most are read-only by design.

ServerWhat it does
Microsoft 365 ManagementTenant administration with preview-then-execute changes
M365 Security InvestigationRead-only sign-in, audit and mailbox forensics for account breach investigations
N-sight RMMRead-only device, check, patch and backup data from N-able N-sight
WatchGuard EPDREndpoint protection status, security events, risk and patch posture
FreshdeskTicket triage, private notes and knowledge base access

Agent skills

  • Freshdesk ticket triage (the team’s most-used): investigates alerts across M365, endpoints, EDR and RMM, writes a clean private note and closes informational tickets
  • M365 breach report: turns an affected account and a rough timeframe into a structured compromise investigation
  • Google Workspace: safe editing of Docs, Sheets and Slides
  • Morning brief and inbox sweep: executive-assistant style daily briefs and inbox triage

Skills ride in the shared vault, so a process written once travels between harnesses instead of being re-taught per tool.

Automation pipelines

  • n8n phishing triage: Freshdesk webhook triggers analysis and filing of reported phishing email
  • Weekly backup review: collects the week’s backup alert tickets from Freshdesk, generates a report and flags the backup jobs that need a closer look
  • AppSheet approval and filing app
  • Composio integration giving each tech’s Claude account API access across the tool stack

Self-hosted agent lab

  • Hermes and OpenClaw agents on a Linux host reachable over Tailscale, with ChatGPT, Claude and OpenCode joining over SSH
  • Seven scheduled jobs: Entra cleanup, PIM digest cleanup, backup reviews, backups and memory backfill
  • GPU-backed memory search and a vault-based portable skill framework
  • Local models with LM Studio, Ollama and CUDA builds of llama.cpp, plus ComfyUI workflows
  • This WordPress site, run in Docker and managed by AI agents with snapshot, change, verify and rollback on every write

Vendor assessment engine

A reusable engine I built with AI agents to answer vendor security questionnaires: the kind a bank sends as a supplier-controls assessment. Instead of re-answering hundreds of controls from scratch, it drafts answers from a stored knowledge base and cites the evidence behind each one.

  • Knowledge base: atomic facts about a client plus a crosswalk mapping any vendor question to the facts that answer it (300+ controls carried over from prior assessments)
  • Semantic matching: local sentence embeddings match a brand-new questionnaire to the known controls and draft an answer
  • Evidence retrieval: the client’s policies and provider attestations (SOC 2, ISO 27001) are indexed so every answer can quote its source
  • Review queue: anything unsupported or unknown is flagged for a person instead of guessed, and reviewer corrections flow back into the knowledge base
  • Multi-client: each client’s facts sit beside a shared crosswalk, so one engine serves many organizations
  • Outputs: a vendor-ready workbook and a PDF evidence pack for the audit trail

Built in R with local embeddings, so client data never leaves the machine. The rule that governs it: it never claims a control it cannot evidence.

Design rule for every agent I build: read-only by default, and nothing changes without a way to undo it.