AI & Automation by OpenCode with DeepSeek v4.1 High

Glowing blue fiber optic strands

I run a small fleet of AI harnesses, and I treat them like a team: one rulebook, one change process, and a way to undo anything they do. This page is the tour. What I run, how it connects to my tools and my Linux lab host, how the self-hosted agents remember, and what I have built with them.

The harnesses I use

I use more than one harness on purpose. They are good at different things, and running the same job through several of them is how I learn which one to reach for. Two of them I host myself on my Linux lab host; the rest are services I use over the web.

HarnessWhat it isWhere it runsWhat I use it for
OpenClawA self-hosted agentMy Linux lab hostExecution-first work: small implementation and cleanup tasks, app actions through Composio, and research sub-agents. When a task turns design-heavy or cross-project it hands the work back instead of improvising.
HermesA self-hosted agentMy Linux lab hostMy everyday agent and the deeper memory of the two: research, notes, scheduled jobs, messaging gateways and agent-to-agent email. It is the one I talk to most.
ChatGPTChat, Work and CodexThe vendor cloudChat for thinking a problem through and drafting, Work for team-facing tasks, and Codex for code.
ClaudeChat, Cowork and CodeThe vendor cloudChat for conversation and analysis, Cowork for longer project work, and Code for code.
OpenCodeA command-line coding harnessFrom my machine, reaching the lab host over SSHRunning the same task with several different models so I can compare them side by side. I have run it with DeepSeek v4.1, Qwen 3.8 Flash and MiMo v2.6 Flash, among others.

The self-hosted agents are not locked to one model. OpenClaw runs a fast model for routine turns and swaps in stronger reasoning models when a task is hard, with choices across DeepSeek, MiniMax, MiMo, GLM, Kimi and Qwen. Hermes runs a Flash-class model day to day and escalates to a Pro-class model for complex work, with a separate local model for search. Both are routed through one model provider, so I can change the model without rewiring the agent.

How they connect

Every harness reaches my tools and my lab host in the same three ways, which keeps the mental model small.

  • Composio for app integrations. Rather than holding an API key for every service, an agent calls Composio, which already holds the authorized connections: ticketing, email, chat, documents and more. Access is scoped to what I have connected, a searchable catalog lets an agent inspect an action before it runs, and a connection can be revoked centrally.
  • Shell access in every harness. Each harness can run a real shell, PowerShell on Windows and bash on Linux. It is the universal fallback for the work no tidy tool covers: git, files, service checks and multi-step scripts. Most of what I automate starts life as a shell command I have run by hand.
  • SSH to the lab host. ChatGPT, Claude and OpenCode reach my Linux lab host over SSH. That is how they administer it, keep it simple to run, and run the backups of the harnesses themselves. OpenClaw and Hermes already live on that host, so they work there directly.

How the self-hosted agents remember

OpenClaw and Hermes both run on the lab host and both carry working memory across sessions, but they solve it differently. OpenClaw behaves like a mesh of specialized stores. Hermes behaves like a session database plus a fact database.

OpenClaw: a memory mesh

OpenClaw keeps several kinds of memory and assembles the right ones for each turn.

  • Live context: what the model is looking at right now. The current conversation, its instructions and recent tool results. It is the fastest and most detailed layer, and it is bounded by the model’s context window.
  • Plain Markdown notes: a curated long-term memory file plus daily notes. Human-readable, easy to audit, and the first place to look at what the agent has been taught to remember.
  • A local search layer: keyword and semantic search over those files, with reranking and query expansion. This is the recall layer for notes.
  • A durable knowledge store: facts, decisions, preferences and instructions, each kept as its own record that survives beyond a single thread.
  • A compaction layer: when a conversation gets long, older turns are compacted into linked summaries instead of being dropped, and the agent can expand a summary back into detail when a question needs it.

A single search can touch all of them: keyword and semantic lookup over the files, direct retrieval of stored knowledge, and expansion of compacted conversation. I keep the layers separate on purpose, because stored does not always mean loaded, and loaded does not always mean stored.

Hermes: always-on notes plus a fact store

Hermes splits always-on context from deep recall, so it pays for only a small amount of memory on every turn.

  • A small pinned notes layer that loads into every session before anything else. It holds stable facts about me and my environment, and it is kept deliberately tiny because it is paid for on every turn.
  • A structured fact store: thousands of facts about me and the environment, each linked to the people, projects and tools it mentions and carrying a trust score. Facts are deduplicated, and trust rises or falls as they prove accurate or stale.
  • Session history: every conversation is stored and indexed, with a summary generated and cached when a session ends. Keyword, partial-word and semantic search all run over it.
  • One search call checks the fact store and the session index together, so a known fact comes back instantly and the relevant past conversations come back alongside it.

Hermes generates its own embeddings for sessions and facts with a local model, and uses local models to expand queries and rerank results. That keeps search fast and keeps the notes on the host.

Backups and recovery

Both agents’ memory is captured by the lab host’s scheduled backups, and a run only counts as complete when it writes a completion marker at the end. That gate matters: without it, a silent failure would look like a good copy.

The databases travel with their own vectors, so a restored machine gets working session search and fact recall without re-embedding anything. Raw transcripts are pruned over time, but the indexes, embeddings and summaries survive, so nothing searchable is lost. I keep a restore drill checklist and use it, because a backup that has never been restored is only a hope.

How I work

One rulebook governs every harness, and it lives in Git so it is versioned and readable by all of them. The rules are the interesting part of this setup, not the models.

  • One rulebook for every harness. The same rules file, read at the start of every session, whether the harness is Claude, ChatGPT, OpenCode, Hermes or OpenClaw.
  • Snapshot before every change. Nothing is written until there is a restorable point to come back to. No exceptions.
  • Typed, one-purpose commits. Each change is small and labelled by type, so the history reads like a log instead of a dump.
  • A Lab Note per session. Every session that changes the site publishes a numbered note in four parts: the ask, what changed, how, and what worked or did not. That is the public changelog.
  • Rollback always available. Any change can be reverted to a snapshot or to any point in the Git history, and every rollback is itself recorded.
  • Read-only by default. Agent actions are classified from read-only up to destructive, and only the highest levels need explicit approval before they run.
  • Bake-offs to compare harnesses and models. The same task, the same rules, one harness at a time, then compare what each built and how carefully it built it. The same harness can enter more than once with a different model.

You can see all of it on the site. The Lab Notes are the public changelog. The Harness Scoreboard ranks each harness on changes, rollback rate and guideline compliance, computed straight from Git and refreshed automatically. The Harness Bake-off runs the same task head to head, most recently across five entries in Round 1.

What I have built with them

Most of the projects below were built with these harnesses doing the heavy lifting while I set the rules, reviewed the work and made the decisions.

MCP servers

Model Context Protocol servers that give structured, scoped access to the tools our techs use every day. Most are read-only by design.

ServerWhat it does
Microsoft 365 ManagementTenant administration with preview-then-execute changes
M365 Security InvestigationRead-only sign-in, audit and mailbox forensics for account breach investigations
N-sight RMMRead-only device, check, patch and backup data from N-able N-sight
WatchGuard EPDREndpoint protection status, security events, risk and patch posture
FreshdeskTicket triage, private notes and knowledge base access

Agent skills

  • Freshdesk ticket triage (the team’s most-used): investigates alerts across M365, endpoints, EDR and RMM, writes a clean private note and closes informational tickets
  • M365 breach report: turns an affected account and a rough timeframe into a structured compromise investigation
  • Google Workspace: safe editing of Docs, Sheets and Slides
  • Morning brief and inbox sweep: executive-assistant style daily briefs and inbox triage

Automation pipelines

  • n8n phishing triage: a Freshdesk webhook triggers analysis and filing of reported phishing email
  • Weekly backup review: collects the week’s backup alert tickets from Freshdesk, generates a report and flags the backup jobs that need a closer look
  • AppSheet approval and filing app
  • Composio integration giving each tech’s Claude account API access across the tool stack

Self-hosted agent lab

  • OpenClaw and Hermes on a Linux host reachable only over Tailscale
  • Seven scheduled jobs: Entra cleanup, PIM digest cleanup, backup reviews, backups and memory backfill
  • GPU-backed memory search and a vault-based portable skill framework
  • Local models with LM Studio, Ollama and CUDA builds of llama.cpp, plus ComfyUI workflows
  • This WordPress site, run in Docker and managed by AI agents with snapshot, change, verify and rollback on every write

Vendor assessment engine

A reusable engine I built with AI agents to answer vendor security questionnaires: the kind a bank sends as a supplier-controls assessment. Instead of re-answering hundreds of controls from scratch, it drafts answers from a stored knowledge base and cites the evidence behind each one.

  • Knowledge base: atomic facts about a client plus a crosswalk mapping any vendor question to the facts that answer it (300+ controls carried over from prior assessments)
  • Semantic matching: local sentence embeddings match a brand-new questionnaire to the known controls and draft an answer
  • Evidence retrieval: the client’s policies and provider attestations (SOC 2, ISO 27001) are indexed so every answer can quote its source
  • Review queue: anything unsupported or unknown is flagged for a person instead of guessed, and reviewer corrections flow back into the knowledge base
  • Multi-client: each client’s facts sit beside a shared crosswalk, so one engine serves many organizations
  • Outputs: a vendor-ready workbook and a PDF evidence pack for the audit trail

Built in R with local embeddings, so client data never leaves the machine. The rule that governs it: it never claims a control it cannot evidence.

Design rule for every agent I build: read-only by default, and nothing changes without a way to undo it.