AI & Automation by ChatGPT with Sol 6.1 High

Connected tools · Searchable memory · Reversible work

I build the plumbing that lets AI do useful work across IT operations: investigate a problem, find the evidence, document the result and make a controlled change.

My setup combines desktop harnesses, self-hosted agents and shared app integrations. I use different tools and models for different jobs, with the same expectation: show what happened, protect private data and leave a way back.

My operating rule
Read-only by default. Verify before claiming success. Make changes small, recorded and reversible.

On this page: The harnesses I use · How they connect · How the agents remember · What I have built · How I compare and control the work

The harnesses I use

A harness is the working environment around a model: its tools, instructions, memory and controls. Changing the model changes one part of that environment. I compare both.

ChatGPT: Chat, Work and Codex

I use Chat for discussion, research and drafting; Work for tasks in my desktop workspace; and Codex for coding and work in a checkout. This family gives me conversational and agent interfaces on my computer, with shell access when the task needs it. I also use it to work on this site through SSH and the shared change process.

Claude: Chat, Cowork and Code

I use Chat for analysis and writing, Cowork for work across files and tools on my Windows PC, and Code for implementation in a working checkout. Claude in Cowork built the first version of this site from my Obsidian profile. My team tooling also gives Claude structured access to ticketing, Microsoft 365 and endpoint data.

OpenCode: one harness, several models

OpenCode is my coding harness in the terminal. I use it for implementation, SSH work and experiments with different models. Round 1 includes OpenCode with DeepSeek, Qwen and MiMo, alongside Claude and ChatGPT. That lets me compare the working result and the process around it, rather than assuming a model name tells me how well it will do.

OpenClaw and Hermes: my self-hosted agents

Both run on my Linux lab host, openclaw01, which I reach over Tailscale. OpenClaw provides an execution worker for smaller implementation and cleanup tasks, with specialized agents available for broader work. Hermes is another execution environment for automation and connected-tool tasks, with persistent conversation and fact memory. Both use bash, Composio and my shared Obsidian knowledge base.

I use models such as MiniMax and DeepSeek for everyday agent work, and GPT models for more complex tasks. The harness gives the model a way to act; my instructions and verification determine whether that action is acceptable.

How they connect

Apps through Composio

Composio is the shared integration layer across my harnesses. It connects tools such as Google Workspace, Microsoft 365, Freshdesk and GitHub through authenticated app actions. The agents can discover a tool, inspect its inputs and use it without rebuilding the integration for each model.

Those actions inherit the connected account’s permissions. I treat access as real authority, and review actions that send information or change an external system.

Shells and SSH

Every harness in my setup has a route to shell tools, using PowerShell or bash. ChatGPT, Claude and OpenCode reach the lab host over SSH to administer it, keep it simple to run and run the backups that protect the harnesses themselves.

For this site, SSH and WP-CLI provide the full change workflow. The WordPress MCP provides another control plane. I choose the route that supports the task and its safeguards.

How the agents remember

Memory is something I store, search and bring back into a conversation. A fact can be saved without being in the model’s current context. I separate working context, durable knowledge and the discussion behind a decision.

OpenClaw: notes, knowledge and conversation recall

  • Live context holds the current conversation, instructions and recent tool results. It is useful for the task in progress, but limited in size.
  • Readable Markdown notes carry daily continuity and curated facts, preferences and decisions. I can inspect and edit what has been deliberately remembered.
  • QMD searches indexed notes using words and meaning, then ranks the relevant passages. It is the search layer for files.
  • clawmem organizes durable knowledge across sessions, including facts, decisions and preferences. It supports keyword and semantic recall.
  • LCM, through lossless-claw, preserves a conversation archive and linked summaries. When a long conversation is compacted, those summaries help locate and recover the original discussion.

Hermes: session history and a separate fact store

Hermes combines small, readable memory notes with two persistent databases. One stores conversations, keyword indexes, semantic embeddings and cached summaries. The other stores durable facts, preferences and decisions through a holographic memory plugin.

Background review and session-end extraction can turn useful context into facts. The fact store supports entity links and feedback scores. A recall query searches facts alongside past sessions, using keyword and meaning-based retrieval. Local models help with embeddings and ranking; cached summaries reduce repeated summarization.

What I remember depends on what I need. A standing preference belongs in curated memory. The reasoning behind a decision belongs in the conversation archive. Search can bring either back, but I still check the evidence before acting on it.

Backing up the ability to remember

The agent backups cover durable memory and the software needed to use it. OpenClaw’s state includes its notes and conversation store. Hermes backups preserve its session and fact databases, memory notes, skills, plugins and custom memory code. My recovery documentation checks completeness and data integrity, with tested reconstruction of the custom Hermes build. I distinguish a copied file from evidence that it can be restored.

What I have built

MCP servers for IT operations

My Model Context Protocol servers give agents structured, scoped access to the tools the team uses. Most are read-only by design.

  • Microsoft 365 Management: tenant administration with preview-then-execute changes.
  • M365 Security Investigation: read-only sign-in, audit and mailbox forensics for account compromise investigations.
  • N-able N-sight RMM: read-only device, check, patch and backup data.
  • WatchGuard EPDR: endpoint protection status, security events, risk and patch posture.
  • Freshdesk: ticket triage, private notes and knowledge base access.

Skills that turn access into a repeatable process

Skills are runbooks for the agent. My Freshdesk triage skill investigates alerts across Microsoft 365, endpoints, EDR and RMM, writes a clear private note and closes informational tickets. The M365 breach-report skill turns an affected account and a timeframe into a structured investigation. Other skills support safe editing in Google Docs, Sheets and Slides, plus morning briefs and inbox triage.

Automation pipelines

  • Phishing triage in n8n: a Freshdesk webhook triggers analysis and filing of reported phishing email.
  • Backup review: gathers backup alert tickets, produces a report and flags jobs for a closer look.
  • AppSheet: an approval and filing app.
  • Composio for the team: connects each tech’s Claude environment to the app stack.

The self-hosted agent lab

The lab brings together OpenClaw, Hermes, GPU-backed memory search and a portable, vault-based skill framework. Scheduled work includes Entra and PIM digest cleanup, backup reviews, backups and memory backfill. I also work with local models through LM Studio, Ollama and CUDA builds of llama.cpp, alongside ComfyUI workflows.

This WordPress site runs in Docker and is part of that lab. It gives me a practical place to test multiple harnesses against the same rules, with a visible record of every change.

An evidence-led vendor assessment engine

I built a reusable engine in R to draft vendor security questionnaire answers from a stored knowledge base. Atomic client facts and a shared control crosswalk support new questionnaires; local sentence embeddings match questions to known controls. Indexed policies and provider attestations supply the evidence behind an answer.

Unsupported or unknown answers go to a human review queue. Corrections feed back into the knowledge base, while each client’s facts remain separate from the shared crosswalk. The outputs are a vendor-ready workbook and a PDF evidence pack. Matching runs locally so client data stays on the machine. My rule is simple: never claim a control I cannot evidence.

How I compare and control the work

I keep one rulebook in Git for every harness working on this site. Before each change, I take a snapshot. I make one logical change, verify it, and record it with a typed commit. Each session ends with a Lab Note describing the ask, the changes, the tools and the limits. Rollback remains available.

Lab Notes: what actually happened

The notes record successes, failed attempts and workarounds. They show how the site moved from Claude’s first build to a shared workflow for multiple harnesses. I want the next session to inherit evidence, not just a claim of success.

Scoreboard and bake-offs: different evidence

The Harness Scoreboard measures recorded process: changes, snapshots, commit types, Lab Notes and rollbacks. Round 1 gave different harnesses and models the same Featured Projects brief. Its comparison includes writing, accuracy, design and accessibility. Process compliance helps me judge how work was done; it does not prove the answer is good.

Across the setup, I start with inspection. System-impacting, destructive and external actions need approval appropriate to their scope. I review drafted communications and verify findings before presenting them or acting on them. A useful agent should make my work easier to understand, maintain and undo.

Based on my existing project page, Lab Notes, shared guidelines, Round 1 records and Obsidian notes on agents, memory and recovery, checked against interviews with OpenClaw and Hermes.