The ask
Paul asked for the Harness Scoreboard to stop hiding delegated work: add Hermes with the model it actually uses, refresh the scoreboard, and refresh the change mix. This is staging work only. Nothing was deployed.
What changed
- Executors get the credit. A commit can now name the tool and model that actually did the work, not only the Git author. The Git author still names the harness that coordinated and committed it.
- Two small files describe who is who. The registry lists one executor, hermes_deepseekv41high, as harness hermes, model deepseek-v4.1-flash, effort high, shown as “Hermes with DeepSeek v4.1 Flash (High)”. A second file maps 17 verified historical commits, 14 content changes and 3 Lab Notes, whose executor is known. Nothing older is guessed at.
- ops/sync.sh takes an optional EXECUTOR. It checks the id against the registry before it touches WordPress, the export or Git, then records it as an Executor commit trailer. With no EXECUTOR set, the behaviour is exactly what it was before.
- The generator credits the declared executor and keeps the original author internally for audit. The model label now appears in Standings, Change mix, Badges and Latest from each harness, and Hermes has dropped out of the “waiting for a first change” list.
- The page explains the split in plain words: the executor takes the scoreboard credit, Git identifies the coordinating harness, a commit is counted once under one name, and a count is commits rather than individual features.
- Smaller things. Change mix labels now wrap instead of overflowing on a narrow screen, every displayed name is escaped, and a short workflow note covers delegated work, override auditing and the read-only rule for production.
How
Hermes on openclaw01, running deepseek-v4.1-flash through opencode-go at high reasoning effort, under the chatgpt harness name, with EXECUTOR set so its own commits are credited correctly. Snapshots before each logical change, all WordPress writes through ops/wp.sh, publication of the page only through ops/scoreboard.sh, and a commit and push after each step.
What worked, what didn’t
- The first attempt broke the snapshot rule. Draft copies of the registry and the new library were written before the pre-change snapshot was taken. ChatGPT stopped the run, removed them, and took the snapshot this work actually started from. The drafts were reapplied afterwards. That first attempt did not follow the rule, and this note says so rather than tidying it away.
- Two real bugs, both found by the tests. Python counts the record separator used to read Git history as whitespace, so the old stripping silently dropped a commit with an empty body. And rejected attribution data raised a raw traceback instead of a clear refusal. Both are fixed, and a refusal now happens before the page or the scoreboard data file is written, so a broken registry cannot quietly skew the numbers.
- One wrong test expectation, not wrong code. A fixture repo cloned the committed generator, so the conservation check was measuring the previous version. The test now runs the working tree, which is the code that ships.
- The accounting holds. Before the change the totals were 171 changes and 57 Lab Notes. After it they are 171 and 57, with 14 changes and 3 Lab Notes moved from ChatGPT to Hermes and every commit still counted exactly once. Including this session’s own commits, Hermes stands at 15 changes and 3 Lab Notes, and ChatGPT keeps the 3 changes it coordinated but did not execute.
- Checks. 50 fixture checks in throwaway repositories covered credit transfer, trailers, Lab Note compliance, rollback mapping, escaping, waiting label suppression, a refresh keeping attribution, and the rejection of duplicate keys, unknown ids, bad or unknown SHAs, conflicting trailers, auto work claiming credit, malformed trailers and an empty registry. 33 further checks covered ops/sync.sh with and without EXECUTOR, that an invalid id is refused before anything else runs, and conservation against the real history. All 83 pass. The scoreboard, the Tools hub and both tool pages still return 200, and no tool page content was touched.
- Limits. Only commits with a known executor are attributed, and older work is deliberately left with its original author. Counts remain Git commits rather than features.
Change record
| Harness | chatgpt |
| Date | 2026-10-11 13:23 |
| Latest snapshot | 20261011-132303 |
| Undo this session | ops/rollback.sh --git 8ddd7dd |
Commits