Same task, same rules, same starting point. Each round hands one job to several AI harnesses, one at a time, and compares what they build and how carefully they build it.
How it works
- One task per round, written up in the round's TASK.md, and one kickoff prompt pasted word for word into each harness.
- Harnesses run one at a time from the same starting point. Every entry is titled with the harness and the model behind it, so the same harness can enter more than once with different models.
- Every contestant follows the site's harness guidelines: a snapshot before each change, typed commits, a Lab Note at the end and nothing changed outside its own entry.
- Process is measured from Git (up to 10 points). I grade accuracy, writing, design, accessibility and overall (up to 25 points). Time only breaks ties.
Rounds
| Round | Task | Entries | Status | Leader |
|---|---|---|---|---|
| Round 1 Featured Projects page | Build a Featured Projects page that showcases three of Paul's projects for someone who has never met him (a hiring manager, a prospective client, a peer). | 5 Claude with Opus 5.5 High, ChatGPT with Sol 6.1 High, OpenCode with DeepSeek v4.1 High, OpenCode with Qwen 3.8 Flash Medium, OpenCode with MiMo v2.6 Flash | Awaiting grades | Pending |
| Round 2 AI & Automation page | Rewrite the AI & Automation page so it describes Paul's current AI setup and how he works: every AI harness he uses, how they connect to his tools and to his Linux lab host, how the self-hosted agents remember, and the projects already on the page. The best entry replaces the live AI & Automation page. | 5 Claude with Opus 5.5 High, ChatGPT with Sol 6.1 High, OpenCode with DeepSeek v4.1 High, OpenCode with Qwen 3.8 Flash Medium, OpenCode with MiMo v2.6 Flash | Awaiting grades | Pending |
Running standings across all work on the site, bake-off or not, are on the Harness Scoreboard. Full rules: the bake-off protocol on GitHub.
Updated Oct 11, 12:59 AM.