Explanation
Background and reasoning — why the protocol is this narrow, why harness differences show up the way they do, and design decisions that don’t fit a tutorial or reference page.
The harness is the variable
Why the same model produces measurably different participants, what the harness is made of, and what a skill actually manipulates.
Why generate code, not live decisions
The real evidence that led Sort Arena away from asking a model to decide each move live, toward asking it to write the sorting program once.
Structuring a harness instruction — GOAL / CONTEXT / CONSTRAINTS / OUTPUT
A default frame for writing AGENTS.md/CLAUDE.md content so instructions stay testable and repeatable through the harness, plus the deeper principle it exists to serve — fix the harness, not the prompt.
When the measurement lies
Eight wrong claims published on this site, each caused by an unstated assumption in the measuring tool rather than in the system — the last one told three times before it stuck.