Sort Arena — Docs
Tutorial

Bring your own participant online

Which tutorial should you be reading? This one and Five seconds to your first sorter both end with a participant that sorts, and they are not alternatives so much as two different questions.

  Five seconds This page
You run four cp commands yourself the sort-arena-harness skill
First result after 5 seconds, no model call one guided conversation
Shows you that harness effects are measurable how the loop that produces them works

If you have never run anything here, start there — it costs five seconds and this page reads better once something of yours is already sorting. If you want to watch the skill drive the loop, and understand what it checks and why, you’re in the right place.

The goal of this tutorial is not “get a model to sort a list.” It’s to understand what has to change in the harness around a model to make the service it produces reliable — not just correct once, but correct the same way every time it’s called. That’s what “reliable” means here, and it’s the whole point of the exercise.

This walkthrough was validated by actually running it, start to finish, on a clean machine with nothing pre-installed — a fresh Ubuntu container, a fresh clone of this repo, nothing assumed. Every command below is the literal command that was typed, and the sample output is real, not paraphrased.

Before you begin

Step 1 — clone the repo and start Claude Code inside it

git clone https://github.com/scimbe/CADS-DEMO-sort.git
cd CADS-DEMO-sort
claude

This matters more than it looks: Claude Code only discovers this repo’s sort-arena-harness skill (.claude/skills/sort-arena-harness/SKILL.md) when it’s running inside a checkout of this repo. Starting claude from your home directory, or any other folder, won’t find it.

Step 2 — invoke the skill

At the prompt, type:

/sort-arena-harness

This is a real slash command — Claude Code resolves it against this repo’s own .claude/skills/ directory the moment you type it, no separate install step. It replies by asking for the three things it needs, one at a time if you don’t already have them ready:

You do not write any code at this step, and you do not hand-edit a prompt. You describe the idea; the skill turns it into a real specification (AGENTS.md) and, from that, into a real program (generated/handler.py).

Both land in your own project directory, outside the clone — the skill asks where you want it and suggests a sibling of the clone named after your participant id. That is deliberate (CADS-DEMO-sort#30): your strategy is your project, not untracked files in someone else’s git tree, and git pull in the clone must never collide with your work. git status in the clone stays empty throughout.

Step 3 — expect the first attempt to be checked, not trusted

The skill generates a program from your spec, then actually verifies it: does it speak the wire contract at all (handler.sh --selftest), does it sort a real array correctly, and does it do so twice on the same input — confirming it’s genuinely the same reliable code both times, not a lucky guess. You are not expected to read the generated code yourself; the checks are what “it works” means here.

Here is the real report from the validation run, verbatim, because seeing the actual shape of a passing report is more useful than a description of one:

Done — `docs-tutorial-test` is built and verified.

## Pass/fail report

**All three contract criteria passed, on the first generation.** No spec-tighten-and-regenerate
cycle was needed here.

| Check | Result |
|---|---|
| `handler.sh --selftest` | **OK** |
| Criterion 1 — format, zero faults | **PASS** (`faults=0`, `wrongDone=0` on every run) |
| Criterion 2 — terminated correctly in budget | **PASS** (`sorted=True` on every run) |
| Criterion 3 — no regression on `correction` | **PASS** (`swap 0,1` without → `compare 0,1` with) |
| Determinism (same array, run twice) | byte-identical output apart from `wallClockMs` |

Main dry run, `--seed 42 --len 8`, array `[82, 15, 4, 95, 36, 32, 29, 18]` → sorted correctly in
**43 rounds** (25 comparisons, 17 swaps), ~3.2s for the whole run. Also clean on: already-sorted,
reversed, heavy duplicates, n=2, and — **with `--budget 600`, because *this* handler is a coached
bubble sort with no shrink optimization, whose worst case is ~(n-1)² rounds** (not because of the
array length as such — see [Check yourself](/CADS-DEMO-sort-docs/tutorials/check-yourself/) Q6,
which retracts the length-based rule; a plain adjacent sorter at n=21 finishes inside 200) (the GUI's own Round Budget
field defaults to 600 for exactly this reason) — n=21/22/24 (301/278/309 rounds). Running any of
these three at the 200 default reports `sorted=False`, budget exhausted, not a broken handler.

Passing on the first try is not the common case worth designing this tutorial around — a real failure teaches you more, and is covered next. But when it does pass first try, this is what that looks like: real numbers, real shapes tested, no hand-waving.

What a real failure looks like, and what actually fixes it

Here is a real one, from this project’s own history: building the bubble-sort-claude participant this repo ships with. The strategy was simple: visit adjacent pairs left to right, swap if out of order, repeat. The first version of the spec was dry-run twice and failed both times:

Run Array Result
1 [5,3,8,1,9,2] budget exhausted, never sorted
2 [18,60,61,29,26,25] budget exhausted, never sorted

Both failed the identical way: at specific positions, the behavior the spec produced repeatedly compared instead of swapping, even though the array plainly showed an out-of-order pair right there. The full round-by-round record is preserved in CADS-DEMO-sort#10 — that issue, not this page, is the primary artifact for this story. The tempting fix is “run it again and hope for a cleaner result.” That fix was not used, because it doesn’t actually address anything — an unreliable spec produces unreliable behavior however many times you re-sample it.

The real fix was to look at why the model kept getting that one decision wrong. The spec said “remember where you are in a pass” — but the handler is invoked fresh every single round with no memory. The instruction was accurate in spirit but didn’t say how to actually reconstruct “where you are” from what a stateless call can actually see. Once the spec was rewritten to say that explicitly — reconstruct the cursor from the single most recent entry in the round history, check the real array directly rather than trying to remember whether a pass had a swap in it — the identical two arrays that failed before both converged cleanly, every time, on retest. The rewritten spec is the one that ships today as participants/bubble-sort-claude/AGENTS.md — its “Where you are” and “Whether you’re done” bullets are the direct fix for this exact failure, so you can read the “after” side of the diff even though the “before” no longer exists (next paragraph).

Don’t expect to reproduce this failure yourself. An earlier version of this page presented it as if you could; CADS-DEMO-sort#22 (item 7) tried, hard, and established three things a reader deserves to know up front:

So treat this section as a documented case study of the diagnose-tighten-retest loop, not as an exercise to replay. The loop is what transfers: when your own generation fails, the failure points at a question your spec left unanswered.

That’s the actual lesson: a generated service is only as reliable as the spec’s answer to “what does the model actually know at the moment it has to decide this?” A vague spec produces code that’s right most of the time and silently wrong at the edges. A spec that states its assumptions explicitly enough to be checked produces code that’s right the same way every time. See “Why generate code, not live decisions” for the full evidence trail this example is drawn from, including what the same model did with no harness at all, and what changed at each step in between.

What you’ll have when this is done

Your own participant directory, beside the clone rather than inside it — AGENTS.md, generated/handler.py, handler.sh and generate.sh — plus a short, real report in the session itself: pass or fail on the contract check, pass or fail on the dry run, and, if either failed, what about your spec the failure points at.

(This page used to promise participants/<your-id>/ inside the clone, and a README file written into it. Neither is what the shipped skill does — measured 2026-08-27, the directory lands outside the clone and no README is written.) When both pass, you have a program that runs in milliseconds and behaves identically on every call, not a live decision you have to hope goes well each round. That reliability is the actual deliverable of this tutorial, not the sorting itself.

On joining the arena: the self-service route in Join as a participant is the only real path today. If any generated file or older note tells you to hand-edit a bridge config instead, it is stale.

Step 4 — watch it sort

Everything above is local verification: --selftest, a dry run, determinism, the correction path. None of it puts your handler on screen — and the visualization is the actual payoff the homepage advertises (“a real, running visualization, not a static mockup”). index.html’s default-selected Solo run tab does exactly this: pick your participant, an array length, a round budget, run it, and watch a live bar visualization, a round timeline colored by compare/swap/fault/done, and a scorecard fill in as your handler actually sorts (CADS-DEMO-sort#22 found this step missing from every tutorial in this repo, despite being the single most convincing thing in the project).

You don’t need the hosted arena or a join request to see this — Run the arena locally gets your own generated handler into that same GUI with zero dependencies, no network, and no operator, using nothing but the node/python3 you already have.

Once your service is reliable: two directions from here

Go deeper into the harness — the natural next step: Evolve the harness — from “it sorts” to “it IS bubble sort, provably” takes the participant you just built and steers it toward a named, mechanically checkable algorithm using dryrun.py’s property checks (--require-adjacent, --require-optimal-swaps) as a goal line and a verification hook as the referee. That’s where the spec-tightening loop you just practiced becomes a real engineering workflow.

Or go live now: turning your generated handler into something sort.bunsenbrenner.org actually calls is self-service — sign in with your Keycloak account (the login is the legitimization), submit, get approved automatically, run the serve command it hands you. See Join as a participant for the full walkthrough.

Found an error, or something that didn't work as documented? Open an issue →