Sort Arena — Docs
How-to

Run it against your own model

The harness, precisely: the move contract, your AGENTS.md, generate.sh itself (including which command it hands your spec to), and the three property gates — laid out in full on The harness is the variable. Read that first if “the harness” isn’t yet a concrete thing to you; everything below assumes it is.

Everything else on this site holds the model constant and varies the harness. This page does the opposite in one respect only: it swaps out where the model lives, and changes nothing else.

That makes it the demanding version of the exercise — the one worth attempting once the rest works, and the most convincing. A participant generated by a model running on someone’s own GPU, passing the same gates and answering real rounds in the hosted arena exactly like one generated by a vendor API, is the proof that this site’s whole argument was never a slogan: the harness is what you were tuning, and it does not care where the tokens came from. This page has been run to completion — generation, all three property gates, and a live round in the arena — and shows you exactly how to do the same.

What has to be true

Two facts meet in one line of generate.sh:

LLM="${CT_LLM_CMD:-claude}"

The generation step shells out to a coding-agent CLI. It does not know or care what that CLI talks to. So pointing the whole pipeline at a self-hosted model is a question about the CLI’s configuration, not about anything in this repo.

Claude Code takes two environment variables for exactly that:

export ANTHROPIC_BASE_URL="https://<your-endpoint>"
export ANTHROPIC_AUTH_TOKEN="<your key>"

With those set, ./generate.sh runs unchanged and the model that writes your sorter is yours.

Get an endpoint and a key

You need an OpenAI- or Anthropic-compatible endpoint serving at least one code-capable model, and a key for it. This page doesn’t ship one — ask whoever runs your target deployment for the endpoint URL and your personal key, the same way you’d ask for any other credential. Treat both the way you’d treat any other secret: set them as environment variables, never commit them, never paste them into AGENTS.md or any file that leaves your machine.

export ANTHROPIC_BASE_URL="<endpoint URL your operator gave you>"
export ANTHROPIC_AUTH_TOKEN="<your key>"

Sanity-check the two variables before touching generate.sh:

curl -s -o /dev/null -w "%{http_code}\n" "$ANTHROPIC_BASE_URL/v1/models" \
  -H "Authorization: Bearer $ANTHROPIC_AUTH_TOKEN"

200 means you’re through. 401 means the key is wrong or missing. Anything else, ask your operator before going further — there’s no point debugging the harness against an endpoint that isn’t answering yet.

Not every coding-agent CLI supports pointing itself at an arbitrary endpoint the same way. As of this writing: Claude Code and opencode do, directly; Codex does via a model_providers block in ~/.codex/config.toml; Gemini CLI does not have a documented way to. Since generate.sh defaults to claude, Claude Code is the short path — set the two variables above and the rest of this page follows.

Why the vendor CLI needs help, and the fix

Reaching a self-hosted endpoint at all is not the same as getting a usable response back for this pipeline’s narrow, code-only prompt. The Claude Code CLI carries its own system prompt and tool-use scaffolding on every request — the part that makes it Claude Code rather than a bare completion client — regardless of which model answers on the other end. A model that was never trained against that framing can respond to it badly: short, generic, or truncated non-answers, even though the same model handles the exact same prompt content correctly when asked directly, with no scaffolding in the way.

The fix is inside what generate.sh already lets you control. It only requires $CT_LLM_CMD to accept -p <prompt> --output-format text … and print the raw response — it never requires that command to be the Claude Code CLI. Point it instead at a small wrapper that skips the CLI entirely and talks straight to your endpoint’s own chat-completions API — no system prompt, no tools, no CLAUDE.md auto-discovery:

#!/usr/bin/env python3
"""CT_LLM_CMD replacement: talks to a self-hosted OpenAI-compatible endpoint directly,
skipping the Claude Code CLI's own system prompt and tool-use scaffolding entirely.
"""
import json, os, sys, urllib.request

def parse_args(argv):
    prompt = None
    i = 0
    while i < len(argv):
        if argv[i] == "-p" and i + 1 < len(argv):
            prompt, i = argv[i + 1], i + 2
        else:
            i += 1  # ignore --output-format, --disallowedTools, --effort, etc.
    return prompt, os.environ.get("CT_LLM_DIRECT_MODEL", "your-model-name")

def main():
    prompt, model = parse_args(sys.argv[1:])
    url = os.environ["ANTHROPIC_BASE_URL"].rstrip("/") + "/v1/chat/completions"
    body = json.dumps({"model": model, "messages": [{"role": "user", "content": prompt}],
                        "max_tokens": int(os.environ.get("CT_LLM_DIRECT_MAX_TOKENS", 4000)),
                        "temperature": float(os.environ.get("CT_LLM_DIRECT_TEMPERATURE", 0))}).encode()
    req = urllib.request.Request(url, data=body, headers={
        "Content-Type": "application/json",
        "Authorization": f"Bearer {os.environ['ANTHROPIC_AUTH_TOKEN']}"})
    with urllib.request.urlopen(req, timeout=120) as resp:
        data = json.loads(resp.read())
    sys.stdout.write(data["choices"][0]["message"]["content"])

if __name__ == "__main__":
    main()

Save that next to your participant’s other scripts (e.g. ct-llm-direct.py), make it executable, and point generate.sh at it:

export CT_LLM_CMD="./ct-llm-direct.py"
export CT_LLM_DIRECT_MODEL="<the model name your endpoint expects>"
./generate.sh

The wrapper defaults to temperature=0, and that default is load-bearing, not cosmetic. Regenerating five times in a row against local-devstral-small2 at temperature=0 produced the same handler byte-for-byte four times out of five — the one outlier differed only in comment wording, not in the logic the property gates actually check. The same experiment without a fixed temperature has no such floor: nothing stops two runs from picking genuinely different (if each individually gate-passing) strategies, which turns “does my spec work against this model” into a moving target. Override it with CT_LLM_DIRECT_TEMPERATURE only if you have a specific reason to want sampling variance back.

One quoting trap worth knowing about in advance: generate.sh invokes "$LLM" -p "$PROMPT" … with $LLM double-quoted. A CT_LLM_CMD containing a space — a wrapper path with a --model flag tacked on, for instance — is executed as one literal command name including the space, not command-plus-arguments, and fails with a bare “command not found” that never reaches the model at all. That’s why the model above goes through the CT_LLM_DIRECT_MODEL environment variable instead of a CLI flag: keep CT_LLM_CMD a single word, always.

If your CLI and model do get along without the wrapper — some do — you can skip it and set ANTHROPIC_BASE_URL/ANTHROPIC_AUTH_TOKEN directly per What has to be true above. Try the plain path first; reach for the wrapper if generation comes back short, generic, or truncated for no reason the specification explains.

Getting there

  1. Make the local tutorial work first, exactly as written in Five seconds to your first sorter. Do not debug two unfamiliar things at once.
  2. Point CT_LLM_CMD at your endpoint — directly, or through the wrapper above — and re-run ./generate.sh. Everything downstream is unchanged, which is the point.
  3. Run the same three gates you already know:
    ./handler.sh --selftest
    python3 "$REPO/dryrun.py" ./handler.sh --seed 42 --len 8 --quiet
    python3 "$REPO/dryrun.py" ./handler.sh --seed 42 --len 8 \
      --require-adjacent --require-optimal-swaps --require-no-wasted-compares --quiet
    

    A working self-hosted setup passes all three, reproducibly — run it more than once before you trust a single clean pass. What that looks like in practice, five back-to-back runs with the wrapper above:

    Run generate.sh All three gates
    1 10 s, compiles clean pass
    2 10 s, compiles clean pass
    3 11 s, compiles clean pass
    4 10 s, compiles clean pass
    5 10 s, compiles clean pass

    Every run: rounds=18 swaps=17 comparisons=0 — the same optimal-adjacent-swap count as the reference handler (swaps == inversions, the bound documented on The harness is the variable). No AGENTS.md rewriting needed to get there — the specification that already worked against a vendor model works here too.

  4. Take it online. From here, going live is identical to every other participant on this site — follow Join as a participant exactly as written. The model behind your handler doesn’t change that flow at all. A real result from doing exactly this, wrapper and temperature=0 included:
    comparisons=0  swaps=24  faults=0  rounds=25  wall=3.0s  sorted=yes
    

    rounds = comparisons + swaps + 1 → 25 = 0 + 24 + 1 — the same accounting identity every participant on this site is measured against. A handler generated entirely by a self-hosted model, through a harness with the vendor CLI’s own framing removed from the loop, sorted a real array correctly in the real hosted arena — and did it again, on a second array with a different inversion count, without touching AGENTS.md between the two runs:

    comparisons=0  swaps=30  faults=0  rounds=31  wall=3.2s  sorted=yes
    

Measuring it without fooling yourself

A self-hosted endpoint is usually one GPU behind a queue, and that changes which numbers mean anything.

None of this makes the exercise pointless — it means the label on your numbers has to say what actually varied. When the measurement lies covers the general discipline in more depth.

Why this is the hard version

Three things get harder at once, and none of them is the model’s fault:

If your strategy survives all three, you’ve shown something stronger than a working sorter: that the specification, the gates and the contract were doing the work — because the one thing you swapped is the part everyone assumes is decisive.

Found an error, or something that didn't work as documented? Open an issue →