Run it against your own model
The harness, precisely: the move contract, your AGENTS.md, generate.sh itself (including
which command it hands your spec to), and the three property gates — laid out in full on
The harness is the variable.
Read that first if “the harness” isn’t yet a concrete thing to you; everything below assumes it is.
Everything else on this site holds the model constant and varies the harness. This page does the opposite in one respect only: it swaps out where the model lives, and changes nothing else.
That makes it the demanding version of the exercise — the one worth attempting once the rest works, and the most convincing. A participant generated by a model running on someone’s own GPU, passing the same gates and answering real rounds in the hosted arena exactly like one generated by a vendor API, is the proof that this site’s whole argument was never a slogan: the harness is what you were tuning, and it does not care where the tokens came from. This page has been run to completion — generation, all three property gates, and a live round in the arena — and shows you exactly how to do the same.
What has to be true
Two facts meet in one line of generate.sh:
LLM="${CT_LLM_CMD:-claude}"
The generation step shells out to a coding-agent CLI. It does not know or care what that CLI talks to. So pointing the whole pipeline at a self-hosted model is a question about the CLI’s configuration, not about anything in this repo.
Claude Code takes two environment variables for exactly that:
export ANTHROPIC_BASE_URL="https://<your-endpoint>"
export ANTHROPIC_AUTH_TOKEN="<your key>"
With those set, ./generate.sh runs unchanged and the model that writes your sorter is yours.
Get an endpoint and a key
You need an OpenAI- or Anthropic-compatible endpoint serving at least one code-capable model, and a
key for it. This page doesn’t ship one — ask whoever runs your target deployment for the endpoint
URL and your personal key, the same way you’d ask for any other credential. Treat both the way you’d
treat any other secret: set them as environment variables, never commit them, never paste them into
AGENTS.md or any file that leaves your machine.
export ANTHROPIC_BASE_URL="<endpoint URL your operator gave you>"
export ANTHROPIC_AUTH_TOKEN="<your key>"
Sanity-check the two variables before touching generate.sh:
curl -s -o /dev/null -w "%{http_code}\n" "$ANTHROPIC_BASE_URL/v1/models" \
-H "Authorization: Bearer $ANTHROPIC_AUTH_TOKEN"
200 means you’re through. 401 means the key is wrong or missing. Anything else, ask your
operator before going further — there’s no point debugging the harness against an endpoint that
isn’t answering yet.
Not every coding-agent CLI supports pointing itself at an arbitrary endpoint the same way. As of
this writing: Claude Code and opencode do, directly; Codex does via a
model_providers block in ~/.codex/config.toml; Gemini CLI does not have a documented way to.
Since generate.sh defaults to claude, Claude Code is the short path — set the two variables above
and the rest of this page follows.
Why the vendor CLI needs help, and the fix
Reaching a self-hosted endpoint at all is not the same as getting a usable response back for this pipeline’s narrow, code-only prompt. The Claude Code CLI carries its own system prompt and tool-use scaffolding on every request — the part that makes it Claude Code rather than a bare completion client — regardless of which model answers on the other end. A model that was never trained against that framing can respond to it badly: short, generic, or truncated non-answers, even though the same model handles the exact same prompt content correctly when asked directly, with no scaffolding in the way.
The fix is inside what generate.sh already lets you control. It only requires $CT_LLM_CMD to
accept -p <prompt> --output-format text … and print the raw response — it never requires that
command to be the Claude Code CLI. Point it instead at a small wrapper that skips the CLI entirely
and talks straight to your endpoint’s own chat-completions API — no system prompt, no tools, no
CLAUDE.md auto-discovery:
#!/usr/bin/env python3
"""CT_LLM_CMD replacement: talks to a self-hosted OpenAI-compatible endpoint directly,
skipping the Claude Code CLI's own system prompt and tool-use scaffolding entirely.
"""
import json, os, sys, urllib.request
def parse_args(argv):
prompt = None
i = 0
while i < len(argv):
if argv[i] == "-p" and i + 1 < len(argv):
prompt, i = argv[i + 1], i + 2
else:
i += 1 # ignore --output-format, --disallowedTools, --effort, etc.
return prompt, os.environ.get("CT_LLM_DIRECT_MODEL", "your-model-name")
def main():
prompt, model = parse_args(sys.argv[1:])
url = os.environ["ANTHROPIC_BASE_URL"].rstrip("/") + "/v1/chat/completions"
body = json.dumps({"model": model, "messages": [{"role": "user", "content": prompt}],
"max_tokens": int(os.environ.get("CT_LLM_DIRECT_MAX_TOKENS", 4000)),
"temperature": float(os.environ.get("CT_LLM_DIRECT_TEMPERATURE", 0))}).encode()
req = urllib.request.Request(url, data=body, headers={
"Content-Type": "application/json",
"Authorization": f"Bearer {os.environ['ANTHROPIC_AUTH_TOKEN']}"})
with urllib.request.urlopen(req, timeout=120) as resp:
data = json.loads(resp.read())
sys.stdout.write(data["choices"][0]["message"]["content"])
if __name__ == "__main__":
main()
Save that next to your participant’s other scripts (e.g. ct-llm-direct.py), make it executable,
and point generate.sh at it:
export CT_LLM_CMD="./ct-llm-direct.py"
export CT_LLM_DIRECT_MODEL="<the model name your endpoint expects>"
./generate.sh
The wrapper defaults to temperature=0, and that default is load-bearing, not cosmetic.
Regenerating five times in a row against local-devstral-small2 at temperature=0 produced the
same handler byte-for-byte four times out of five — the one outlier differed only in comment
wording, not in the logic the property gates actually check. The same experiment without a fixed
temperature has no such floor: nothing stops two runs from picking genuinely different (if each
individually gate-passing) strategies, which turns “does my spec work against this model” into a
moving target. Override it with CT_LLM_DIRECT_TEMPERATURE only if you have a specific reason to
want sampling variance back.
One quoting trap worth knowing about in advance: generate.sh invokes "$LLM" -p "$PROMPT" …
with $LLM double-quoted. A CT_LLM_CMD containing a space — a wrapper path with a --model flag
tacked on, for instance — is executed as one literal command name including the space, not
command-plus-arguments, and fails with a bare “command not found” that never reaches the model at
all. That’s why the model above goes through the CT_LLM_DIRECT_MODEL environment variable instead
of a CLI flag: keep CT_LLM_CMD a single word, always.
If your CLI and model do get along without the wrapper — some do — you can skip it and set
ANTHROPIC_BASE_URL/ANTHROPIC_AUTH_TOKEN directly per What has to be true
above. Try the plain path first; reach for the wrapper if generation comes back short, generic, or
truncated for no reason the specification explains.
Getting there
- Make the local tutorial work first, exactly as written in Five seconds to your first sorter. Do not debug two unfamiliar things at once.
- Point
CT_LLM_CMDat your endpoint — directly, or through the wrapper above — and re-run./generate.sh. Everything downstream is unchanged, which is the point. - Run the same three gates you already know:
./handler.sh --selftest python3 "$REPO/dryrun.py" ./handler.sh --seed 42 --len 8 --quiet python3 "$REPO/dryrun.py" ./handler.sh --seed 42 --len 8 \ --require-adjacent --require-optimal-swaps --require-no-wasted-compares --quietA working self-hosted setup passes all three, reproducibly — run it more than once before you trust a single clean pass. What that looks like in practice, five back-to-back runs with the wrapper above:
Run generate.shAll three gates 1 10 s, compiles clean pass 2 10 s, compiles clean pass 3 11 s, compiles clean pass 4 10 s, compiles clean pass 5 10 s, compiles clean pass Every run:
rounds=18 swaps=17 comparisons=0— the same optimal-adjacent-swap count as the reference handler (swaps == inversions, the bound documented on The harness is the variable). NoAGENTS.mdrewriting needed to get there — the specification that already worked against a vendor model works here too. - Take it online. From here, going live is identical to every other participant on this site —
follow Join as a participant exactly as
written. The model behind your handler doesn’t change that flow at all. A real result from doing
exactly this, wrapper and
temperature=0included:comparisons=0 swaps=24 faults=0 rounds=25 wall=3.0s sorted=yesrounds = comparisons + swaps + 1→25 = 0 + 24 + 1— the same accounting identity every participant on this site is measured against. A handler generated entirely by a self-hosted model, through a harness with the vendor CLI’s own framing removed from the loop, sorted a real array correctly in the real hosted arena — and did it again, on a second array with a different inversion count, without touchingAGENTS.mdbetween the two runs:comparisons=0 swaps=30 faults=0 rounds=31 wall=3.2s sorted=yes
Measuring it without fooling yourself
A self-hosted endpoint is usually one GPU behind a queue, and that changes which numbers mean anything.
- Per-request latency stays meaningful. Median and spread of a single generation are a property of the model and the hardware, comparable to a vendor API.
- Throughput does not, if your endpoint serialises requests through a single GPU. Run five generations at once and you’ve measured the queue, not the model. Compare one at a time, and say which you did.
- The model may not be in memory when you ask. Idle-unload timeouts mean the first request after a pause pays a load cost on top of generation. Send a throwaway warm-up request before a series and discard it; log cold starts separately rather than folding them into a median.
- Sampling settings decide rule-following rates before anything else does. If you’re measuring
how often a generated handler violates a contract rule, that rate moves with
temperatureandtop_pbefore it moves with the model. The Claude Code CLI sends no sampling parameters at all — both sides run on their own server-side defaults, which you can’t equalise through this path. Label results as a statement about this deployment, not this model. The wrapper above setstemperature=0by default for exactly this reason — pin it before you draw any conclusion from a handful of runs, or every “it worked” / “it didn’t” you observe is entangled with which sample you happened to draw.
None of this makes the exercise pointless — it means the label on your numbers has to say what actually varied. When the measurement lies covers the general discipline in more depth.
Why this is the hard version
Three things get harder at once, and none of them is the model’s fault:
- The gates stop being decoration. A hosted frontier model tends to pass
--selfteston the first try. A smaller local one may not — and then--require-adjacent,--require-optimal-swapsand--require-no-wasted-comparesare what tells you which part of your specification was too vague, instead of leaving you with “it didn’t work”. - The specification may need to carry more — but check which layer actually failed before you
assume that. A failure that looks like a specification problem can just as easily belong to the
generation step (the CLI wrapper, in particular) instead. A control-arm test — send the same
prompt content directly to the endpoint, no CLI in the way — tells you which one it actually is
before you spend a round of
AGENTS.mditeration on the wrong layer. - Failure moves to where you can see it. A generation that fails against your own endpoint fails at build time, with a compile error you can read, on hardware you control.
If your strategy survives all three, you’ve shown something stronger than a working sorter: that the specification, the gates and the contract were doing the work — because the one thing you swapped is the part everyone assumes is decisive.