Local LLM Demo — Docs
Tutorial

Call the API and point Claude Code at it

Every command and response below is real — copy-pasted from an actual session against the live endpoint at https://llm-34a13a96.bunsenbrenner.org, not paraphrased.

1. Get a scoped key

Ask the operator to generate one (or, if you hold the LiteLLM master key yourself):

curl -s http://127.0.0.1:4001/key/generate \
  -H "Authorization: Bearer $LITELLM_MASTER_KEY" -H "Content-Type: application/json" \
  -d '{"models": ["local-qwen3-coder"], "user_id": "you@example.com", "key_alias": "llm-demo-you"}'

You get back a key like sk-REPLACE_WITH_YOUR_KEY, scoped to exactly one model. Try it against a model it’s not allowed to use and confirm it’s rejected — this is the access-control guarantee the whole setup rests on:

curl -s https://llm-34a13a96.bunsenbrenner.org/v1/chat/completions \
  -H "Authorization: Bearer sk-REPLACE_WITH_YOUR_KEY" -H "Content-Type: application/json" \
  -d '{"model":"cf-llama-70b","messages":[{"role":"user","content":"hi"}]}'
{"error":{"message":"key not allowed to access model. This key can only access models=['local-qwen3-coder']. Tried to access cf-llama-70b","type":"key_model_access_denied","param":"model","code":"401"}}

measured Reproduced 2026-08-29: a scoped key is refused for cf-llama-70b with this exact key_model_access_denied message. Note the proxy now returns HTTP 403 for this case; the captured body above shows the older 401.

2. Make a real call

curl -s https://llm-34a13a96.bunsenbrenner.org/v1/chat/completions \
  -H "Authorization: Bearer sk-REPLACE_WITH_YOUR_KEY" -H "Content-Type: application/json" \
  -d '{"model":"local-qwen3-coder","messages":[{"role":"user","content":"Reply with exactly the word: pong"}],"max_tokens":20}'
{"id":"chatcmpl-3980139d-136f-404f-a82b-aaf5165425a9","created":1786787743,"model":"local-qwen3-coder","object":"chat.completion","choices":[{"finish_reason":"stop","index":0,"message":{"content":"Hi!","role":"assistant"}}],"usage":{"completion_tokens":3,"prompt_tokens":18,"total_tokens":21}}

That’s a real request, over the real tunnel, hitting real local GPU inference (qwen3-coder:30b on the origin’s TITAN RTX). Round-trip latency for a short reply like this is typically well under a second — see capability-probe/transcripts/ in the code repo for a wider sample.

audited This completion — and the Claude Code run below — are quoted verbatim from a real session against the live model, not re-issued here. The latency figure is the maintainer's observation, not independently re-timed.

3. Point Claude Code at it

The endpoint also answers Anthropic’s /v1/messages shape (LiteLLM translates it), so Claude Code itself can talk to it directly:

ANTHROPIC_BASE_URL="https://llm-34a13a96.bunsenbrenner.org" \
ANTHROPIC_AUTH_TOKEN="sk-REPLACE_WITH_YOUR_KEY" \
CLAUDE_CODE_DISABLE_UNKNOWN_MODEL_WINDOW_ENFORCEMENT=1 \
  claude -p "Reply with exactly the word: pong" --model local-qwen3-coder

Two real things happened here that are worth pausing on:

What you’ve verified

Next: add another authorized person or read the model and endpoint reference.

Found an error, or something that didn't work as documented? Open an issue →