Model · s1-llm-auto-router
Classify a request before you route it
s1-llm-auto-router is a small CPU model that answers the seven questions an LLM router needs: what kind of work a request is, how hard it is, how costly a mistake would be, and whether it needs tools, vision, long context or the previous turn. All seven answers come back from one request in about 14 ms (p50, on the node). It costs $0.001 (€0.001) per 1M input tokens in both tiers, and output is free.
This model answers only the seven questions below, using these exact option sets. The model does not read the question text. A request with other question names, types or options returns 400 invalid_request. For general typed decisions, use s1-fast or s1-pro. This is the model with seven fixed routing questions; s1-pro also accepts up to three text questions.
Questions
| Name | Type | Options |
|---|---|---|
category | choice | coding, agentic, math, knowledge, long_context, tool_use, design, summarisation, general (all nine keys required) |
difficulty | score | 5 levels: trivial, easy, moderate, hard, frontier |
stakes | score | 4 levels: negligible, low, medium, high |
needs_tools, needs_vision, needs_long_context, follow_up | noul | yes/no probability |
Put the user turn in state.request and an optional short summary of the conversation in state.context. The model reads up to 512 tokens and the first 600 characters of context. You can send any subset of the seven questions; the price is per input token, not per question.
From signals to a concrete model
The classifier returns task signals instead of model names so it works with any model catalog. When models, prices or availability change, update your routing configuration without retraining the classifier.
Each example below makes two real API calls: classify with System1 Models, then generate an answer with an independently configured TensorX chat model. Set S1M_API_KEY and TENSORX_API_KEY in your environment; each provider bills its own call. The IDs below were checked against https://api.tensorx.ai/v1/models on 2 October 2026; replace them with models available in your own catalog.
| Category | Expected difficulty | Model ID (TensorX) |
|---|---|---|
| coding / general | below 2 (easy bucket) | deepseek/deepseek-v3.2 |
| coding / general | 2 or above (hard bucket) | deepseek/deepseek-v4.1-flash |
| Any other category or missing table entry | any | deepseek/deepseek-v4.1-flash (fallback) |
This small table demonstrates a configurable policy. Low category confidence (below 0.6) or expected stakes of 2 or above also selects the fallback; this is an example policy you can tune.
Read the probabilities correctly
answers.category.choice is the most likely category; probabilities contains the full distribution and confidence is that choice's probability. Difficulty and stakes score are expected ordinal levels (0–4 and 0–3), rather than probabilities or rounded labels: their probabilities objects use string level keys, and legend maps those keys to labels. A flag's noul is the probability that the flag is true (0–1), with false probability 1 - noul. Each snippet logs all these signals, the request ID, selected model and reason, without recording the prompt or API keys.
The examples handle a new text turn. If tools, vision, long context or follow-up probability reaches 0.5, they log the decision and stop before forwarding; configure a capable model, tools, images and conversation history for those routes. The classifier itself accepts text only; it cannot inspect an image.
curl + jq
Save as route.sh and run bash route.sh. Requires bash, curl and jq. Download exact script.
#!/usr/bin/env bash
# Requires bash, curl, jq, S1M_API_KEY and TENSORX_API_KEY.
set -euo pipefail
: "${S1M_API_KEY:?Set S1M_API_KEY}" "${TENSORX_API_KEY:?Set TENSORX_API_KEY}"
request='Write a Python function that returns the square of a number.'
fast='deepseek/deepseek-v3.2'
strong='deepseek/deepseek-v4.1-flash'
# Edit this category x difficulty table and fallback for your own catalog.
routes=$(jq -n --arg fast "$fast" --arg strong "$strong" '{coding:{easy:$fast,hard:$strong},general:{easy:$fast,hard:$strong}}')
classifier_payload=$(jq --arg request "$request" '.state.request = $request' <<'CLASSIFIER_JSON'
{"model":"s1-llm-auto-router","state":{"request":"Write a Python function that returns the square of a number.","context":"(new conversation)"},"questions":{"category":{"type":"choice","criteria":{"coding":"","agentic":"","math":"","knowledge":"","long_context":"","tool_use":"","design":"","summarisation":"","general":""}},"difficulty":{"type":"score","criteria":["trivial","easy","moderate","hard","frontier"]},"stakes":{"type":"score","criteria":["negligible","low","medium","high"]},"needs_tools":{"type":"noul"},"needs_vision":{"type":"noul"},"needs_long_context":{"type":"noul"},"follow_up":{"type":"noul"}}}
CLASSIFIER_JSON
)
result=$(curl --fail-with-body --silent --show-error --max-time 60 \
https://api.system1models.ai/v1/systemone \
-H "Authorization: Bearer $S1M_API_KEY" -H 'Content-Type: application/json' \
--data-binary "$classifier_payload")
decision=$(jq --argjson routes "$routes" --arg fallback "$strong" '
.answers as $a | (if $a.difficulty.score >= 2 then "hard" else "easy" end) as $level |
([$a | to_entries[] | select(.key == "needs_tools" or .key == "needs_vision" or
.key == "needs_long_context" or .key == "follow_up") | select(.value.noul >= 0.5) | .key]) as $unsupported |
{request_id:.id, signals:$a, unsupported:$unsupported,
model:(if $a.category.confidence < 0.6 or $a.stakes.score >= 2 then $fallback
else ($routes[$a.category.choice][$level] // $fallback) end),
reason:(if $a.category.confidence < 0.6 or $a.stakes.score >= 2 then
"low confidence or high stakes" else "category/difficulty table" end)}' <<<"$result")
printf '%s\n' "$decision" # distributions + flag probabilities; no prompts/keys
if [ "$(jq '.unsupported | length' <<<"$decision")" != 0 ]; then
echo 'Configure capability support / conversation history before forwarding' >&2; exit 1
fi
model=$(jq -r '.model' <<<"$decision")
payload=$(jq -n --arg model "$model" --arg request "$request" '{model:$model,messages:[{role:"user",content:$request}],max_tokens:80,temperature:0}')
curl --fail-with-body --silent --show-error --max-time 60 https://api.tensorx.ai/v1/chat/completions -H "Authorization: Bearer $TENSORX_API_KEY" -H 'Content-Type: application/json' --data-binary "$payload" | jq -r '.choices[0].message.content'
Python
Save as route.py and run python3 route.py. Python 3.10+, standard library only. Download exact script.
import json, os, urllib.request
# Requires S1M_API_KEY and TENSORX_API_KEY; Python 3.10+ (standard library only).
REQUEST = "Write a Python function that returns the square of a number."
QUESTIONS = {'category': {'type': 'choice', 'criteria': {'coding': '', 'agentic': '', 'math': '', 'knowledge': '', 'long_context': '', 'tool_use': '', 'design': '', 'summarisation': '', 'general': ''}}, 'difficulty': {'type': 'score', 'criteria': ['trivial', 'easy', 'moderate', 'hard', 'frontier']}, 'stakes': {'type': 'score', 'criteria': ['negligible', 'low', 'medium', 'high']}, 'needs_tools': {'type': 'noul'}, 'needs_vision': {'type': 'noul'}, 'needs_long_context': {'type': 'noul'}, 'follow_up': {'type': 'noul'}}
FAST = "deepseek/deepseek-v3.2"
STRONG = "deepseek/deepseek-v4.1-flash"
ROUTES = {"coding": {"easy": FAST, "hard": STRONG},
"general": {"easy": FAST, "hard": STRONG}}
FALLBACK = STRONG
def post(url, key, body):
req = urllib.request.Request(url, data=json.dumps(body).encode(), headers={
"Authorization": "Bearer " + key, "Content-Type": "application/json"})
with urllib.request.urlopen(req, timeout=60) as response:
return json.load(response)
result = post("https://api.system1models.ai/v1/systemone", os.environ["S1M_API_KEY"],
{"model": "s1-llm-auto-router", "state": {"request": REQUEST,
"context": "(new conversation)"}, "questions": QUESTIONS})
a = result["answers"]
# score is an expected LEVEL (0..4 / 0..3), not a probability.
difficulty = "hard" if a["difficulty"]["score"] >= 2 else "easy"
model = ROUTES.get(a["category"]["choice"], {}).get(difficulty, FALLBACK)
reason = "category/difficulty table"
if a["category"]["confidence"] < 0.6 or a["stakes"]["score"] >= 2:
model, reason = FALLBACK, "low confidence or high stakes"
# This demo handles a new text turn. Never silently lose required capabilities/context.
unsupported = [f for f in ("needs_tools", "needs_vision", "needs_long_context", "follow_up")
if a[f]["noul"] >= 0.5]
decision = {"request_id": result["id"], "model": model, "reason": reason,
"signals": a, "unsupported": unsupported}
print(json.dumps(decision)) # All distributions and flag probabilities; no prompts or keys.
if unsupported:
raise SystemExit("Configure capability support / conversation history before forwarding")
completion = post("https://api.tensorx.ai/v1/chat/completions", os.environ["TENSORX_API_KEY"],
{"model": model, "messages": [{"role": "user", "content": REQUEST}],
"max_tokens": 80, "temperature": 0})
print(completion["choices"][0]["message"]["content"])
TypeScript
Save as route.ts and run node route.ts with Node 22.18+ (built-in type stripping; no dependencies). Download exact script.
// Node 22.18+ supports: node route.ts (no dependencies).
const REQUEST = "Write a Python function that returns the square of a number.";
const QUESTIONS = {"category": {"type": "choice", "criteria": {"coding": "", "agentic": "", "math": "", "knowledge": "", "long_context": "", "tool_use": "", "design": "", "summarisation": "", "general": ""}}, "difficulty": {"type": "score", "criteria": ["trivial", "easy", "moderate", "hard", "frontier"]}, "stakes": {"type": "score", "criteria": ["negligible", "low", "medium", "high"]}, "needs_tools": {"type": "noul"}, "needs_vision": {"type": "noul"}, "needs_long_context": {"type": "noul"}, "follow_up": {"type": "noul"}};
const FAST = "deepseek/deepseek-v3.2";
const STRONG = "deepseek/deepseek-v4.1-flash";
const ROUTES: Record<string, Record<string, string>> = {
coding: {easy: FAST, hard: STRONG}, general: {easy: FAST, hard: STRONG}
};
const FALLBACK = STRONG;
function key(name: string): string {
const value = process.env[name]; if (!value) throw new Error(`Set ${name}`); return value;
}
async function post(url: string, token: string, body: unknown): Promise<any> {
const response = await fetch(url, {method: "POST", headers: {
Authorization: `Bearer ${token}`, "Content-Type": "application/json"},
body: JSON.stringify(body), signal: AbortSignal.timeout(60000)});
if (!response.ok) throw new Error(`API returned HTTP ${response.status}`);
return response.json();
}
const result = await post("https://api.system1models.ai/v1/systemone", key("S1M_API_KEY"),
{model: "s1-llm-auto-router", state: {request: REQUEST, context: "(new conversation)"},
questions: QUESTIONS});
const a = result.answers;
// score is an expected LEVEL (0..4 / 0..3), not a probability.
const difficulty = a.difficulty.score >= 2 ? "hard" : "easy";
let model = ROUTES[a.category.choice]?.[difficulty] ?? FALLBACK;
let reason = "category/difficulty table";
if (a.category.confidence < 0.6 || a.stakes.score >= 2) {
model = FALLBACK; reason = "low confidence or high stakes";
}
const unsupported = ["needs_tools", "needs_vision", "needs_long_context", "follow_up"]
.filter(f => a[f].noul >= 0.5);
console.log(JSON.stringify({request_id: result.id, model, reason, signals: a, unsupported}));
if (unsupported.length) throw new Error("Configure capability support / conversation history before forwarding");
const completion = await post("https://api.tensorx.ai/v1/chat/completions", key("TENSORX_API_KEY"),
{model, messages: [{role: "user", content: REQUEST}], max_tokens: 80, temperature: 0});
console.log(completion.choices[0].message.content);
export {};
Use it from auto-model-router
auto-model-router uses the same seven signals with its catalog, capability, cost and quota policy. The local CPU backend s1-llm-auto-router and hosted backend s1-llm-auto-router-api are in PR #5 (pending merge, checked on 2 October 2026); install the pinned revision below until that PR is released. Open weights are on Hugging Face under the MIT licence.
git clone https://github.com/fstandhartinger/auto-model-router.git
cd auto-model-router
git fetch origin pull/5/head
git checkout a92814ecd17cb98f3f04c1f866bde8f3592f0b9c
python3 -m venv .venv
. .venv/bin/activate
pip install -e .
cp examples/models.yaml router.local.yaml
# Edit providers/models in router.local.yaml for your own catalog and keys.
# Set S1M_API_KEY in the environment, then put the YAML below under policy.
export AUTO_ROUTER_CONFIG="$PWD/router.local.yaml"
uvicorn auto_router.server:app --host 127.0.0.1 --port 8787policy:
name: F_expected
classifier:
backend: s1-llm-auto-router-api
api_key_env: S1M_API_KEY
verify:
enabled: false # start with selection only; configure an answer judge separately
# Keep the providers/models blocks from your own edited examples/models.yaml.curl --fail-with-body http://127.0.0.1:8787/v1/chat/completions -H 'Content-Type: application/json' -d '{"model":"auto","messages":[{"role":"user","content":"Write a Python function that returns the square of a number."}],"max_tokens":80}'
curl --fail-with-body http://127.0.0.1:8787/v1/router/decisionsmodel: auto lets the router choose among your configured models; /v1/router/decisions returns its decision records. The hosted backend returns classifier signals to the existing policy; it does not replace your model catalog. To run locally instead, install pip install -e ".[s1-router]" at the same revision and set policy.classifier.backend: s1-llm-auto-router.
Measured quality
We measured it on a held-out set of 2,175 synthetic routing requests in 22 languages, including SAP ABAP and PL/I/COBOL. Category accuracy was 89.8 % (hosted Jev 1.13: 84.5 %; Winnow-12B: 84.1 %). In 74.2 % of cases the router picked the same model from its answers as from the gold labels (Jev: 63.0 %). On 100 requests labelled independently by Claude Opus, category accuracy was 0.88 (Jev: 0.80). Expected calibration error was 0.018 for category.
Limits
- The training requests are synthetic, and the labels come from an ensemble of open LLMs. Real traffic may differ.
- It is weakest on Greek, Korean and Italian, at about 0.73–0.78 category accuracy.
- Text only; no images.
- It is a System1 Models model. If we benchmark it, it is labelled as ours.
It runs on dedicated CPU cores in Finland (EU) and is available in the Peer-to-Peer and EU tiers. Prompts and answers are not stored.