I ran out of usage three days in a row and got annoyed enough to actually measure it. A week of logging every call, tagged by what the task was. The split was not close.
Roughly 60% of what I was sending to a frontier model was mechanical: rename these symbols, convert this JSON to a type, extract the URLs from this blob, summarise this log, classify this comment. Work where there is one correct answer and finding it needs no judgment about my codebase at all.
The rule that decides where a call goes
Forget benchmarks for this. The only question that matters is practical:
Does this task need to hold my whole system in its head, or is it a local transform where the input contains everything needed to produce the output?
Local transform, small model. Needs to reason across files, weigh trade-offs, or decide something you'd have to think about yourself? Frontier model. That's the whole rule and it's right almost every time.
What that looks like concretely
Route down to a cheap or local model:
- Format conversion: JSON to types, CSV to SQL, YAML to env
- Extraction: pull the errors out of this build log
- Classification: is this comment a question or praise
- Mechanical renames and single-file boilerplate
- First-pass summaries you're going to read anyway
Keep on the frontier model:
- Anything touching more than two files at once
- Debugging where the cause isn't in the file with the symptom
- Architecture and “should we even do it this way”
- Code review, because catching the subtle one is the entire point
- Anything where a plausible wrong answer costs you an hour
Wiring it up
The router is a config, not a service. Point the cheap tier at a local model or a free API, keep the default on your main model, and give it a task-type hint:
{
"default": "claude",
"tiers": {
"claude": { "provider": "anthropic", "model": "claude-sonnet-5" },
"local": { "provider": "ollama", "model": "llama3.2:3b" }
},
"routes": [
{ "task": "extract", "tier": "local" },
{ "task": "classify", "tier": "local" },
{ "task": "convert", "tier": "local" },
{ "task": "summarize", "tier": "local" },
{ "task": "review", "tier": "claude" },
{ "task": "debug", "tier": "claude" },
{ "task": "design", "tier": "claude" }
]
}The part people skip: log which tier served each call and what it cost. Without that you're guessing about your own usage, which is exactly how you got here.
export function pickTier(task: string, input: string) {
// A local transform can't be local if the input doesn't fit.
if (estimateTokens(input) > 8_000) return "claude";
const rule = config.routes.find((r) => r.task === task);
return rule?.tier ?? config.default;
}The failure mode to watch for
Small models fail quietly. A frontier model that can't do something tends to say so. A 3B model hands you a confident, clean, wrong answer, and because the task was “just” an extraction, you don't check it.
So route by task type, but never route work whose output goes straight into something else unchecked. If nothing downstream would catch the error, it stays on the good model regardless of how mechanical it looks.
What actually changed
I stopped hitting the limit. Not because I use less AI (I use noticeably more) but because the frontier model is now spending its budget on the work that needed it. The mechanical half runs on my own machine for nothing, and honestly returns faster.