The Fix

Stop routing everything to Claude

You hit your usage limit at 2pm and blame the limit. I logged a week of my own calls instead. Over half of them were mechanical work that any small model does identically, for free.

E
Endi
Engineer · Buka.labs
14 min read

I ran out of usage three days in a row and got annoyed enough to actually measure it. A week of logging every call, tagged by what the task was. The split was not close.

Roughly 60% of what I was sending to a frontier model was mechanical: rename these symbols, convert this JSON to a type, extract the URLs from this blob, summarise this log, classify this comment. Work where there is one correct answer and finding it needs no judgment about my codebase at all.

The rule that decides where a call goes

Forget benchmarks for this. The only question that matters is practical:

Does this task need to hold my whole system in its head, or is it a local transform where the input contains everything needed to produce the output?

Local transform, small model. Needs to reason across files, weigh trade-offs, or decide something you'd have to think about yourself? Frontier model. That's the whole rule and it's right almost every time.

What that looks like concretely

Route down to a cheap or local model:

  • Format conversion: JSON to types, CSV to SQL, YAML to env
  • Extraction: pull the errors out of this build log
  • Classification: is this comment a question or praise
  • Mechanical renames and single-file boilerplate
  • First-pass summaries you're going to read anyway

Keep on the frontier model:

  • Anything touching more than two files at once
  • Debugging where the cause isn't in the file with the symptom
  • Architecture and “should we even do it this way”
  • Code review, because catching the subtle one is the entire point
  • Anything where a plausible wrong answer costs you an hour

Wiring it up

The router is a config, not a service. Point the cheap tier at a local model or a free API, keep the default on your main model, and give it a task-type hint:

router.json
{
  "default": "claude",
  "tiers": {
    "claude": { "provider": "anthropic", "model": "claude-sonnet-5" },
    "local":  { "provider": "ollama",    "model": "llama3.2:3b" }
  },
  "routes": [
    { "task": "extract",   "tier": "local" },
    { "task": "classify",  "tier": "local" },
    { "task": "convert",   "tier": "local" },
    { "task": "summarize", "tier": "local" },
    { "task": "review",    "tier": "claude" },
    { "task": "debug",     "tier": "claude" },
    { "task": "design",    "tier": "claude" }
  ]
}

The part people skip: log which tier served each call and what it cost. Without that you're guessing about your own usage, which is exactly how you got here.

route.ts
export function pickTier(task: string, input: string) {
  // A local transform can't be local if the input doesn't fit.
  if (estimateTokens(input) > 8_000) return "claude";

  const rule = config.routes.find((r) => r.task === task);
  return rule?.tier ?? config.default;
}

The failure mode to watch for

Small models fail quietly. A frontier model that can't do something tends to say so. A 3B model hands you a confident, clean, wrong answer, and because the task was “just” an extraction, you don't check it.

So route by task type, but never route work whose output goes straight into something else unchecked. If nothing downstream would catch the error, it stays on the good model regardless of how mechanical it looks.

What actually changed

I stopped hitting the limit. Not because I use less AI (I use noticeably more) but because the frontier model is now spending its budget on the work that needed it. The mechanical half runs on my own machine for nothing, and honestly returns faster.

The FixClaude CodeBuka.labs