One endpoint, three models. Every request goes to a judge model that estimates whether a cheap model can handle it — easy work stays on a 3B local model, harder work runs on a 35B local model, and only genuinely hard requests escalate to Claude Opus. Below is a real recorded run: each card shows the question, the router's decision, and which model actually answered.