There are three honest answers to “which model should serve this request”.
Pin one model and send everything to it. Orchestrate across a pool and match the model to the request. Or send everything to a frontier model and stop thinking about it.
Most systems I see have picked the third without deciding to. It is the default that emerges when nobody owns the question. Over an evening of testing I costed all three from billed data, and the gap between the considered answers and the unconsidered one is two orders of magnitude.
I build an orchestration platform, so treat the recommendation at the end accordingly. The numbers are from my own meter and I have shown the workings.
Key Takeaways
- A frontier-by-default routing strategy is usually not a decision at all, it is what happens when nobody owns the model selection question, and it cost about 124 times more than the small-model path on matched work.
- Pricing across models varies by a 460 fold spread end to end, and by five fold inside the small model tier alone, so “use a small model” is the start of a decision, not the end of one.
- A pinned model gives predictable cost and tight latency, but it is a bet on traffic shape that quietly goes wrong as pricing and models change underneath it.
- An orchestrated pool matched a hand pinned model on accuracy while adding automatic failover, so a struggling or unavailable provider gets substituted without paging anyone.
- In this dataset, frontier model calls were 7% of total traffic but 47% of total spend, which is usually the single largest line item hiding in an AI bill.
- What the Billed Data Says
- The Interesting Number Is Inside the Small Tier
- Pinned: Predictable, and Correct on the Day You Picked It
- Orchestrated: the Right Model Per Request, Without Owning the Problem
- The Axis I Did Not Test, and Should Have
- Frontier by Default: the Expensive Non-Decision
- Latency Has Two Separate Stories and Mixing Them Misleads
- Enforce the Schema at the Boundary
- Where I Would Put Each One
- A Rule to Build Against
- Limits
- Where WrangleAI Fits
- FAQs
What the Billed Data Says
384 metered gateway calls on 19 September, plus TypeSafe usage tokens priced at their published rate. Not pricing-page arithmetic.
| Calls | Total $ | Avg $/call | |
| Small-model tier | 265 | $0.002083 | $0.0000079 |
| Typed judgment (Jev) | 242 | $0.004238 | $0.0000175 |
| Frontier (Opus 4.6, Sonnet 4.5) | 26 | $0.020777 | $0.000799 |
A small model costs about 2.2 times less per call than a typed judgment API, and about 100 times less than a frontier call. On the matched twelve-ticket department window, where every arm did identical work, frontier came to roughly 124 times the small-model cost for the same labels.
That 100-fold gap is the headline, but it is not the interesting number.
The Interesting Number Is Inside the Small Tier
Blended dollars per million tokens, measured from the meter across every model my gateway touched that week:
| Model | Blended $/MTok |
| Mistral-Nemo | $0.023 |
| Llama-3.1-8B-Turbo | $0.025 |
| gemma-3-4b | $0.058 |
| Mistral-Small-24B | $0.059 |
| gpt-oss-20b | $0.082 |
| Qwen3-32B | $0.105 |
| phi-4 | $0.135 |
| gemini-2.5-flash-lite | $0.140 |
| Qwen2.5-72B | $0.366 |
| gpt-5-mini | $0.905 |
| gemini-2.5-flash | $1.232 |
| gemini-2.5-pro | $4.871 |
| claude-opus-4.6 | $7.608 |
| claude-sonnet-4.5 | $10.774 |
A 460-fold spread end to end. A five-fold spread inside the small tier alone, where Llama-3.1-8B came in at less than half the cost of the smaller gemma-3-4b, because provider economics do not track parameter count in any way you can reason about from the outside.
This is why “use a small model” is not an architectural decision. It is the beginning of one. The decision is which small model for which request shape, and that answer has a shelf life measured in weeks.
Quick link: Top AI Observability Platforms of All Time
Pinned: Predictable, and Correct on the Day You Picked It
Pinning gemma-3-4b for department triage gave 11 out of 12 accuracy at 960 milliseconds average, 100% parse success, and $0.0000067 per call on the matched window. Same model every time, so the cost is predictable to the sixth decimal place and the latency distribution is tight.
That predictability is the whole argument, and it is a good one for a single high-volume path with one known request shape and a tight latency SLO.
Two things go wrong with it over time. The first is that pinning is a bet on the shape of your traffic: the moment a request arrives that a 4B model cannot handle, you have no graceful answer, so you either accept the degraded result or you hand-build escalation logic, which is orchestration, badly. The second is the pricing table above. You pinned the right model in September. Nobody will tell you in December that it stopped being the right one, because nothing will break. The bill just drifts.
Orchestrated: the Right Model Per Request, Without Owning the Problem
The orchestrated arm matched the pinned model on accuracy, 11 out of 12, with 100% parse success, and its single miss was the same ticket the pinned model missed: a sales upgrade classified as billing. Valid JSON, correct shape, wrong label. That is generative drift, identical in both arms, and not a routing fault.
What orchestration buys is reach and maintenance. The gateway selected across gemma, Qwen and Flash depending on the request, which means a workload with genuinely mixed difficulty gets matched rather than compromised. A pinned 4B cannot do that. Pinning a frontier model to get it costs 124 times more on the easy majority. And when a provider reprices or a new checkpoint lands, the selection layer absorbs it rather than your backlog.
The cost is latency variance: 4 seconds average against the pinned model’s 0.96, with tails when the router reaches for a larger reasoning model. You are trading a tight latency distribution for coverage across difficulty. That is the right trade for a mixed workload and the wrong one for a single uniform high-volume path. Both are legitimate, they are just different problems.

The Axis I Did Not Test, and Should Have
My harness ran for an evening against providers that stayed up. So this next part is an architectural argument rather than a measurement, and I am labelling it as such.
A pinned model is a single point of failure with a vendor’s uptime attached to it. When that provider has a bad hour, rate-limits you, deprecates the checkpoint, or degrades quietly under load, your gate layer goes down with it. The usual response is a hand-rolled fallback: a try or except block, a hard-coded second model, a retry policy someone wrote in a hurry.
That code is easy to write and hard to keep correct, because it needs a current view of which models are near-peers on quality, which are available right now, and which of your policy constraints still hold when you substitute one for another. It is also the code that gets tested least, because it only runs when something is already going wrong.
Orchestration turns that from application code into infrastructure. If a model is unavailable, the selection layer substitutes a near-peer from the pool and the request completes. Your application never sees the failure, and nobody is woken up to edit a constant and redeploy.
There is a second-order version of this that matters more than uptime. When a model call fails mid-flow, the common failure is not a clean error, it is a partially processed request: a payload that has left your estate, a gate that never returned a verdict, a record written in an inconsistent state. If your fallback path is a hastily written retry, it is also the path where data handling rules are most likely to be skipped, because the redaction and policy checks live in the happy path somebody wrote first.
Failover done properly applies the same DLP and policy controls to the substitute call as to the original. Failover done as an afterthought is precisely where the payload you were careful about leaks to a provider you never assessed.
Three things a selection layer is buying you, then, and only one of them is cost: engineering time you do not spend building and maintaining fallback logic, downtime you do not take when one provider has a bad day, and the incident you do not have to explain to a regulator because the substitute call was subject to the same controls as the original.
Frontier by Default: the Expensive Non-Decision
Opus scored 12 out of 12 and averaged $0.000799 per call against $0.0000079 for the small tier, at 2.3 seconds against 0.96.
Ceiling quality on hard cases is real. The question is what share of your traffic is actually hard. In this bakeoff, none of it was: every ticket was routed correctly by a logistic regression running in two milliseconds for free.
If you cannot state what fraction of your requests need frontier reasoning, you are not choosing frontier. You are defaulting to it, and in my logs the frontier rows were 7% of the calls and 47% of the spend.
Latency Has Two Separate Stories and Mixing Them Misleads
One judgment. The typed API and a pinned small model are usually within 15% of each other: 647 against 703 milliseconds on baseline, 671 against 686 on contested, 656 against 600 on the stress suite where the small model was faster. Peers.
Eight judgments about the same input. One typed call carrying all eight questions returns in about 751 milliseconds. Eight sequential small-model calls take about 22 seconds, roughly 29 times slower. Batching into one prompt brings it back to 1.3 seconds and makes all eight answers dependent on a single blob parsing cleanly.
Whichever story you are telling, say which. A gate layer that needs six or eight signals per request has a completely different latency profile from one that needs a single label, and a benchmark quoting single-call medians tells you nothing useful about the first case.
Enforce the Schema at the Boundary
The consumer of a routing decision is code, not a human, and that changes what you need from the output.
Score parse failures as failures in your own benchmarks. Do not repair the output before scoring it, because the repair layer is a real thing you will write and maintain, and hiding it in the harness means you never see its cost. Set the response format explicitly rather than relying on prompt instructions. Log the finish reason on every call, because a truncated response and a wrong response look identical downstream and need different fixes.
Across all five arms in my run, parse success was 100%. That was a property of enforcing the schema, not of the models being reliable by nature.
Where I Would Put Each One
| Pinned small model | Orchestrated | Typed judgment | Frontier | |
| Best for | One known shape, very high volume | Mixed difficulty, no appetite to own model selection | Policy gates, multi-judgment packs | Genuinely hard cases, low volume |
| Cost per call | $0.0000079 | Between pinned and frontier | $0.0000175 | $0.000799 |
| Latency profile | Tight, ~1 s | Wider, ~4 s with tails | Tight, ~0.7 s, flat under fan-out | ~2.3 s |
| Maintenance | You re-benchmark when prices move | Absorbed by the selection layer | Edit criteria, no retraining | None, you just pay |
| If the provider is down | You are down, or your fallback code runs | Near-peer substituted automatically | Vendor uptime | You are down |
| Fails by | Silent wrong label | Silent wrong label | Wrong typed answer, no parse risk | Cost |
A Rule to Build Against
Build the gate layer first and price it from your own billed data rather than a pricing page. Pin the model for any path with a single known shape and real volume, and diarise a re-benchmark. Orchestrate when difficulty genuinely varies, or when nobody on the team wants to own the selection problem permanently. Reserve the frontier model for cases you can name in advance as hard, and measure what share of traffic actually meets that bar.
Then check that number. If it turns out to be close to zero, as it was here, you have just found the largest line item in your inference bill.
Limits
Small gold sets, many clear-cut cases, twelve tickets in the bakeoff. Synthetic tickets rather than production drift. Jev costs derived from tokens times a published rate rather than billed. A handful of frontier rows showed $0, which I read as incomplete metering. Every row was marked optimised: false, so none of these figures reflect a caching or optimisation layer.
Directionally I would stand behind all of it.
Where WrangleAI Fits
This post is the problem I built the company around, so here it is plainly.
The pricing table above is the reason. Fourteen models, a 460-fold cost spread, and the cheapest competent option for a given request changing every few weeks. No engineering team should be carrying that as a standing maintenance job, and most teams resolve it by pinning one model and never revisiting it, or by defaulting to frontier and absorbing a bill they cannot explain.
SmartRouter is the alternative. One API, a managed pool of small models behind it, semantic classification choosing which one handles each request, and escalation to larger models only where the request warrants it. You stop making model-selection decisions and you stop re-making them. In this experiment the orchestrated arm matched a hand-pinned model on accuracy with full parse reliability, while retaining reach into the mid tier that a pinned 4B does not have.
Resilience comes with it rather than being a project of its own. If a model is unavailable or degraded, SmartRouter substitutes a near-peer from the pool in-flight, at speed, with no fallback logic in your application and no pager going off. The controls travel with the substitution: hard budget stops that cannot be exceeded, DLP redaction before a payload leaves your estate, and policy rules applied ahead of the model rather than after it. That last point is the one worth dwelling on, because the fallback path is exactly where hand-rolled retry code tends to skip the checks that the happy path applies.
Saved developer time, avoided downtime, and a payload that never reaches an unassessed provider because a retry went sideways at three in the morning. The first two show up in your roadmap. The third only shows up when it goes wrong, which is the worst possible time to discover you built the fallback in a hurry.
Meter produced every number in this series. Billed cost per call, per model, per cost centre, with forecasting on top. The 47%-of-spend-on-7%-of-calls figure above took one query. Without that view, the frontier default stays invisible, which is exactly why it persists.
If this series has convinced you that model choice is a real architectural decision rather than a config value, the follow-on question is who maintains it. That is us. See the platform or talk to me.
FAQs
What is the difference between a pinned model and an orchestrated one?
A pinned model uses one fixed model for every request, which gives predictable cost and tight latency but no graceful answer when a request does not fit it. An orchestrated setup selects from a pool per request, matching mixed difficulty and absorbing pricing or availability changes automatically.
Why is defaulting to a frontier model expensive?
In this dataset, frontier calls made up 7% of traffic but 47% of total spend, and cost roughly 124 times more per call than a small model on matched work. Ceiling quality is real for genuinely hard cases, but most traffic does not need it.
What happens if a pinned model’s provider goes down?
A pinned model is a single point of failure tied to one vendor’s uptime, so an outage takes the gate layer down with it unless hand-built fallback logic exists. An orchestrated pool substitutes a near-peer model automatically when one becomes unavailable.
Does model routing affect latency?
Yes. A pinned small model averaged about 1 second per call, while an orchestrated pool averaged closer to 4 seconds with occasional longer tails, since it sometimes reaches for larger models. That trade buys coverage across difficulty rather than a single tight latency distribution.
How much can AI cost vary between models?
In this data, blended price per million tokens varied by a 460 fold spread across fourteen models, and by five fold inside the small model tier alone, which is why picking one model and never revisiting the choice tends to drift into an expensive default over time.




