AI Model

Five Ways to Make a Decision in an AI System, and How to Choose Between Them

Most agentic systems contain far more decisions than generations. Before a single token of useful output appears, something has to work out what the user is asking for, whether they are allowed to ask it, which model should handle it, whether the payload contains secrets, and whether a human needs to see it. Only then does anything get written.

That decision layer is where the cost, the latency and most of the failure modes live. It is also the part people think about least, because the generation step is the bit that looks like AI. So instead of watching the new series of MobLand, I built a harness and ran five approaches against the same tasks. This post covers how to choose between them.

Key Takeaways

  • Four of the five approaches landed within one ticket of each other on accuracy, so cost and operational fit, not accuracy, is usually the real decision.
  • A pinned small model cost about 2.2 times more per call than a typed judgment API, and about 100 times less than a pinned frontier model, in this experiment’s billed data.
  • Blended price per million tokens varied by 300 fold across small models alone, so “use a small model” is the start of a decision, not the end of one.
  • Orchestrated model selection matched a hand pinned model on accuracy with full parse reliability, without needing to be re-tuned as provider pricing shifted.
  • Every ticket in the five way bakeoff was correctly routed by a logistic regression running in two milliseconds for free, which is worth sitting with before defaulting to a frontier model.

The Five Options

In a sales conversation these all flatten into “AI classification”. In an architecture they are five different components with different cost curves and different operational burdens.

Classical ML

TF-IDF plus logistic regression, trained on labelled examples, running locally. No API call.

A Typed Judgment API

TypeSafe’s Jev is the example I used. You send application state and narrow typed questions, a boolean, a choice from a set, a score, and you get answers with probabilities your code gates on. It does not write prose. It cannot.

A Small Model, Pinned

One SLM doing the work. In my harness this was gemma-3-4b, reached through a gateway but with the model fixed rather than selected.

A Small Model, Orchestrated

The same class of model, but with the selection and management handled for you: the gateway picks from a pool per request. This is WrangleAI’s SmartRouter and it is the layer I build, so read the relevant sections with that in mind.

A Pinned Frontier Model

Claude Opus, called with the same classifier prompt as the other chat arms.

The distinction between the third and fourth is the one people collapse most often, and it is the one that turned out to matter most in the cost data. More on that below.

How I Set It Up

Two programmes running in parallel. A multi-suite peer pack comparing Jev against a pinned small model across baseline judgments, deliberately contested judgments, stress cases, and gateway-shaped gates covering routing, policy and DLP. And a five-way bakeoff putting all five options above on the same twelve held-out support tickets, choosing between billing, technical and sales.

Three rules to keep it fair. Every chat arm got the same classifier prompt, so nobody won on wording. Any arm that returned something unparseable scored as wrong, because a production router cannot act on prose. And I costed everything from actual billed exports, 384 metered gateway calls plus TypeSafe’s usage tokens, rather than from pricing pages.

I also refused to score systems on jobs they are not built for. Jev is marked out of scope for generation rather than tortured into producing a paragraph through a chain of choices.

Quick link: WrangleAI vs Arize AI

The Five-Way Result

SystemAccuracyAvg latencyp50 latencyParse OK
Classical ML (sklearn)12/122.1 ms2.3 ms100%
Jev (typed judgment)12/12678 ms691 ms100%
Small model, pinned11/12960 ms798 ms100%
Small model, orchestrated11/124,060 ms2,127 ms100%
Opus (frontier)12/122,293 ms2,324 ms100%

Four of the five land within one ticket of each other. If accuracy were the whole story this would be a coin toss decided on price.

It is not the whole story, and the interesting thing about that table is how little it tells you. The real differences are in what happens when the taxonomy changes, when you need eight judgments instead of one, and when the volume goes up by three orders of magnitude.

Classical ML: Still the Right Answer More Often Than People Admit

2.1 milliseconds. Zero marginal cost. Perfect accuracy on the eval set. Nothing else came within two orders of magnitude on either axis.

The cost is entirely in the setup and the maintenance. It needed 105 labelled training tickets. Add a fourth department and you retrain. Change what “technical” means and you retrain. It has no notion of a criterion you can edit, only a decision surface frozen at training time.

My eval was synthetic and clean, which flatters this arm more than the others. Production text drifts, and classical ML feels drift hardest because it cannot be told about the change, only shown.

Choose it when: your categories are stable, you have labels or can generate them, and your volume is high enough that per-call API cost matters.

Avoid it when: the taxonomy is still moving, or you have no labelled corpus and no cheap way to build one.

Small Models: The Economical Workhorse

A pinned gemma-3-4b scored 88% on baseline and contested judgment suites, 92% on department routing, and got through twelve breadth tasks (coding, summarisation, formatting, grammar, copy, reasoning, creative writing) at 12 out of 12.

The billed numbers across the whole experiment: 265 small-model calls, $0.0000079 per call. That is about 2.2 times cheaper per call than the typed judgment API and about 100 times cheaper than the frontier calls in the same logs.

The honest weakness is silent error. The one department miss was a sales upgrade ticket labelled billing. Valid JSON, correct shape, wrong answer, no parse error to catch it and no confidence signal low enough to flag it. It went quietly to the wrong queue. That is the characteristic failure of a generative classifier and it is harder to monitor than a crash.

Choose it when: you need bulk classification and can tolerate a small silent error rate or catch it downstream.

Avoid it when: a wrong answer is expensive and you have nothing to catch it with.

The Bit Everyone Skips: Which Small Model

Here is the number that surprised me most, and it is not about accuracy at all. These are blended dollars per million tokens, measured from the meter rather than quoted from a pricing page.

ModelBlended $/MTok
Mistral-Nemo$0.023
Llama-3.1-8B-Turbo$0.025
gemma-3-4b$0.058
Mistral-Small-24B$0.059
gpt-oss-20b$0.082
Qwen3-32B$0.105
phi-4$0.135
Qwen2.5-72B$0.366
gpt-5-mini$0.905
gemini-2.5-flash$1.232
claude-opus-4.6$7.608

That is a 300-fold spread between the cheapest small model and the frontier tier, and a 5-fold spread inside the small tier alone. Llama-3.1-8B cost me less than half what gemma-3-4b cost for comparable work, despite being the larger model, because the per-token economics of these providers do not track parameter count in any way you can reason about from the outside.

So “use a small model” is not a decision. It is the start of one. The real decision is which small model for which request shape, and that is a question whose answer changes every time a provider adjusts pricing or ships a new checkpoint. Pinning one model means you got that answer right on the day you pinned it, and you will not notice when it stops being right.

This is where orchestration stops being a convenience and starts being the actual product. In the bakeoff, the orchestrated arm matched the pinned arm on accuracy, 11 out of 12, with 100% parse success and the same single miss, while reaching mid-tier models for the requests that warranted them. You trade some latency for coverage across difficulty and for never having to re-run this pricing exercise by hand.

Quick link: EU AI Act Compliance Checklist for Enterprises Using GPT, Claude, and Gemini

Frontier Models: Ceiling Quality, Priced Accordingly

Opus scored 12 out of 12. It also averaged $0.000799 per call across the experiment, against $0.0000079 for the small tier. On the matched twelve-ticket window it worked out at roughly 124 times the small-model cost for the same labels, at 2.3 seconds per call against 0.96.

Ceiling quality on genuinely hard cases is real and worth paying for. The question is what share of your traffic is genuinely hard. In this bakeoff the answer was none of it: every ticket was routed correctly by a logistic regression running in two milliseconds for free.

If you cannot state what fraction of your requests need frontier reasoning, you are not choosing frontier. You are defaulting to it, and that default was the single largest line item in this experiment.

Typed Judgment: The One Most People Have Not Evaluated

Jev scored 100% on baseline, contested and stress suites, and 94% on the gateway-shaped policy gates, the highest of anything I tested on decision-shaped work. It costs $0.0000175 per call, so roughly twice a small model and about 2% of a frontier call.

Accuracy is not the interesting property. The interesting property is that there is no parse step between the model and your if statement, and no retraining step when a criterion changes. Part 2 is a proper review of it, including where it does not fit.

The Architecture All of This Points At

decision_gate_architecture

Decide first, generate second. The gates are cheap, fast and typed, and they resolve most requests without ever reaching a generative model. What survives them goes to a model that writes something, and by then you know enough about the request to pick the right one.

The diagram is not novel. What the evidence adds is the price of skipping it. A system that sends everything to a frontier model because nobody built the gate layer pays roughly a hundred times more on the traffic a classifier or a small model should have absorbed, and gets no better answers on that traffic.

The Short Version

Your situationUse
Stable taxonomy, labelled corpus, high volumeClassical ML
Criteria change often, several judgments per request, typed gatesTyped judgment API
One known request shape, very high volume, tight latency SLOA pinned small model
Mixed difficulty, or you do not want to own model selectionOrchestrated small models
Genuinely ambiguous or high-stakes, low volumeFrontier model

What I Am Not Claiming

Several of these gold sets are small and many of the cases are clear-cut. Twelve tickets is a signal, not a proof. The tickets are synthetic, which flatters classical ML in particular and understates production drift generally. Jev cost figures are derived from token counts times a published rate rather than a billed total. And a handful of frontier rows in the export showed $0, which I read as incomplete metering rather than free inference.

The direction of these findings is more robust than any single number. The direction is what should change your architecture.

Where WrangleAI Fits

The 300-fold price spread in that table is the problem I built WrangleAI to solve.

Nothing in this post is hard to do once. You can benchmark eleven models, read your meter, pick the cheapest one that clears your quality bar, and pin it. What you cannot do once is keep that decision correct. Providers reprice, new checkpoints land, your traffic mix shifts, and the model you pinned in September is the wrong answer by December without anything visibly breaking.

SmartRouter takes model selection off your team. One API, a pool of small models behind it, semantic classification deciding which one handles each request, and a fallback to larger models only for the requests that need them. In this experiment that arm matched a hand-pinned model on accuracy with full parse reliability, and it does not need re-tuning when the pricing table moves under it. It also carries the gate-layer controls this post argues for: hard budget stops, DLP redaction, and policy rules applied before a request reaches any model.

Meter is where the numbers in this post came from. Every figure here is billed cost per call from our own usage export, broken down by model and mapped to a cost centre. That is the difference between knowing your inference spend and estimating it, and it is what let me find the 124-fold gap between the frontier default and the small-model path.

If you are building the gate-then-generate architecture above and do not want to own the model selection problem underneath it, that is exactly the layer we provide. Get in touch or see the platform.

CTA

FAQs

What is the cheapest way to classify AI requests?

In this experiment, classical machine learning was the cheapest option by a wide margin, running at 2.1 milliseconds with zero marginal cost once trained. It works best when the categories are stable and a labelled corpus already exists.

Why not just always use a frontier model like Claude Opus?

Frontier models cost roughly 100 times more per call than a small model in this dataset, yet every ticket in the bakeoff was routed correctly by a two millisecond logistic regression. Frontier reasoning is worth paying for on genuinely hard cases, but defaulting to it for everything is usually the single largest avoidable cost in a decision layer.

What is the difference between a pinned small model and an orchestrated small model?

A pinned small model uses one fixed model for every request. An orchestrated small model, such as WrangleAI’s SmartRouter, selects from a pool of models per request, which matched pinned accuracy in this test while removing the need to manually re-tune the choice as pricing and models change.

Does typed judgment replace a language model entirely?

Not entirely. A typed judgment API such as TypeSafe’s Jev answers narrow, typed questions with probabilities rather than writing prose, which suits decision gates well, but it is out of scope for actual text generation.

How much can model routing save on AI costs?

In this experiment, blended price per million tokens varied by roughly 300 fold across small models alone, and by around 124 fold between the small-model path and the frontier default on matched work. The exact saving depends on traffic mix, but the gap is large enough to be the biggest line item in many AI systems.

Scroll to Top
Contact Form Demo