Three decision models are now available through the Requesty router: sference/clef, cloudflare/clef and perplexity/pplx-decider-v1-27b. They join TypeSafe Jev, which we covered in TypeSafe Jev explained. All four speak the same System One API: you send a state and a set of typed questions, and get back probabilities instead of text. [1] [2] [3]
Before recommending any of them we wanted answers to three questions:
- Sference says its Clef endpoint serves Cloudflare's open-weight
Cloudflare/clef. Does it behave like the same model? - On public benchmarks with gold labels, how do the three decision models compare with each other and with the published Jev numbers?
- On the kind of decisions a gateway user makes every day (routing, moderation, triage, tool-call gating, response preference), how do they compare with LLMs prompted for JSON, on accuracy, calibration, latency, cost and failure rate?
So we ran about 22,000 decisions through the production router on October 5, 2026, with the same request bodies to every model. This post is the write-up. All charts are from that run.
TL;DR
- Sference Clef and Cloudflare Clef are the same model, served differently. 97.3% identical answers on 7,137 shared items, median probability difference 0.010, 100% agreement once the model's own margin is above 0.5. A different model (the Perplexity decider) agrees with Clef 82.1% of the time with a median difference of 0.074.
- Clef leads on fine-grained intent. 94.3% on Banking77 and 89.3% on MASSIVE, against 79.5% and 69.7% for the Perplexity decider and 80.0% and 67.7% published for Jev. On the typed-decisions set and on Emotion all four are tied within noise.
- Decision models are well calibrated. Clef's expected calibration error is 0.02 to 0.04 on four of five datasets; the published Jev range is 0.04 to 0.28.
- On human-labeled gateway tasks, Clef averages 83.2%. That is above every cheap LLM we tried, 3.4 points below Claude Opus 5.5, at 3% of Opus's cost and a quarter of its latency. The Perplexity decider averages 81.8% at $0.02 per 1,000 items.
- Decision models do not refuse. 0 unusable answers in 4,000. The Azure-hosted GPT models could not answer 33% to 36% of the toxic-chat prompts.
- The weak spot is multi-level scoring. "Should the agent call a tool" and three-level complexity are where Clef trails Opus by 20 points or more.
Setup
Every request went through router.requesty.ai with the same body shape the decision models receive in production: one System One state, one or more typed questions, noul for yes/no probabilities, choice for pick-one, score for an ordered rubric. Decision models return a probability per option; we take the argmax as the answer and the probability of the argmax as its confidence. LLM baselines received the same state with the same question text, and were asked for a JSON object with one key per question. Latency is wall-clock from the client, so it includes the router hop. Cost is the amount Requesty billed for the run.
Models in the run:
| Model on Requesty | Type | Input price | Output | Context |
|---|---|---|---|---|
sference/clef | decision, 27B | $0.24 / M | free | 64K |
cloudflare/clef | decision, 27B | $0.24 / M | free | 64K |
perplexity/pplx-decider-v1-27b | decision, 27B | $0.04 / M | free | 262K |
anthropic/claude-opus-5-5, azure/gpt-5.5 | frontier LLM | |||
anthropic/claude-haiku-4-5, azure/gpt-5.6-luna, azure/gpt-5-nano, deepinfra/deepseek-v4.1-flash, google/gemini-3.5-flash-lite | small LLM |
One caveat up front: our Cloudflare account hit its daily Workers AI quota midway through the open benchmark, so cloudflare/clef has full coverage on typed decisions, Banking77 and Emotion, only 62 AG News items, and no MASSIVE run. Where its bars are missing or marked n=62, that is why.
Part 1: is Clef on Sference the same model as Clef on Cloudflare?
Sference lists its endpoint as serving Cloudflare/clef, the open-weight 27B model Cloudflare released on Workers AI. [2] [4] The cheap way to check is to send identical requests to both and compare the probabilities. If the weights are the same, the numbers should match to within serving precision. If they are different models, they should disagree the way two different models do.
We had 7,137 items where both endpoints returned a result. The same-answer rate was 97.3%. Most of the disagreements sit where the model itself is unsure:

When Clef's top probability beats the runner-up by less than 0.1, the two deployments pick the same answer 66% of the time: a coin flip that tips differently in two numerics stacks. Above a margin of 0.5 agreement is 99% to 100%. The Perplexity decider, a different 27B model with the same API, never gets above 92% agreement even on Clef's most confident items.
The probabilities themselves tell the same story. Plot each item's top-answer probability from Sference against the probability the other endpoint gives the same answer:

Correlation 0.968 against Cloudflare, 0.542 against Perplexity. The distribution of the largest per-item probability difference makes the gap obvious on a log scale:

Median difference 0.010 between the two Clef deployments, 0.074 against Perplexity. Per dataset the Clef agreement is 94.5% on typed decisions, 99.1% on Banking77, 100% on AG News and 97.2% on Emotion; against Perplexity it is 80% to 88%:

Three more facts line up. Both deployments are deterministic (six repeats of the same request, six identical outputs on each). Both report identical input token counts for every request, which means the same tokenizer. And on every dataset where both have full coverage, accuracy is within 0.2 points.
The outputs are never bit-identical, which is what you expect from one set of weights behind two inference stacks with different kernels or quantization. We cannot see the checkpoint bytes, so "same weights, different serving" is a conclusion from behaviour, not a proof. We measured the same phenomenon across providers of a single open-weight LLM in same weights, twelve providers.
One API field that is not portable: confidence
The System One response carries a confidence field next to the probabilities on choice and score questions. Sference returns the maximum probability. Cloudflare returns something else, consistently lower. Perplexity returns a third thing that equals the max probability on a quarter of items:

If your code thresholds on confidence, it will behave differently when you switch providers. Threshold on the probabilities instead; those are portable.
Part 2: the open decision benchmark
For gold-label results we used the public decision benchmark manifests that ship with published Jev results: 2,000 typed decisions across real-world workflows (600 noul, 600 choice, 800 score), Banking77 (3,076 items, 77 intents), AG News (2,000 items), Emotion (2,000 items) and a 1,000-item subset of MASSIVE (60 intents). [5] Every item is one question in one call, so the numbers are directly comparable with the published figures.

Two patterns. On typed decisions and on Emotion, all four models land on the same number: 72.8% to 73.2% and 58.5% to 59.9%. Emotion is a noisy six-class dataset where few models do well; typed decisions look like they share a ceiling set by label ambiguity rather than by model quality.
On fine-grained intent classification Clef pulls away: 94.3% on Banking77 against 79.5% for the Perplexity decider and 80.0% published for Jev; 89.3% on MASSIVE against 69.7% and 67.7%. Macro F1 says the same thing, so the gain is not concentrated in a few frequent classes:

Per intent, Clef is at or above 90% on 66 of the 77 Banking77 intents. The Perplexity decider's median intent is at 88%, with a long tail of intents under 50%:

The hardest intents for Clef are the ones that are hard for humans too: "topping up by card" vs "pending top up", "declined transfer" vs "failed transfer", "balance not updated after bank transfer" vs "transfer not received by recipient":

Question type matters more than model
Within the 2,000 typed decisions, the three models are within 2.7 points of each other on every question type, and the spread between question types is much bigger than the spread between models:

noul (a yes/no probability) is at 79% to 82%. choice and score are at 69% to 72%. If you are designing questions for a decision model, a stack of binary noul questions is more reliable than one multi-level score.
Calibration
A decision model's probabilities are only useful if 0.8 means 80%. Expected calibration error with 15 bins:

Clef is at 0.017 to 0.040 on four of the five datasets. The Perplexity decider is at 0.031 to 0.096, published Jev at 0.038 to 0.135. Emotion is badly calibrated for everyone (0.17 to 0.28), which fits a dataset where the labels themselves are inconsistent. The reliability diagrams show the shape: Clef on Banking77 sits on the diagonal, the Perplexity decider is over-confident in the middle of its range, and every model is over-confident on Emotion:

Latency on single-question calls
Median latency per call through the router, with 4 to 8 requests in flight:

Perplexity is the fastest at about 385 ms regardless of dataset. Sference Clef is 407 to 566 ms, Cloudflare Clef 626 to 780 ms. These are deployment numbers for one client location on one day, not model properties: the same weights are 200 ms apart depending on who serves them.
Part 3: the Requesty decision benchmark
Public datasets do not look like gateway traffic, so we built a 2,000-item set from the decisions a Requesty user makes around an LLM call:
| Task | Items | Source | Questions | Reference label |
|---|---|---|---|---|
| Gateway routing | 400 | WildChat prompts | task type (7 options), complexity 0/1/2, needs tools, long output, contains PII | Claude Opus 5.5 |
| Toxic prompt and jailbreak | 300 | lmsys toxic-chat | two noul | human |
| Unsafe prompt | 300 | NVIDIA Aegis 2.0 | one noul | human |
| Response preference | 400 | HelpSteer3 and Arena pairs | which response is better, A or B | human |
| Tool-call gating | 300 | Glaive function-calling | should the agent call a tool | derived from the trace |
| Support triage | 300 | synthetic support tickets | queue, priority, ticket type | dataset labels |
Two label types are deliberately kept apart. Safety, preference and tool-call labels are human or derived from the source trace; they are the ground truth. The gateway routing labels were produced by Claude Opus 5.5, because no public dataset labels prompts with "needs a frontier model" or "will produce a long output". Agreement with Opus labels measures how closely a model matches Opus's judgement, not accuracy against truth, and Opus itself is excluded from that comparison.
Human-labeled tasks

Averaged over the five human and derived fields:

GPT-5.5 (87.0%) and Opus 5.5 (86.6%) lead. Clef is at 83.2%, above DeepSeek v4.1 Flash, the Perplexity decider, Gemini 3.5 Flash Lite and Haiku 4.5 (80.6% to 82.0%). GPT-5.6 Luna and GPT-5 Nano post 83.5% and 82.2%, but only on the 93% to 94% of items they answered; their refusals are concentrated on the hardest moderation prompts, which flatters their accuracy.
The field-level view explains the averages. On moderation Clef is the best or near the best of anything cheaper than a frontier model: 90.3% on toxic prompts, 96.0% on jailbreaks, 87.0% on Aegis, which is 2.5 points above Opus on Aegis. On response preference everyone is at 69% to 76%, Opus included; pairwise preference is a hard, noisy task. Ticket type sits at 66% to 77% for every model including Opus, which says more about the dataset's labels than about the models.
The one place the decision models lose clearly is tool-call gating: 68.7% for both Clef and the Perplexity decider against 90.0% for Opus and 85.7% for GPT-5.5. Deciding whether a user turn warrants a function call needs the tool schema and some reasoning about it, and that is what a 27B decision model does not do.
Routing fields, measured against Opus 5.5

Clef agrees with Opus on 93.2% of task types, 97.5% of needs-tools calls, 90.2% of long-output calls and 99.0% of PII flags: as close as GPT-5.5, and closer than any of the small LLMs on three of the four. The exception is three-level complexity, where Clef agrees with Opus 51.9% of the time, the Perplexity decider 56.1%, and GPT-5.5 81.7%. It is the same pattern as the score questions in the open benchmark: ordered multi-level rubrics are where decision models struggle.
Clef vs the Perplexity decider
Field by field, the two decision models are close. Clef is 4.7 points ahead on toxicity, 2 points ahead on Aegis and preference, 1 to 2 points ahead on three of the routing fields, 4.3 points behind on complexity and 1.3 behind on jailbreak:

The Perplexity decider is a sixth of the price per token and about 70 ms faster. On the open benchmark the gap was wide (Banking77, MASSIVE); on our gateway tasks it is narrow. Which one to pick depends on how many fine-grained options your questions have.
Latency and cost
Each item is one request that answers all of its questions (one to five) at once, so these are per-item numbers:

Clef: p50 0.45 s, p95 0.53 s. Perplexity decider: 0.39 s and 0.45 s. The gap between p50 and p95 is what stands out. The decision models have about 70 ms between p50 and p95; the LLMs have between 0.2 s and 2.7 s, because output length varies and generation is sequential:

Cost per 1,000 items, as billed:

Clef at $0.13 is in the same bracket as the small LLMs ($0.07 to $0.59) because it charges $0.24 per million input tokens and nothing for output; the Perplexity decider at $0.02 is the cheapest thing in the run by a factor of three. Opus and GPT-5.5 are $3.80 and $4.13. Cost against accuracy on the human-labeled fields:

Failure rate
Decision models cannot refuse, hallucinate a format or hit a content filter, because they do not generate text. Over 2,000 items each, Clef and the Perplexity decider returned 2,000 usable answers:

GPT-5.5, GPT-5.6 Luna and GPT-5 Nano on Azure could not answer 33% to 36% of the toxic-chat prompts: the Azure content filter rejected the request before the model saw it, or the model refused to classify. Opus 5.5 hit content_filter on 14 items. If you are building a moderation step, the model that classifies the prompt has to be able to read it; that is the strongest argument for a decision model we found in this run.
Using the probabilities: escalate the uncertain ones
Because the outputs are calibrated probabilities, the natural way to run a decision model is to act on the confident decisions and route the rest to a human or a larger model. Accuracy on the kept decisions, as a function of how many you keep, over the 1,600 human-labeled decisions:

Keep the 80% of decisions Clef is most confident about and accuracy on them is 82%; keep the top half and it is 86%. The Perplexity decider's curve is steeper: 80% at 80% coverage, 88% at 50%. At 50% coverage the $0.02 model is at Opus level on what it keeps, and only the other half of the traffic is billed at Opus prices.
What we take from it
sference/clefandcloudflare/clefare interchangeable at the model level; pick on price, latency from your region, and limits. Sference caps a request at 16 questions, Cloudflare at 64, Perplexity at 128. Do not port aconfidencethreshold between them.- For intent classification with many options, Clef is a large step up from Jev and the Perplexity decider. For a handful of binary questions, the two decision models are close and the Perplexity decider is cheaper.
- Prefer
noulquestions over multi-levelscorequestions where you can. Decision models are 10 points weaker on ordered rubrics across both benchmarks. - Decision models are the right tool for moderation and routing steps that need an answer every time, in under half a second, at LLM-Flash prices. They are the wrong tool for anything that needs the tool schema or reasoning, like tool-call gating.
- Use the probabilities. A confidence threshold plus a frontier fallback gives Opus-level accuracy on the easy majority for a few cents.
How to call them
All three models take a standard chat completions, responses or messages request with a questions response format. Only the model name changes. Details in the Decisions documentation.
curl https://router.requesty.ai/v1/chat/completions \
-H "Authorization: Bearer $REQUESTY_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "sference/clef",
"messages": [{"role": "user", "content": "Our checkout returns 500s and orders are blocked."}],
"response_format": {
"type": "questions",
"questions": {
"department": {"type": "choice", "instructions": "Which team should handle this?",
"criteria": {"billing": "Payments or invoices", "technical": "Bugs or outages"}},
"outage": {"type": "noul", "instructions": "Is a service down?"}
}
}
}'The response is one assistant message whose content is a JSON object: {"department": {"choice": "technical", "probabilities": {...}}, "outage": {"noul": 0.97}}. Swap sference/clef for cloudflare/clef or perplexity/pplx-decider-v1-27b and the request is otherwise unchanged.
Method notes and limits
- One run, one day, one client location. Latency numbers describe the deployments on October 5, 2026, through our router, and will move.
- Cloudflare coverage on the open benchmark is partial (see Setup). Its same-model comparison uses the 7,137 items both endpoints completed.
- The gateway routing labels are Opus 5.5 labels. Agreement with them is agreement with Opus, and a model that reasons like Opus will score well on them for that reason.
- The support triage dataset's ticket-type labels appear inconsistent (every model including Opus lands at 66% to 77%); treat that column as a dataset property.
- LLM baselines were prompted once, for JSON, with no few-shot examples and no retries on refusals. A tuned prompt would move them, probably upward; we wanted the comparison a gateway user gets out of the box.
- Accuracy for LLMs is on the items they answered. The failure rate chart shows what they did not answer.
Sources
Frequently asked questions
- Is Clef on Sference the same model as Clef on Cloudflare Workers AI?
- The evidence says yes. On 7,137 identical requests the two returned the same answer 97.3% of the time, the median probability difference was 0.010, both are deterministic, token counts match, and accuracy on every dataset is within 0.2 points. The outputs are not bit-identical, which is what you expect from two serving stacks running the same weights. We cannot verify the checkpoint bytes from the outside.
- How accurate are decision models compared with LLMs?
- On our human-labeled tasks Clef averaged 83.2% and the Perplexity decider 81.8%, versus 86.6% for Claude Opus 5.5 and 80.6% to 83.5% for the cheaper LLMs. The decision models answered every item; the Azure-hosted GPT models returned no usable answer on 6% to 7% of items, mostly content filtering on the moderation prompts.
- What do decision models cost through Requesty?
- Clef is $0.24 per million input tokens on Sference and Cloudflare, the Perplexity decider is $0.04. Output tokens are not billed. On our 2,000-item benchmark that came to $0.13 and $0.02 per 1,000 items, against $3.80 for Opus 5.5 and $0.07 to $0.59 for the small LLMs.
- How do I call a decision model through Requesty?
- Send a normal chat completions, responses or messages request with the model set to sference/clef, cloudflare/clef or perplexity/pplx-decider-v1-27b and a response_format of type questions. The answer comes back as a JSON object with a probability for every question.
- SEP '26
TypeSafe Jev explained: how it works, LLM differences and API pricing
A practical guide to TypeSafe's decision model, what its probabilities mean, and how to access Jev through Requesty.
- JUL '26
Same weights, twelve providers, 4.8x price gap: the provider variance problem
On July 20 the same DeepSeek V4 Pro weights were listed across twelve providers from $0.44 to $2.10 per million tokens. Open weights did not commoditize inference, they moved the variance from the model to the provider. Here is how to route for it.
- AUG '26
36x the price for 22% more quality: the Pareto data that should decide your model mix
Glean benchmarked 37 models and reasoning configurations across 1,000 enterprise tasks. The gap between the cheapest and the priciest frontier model was 36x on price and 22% on quality. Add an 11x provider spread and a 2x reasoning effort penalty on the same model, and single model deployments start to look like the most expensive decision in the stack.
