Requesty
Back|OCT '26AI MODELS / ROUTING
14 MIN READ|

We benchmarked the decision models: Clef on Sference vs Cloudflare, Perplexity's decider, and eight LLMs on 22,000 decisions

Last updated

Three decision models are now available through the Requesty router: sference/clef, cloudflare/clef and perplexity/pplx-decider-v1-27b. They join TypeSafe Jev, which we covered in TypeSafe Jev explained. All four speak the same System One API: you send a state and a set of typed questions, and get back probabilities instead of text. [1] [2] [3]

Before recommending any of them we wanted answers to three questions:

  1. Sference says its Clef endpoint serves Cloudflare's open-weight Cloudflare/clef. Does it behave like the same model?
  2. On public benchmarks with gold labels, how do the three decision models compare with each other and with the published Jev numbers?
  3. On the kind of decisions a gateway user makes every day (routing, moderation, triage, tool-call gating, response preference), how do they compare with LLMs prompted for JSON, on accuracy, calibration, latency, cost and failure rate?

So we ran about 22,000 decisions through the production router on October 5, 2026, with the same request bodies to every model. This post is the write-up. All charts are from that run.

TL;DR

  • Sference Clef and Cloudflare Clef are the same model, served differently. 97.3% identical answers on 7,137 shared items, median probability difference 0.010, 100% agreement once the model's own margin is above 0.5. A different model (the Perplexity decider) agrees with Clef 82.1% of the time with a median difference of 0.074.
  • Clef leads on fine-grained intent. 94.3% on Banking77 and 89.3% on MASSIVE, against 79.5% and 69.7% for the Perplexity decider and 80.0% and 67.7% published for Jev. On the typed-decisions set and on Emotion all four are tied within noise.
  • Decision models are well calibrated. Clef's expected calibration error is 0.02 to 0.04 on four of five datasets; the published Jev range is 0.04 to 0.28.
  • On human-labeled gateway tasks, Clef averages 83.2%. That is above every cheap LLM we tried, 3.4 points below Claude Opus 5.5, at 3% of Opus's cost and a quarter of its latency. The Perplexity decider averages 81.8% at $0.02 per 1,000 items.
  • Decision models do not refuse. 0 unusable answers in 4,000. The Azure-hosted GPT models could not answer 33% to 36% of the toxic-chat prompts.
  • The weak spot is multi-level scoring. "Should the agent call a tool" and three-level complexity are where Clef trails Opus by 20 points or more.

Setup

Every request went through router.requesty.ai with the same body shape the decision models receive in production: one System One state, one or more typed questions, noul for yes/no probabilities, choice for pick-one, score for an ordered rubric. Decision models return a probability per option; we take the argmax as the answer and the probability of the argmax as its confidence. LLM baselines received the same state with the same question text, and were asked for a JSON object with one key per question. Latency is wall-clock from the client, so it includes the router hop. Cost is the amount Requesty billed for the run.

Models in the run:

Model on RequestyTypeInput priceOutputContext
sference/clefdecision, 27B$0.24 / Mfree64K
cloudflare/clefdecision, 27B$0.24 / Mfree64K
perplexity/pplx-decider-v1-27bdecision, 27B$0.04 / Mfree262K
anthropic/claude-opus-5-5, azure/gpt-5.5frontier LLM
anthropic/claude-haiku-4-5, azure/gpt-5.6-luna, azure/gpt-5-nano, deepinfra/deepseek-v4.1-flash, google/gemini-3.5-flash-litesmall LLM

One caveat up front: our Cloudflare account hit its daily Workers AI quota midway through the open benchmark, so cloudflare/clef has full coverage on typed decisions, Banking77 and Emotion, only 62 AG News items, and no MASSIVE run. Where its bars are missing or marked n=62, that is why.

Part 1: is Clef on Sference the same model as Clef on Cloudflare?

Sference lists its endpoint as serving Cloudflare/clef, the open-weight 27B model Cloudflare released on Workers AI. [2] [4] The cheap way to check is to send identical requests to both and compare the probabilities. If the weights are the same, the numbers should match to within serving precision. If they are different models, they should disagree the way two different models do.

We had 7,137 items where both endpoints returned a result. The same-answer rate was 97.3%. Most of the disagreements sit where the model itself is unsure:

Same answer rate by decision margin
Same answer rate by decision margin

When Clef's top probability beats the runner-up by less than 0.1, the two deployments pick the same answer 66% of the time: a coin flip that tips differently in two numerics stacks. Above a margin of 0.5 agreement is 99% to 100%. The Perplexity decider, a different 27B model with the same API, never gets above 92% agreement even on Clef's most confident items.

The probabilities themselves tell the same story. Plot each item's top-answer probability from Sference against the probability the other endpoint gives the same answer:

Top-answer probability scatter
Top-answer probability scatter

Correlation 0.968 against Cloudflare, 0.542 against Perplexity. The distribution of the largest per-item probability difference makes the gap obvious on a log scale:

Probability difference per item
Probability difference per item

Median difference 0.010 between the two Clef deployments, 0.074 against Perplexity. Per dataset the Clef agreement is 94.5% on typed decisions, 99.1% on Banking77, 100% on AG News and 97.2% on Emotion; against Perplexity it is 80% to 88%:

Same answer rate per dataset
Same answer rate per dataset

Three more facts line up. Both deployments are deterministic (six repeats of the same request, six identical outputs on each). Both report identical input token counts for every request, which means the same tokenizer. And on every dataset where both have full coverage, accuracy is within 0.2 points.

The outputs are never bit-identical, which is what you expect from one set of weights behind two inference stacks with different kernels or quantization. We cannot see the checkpoint bytes, so "same weights, different serving" is a conclusion from behaviour, not a proof. We measured the same phenomenon across providers of a single open-weight LLM in same weights, twelve providers.

One API field that is not portable: confidence

The System One response carries a confidence field next to the probabilities on choice and score questions. Sference returns the maximum probability. Cloudflare returns something else, consistently lower. Perplexity returns a third thing that equals the max probability on a quarter of items:

Confidence field by provider
Confidence field by provider

If your code thresholds on confidence, it will behave differently when you switch providers. Threshold on the probabilities instead; those are portable.

Part 2: the open decision benchmark

For gold-label results we used the public decision benchmark manifests that ship with published Jev results: 2,000 typed decisions across real-world workflows (600 noul, 600 choice, 800 score), Banking77 (3,076 items, 77 intents), AG News (2,000 items), Emotion (2,000 items) and a 1,000-item subset of MASSIVE (60 intents). [5] Every item is one question in one call, so the numbers are directly comparable with the published figures.

Open benchmark accuracy
Open benchmark accuracy

Two patterns. On typed decisions and on Emotion, all four models land on the same number: 72.8% to 73.2% and 58.5% to 59.9%. Emotion is a noisy six-class dataset where few models do well; typed decisions look like they share a ceiling set by label ambiguity rather than by model quality.

On fine-grained intent classification Clef pulls away: 94.3% on Banking77 against 79.5% for the Perplexity decider and 80.0% published for Jev; 89.3% on MASSIVE against 69.7% and 67.7%. Macro F1 says the same thing, so the gain is not concentrated in a few frequent classes:

Open benchmark macro F1
Open benchmark macro F1

Per intent, Clef is at or above 90% on 66 of the 77 Banking77 intents. The Perplexity decider's median intent is at 88%, with a long tail of intents under 50%:

Banking77 per-intent accuracy curve
Banking77 per-intent accuracy curve

The hardest intents for Clef are the ones that are hard for humans too: "topping up by card" vs "pending top up", "declined transfer" vs "failed transfer", "balance not updated after bank transfer" vs "transfer not received by recipient":

Banking77 hardest intents
Banking77 hardest intents

Question type matters more than model

Within the 2,000 typed decisions, the three models are within 2.7 points of each other on every question type, and the spread between question types is much bigger than the spread between models:

Typed decisions by question type
Typed decisions by question type

noul (a yes/no probability) is at 79% to 82%. choice and score are at 69% to 72%. If you are designing questions for a decision model, a stack of binary noul questions is more reliable than one multi-level score.

Calibration

A decision model's probabilities are only useful if 0.8 means 80%. Expected calibration error with 15 bins:

Calibration ECE
Calibration ECE

Clef is at 0.017 to 0.040 on four of the five datasets. The Perplexity decider is at 0.031 to 0.096, published Jev at 0.038 to 0.135. Emotion is badly calibrated for everyone (0.17 to 0.28), which fits a dataset where the labels themselves are inconsistent. The reliability diagrams show the shape: Clef on Banking77 sits on the diagonal, the Perplexity decider is over-confident in the middle of its range, and every model is over-confident on Emotion:

Reliability diagrams
Reliability diagrams

Latency on single-question calls

Median latency per call through the router, with 4 to 8 requests in flight:

Open benchmark latency
Open benchmark latency

Perplexity is the fastest at about 385 ms regardless of dataset. Sference Clef is 407 to 566 ms, Cloudflare Clef 626 to 780 ms. These are deployment numbers for one client location on one day, not model properties: the same weights are 200 ms apart depending on who serves them.

Part 3: the Requesty decision benchmark

Public datasets do not look like gateway traffic, so we built a 2,000-item set from the decisions a Requesty user makes around an LLM call:

TaskItemsSourceQuestionsReference label
Gateway routing400WildChat promptstask type (7 options), complexity 0/1/2, needs tools, long output, contains PIIClaude Opus 5.5
Toxic prompt and jailbreak300lmsys toxic-chattwo noulhuman
Unsafe prompt300NVIDIA Aegis 2.0one noulhuman
Response preference400HelpSteer3 and Arena pairswhich response is better, A or Bhuman
Tool-call gating300Glaive function-callingshould the agent call a toolderived from the trace
Support triage300synthetic support ticketsqueue, priority, ticket typedataset labels

Two label types are deliberately kept apart. Safety, preference and tool-call labels are human or derived from the source trace; they are the ground truth. The gateway routing labels were produced by Claude Opus 5.5, because no public dataset labels prompts with "needs a frontier model" or "will produce a long output". Agreement with Opus labels measures how closely a model matches Opus's judgement, not accuracy against truth, and Opus itself is excluded from that comparison.

Human-labeled tasks

Custom benchmark, human labels
Custom benchmark, human labels

Averaged over the five human and derived fields:

Overall ranking on human-labeled decisions
Overall ranking on human-labeled decisions

GPT-5.5 (87.0%) and Opus 5.5 (86.6%) lead. Clef is at 83.2%, above DeepSeek v4.1 Flash, the Perplexity decider, Gemini 3.5 Flash Lite and Haiku 4.5 (80.6% to 82.0%). GPT-5.6 Luna and GPT-5 Nano post 83.5% and 82.2%, but only on the 93% to 94% of items they answered; their refusals are concentrated on the hardest moderation prompts, which flatters their accuracy.

The field-level view explains the averages. On moderation Clef is the best or near the best of anything cheaper than a frontier model: 90.3% on toxic prompts, 96.0% on jailbreaks, 87.0% on Aegis, which is 2.5 points above Opus on Aegis. On response preference everyone is at 69% to 76%, Opus included; pairwise preference is a hard, noisy task. Ticket type sits at 66% to 77% for every model including Opus, which says more about the dataset's labels than about the models.

The one place the decision models lose clearly is tool-call gating: 68.7% for both Clef and the Perplexity decider against 90.0% for Opus and 85.7% for GPT-5.5. Deciding whether a user turn warrants a function call needs the tool schema and some reasoning about it, and that is what a 27B decision model does not do.

Routing fields, measured against Opus 5.5

Gateway routing fields vs Opus labels
Gateway routing fields vs Opus labels

Clef agrees with Opus on 93.2% of task types, 97.5% of needs-tools calls, 90.2% of long-output calls and 99.0% of PII flags: as close as GPT-5.5, and closer than any of the small LLMs on three of the four. The exception is three-level complexity, where Clef agrees with Opus 51.9% of the time, the Perplexity decider 56.1%, and GPT-5.5 81.7%. It is the same pattern as the score questions in the open benchmark: ordered multi-level rubrics are where decision models struggle.

Clef vs the Perplexity decider

Field by field, the two decision models are close. Clef is 4.7 points ahead on toxicity, 2 points ahead on Aegis and preference, 1 to 2 points ahead on three of the routing fields, 4.3 points behind on complexity and 1.3 behind on jailbreak:

Clef vs Perplexity decider head to head
Clef vs Perplexity decider head to head

The Perplexity decider is a sixth of the price per token and about 70 ms faster. On the open benchmark the gap was wide (Banking77, MASSIVE); on our gateway tasks it is narrow. Which one to pick depends on how many fine-grained options your questions have.

Latency and cost

Each item is one request that answers all of its questions (one to five) at once, so these are per-item numbers:

Latency p50 and p95
Latency p50 and p95

Clef: p50 0.45 s, p95 0.53 s. Perplexity decider: 0.39 s and 0.45 s. The gap between p50 and p95 is what stands out. The decision models have about 70 ms between p50 and p95; the LLMs have between 0.2 s and 2.7 s, because output length varies and generation is sequential:

Latency distribution
Latency distribution

Cost per 1,000 items, as billed:

Cost per 1,000 items
Cost per 1,000 items

Clef at $0.13 is in the same bracket as the small LLMs ($0.07 to $0.59) because it charges $0.24 per million input tokens and nothing for output; the Perplexity decider at $0.02 is the cheapest thing in the run by a factor of three. Opus and GPT-5.5 are $3.80 and $4.13. Cost against accuracy on the human-labeled fields:

Cost vs accuracy
Cost vs accuracy

Failure rate

Decision models cannot refuse, hallucinate a format or hit a content filter, because they do not generate text. Over 2,000 items each, Clef and the Perplexity decider returned 2,000 usable answers:

Failure rate
Failure rate

GPT-5.5, GPT-5.6 Luna and GPT-5 Nano on Azure could not answer 33% to 36% of the toxic-chat prompts: the Azure content filter rejected the request before the model saw it, or the model refused to classify. Opus 5.5 hit content_filter on 14 items. If you are building a moderation step, the model that classifies the prompt has to be able to read it; that is the strongest argument for a decision model we found in this run.

Using the probabilities: escalate the uncertain ones

Because the outputs are calibrated probabilities, the natural way to run a decision model is to act on the confident decisions and route the rest to a human or a larger model. Accuracy on the kept decisions, as a function of how many you keep, over the 1,600 human-labeled decisions:

Confidence coverage curve
Confidence coverage curve

Keep the 80% of decisions Clef is most confident about and accuracy on them is 82%; keep the top half and it is 86%. The Perplexity decider's curve is steeper: 80% at 80% coverage, 88% at 50%. At 50% coverage the $0.02 model is at Opus level on what it keeps, and only the other half of the traffic is billed at Opus prices.

What we take from it

  • sference/clef and cloudflare/clef are interchangeable at the model level; pick on price, latency from your region, and limits. Sference caps a request at 16 questions, Cloudflare at 64, Perplexity at 128. Do not port a confidence threshold between them.
  • For intent classification with many options, Clef is a large step up from Jev and the Perplexity decider. For a handful of binary questions, the two decision models are close and the Perplexity decider is cheaper.
  • Prefer noul questions over multi-level score questions where you can. Decision models are 10 points weaker on ordered rubrics across both benchmarks.
  • Decision models are the right tool for moderation and routing steps that need an answer every time, in under half a second, at LLM-Flash prices. They are the wrong tool for anything that needs the tool schema or reasoning, like tool-call gating.
  • Use the probabilities. A confidence threshold plus a frontier fallback gives Opus-level accuracy on the easy majority for a few cents.

How to call them

All three models take a standard chat completions, responses or messages request with a questions response format. Only the model name changes. Details in the Decisions documentation.

Shell
curl https://router.requesty.ai/v1/chat/completions \
  -H "Authorization: Bearer $REQUESTY_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "sference/clef",
    "messages": [{"role": "user", "content": "Our checkout returns 500s and orders are blocked."}],
    "response_format": {
      "type": "questions",
      "questions": {
        "department": {"type": "choice", "instructions": "Which team should handle this?",
                       "criteria": {"billing": "Payments or invoices", "technical": "Bugs or outages"}},
        "outage": {"type": "noul", "instructions": "Is a service down?"}
      }
    }
  }'

The response is one assistant message whose content is a JSON object: {"department": {"choice": "technical", "probabilities": {...}}, "outage": {"noul": 0.97}}. Swap sference/clef for cloudflare/clef or perplexity/pplx-decider-v1-27b and the request is otherwise unchanged.

Method notes and limits

  • One run, one day, one client location. Latency numbers describe the deployments on October 5, 2026, through our router, and will move.
  • Cloudflare coverage on the open benchmark is partial (see Setup). Its same-model comparison uses the 7,137 items both endpoints completed.
  • The gateway routing labels are Opus 5.5 labels. Agreement with them is agreement with Opus, and a model that reasons like Opus will score well on them for that reason.
  • The support triage dataset's ticket-type labels appear inconsistent (every model including Opus lands at 66% to 77%); treat that column as a dataset property.
  • LLM baselines were prompted once, for JSON, with no few-shot examples and no retries on refusals. A tuned prompt would move them, probably upward; we wanted the comparison a gateway user gets out of the box.
  • Accuracy for LLMs is on the items they answered. The failure rate chart shows what they did not answer.

Sources

  1. TypeSafe: System One concepts
  2. Cloudflare Workers AI: Clef model page
  3. Perplexity: Decisions API
  4. Hugging Face: Cloudflare/clef model card
  5. Hugging Face: decision benchmark manifests and published Jev results
  6. Requesty: Decisions documentation
Frequently asked questions
Is Clef on Sference the same model as Clef on Cloudflare Workers AI?
The evidence says yes. On 7,137 identical requests the two returned the same answer 97.3% of the time, the median probability difference was 0.010, both are deterministic, token counts match, and accuracy on every dataset is within 0.2 points. The outputs are not bit-identical, which is what you expect from two serving stacks running the same weights. We cannot verify the checkpoint bytes from the outside.
How accurate are decision models compared with LLMs?
On our human-labeled tasks Clef averaged 83.2% and the Perplexity decider 81.8%, versus 86.6% for Claude Opus 5.5 and 80.6% to 83.5% for the cheaper LLMs. The decision models answered every item; the Azure-hosted GPT models returned no usable answer on 6% to 7% of items, mostly content filtering on the moderation prompts.
What do decision models cost through Requesty?
Clef is $0.24 per million input tokens on Sference and Cloudflare, the Perplexity decider is $0.04. Output tokens are not billed. On our 2,000-item benchmark that came to $0.13 and $0.02 per 1,000 items, against $3.80 for Opus 5.5 and $0.07 to $0.59 for the small LLMs.
How do I call a decision model through Requesty?
Send a normal chat completions, responses or messages request with the model set to sference/clef, cloudflare/clef or perplexity/pplx-decider-v1-27b and a response_format of type questions. The answer comes back as a JSON object with a probability for every question.
Related reading

Start building with Requesty

One line of code. 600+ models. Full control.

Speak to founders