Updated 06:47
Clef wins seven of ten, and loses the one its guardrail pitch needs
Cloudflare Clef (27B) and Clef-flash (9B) on Workers AI: the System One API, the price per million tokens, and the three benchmarks where Jev or Cloudflare's own prototype still leads.
Cloudflare released its first self-trained models on October 1: Clef and Clef-flash, two "decision models" built to answer questions with probabilities instead of prose. The launch materials make the obvious headline easy to write: Clef scores highest on seven of ten benchmarks against Typesafe's Jev. The more useful story for anyone thinking of putting one into an agent loop is which three it didn't win, and what Cloudflare suggests using it for anyway.
What a decision model is, in API terms
According to Cloudflare's changelog, Clef "reads an input state and a set of typed questions, then returns a probability for every allowed answer." It doesn't generate free text. You send a state (a string or structured data) and up to 64 questions in three types:
noul: yes/no, "Returns the probability that the answer is yes."choice: pick one option from a set you define; returns the chosen option, a probability for each option, and a confidence value.score: rate against an ordered rubric; returns a probability-weighted score and a probability per level.
The request shape from the launch post, abridged:
{
"model": "clef",
"state": "Checkout has been failing for every customer for the last hour.",
"questions": {
"urgent": { "type": "noul", "instructions": "Is this support request urgent?" },
"severity": {
"type": "score",
"instructions": "How severe is the customer impact?",
"criteria": ["No impact", "Minor", "Major", "Critical"]
}
}
}
Cloudflare says Clef "follows the System One API," Jev's interface, "so you can switch an existing Jev integration to Clef by changing the endpoint and model." The model pages list @cf/cloudflare/clef (27B) at $0.24 per million input tokens and @cf/cloudflare/clef-flash (9B) at $0.09, both with a 65,536-token context window. Clef also accepts up to four images (PNG, JPEG or WebP, embedded, not by URL), which Cloudflare's docs call a "Clef extension to the System One API." The weights are on Hugging Face under Apache 2.0.
Under the hood, according to the blog post, Cloudflare froze Qwen3.8-27B (Clef) and Qwen3.5-9B (Clef-flash) and trained a routing head with rank-256 low-rank adapters. Inference is "a prefill-only pass" that "then scores the valid schema choices in parallel," so there are no tokens to generate one at a time.
The seven wins
The full table in the launch post covers six models across ten benchmarks. Counting the top score in each row backs up the changelog's claim that "a Clef model scores highest on 7." Some of the wins are large. On the Home appliances benchmark (case exact), Clef-flash scores 97.73 against Jev's 52.27. On BANKING77 (macro-F1), Clef scores 94.20 against 79.74.
Latency is the other headline. Across 43 benchmark runs, Cloudflare reports median latency of 209.3 ms for Clef and 38.8 ms for Clef-flash, against Jev's 524.1 ms. That is the source of its "2.5x" and "13x faster" figures.
The three losses
These are the rows where neither Clef model came first:
| Benchmark | Clef | Clef-flash | Winner | |---|---|---|---| | When2Call · accuracy | 72.37 | 65.58 | Jev, 80.97 | | BRIGHT · nDCG@10 | 45.91 | 39.26 | Jev, 47.52 | | PhishNChips · accuracy | 79.60 | 75.05 | DiffusionGemma Jev, 85.35 |
When2Call is the one that matters for agents. NVIDIA's README for the benchmark says it measures "tool-calling decision-making: when to generate a tool call, when to ask follow-up questions and when to admit the question can't be answered with the tools provided." That is close to the use case Cloudflare's changelog promotes under "Agent guardrails": "Let an agent check 'should I take this action?' in tens of milliseconds before calling a tool." On the benchmark nearest that pitch, Jev leads Clef by 8.6 points and Clef-flash by 15.39.
That doesn't make the guardrail idea wrong. A gate that answers in 38.8 ms at the median can sit in a hot path where a 524 ms gate cannot, and Cloudflare is upfront about the speed-quality trade elsewhere in the post. But a gate that is fast and wrong is still wrong, and When2Call is the closest public evidence of how often that happens.
The PhishNChips row is interesting for a different reason. The model that beats both Clefs there, DiffusionGemma Jev, is Cloudflare's own earlier experiment. The blog post describes adapting DiffusionGemma "to output deterministic probabilities by exposing the logprobs." On a phishing benchmark, which is the one security-shaped classifier in the table, Cloudflare's prototype outscored the product it shipped, 85.35 to 79.60.
There's also a small inconsistency in the presentation. The changelog's "highlights" table includes CLINC150+OOS, where Clef's 97.43 leads. In the same row, Clef-flash scores 66.77, more than 22 points behind Jev's 89.27. Clef-flash is the model Cloudflare recommends for "latency-critical, hot-path decisions," so speed isn't the only thing that differs between the two tiers.
How to read the claims
These are vendor benchmarks, chosen by the vendor, with a competitor's API as the target to beat. Cloudflare says it "shortlisted some evaluations" that matter "as defined by the Jev Decision Index." It also ran Typesafe's own workflow suite, where Clef beat Jev in three of four areas and lost "Agent trace observability" (68.5 to Jev's 71.6, with Clef-flash at 69.8).
The practical takeaway is simple. Clef is cheap, fast, open-weight and a drop-in replacement for an existing Jev integration, so it's easy to test. If you plan to use it as a tool-call gate, test it on your own data first: on the public benchmark closest to that job, it's the one that loses.
Cloudflare's fine-tuning offer goes directly at that gap. A reinforcement learning service starts "as a hands-on partner with our forward-deployed engineer (FDE) team," with a self-serve platform to follow. No pricing has been published for either.
Primary sources: Cloudflare blog, "Introducing Clef", Workers AI changelog, October 1, 2026, Clef model page, Clef-flash model page, NVIDIA When2Call README, read 2026-10-06.