🇫🇷 Cet article est aussi disponible en français.

Cloudflare's Clef Brings Decision Models to the Agent Hot Path

Cloudflare's Clef Brings Decision Models to the Agent Hot Path
🇫🇷 Cet article est aussi disponible en français.
📑 Table of Contents

TL;DR

  • Cloudflare released Clef (27B) and Clef-flash (9B), the first models trained by its Workers AI team, open-sourced under Apache 2.0 on Hugging Face.
  • A decision model reads an input state plus typed questions and returns a probability for every allowed answer — no free-form text, no reasoning tokens — aimed at the “hot path” of agent pipelines.
  • Across 43 benchmark runs, Clef posts a 209.3 ms median latency against Jev’s 524.1 ms (2.5x faster), while Clef-flash hits 38.8 ms (13x faster).
  • Clef is drop-in compatible with Typesafe’s Jev via the System One API and adds a vision encoder, though it still trails Jev on a few evals.

A smaller model that only makes decisions

For two years the default answer to “how should an agent decide something?” has been to ask an LLM. But that is slow and non-deterministic: the model emits tokens one at a time, and by the time your code has parsed free-form text the moment for a routing decision may have passed. Cloudflare’s Workers AI team shipped a different bet on October 1. Clef and Clef-flash are its first in-house trained models, positioned not as LLMs but as decision models, “in the same family as Typesafe’s Jev” — they read an input state plus typed questions and return a probability for every allowed answer (Source : Cloudflare — Introducing Clef).

This is the System One / System Two split applied to infrastructure. A decision model handles the fast, bounded reflex — route this ticket, block this request, escalate to a human — while a general LLM handles open-ended reasoning and tool calls. Cloudflare frames the two as complements, not rivals. The decision model sits in the request path; the LLM takes the action (Source : Cloudflare Changelog — Introducing Clef).

Typed questions, parseable answers

The interface is small and explicit. Each request carries a state plus as many as 64 questions, in three types (Source : Cloudflare Changelog — Introducing Clef):

  • noul — a yes/no question. Returns the probability that the answer is yes.
  • choice — pick one option from a set you define. Returns the chosen option, a probability per option, and a confidence value.
  • score — rate against an ordered rubric. Returns a probability-weighted score and a probability per level.

Because the output is strictly typed and the allowed answers are bounded by the schema, there is no text to parse and no chain of reasoning to wait on. Cloudflare argues this is the point: a decision model “produces bounded structured outputs cheaply, quickly and consistently” where an LLM is “largely non-deterministic” (Source : Cloudflare — Introducing Clef).

The latency and accuracy numbers

Cloudflare’s headline claim is speed. Across the 43 eval benchmarks it ran, Clef measured a 209.3 ms median latency (238.6 ms at p95) against Jev’s 524.1 ms median (536.0 ms at p95). Clef-flash dropped to 38.8 ms median and 122.4 ms at p95. That makes Clef 2.5x faster than Jev at the median and Clef-flash 13x faster (Source : Cloudflare Changelog — Introducing Clef).

The accuracy picture is more mixed and worth reading carefully. Across 10 decision benchmarks, a Clef model scores highest on 7 (Source : Cloudflare Changelog — Introducing Clef). On BFCL case-exact, Clef scores 98.47 and Clef-flash 98.76 against Jev’s 95.75. On BANKING77 macro-F1, Clef leads at 94.20 versus Jev’s 79.74. On CLINC150+OOS macro-F1, Clef reaches 97.43 while Clef-flash collapses to 66.77 — below Jev’s 89.27. The pattern cuts the other way on home appliances, where Clef-flash scores 97.73 against Clef’s 82.95 and Jev’s 52.27 (Source : Cloudflare Changelog — Introducing Clef).

On Typesafe’s own workflow evaluations, Clef beats Jev in 3 of 4 areas — invoice processing (64.7 vs 61.8), customer service (76.3 vs 76.0), and security incidents (62.9 vs 61.7). Jev still wins agent trace observability, 71.6 to Clef-flash’s 69.8 (Source : Cloudflare — Introducing Clef).

The takeaway is that Clef-flash is a genuine trade-off, not a free lunch: it wins latency by an order of magnitude but can lose substantial accuracy on some classification tasks, so task selection matters.

Under the hood, and in the hot path

Clef uses Qwen as a backbone, post-trained for decision use cases. During inference it runs a prefill-only pass, then scores valid schema choices in parallel — the decision step is non-autoregressive, so no text is generated token by token (Source : Cloudflare — Introducing Clef). Cloudflare reports that training freezes the backbone and jointly optimizes a routing head with rank-256 low-rank adapters, calibrating with a Brier loss and a “Reinforcement Learning for Calibrated Decisions” objective (Source : Cloudflare — Introducing Clef).

Clef (27B) is the highest-precision option; Clef-flash (9B) is for latency-critical decisions. Both carry a 64K-token context window, double Jev’s 32K (Source : Cloudflare — Introducing Clef). Clef also has a vision encoder and accepts up to four images alongside the state, unlike text-only decision models like Jev (Source : Cloudflare Changelog — Introducing Clef).

Because the models run on Workers AI GPUs across Cloudflare’s network, the network round trip stays short, which is what makes the hot-path pitch plausible (Source : Cloudflare Changelog — Introducing Clef). Cloudflare’s own threat-intelligence team uses it to classify domains: paired with Browser Run, Clef fetched, rendered, and classified a site in 2.2 seconds versus 4.7 seconds for gpt-oss-120b (Source : Cloudflare — Introducing Clef).

Open weights and an RL fine-tuning play

The weights are on Hugging Face under Apache 2.0, and Clef is reachable through the Workers AI binding (env.AI.run()) or the REST API at /ai/run, and through AI Gateway (Source : Cloudflare Changelog — Introducing Clef).

Alongside the models, Cloudflare launched a reinforcement-learning fine-tuning service, starting with its forward-deployed engineering team and intended to become self-serve, letting customers capture data via AI Gateway, generate rollouts on Workers AI, score actions in Containers-based RL sandboxes, retrain, and redeploy (Source : Cloudflare — Introducing Clef).

FAQ

What is a decision model?

A model that reads an input state plus typed questions and returns a probability for every allowed answer, rather than generating free-form text. It makes bounded, programmatic choices inside an agent workflow.

How much faster is Clef than Jev?

Across 43 benchmark runs, Clef’s median latency was 209.3 ms versus Jev’s 524.1 ms — 2.5x faster. Clef-flash came in at 38.8 ms median, 13x faster (Source : Cloudflare Changelog — Introducing Clef).

Does Clef replace LLMs?

No, Cloudflare positions it as a complement. Clef handles fast decisions in the request path; an LLM on Workers AI then takes the action, so the two are chained rather than swapped (Source : Cloudflare Changelog — Introducing Clef).

Can I switch an existing Jev integration to Clef?

Yes. Clef follows the System One API, so switching means changing the endpoint and the model name (Source : Cloudflare Changelog — Introducing Clef).

Further Reading

— The Agent Report