# Clef and Clef-Flash: Cloudflare’s Open Decision Models for Structured AI Responses
Cloudflare has entered the frontier of specialized AI models with the release of **Clef** and **Clef-flash**, the first models developed by its Workers AI research team. Unlike the large language models most users are familiar with, these are *decision models* — purpose-built systems that read an input state alongside a structured schema of typed questions and return a probability distribution over every permitted answer. There is no free-form text generation, no token-by-token sampling, and no conversational preamble. The result is a deterministic, typed output that is immediately usable in production pipelines.
Both models are open-weight, released under the **Apache 2.0** license, and are available for deployment on Cloudflare’s Workers AI platform today. The trained weights are also accessible on Hugging Face for those who prefer self-hosting. Crucially, Clef and Clef-flash are compatible with **TypeSafe AI’s Jev API**, meaning developers already working within the System One ecosystem can integrate them by simply changing the endpoint and model name.
—
## Why Decision Models Matter
Traditional large language models generate text one token at a time. Even when a developer asks a yes/no question, the model must produce a sequence of tokens, and the downstream application must parse and interpret that sequence to extract a structured answer. This introduces latency, parsing complexity, and ambiguity — the model might say “Yes, absolutely,” or “The answer is affirmative,” and the caller must handle both.
A decision model eliminates this ambiguity entirely. It accepts a fixed schema of questions, each with a defined set of allowed answers, and outputs a calibrated probability for every single option in a single forward pass. There is no generation step, no sampling, and no text to interpret. The output is a set of typed probabilities ready for direct use in routing logic, guardrails, or automated decision pipelines.
### Three Question Types Supported
Clef supports three distinct question formats, each tailored to a common pattern in production AI systems:
– **`noul` (yes/no):** Returns the probability that the answer is “yes.” Useful for binary classification tasks like flagging urgent tickets, checking whether a transaction exceeds a threshold, or verifying whether a system state is safe.
– **`choice`:** Selects exactly one option from a named list of candidates, providing per-option probabilities along with an overall confidence value. This maps naturally to intent classification, routing decisions, and categorical labeling.
– **`score`:** Rates an input against an ordered rubric, returning a probability-weighted score. This is ideal for severity assessments, priority rankings, or any scenario where answers exist on a spectrum from low to high.
A single request can carry up to **64 questions** and up to **4 images**, making the model suitable for rich, multimodal decision workloads.
—
## How Clef Is Built
### Architecture
Clef is built on top of **Qwen3.8-27B**, while Clef-flash uses **Qwen3.5-9B** as its foundation. In both cases, the backbone language and vision encoder are kept frozen after pre-training. Only a small additional component — called the **joint schema head** — is trained on top.
The inference pipeline operates in two stages:
1. **Prefill Pass:** The frozen Qwen backbone processes the entire input — the state description, the schema of questions, and any attached images — in a single forward pass. No tokens are generated during this stage.
2. **Schema Head Processing:** The joint schema head reads the final hidden states from the backbone. It routes relevant evidence from the input to each individual question, allows cross-attention between different question fields, and scores all candidate options jointly. A per-question softmax then converts raw logits into well-calibrated probabilities.
### Training Approach
During training, both the backbone and the schema head were optimized jointly, but the backbone itself was frozen. The schema head was trained using **rank-256 low-rank adapters**, keeping the parameter count modest while enabling effective specialization.
The training loss combined three components:
– **Label-smoothed cross-entropy** for primary classification accuracy.
– **Brier loss** to encourage probability calibration — ensuring that when the model says an answer has 90% probability, it is correct roughly 90% of the time.
– **Reinforcement Learning for Calibrated Decisions (RLCD)**, a secondary objective that gives partial credit to adjacent ordinal choices in the `score` question type, acknowledging that a near-miss on a ranked scale still carries meaningful signal.
—
## Deployment and API Compatibility
Both Clef and Clef-flash are available on Cloudflare’s Workers AI platform, where they can be called alongside other AI models with the same tooling and infrastructure developers already use. The models also follow the **System One API** standard introduced by TypeSafe AI, which means they share the same request/response format as the Jev model family.
Developers working with the Jev API can switch to Clef or Clef-flash by updating the model name and endpoint. This compatibility lowers the barrier to adoption and enables teams to compare performance across models without rewriting their integration code.
For teams that prefer to run models on their own infrastructure, the full model weights are hosted on Hugging Face under open licenses, supporting self-hosted deployment.
—
## Performance and Benchmarks
Clef and Clef-flash have been benchmarked across a diverse set of evaluation datasets covering financial text classification, intent detection, tool-use evaluation, and phishing detection. Across these benchmarks, Clef consistently outperforms or matches larger models on structured decision tasks, often with significantly lower latency.
In head-to-head comparisons on the Workers AI platform, Clef-flash delivers median latency under 40 milliseconds, while Clef achieves competitive accuracy at higher latency. Both are dramatically faster than the Jev model family on structured decision workloads, and even the lighter-weight alternatives in the ecosystem.
The benchmark results highlight a key takeaway: for tasks that require typed, structured answers rather than free-form generation, purpose-built decision models can outperform general-purpose LLMs while using fewer compute resources.
—
## FAQ
**Q: What is a decision model, and how is it different from a chatbot?**
A: A decision model does not generate free-form text. Instead, it reads an input state and a schema of typed questions, then returns a probability for every allowed answer. There is no token generation, no conversation, and no interpretation layer needed on the caller side.
**Q: Can Clef handle images as input?**
A: Yes. Each request supports up to four images alongside text input, leveraging the vision encoder from the underlying Qwen backbone.
**Q: What licenses are Clef and Clef-flash released under?**
A: Both models are open-weight under the Apache 2.0 license, permitting commercial use, modification, and distribution.
**Q: Where can I access the model weights?**
A: The weights are available on Hugging Face for self-hosting, and the models are deployable directly on Cloudflare’s Workers AI platform.
**Q: How does Clef compare to the Jev API from TypeSafe AI?**
A: Clef uses the same System One API format as Jev, making integration a drop-in replacement at the endpoint level. Performance differences vary by task, but Clef is designed to offer competitive or superior accuracy on structured decision benchmarks with lower latency.
**Q: What question types does Clef support?**
A: Clef supports three types: `noul` (yes/no with probability), `choice` (one option from a list with per-option probabilities), and `score` (ordered rubric with probability-weighted scoring).
**Q: What is the maximum request size?**
A: A single request supports up to 64 questions and up to 4 images.
**Q: Is Clef-flash a smaller, faster version of Clef?**
A: Yes. Clef-flash is built on Qwen3.5-9B and delivers significantly lower latency while maintaining strong accuracy on decision benchmarks. Clef is built on Qwen3.8-27B and offers higher accuracy at the cost of increased latency.
**Q: What is RLCD?**
A: RLCD stands for Reinforcement Learning for Calibrated Decisions. It is a training objective that provides partial credit for adjacent ordinal answers in `score`-type questions, helping the model learn that answers close to the correct one on a rubric are still informative.
**Q: Can I use Clef for tasks beyond the three question types?**
A: The current model is designed specifically for the `noul`, `choice`, and `score` question types. For tasks requiring free-form generation or more complex output structures, a traditional LLM may still be more appropriate.
—
## Conclusion
Cloudflare’s Clef and Clef-flash represent a meaningful shift in how AI models can be applied in production environments. By moving away from open-ended text generation and toward structured, typed, probabilistic outputs, decision models like Clef address a persistent pain point in AI integration: the gap between what a model outputs and what an application can reliably consume.
The combination of Apache 2.0 licensing, API compatibility with the System One ecosystem, deployment on Workers AI, and open weights on Hugging Face makes these models highly accessible. For teams building support triage systems, agent guardrails, invoice processors, or any pipeline that requires reliable, structured classification from rich inputs, Clef and Clef-flash offer a purpose-built solution that outperforms general-purpose models on both accuracy and latency.
As the AI landscape matures, the trend is clearly moving toward specialization — models that do fewer things but do them with precision, speed, and reliability. Clef and Clef-flash are a strong statement of intent from Cloudflare’s AI team in that direction.
Thank you for reading



