# Clef and Clef-flash: Cloudflare’s New Family of Decision Models for Agentic AI
## Introduction
The artificial intelligence landscape is evolving rapidly, and a new category of models has emerged to fill a critical gap: decision models. Unlike traditional classifier models that require constant retraining to accommodate new categories, or large language models (LLMs) that produce open-ended, non-deterministic outputs, decision models are designed to produce structured, probabilistic classifications quickly and consistently. They sit at the heart of agentic workflows, enabling AI systems to make programmatic decisions without always requiring human oversight.
Cloudflare has entered this space with two new models — **Clef** and **Clef-flash** — now available on Workers AI and fully open-sourced on Hugging Face under an Apache 2.0 license. These models represent a significant step forward in making decision AI accessible, fast, and practical for real-world workloads.
## Understanding Decision Models
A decision model operates by accepting inputs — such as a customer support message, a website URL, or an image — and producing typed outputs with associated probabilities. For instance, you might pass a support ticket into a decision model and ask it to determine whether the request is urgent, which team should handle it, and how severe the impact is. The model returns structured answers with confidence scores that your code can then use to route the ticket, escalate the issue, or defer to a human agent.
This capability fundamentally changes how agents operate. Instead of relying on an LLM to reason through every decision in a slow, non-deterministic fashion, agents can use decision models to make fast, reliable classifications that feed directly into automated workflows. The result is a system that can gather context, classify situations, and take action — or flag things for human review — all at remarkable speed.
### A Real-World Example
Cloudflare’s Threat Intelligence team has been testing Clef on a practical problem: classifying website domains to determine whether they are legitimate or malicious. By giving Clef a domain to analyze, the model can quickly identify categories the domain falls into — for example, a 95% probability it is a fashion website, an 85% probability it is an e-commerce site, and less than 1% probability it is a phishing domain. This entire process took Clef just 2.2 seconds, compared to 4.7 seconds for a comparable LLM that returned far fewer classifications.
This kind of speed advantage, when multiplied across thousands or millions of domains, creates a meaningful operational edge for security teams and other decision-heavy workflows.
## Clef vs. Clef-flash: Two Models for Different Needs
Cloudflare released two variants to serve different use cases:
– **Clef** is the larger, more powerful model designed for precision classification tasks where accuracy is paramount. It currently leads the Jev Decision Index benchmarks and features a vision encoder that can process images alongside text inputs.
– **Clef-flash** is a leaner, faster variant optimized for latency-critical decisions where speed matters most. Despite being significantly smaller, it delivers competitive accuracy and can make decisions in under 40 milliseconds at the median.
Both models share a common API surface, making it straightforward to swap between them depending on the demands of your workflow.
## What Sets Clef Apart
Several features distinguish Clef from other decision models available in the market today.
### Vision Capabilities
Clef includes a vision encoder that allows it to take in images and classify visual content. This is a notable advantage over models that are limited to text-only classification, opening up use cases in image-based decision making — from analyzing screenshots to processing visual inputs in agentic pipelines.
### Extended Context Window
Clef offers a 64,000-token context window, compared to 32,000 tokens for many competing decision models. This larger window allows users to feed more input state into the model, enabling richer and more informed classifications.
### Speed and Efficiency
The architecture behind Clef is designed for speed. Rather than generating text token by token in an autoregressive fashion, Clef uses a non-autoregressive approach. It processes the input through a prefill pass with its Qwen-based backbone, then scores valid schema choices in parallel. There is no intermediate text generation, which dramatically reduces latency while maintaining high accuracy.
### Competitive Benchmarks
Across 43 evaluation benchmarks, Clef models consistently outperformed competing decision models on latency while maintaining or exceeding accuracy. Clef-flash, in particular, demonstrated that a smaller model can deliver exceptional value when speed is the priority. In the Typesafe internal evaluation suite, Clef beat Jev in three out of four workflow categories, and Clef-flash led in one of them.
| Benchmark Area | Clef | Clef-flash | Jev |
|—|—|—|—|
| BFCL Case Exact | 98.47 | 98.76 | 95.75 |
| ToolRet nDCG@10 | 69.19 | 66.43 | 65.28 |
| API-Bank Accuracy | 91.93 | 93.11 | 88.19 |
| Median Latency (ms) | 209.3 | 38.8 | 524.1 |
| p95 Latency (ms) | 238.6 | 122.4 | 536.0 |
## How Clef Was Trained
Clef builds on prior research into making large language models produce deterministic probability outputs. The team adapted a DiffusionGemma-style approach but used Qwen as the base model instead. During inference, Clef performs a prefill-only pass with Qwen, then routes each valid schema choice through a specialized attention mechanism that allows individual classification fields to cross-attend with each other and with the original input payload.
The training process used label-smoothed cross-entropy for valid schema outputs, paired with a Brier loss to refine probability calibration. A proprietary training technique called Reinforcement Learning for Calibrated Decisions (RLCD) was also employed, which grants partial credit to adjacent ordinal choices, rewards fully precise outputs, and applies penalties to prevent distribution shift. The models were trained on internal synthetic datasets that permuted field orders, prompts, and schema structures to ensure robust generalization.
## Fine-Tuning and the Reinforcement Learning Service
Recognizing that many production use cases require domain-specific classifiers, Cloudflare is offering a fine-tuning service for Clef. Internal teams at Cloudflare have already identified use cases such as evaluating Trust & Safety submissions, triaging support requests, and classifying bots in their security products.
The fine-tuning offering begins with a hands-on partnership with Cloudflare’s Forward Deployed Engineer (FDE) team, with plans to eventually transition to a self-serve platform where customers can capture their own data, fine-tune Clef, and redeploy it on Cloudflare’s infrastructure.
### The New RL Product
Cloudflare is also debuting a reinforcement learning service that allows customers to fine-tune Clef models for their specific workloads. This service leverages several primitives already built on the Cloudflare platform:
– **AI Gateway** captures all AI traffic to automatically build datasets for your use case
– **Workers AI** generates rollouts against the base Clef model
– **Containers** provides an RL sandbox for scoring and replaying agent actions
– **Trainer** updates the weights of the fine-tuned model
– **Workers AI with BYO Model** redeploys the fine-tuned model on edge infrastructure
This combination of tools creates a complete loop for continuous model improvement, from data collection to training to deployment — all on Cloudflare’s platform.
## Privacy and Enterprise Readiness
Both Clef and Clef-flash come with Cloudflare’s standard enterprise guarantees. The company does not read, store, or train on your requests or responses unless you explicitly opt into the fine-tuning service. This ensures that sensitive data remains protected while still benefiting from the power of AI-driven decision making.
## Getting Started
Clef and Clef-flash are available now on Workers AI, and the model weights are fully open-sourced on Hugging Face under an Apache 2.0 license. Developers can experiment with the hosted models through the Workers AI API or run them locally using the open-source weights. Detailed documentation is available for those looking to integrate decision models into their agentic workflows.
For organizations with specific fine-tuning needs, Cloudflare’s FDE team is available for early partnerships, and a self-serve fine-tuning platform is in development.
—
## Frequently Asked Questions
### What exactly is a decision model?
A decision model is a type of AI model designed to produce structured, probabilistic classifications based on inputs. Unlike LLMs that generate free-form text, decision models return typed outputs with confidence scores for each possible classification. These outputs can be used directly by code to make automated decisions, route tasks, or trigger actions in an agentic workflow.
### How do Clef and Clef-flash differ?
Clef is the larger, more powerful variant optimized for accuracy and precision across a wide range of classification tasks. It includes a vision encoder and a 64k context window. Clef-flash is a smaller, faster model optimized for low-latency decision making, ideal for scenarios where speed is more critical than the absolute highest accuracy.
### Can Clef process images?
Yes. Clef includes a vision encoder that allows it to take in images and classify visual content alongside text inputs. This is a distinguishing feature compared to many other decision models that are text-only.
### Is Clef API-compatible with other decision models?
Yes. Clef produces strictly typed outputs and is fully API-compatible with the Jev decision model format, making it easy to swap Clef into existing workflows that use other decision models.
### Where can I try Clef?
Clef is hosted on Workers AI and can be accessed through the Cloudflare API. The model weights are also available on Hugging Face under an Apache 2.0 license for local experimentation and deployment.
### What is the reinforcement learning service?
The RL service allows customers to fine-tune Clef models for their specific use cases. It leverages Cloudflare’s AI Gateway to capture traffic, Workers AI for rollouts, and Containers for RL sandboxing, creating a complete feedback loop for model improvement.
### How does Clef compare to LLMs for decision tasks?
Clef is significantly faster than general-purpose LLMs for classification tasks. In one test, Clef classified a domain in 2.2 seconds compared to 4.7 seconds for a large LLM — more than double the speed. Additionally, Clef produces structured probabilistic outputs rather than open-ended text, making its results easier to integrate into automated workflows.
### Is my data safe when using Clef?
Yes. Cloudflare does not read, store, or train on your requests or responses by default. If you opt into the fine-tuning service, your data is used only for that specific purpose under your control.
—
## Conclusion
The introduction of Clef and Clef-flash represents a meaningful advancement in the decision model space. By combining vision capabilities, large context windows, speed, and open-source accessibility, Cloudflare has created tools that can slot directly into agentic workflows and deliver measurable improvements in latency and accuracy.
As agentic AI continues to grow in adoption, the need for fast, reliable, and structured decision-making models will only increase. Clef and Clef-flash position Cloudflare at the forefront of this emerging category, and the upcoming fine-tuning and reinforcement learning services promise to make these models even more adaptable to specific real-world needs.
For developers and teams looking to build faster, more autonomous AI systems, Clef offers a compelling option — whether deployed at the edge through Workers AI or run locally with the open-source weights.
Thank you for reading



