Cloudflare Releases Open-Weight Clef Decision Models
Post-trained from Qwen backbones, Clef and Clef-flash bring multimodal, non-autoregressive decision scoring to edge infrastructure.

Key takeaways
- Cloudflare released Clef (27B) and Clef-flash (9B) under the Apache 2.0 license on Hugging Face and Workers AI.
- Unlike standard text-generating LLMs, the models run a non-autoregressive prefill pass and a joint schema head to output typed probabilities.
- Both models accept text, JSON, images, and video across a 64,000-token context window.
- Clef-flash clocked a median latency of 38.8 ms compared to 524.1 ms for TypeSafe AI's Jev in internal benchmark runs.
Cloudflare has launched Clef and Clef-flash, two open-weight decision models designed to deliver structured, typed probabilities rather than free-form conversational text. Announced on October 1, 2026, the releases represent the first machine learning models post-trained by the company's Workers AI team. Both architectures are available under an open-source Apache 2.0 license for self-hosting and are deployed across Cloudflare's edge infrastructure.
According to the Cloudflare announcement, the models target automated agent workflows that require deterministic, bounded outputs to make programmatic choices. Instead of using traditional autoregressive token generation that demands subsequent parsing, the Clef models calculate decision probabilities across structured question schemas in a single forward pass.

Architecture and Supported Schema Types
The Clef model family includes two sizes: the 27-billion-parameter Clef, post-trained from Qwen/Qwen3.8-27B, and the smaller 9-billion-parameter Clef-flash, derived from Qwen/Qwen3.5-9B. Both models retain their base architecture's vision encoder, allowing them to evaluate multimodal states including images, video frame arrays, JSON payloads, and text across a 64,000-token context window.
Inference operates in two distinct stages. The frozen base backbone first processes a single prefill-only pass over the input state and question schema. A specialized transformer module called the joint schema head then inspects the backbone's final hidden states, cross-attends between fields and original input context, and jointly scores all allowed choices before a per-question softmax converts logits into calibrated probabilities.
As detailed in the Clef-flash repository, the models support three distinct question types:
- noul: True/false evaluations returning the exact probability of true.
- choice: Mutually exclusive options from a named set, providing per-option probabilities and confidence ratings.
- score: Ordered rubrics evaluated along an indexed scale to return probability-weighted scores.
Cloudflare trained the joint schema head alongside rank-256 low-rank adapters using synthetic permutations, label-smoothed cross-entropy, Brier calibration losses, and a Reinforcement Learning for Calibrated Decisions (RLCD) objective to penalize distribution shift while rewarding ordinal precision.

Performance and Benchmark Results
Cloudflare evaluated the models on its internal run of the Decision Index 0.2.1 suite and workflow benchmarks. On speed, internal benchmark data shows Clef-flash achieving a median latency of 38.8 milliseconds (p95 latency of 122.4 ms), while the larger Clef recorded a median latency of 209.3 milliseconds (p95 latency of 238.6 ms). By comparison, TypeSafe AI's Jev registered a median latency of 524.1 milliseconds (p95 latency of 536.0 ms).
Across classification and reasoning evaluations, Clef scored 98.5% on BFCL case-exact accuracy, 94.2 macro-F1 on BANKING77, and 97.4 macro-F1 on CLINC150+OOS. Clef-flash delivered 98.8% on BFCL, 97.7% on the Home appliance simulator, and a 10.6 Brier score on ForecastBench (where lower numbers represent better calibration).
However, Jev retained significant leads in knowledge-intensive and symbolic reasoning benchmarks, scoring 78.3% on GPQA Diamond (compared to 48.0% for Clef and 51.0% for Clef-flash), 82.7% on MMLU-Pro, and 92.9% on Big-Bench Hard (BBH).
On TypeSafe's four end-to-end workflow evaluations, Clef outperformed Jev on invoice processing (64.7% versus 61.8% exact actions), customer service (76.3% versus 76.0%), and security incident response (62.9% versus 61.7%), while Jev led agent trace observability (71.6% versus 68.5% primary action accuracy).
Edge Deployment and Fine-Tuning Service
Clef and Clef-flash are compatible with the Jev and SystemOne API specifications, meaning existing applications can switch models by adjusting request endpoints. The models are available on Cloudflare Workers AI priced at $0.24 per million input tokens for Clef and $0.09 per million input tokens for Clef-flash, as reported by MarkTechPost.
For on-premises and private infrastructure, the safetensor weights, processor configurations, and schema modules are downloadable on Hugging Face, having been tested with PyTorch 2.11 and Hugging Face Transformers 5.10.2 on a single H200 accelerator.
Alongside the open-weight release, Cloudflare unveiled a reinforcement learning fine-tuning platform. The workflow combines Cloudflare AI Gateway for logging requests, Workers AI for generating rollouts, Cloudflare Containers as an execution sandbox, and a new Trainer component to update model weights. Cloudflare is rolling out the service first through its forward-deployed engineering team before launching a self-serve platform.
Frequently asked questions
What is the difference between an LLM and a decision model?
A decision model does not generate conversational text token by token. Instead, it processes an input state alongside a schema of questions to calculate mathematical probabilities for predefined options in a single prefill pass.
What licenses govern Clef and Clef-flash?
Both Clef and Clef-flash are released under the open-source Apache-2.0 license and are available on Hugging Face for self-hosting.
Can Clef process images and video?
Yes. Both Clef models inherit the vision encoder from their underlying Qwen base architectures, allowing them to accept text, JSON, images, and video frames within a 64,000-token context window.
Sources
- Cloudflare/clef-flashHugging Face · Sep 30, 2026 · Official
- Cloudflare/clefHugging Face · Sep 30, 2026 · Official
- Cloudflare Releases Clef and Clef-flash: Open-Weight Decision Models That Return Typed Probabilities Instead of TextMarkTechPost · Oct 2, 2026
- Introducing Clef: our open-source decision models, and new RL fine-tuning platformCloudflare Blog · Oct 1, 2026
- Cloudflare Open-Sources Clef, Beats Jev Latency on Edge GPUsAI Weekly · Oct 1, 2026
How this story was made: the newsroom picked it up from blog.cloudflare.com, Google News and Hugging Face, gathered the full text of the sources above, and drafted it with AI assistance. Every factual claim was then checked against those sources before publishing (25 claims checked). Illustrations marked as AI-generated are not photographs. Spotted an error? Tell us.
Published October 2, 2026 at 01:36 UTC


