Cloudflare Clef: Open-Weight Decision Models & RL Fine-Tuning Platform
AI News

Cloudflare Clef: Open-Weight Decision Models & RL Fine-Tuning Platform

5 min
10/2/2026
Cloudflare Clefdecision modelsopen-weight AIreinforcement learning

Cloudflare Enters the Decision Model Arena with Clef

Cloudflare has officially entered the decision model space with the release of Clef and Clef-flash, two open-weight models designed for fast, structured classification tasks. Unlike traditional large language models (LLMs) that generate open-ended text, decision models return typed probabilities for a set of predefined choices, making them ideal for agentic workflows where speed and determinism are critical. The models are now hosted on Workers AI and available under an Apache 2.0 license on Hugging Face, signaling Cloudflare's commitment to open-source AI infrastructure.

This launch comes at a time when decision models are gaining traction, with TypeSafe AI's Jev being a notable predecessor. Cloudflare's entry into this niche is not just about adding another model—it's about integrating decision-making capabilities directly into its edge network, reducing latency for AI agents that need to make real-time choices. The company also announced a reinforcement learning (RL) fine-tuning service, positioning itself as a full-stack AI provider for enterprise workflows.

What Makes Clef Different?

Clef distinguishes itself from existing decision models like Jev in several key ways. First, it incorporates a vision encoder, allowing it to process and classify visual content, a capability Jev currently lacks. This expands its utility beyond text-based inputs to include images, which is crucial for applications like content moderation or visual quality checks. Second, Clef offers a 64k context window, double that of Jev's 32k, enabling it to analyze larger sets of input data in a single pass.

Third, Clef is engineered for speed. The inference process uses a prefill-only pass with parallel scoring of valid schema choices, bypassing the autoregressive token-by-token generation that slows down LLMs. This non-autoregressive approach allows Clef to deliver median latency of 209.3 ms (and just 38.8 ms for Clef-flash), compared to Jev's 524.1 ms. When combined with Cloudflare's edge GPUs, this makes Clef suitable for hot-path decisions in production environments.

Benchmark Performance: Clef vs. Jev and Others

Cloudflare reports that Clef leads the Jev Decision Index on several key benchmarks. In the table below, Clef and Clef-flash outperform Jev on tasks like BFCL (case exact), ToolRet (nDCG@10), and BANKING77 (macro-F1). However, Jev still holds an edge in some areas like When2Call and BRIGHT, indicating that the models have complementary strengths depending on the use case.

  • BFCL (case exact): Clef 98.47, Jev 98.76
  • ToolRet (nDCG@10): Clef 69.19, Jev 66.43
  • BANKING77 (macro-F1): Clef 94.20, Jev 90.93
  • When2Call (accuracy): Clef 72.37, Jev 80.97

On Typesafe's own WorkflowEvals suite, Clef and Clef-flash beat Jev in three out of four workflows, including invoice processing and customer service. While Jev remains strong in agent trace observability, Cloudflare's models demonstrate competitive performance across a broad range of decision-making tasks, making them a viable choice for enterprises.

continue reading below...

Technical Architecture: Built on Qwen with RLCD

Clef is built on top of Qwen, with Clef using the Qwen3.8-27B backbone and Clef-flash using Qwen3.5-9B. The models freeze the base and add rank-256 low-rank adapters, optimizing a routing head that scores schema choices in parallel. This architecture, combined with a two-stage attention routing process, allows each field to cross-attend with others, preserving semantic intent across options.

Training involved label-smoothed cross-entropy and a Brier loss for calibration, plus a novel Reinforcement Learning for Calibrated Decisions (RLCD) objective. RLCD grants partial credit to adjacent ordinal choices, rewards precise outputs, and applies a reference penalty to prevent distribution shift. This results in better accuracy and generalization, according to Cloudflare's engineering team.

The New RL Fine-Tuning Service

Cloudflare is also launching a reinforcement learning fine-tuning service, initially offered hands-on with its forward-deployed engineer (FDE) team. The service leverages existing Cloudflare primitives: AI Gateway to capture request data, Workers AI for rollouts, Containers for RL sandboxes, and a new Trainer component to update model weights. This pipeline allows customers to fine-tune Clef on their own data and redeploy it on Workers AI via BYO Model.

This move positions Cloudflare not just as a model provider but as a platform for custom AI decision-making. Internal use cases include triaging support requests, detecting bad bots, and evaluating Trust & Safety submissions. By offering fine-tuning, Cloudflare aims to address highly specific enterprise needs where generic models fall short.

Deployment and Compatibility

Both models are fully compatible with the Jev API, making migration straightforward for existing users. They can be accessed via the Workers AI REST API, binding, or AI Gateway. For self-hosting, the models have been tested on a single H200 with BF16 weights. Cloudflare emphasizes that it does not read, store, or train on customer requests unless they opt into the fine-tuning service, addressing enterprise data privacy concerns.

Why This Matters

Cloudflare's entry into decision models with Clef and its RL fine-tuning platform is a significant step toward making AI agents more efficient and deterministic. By open-sourcing the models and integrating them into its edge network, Cloudflare is democratizing access to high-performance decision-making AI. The low latency, vision capabilities, and fine-tuning options make Clef a compelling choice for developers building agentic workflows that require quick, reliable decisions.

As the AI landscape evolves, the ability to make fast, accurate decisions will become as important as raw generative power. With Clef, Cloudflare is betting that the future of AI lies not just in generating text, but in making the right choices—quickly and reliably.