Back to the blog
BunAICloudflareOpen Source

Cloudflare Clef: decision models in your Bun backend

October 03, 2026·6 min read·Diego Horvatti

Count how many LLM calls in your backend exist only to pick one option from a list. Classifying a ticket. Deciding if a comment is spam. Choosing which queue a task goes to. Cloudflare announced Clef, a family of open-weight decision models, along with a reinforcement learning (RL) fine-tuning platform. It targets exactly this kind of call. My everyday backend runs on Bun, so I want to look at the announcement from that angle. What changes for someone who writes TypeScript and just wants the decision to come out right, fast and cheap?

The original announcement is on the Cloudflare blog. You'll find the details on size, license and benchmarks there. Here is my take as a dev.

What Cloudflare launched with Clef

There are two pieces.

The first is the models. They are open-weight. That means you can download the weights and run them wherever you want, without depending on anyone's API. The focus isn't chatting or writing essays. The focus is making a decision from an input and a set of options.

The second is the reinforcement fine-tuning platform. You don't build a huge dataset of "input → correct answer". Instead, you teach the model by scoring what it decided. Got it right, it gets a reward. Got it wrong, it loses. Over time the model adapts to your domain.

The combination is what matters. Open models alone are everywhere. An RL platform alone is ML team stuff. Together, packaged for people who already use Workers, they become a product tool.

What a decision model is, in practice

Think of a support triage endpoint. Today a lot of people solve it like this. They send the ticket text to a big LLM with a prompt like "answer only with one of these categories: bug, billing, feature, spam". Then they parse the response. They hope it doesn't come back as "Sure! The category is bug." And they handle the case where the model invents a fifth category.

It works. But you're paying for a model that can write a sonnet just to get one word back.

A decision model flips this. It takes the input and the options and returns a choice. The output is structured from the start. No free-text parsing. No creative answers. No "as a language model...".

Most LLM calls in production are a decision dressed up as text.

That's my strong opinion of the day. If you open your app's logs and count, I bet more than half of the calls end in a switch or an if. For those cases, generating text wastes latency, tokens and patience.

Why RL fine-tuning matters to people who don't do ML

Traditional fine-tuning is scary because it needs a labeled dataset. You need thousands of examples with the right answer. Clean. Balanced. Almost no product team has that ready.

RL changes the question. You don't need to know the right answer in advance. You need to be able to say, afterwards, whether the decision was good. And your system often already knows that:

  • Did the agent move the ticket to another queue? The decision was bad.
  • Did the user click the recommendation? It was good.
  • Was the comment flagged as spam restored by a moderator? It was bad.

That signal is already in your database. It just isn't being used to train anything. With a managed RL platform, that history becomes raw material. For a full stack dev, this is what takes the news from "cool, another model" to "ok, I can actually use this".

How this fits into a Bun backend

I'll show the flow with a triage example. The exact payload format and model name are in the Cloudflare docs. The code here is illustrative, just to show how it fits.

First, the call. On the Bun side it's just a fetch to the Workers AI REST API:

// triage.ts
const options = ["bug", "billing", "feature", "spam"] as const
type Choice = (typeof options)[number]

const url = `https://api.cloudflare.com/client/v4/accounts/${Bun.env.CF_ACCOUNT}/ai/run/${Bun.env.DECISION_MODEL}`

export async function triage(ticket: string): Promise<Choice> {
  const res = await fetch(url, {
    method: "POST",
    headers: { Authorization: `Bearer ${Bun.env.CF_TOKEN}` },
    body: JSON.stringify({ input: ticket, options }), // illustrative format
  })
  if (!res.ok) throw new Error(`decision failed: ${res.status}`)
  const { result } = await res.json()
  return result.choice
}

Notice there's no regex to clean up the response. No JSON.parse inside a try while you pray. The Choice type seals the contract.

Second, and this is where most people will skip a step: store the decision and the outcome. Without that, RL has nothing to learn from. With native Bun.sql and Postgres, it stays short:

import { sql } from "bun"

// when deciding
const [row] = await sql`
  insert into decisions (input, choice, model)
  values (${ticket}, ${choice}, ${Bun.env.DECISION_MODEL})
  returning id
`

// when a human corrects the queue (or doesn't)
await sql`
  update decisions
  set reward = ${finalQueue === choice ? 1 : 0}
  where id = ${row.id}
`

It takes less planning than it seems. One reward column and you already have the skeleton of a continuous improvement loop. A month from now, that history feeds the fine-tuning. You swap DECISION_MODEL for the tuned version and compare the accuracy rate. Environment variable, deploy, done.

And since the weights are open, there's a plan B. If Cloudflare ever changes pricing or terms, you take the tuned model and run it somewhere else. That's not a detail. I've seen teams rewrite half their backend because the AI provider changed its API from one month to the next.

Where I'd stay cautious

It's not all good news. A few things I'd check before putting this in production:

A bad reward signal makes a bad model. RL optimizes exactly what you measure. If your reward is "the user clicked", the model learns to generate clicks, not value. Classic. Think twice about what "getting it right" means in your domain.

Volume matters. If your endpoint makes 30 decisions a day, the history will take a long time to teach anything. For low volume, a well-written prompt on a generic model still does the job, with fewer moving parts.

Hidden lock-in. The weights are open, but the RL platform belongs to Cloudflare. If your training pipeline depends on it, the most valuable part (the improvement process) is still locked in. Check whether you can export the tuned model and the training data without pain.

Decisions aren't everything. If the task needs an explanation, a summary or text for the end user, a decision model won't help. It's a scalpel, not a Swiss Army knife. And that's fine. Just don't try to make it write the reply email to the customer.

Is Clef worth testing?

For me, yes, with a clear scope. Pick one decision in your system that runs at decent volume and already has a natural right-or-wrong signal. Triage, moderation, routing. Put the decision model side by side with what you use today. Log everything. Compare accuracy, latency and cost per thousand calls for two weeks.

If it ties on accuracy and wins on latency and cost, you just pulled a giant model out of a place it never should have been. If it loses, you've gained a decisions table full of real data. That's already better than the guesswork you had before.

What I like most about this announcement is that it pushes the AI conversation somewhere more honest. Less "agent that does everything", more "function that makes one decision well". It's the kind of AI that fits in a real backend, with types, tests and metrics. If you want to see how I usually structure this kind of integration in real projects, take a look at my projects.

LinkedIn summary

Most LLM calls in production are just a decision dressed up as text.

Classify a ticket, flag spam, pick a queue. We pay for a model that can write a sonnet just to get one word back.

Cloudflare launched Clef: open-weight decision models with reinforcement fine-tuning.

The good part is you don't need a labeled dataset. The success signal is already in your database: the ticket that changed queues, the comment a moderator restored.

In my Bun backend, this comes down to a typed fetch and a reward column in Postgres.

Less "agent that does everything", more "function that makes one decision well". I wrote up my full take on the blog, with code and the points where I'd stay cautious.

#Cloudflare #AI #TypeScript #Bun #Backend