AgentryBook a call
← All posts

Decision Models: Cheaper AI Agents That Actually Ship

Agentry#aiagents#llmfine-tuning#reinforcementlearning#agentarchitecture
Decision Models: Cheaper AI Agents That Actually Ship

Most agent failures aren't a reasoning problem. They're a cost and latency problem — you picked a model too big for the job.

Cloudflare's Clef project just dropped open-weight decision models trained with reinforcement learning, paired with a platform for RL fine-tuning. It's a quiet release, but it points at something practitioners have been working around for a while.

What a Decision Model Actually Is

A decision model is a small, specialized model trained to make one class of choice well: route this request, classify this input, pick the next tool call. It doesn't need to write prose or summarize documents. It needs to be fast and right.

Most teams default to throwing GPT-4-class models at every step in their agent pipeline. That works in a demo. In production, it's slow and expensive, and the cost compounds across every workflow run.

Decision models are the alternative: train a smaller model on the specific decisions your agent makes, and swap the big model out for the narrow one wherever reasoning depth isn't required.

Why RL Fine-Tuning Changes the Calculus

The traditional fine-tuning playbook is supervised: collect labeled examples, train on them, evaluate. It works, but building a labeled dataset for agent decisions is painful. What counts as the "correct" tool call in a branching workflow?

Reinforcement learning sidesteps the labeling problem by training on outcomes instead. The model tries things, gets a reward signal based on whether it worked, and adjusts. For decision tasks with clear success criteria (did the right tool get called? did the router send traffic to the right handler?), RL is a much cleaner fit than supervised fine-tuning.

Cloudflare building an RL fine-tuning platform into Clef means that loop — train, evaluate on real outcomes, iterate — gets easier to run without standing up your own training infrastructure.

What This Means for Agent Architecture

The practical pattern this enables is a two-tier agent: a small, fast decision model handles routing and tool selection; a larger model steps in only when generation quality actually matters (writing a response, drafting a document, synthesizing a long context).

This isn't new as an idea. Teams at scale have been doing it manually. What Clef does is lower the floor: open weights mean you can run and fine-tune these models without a vendor relationship, and the RL platform means you don't need an ML team to close the training loop.

For a mid-size team shipping an internal agent, that's the difference between "we could do this" and "we can actually do this."

Where Teams Usually Get This Wrong

The failure mode is treating every step of your agent pipeline as equally complex. It isn't. Classifying which department owns an incoming support ticket is a different problem from generating the reply. They shouldn't use the same model.

The teams that build cost-efficient agents in production draw a map of their agent's decision points, sort them by reasoning complexity, and match model size to task. Decision models fit in the bottom tier: high volume, low complexity, clear success criteria.

If you haven't mapped those decision points for your own workflows yet, our free AI Opportunity Audit will pull your three highest-impact automations from your website alone. It's a fast way to see where the routing and classification work actually lives before you start building.

Open Weights Matter More Than They Look

The open-weight piece of the Clef release is worth flagging separately. Proprietary decision models lock you into one vendor's pricing and uptime. Open weights mean you can self-host, fine-tune further, and move the model if something cheaper comes along.

For agent infrastructure that handles high-volume internal decisions, vendor lock-in at the model layer is a real operational risk. A self-hosted decision model at the routing layer, with a proprietary model handling only the generation steps, is a more defensible architecture.

The RL fine-tuning platform on top of that is what makes it practical. Without tooling to close the training loop, open weights are mostly theoretical. With it, you can actually iterate toward a model that knows your specific workflows.

Closing

If this kind of tiered agent architecture is what you want built for your team, book a call and we can talk through what it would look like for your workflows.

Want an agent like this built for your business?

Agentry ships production AI agents in weeks. See where they'd help you first with the free AI Opportunity Audit or the other tools, then book a call to scope it.

Book a call →