AgentryBook a call
← All posts

When AI Vendors Oversell: A Builder's Reality Check

Agentry#aiagents#llmevaluation#vendorclaims#aicoding
When AI Vendors Oversell: A Builder's Reality Check

A credible technical founder publicly called out Anthropic this week for blowing smoke. That's worth sitting with — not as drama, but as a signal about where we actually are with AI agents in production.

What happened and why it matters

Nathan Sobo, creator of the Zed code editor, pushed back on Anthropic's framing around AI coding capabilities. The specifics matter less than the pattern: a practitioner who ships real software, building tools real developers use, looked at vendor claims and said they don't match reality.

This happens constantly in AI right now. Benchmarks get cherry-picked. Demo conditions don't reflect production. "State of the art" means something very specific that doesn't transfer to your actual codebase or your actual workflow.

For anyone evaluating AI agents for their business, this is the core problem: the gap between what a model can do on a controlled task and what an agent built on that model will do reliably in your environment.

The benchmark trap

Vendors optimize for benchmark performance because benchmarks are measurable and publishable. Production reliability is messy, context-dependent, and hard to market.

Here's what that gap looks like in practice:

  • A model scores well on coding benchmarks with self-contained problems. Your codebase has 8 years of context, custom abstractions, and implicit conventions no benchmark captures.
  • An agent demos beautifully on a 5-step task. Your real workflow has 23 steps, several of which require judgment calls the agent hasn't been trained to flag.
  • Latency is fine in the demo. Under real load, with real data sizes, it falls apart.

None of this means AI agents aren't useful. They are — we build them for a living. It means the evaluation has to happen in your environment, on your tasks, not on a vendor's slide deck.

How to actually evaluate an AI agent claim

When a vendor or a demo makes a capability claim, run it through these questions before building anything on top of it:

What's the task distribution? If they tested on 100 examples, what were those examples? Are they representative of what you'd actually throw at it?

What does failure look like? Good agents fail gracefully. They surface uncertainty, hand off to humans, log the gap. Bad agents hallucinate confidently and you find out three steps later.

What's the latency and cost at your volume? A model that's fast and cheap at 10 calls a day may be neither at 10,000.

Can you run a blind test? Take 20 real examples from your workflow. Don't tell the vendor which ones are hard. See what happens.

This is exactly the kind of audit that saves teams from building on a foundation that cracks under real load. If you want a faster starting point, our free AI Opportunity Audit looks at your actual business context and surfaces where agents are likely to hold up — and where they won't.

What credible pushback from builders means for the field

Sobo's public criticism is actually healthy. The AI space needs more of it.

When credible technical people — people building tools others depend on — start saying "this doesn't match what I'm seeing," it forces a more honest conversation. That conversation is better for everyone who wants to actually use AI agents to get work done, rather than just talk about it.

The studios and teams doing serious agent work already know the limits. They build around them: smaller, well-scoped agents with clear handoff conditions, evals that run on real data, monitoring that catches drift early. The hype doesn't help them. What helps is an honest map of where the technology is right now.

We're still early. The agents that work reliably in production today are narrower than the marketing suggests. They're also genuinely useful — often dramatically so — within those narrower bounds.

The takeaway for operators and founders

Don't let vendor benchmarks set your expectations. Set your expectations from a small, controlled pilot on your own data and your own workflow. Define what "good" looks like before you run it, not after.

That discipline is what separates teams that get real ROI from agents from teams that run a pilot, get burned by the gap between demo and reality, and write off the whole category too early.

The technology is worth using. It just needs honest evaluation, not slide decks.

Want this built instead of researched?

If you're ready to move past evaluation and actually ship an agent for your workflow, book a call and we can scope it. We build narrow, production-ready agents — no hype, just working software.

Want an agent like this built for your business?

Agentry ships production AI agents in weeks. See where they'd help you first with the free AI Opportunity Audit or the other tools, then book a call to scope it.

Book a call →