Perspective

Agents Are Having a Moment. Reliability Is Having a Reckoning.

July 9, 2026·6 min read·By Ricky Grannis-Vu

Agents are having a moment. Reliability is having a reckoning.

We are somewhere near the top of the agent hype cycle. Every product has an agent now. Every demo shows a model booking the travel, reconciling the books, resolving the ticket, all on its own. The demos are genuinely impressive, and I don't say that dismissively, because the underlying models really have gotten remarkable at understanding what you want and stringing together a plausible plan to get it.

But there's a widening gap between what these systems can demonstrate and what teams can actually depend on, and I think 2026 is the year that gap stops being a footnote and becomes the whole conversation. The excitement was about capability. The reckoning is about reliability.

The demo-to-production cliff

Here's the pattern I've watched play out over and over, at Orby and since. A team adopts an agentic tool. It dazzles in the pilot. It breezes through the happy path, handles the common cases, saves obvious time, and everyone's thrilled. Then it meets reality: the unusual invoice, the vendor whose name is spelled three different ways, the edge case that only appears at month-end. And the system that felt magical in the demo does something subtly, expensively wrong. Worse, it does it differently each time, so nobody can quite reproduce it or trust the fix. The hours saved in the morning get spent in the afternoon cleanup. Confidence erodes. The tool gets quietly demoted to the tasks that didn't really matter anyway.

This isn't a story about bad models. It's a story about asking a probabilistic system to be the deterministic core of a process that can't tolerate variance.

A simple chart making the point that a 95 percent reliable step, run 1,000 times a day, produces about 50 failures a day. As the same step is chained across a multi-step workflow, the share of runs that succeed end to end falls sharply.

Why "just wait for a better model" doesn't close the gap

The instinct is to wait for the next model. Surely at some accuracy threshold the reliability problem simply dissolves. I don't think it does, and the reason is structural rather than a matter of scale.

Language models are probabilistic by design. That's the source of their flexibility and their fluency, and it is also, inescapably, the source of their variance. You can push the error rate down, but you cannot make a fundamentally probabilistic process return identical results on identical inputs every single time, which happens to be the exact guarantee that serious operational work requires. Do the arithmetic and it gets uncomfortable fast. A step that's 95 percent reliable sounds great until you run it a thousand times a day, at which point it's failing fifty times a day. Chain a few such steps together and the odds that a whole run comes out clean end to end start dropping toward a coin flip. Payroll can't be approximately right. A ledger entry can't be a good guess. Compliance can't hold most of the time.

So the answer isn't a better guesser. It's to stop asking the model to guess at the part of the job that was never supposed to be a guess.

The reframe: intelligence at the edges, determinism at the core

The teams getting durable value from AI right now, and I mean durable value rather than demo value, have mostly converged on the same architecture, whether or not they'd describe it that way. They use models for what models are uniquely good at, which is understanding a fuzzy request, reading a messy document, proposing a plan. And they refuse to let a model be the thing that executes irreversible work at runtime.

In that design, models do what they're genuinely good at while staying well clear of the execution path. What actually runs is a compiled, explicit sequence of steps: deterministic, inspectable, the same every time. The intelligence lives at the edges, at the specific moments where reading and judgment are truly needed, and the core stays boring on purpose. When something goes wrong there's an audit trail that explains exactly what happened, and when you fix it the fix stays fixed, because the execution isn't being re-improvised on every run. This is less exciting than "the agent does everything." It's also the version a finance team can actually put in front of an auditor, which is the version that survives contact with a real business.

What the reckoning actually clears the way for

I want to be clear that this is an optimistic argument, not a skeptical one. The point isn't that AI can't be trusted with important work. It's that trust has to be engineered, and the industry is finally getting serious about engineering it. That's progress, not retreat.

The winners of the next few years won't be whoever has the most autonomous-looking demo. They'll be whoever makes AI dependable enough that a careful person is comfortable putting their name on its output and being on the hook for it. That's a higher bar than impressive, and it's the bar that unlocks the actual prize, which was never automating the easy ninety percent that was never the problem. It was earning enough trust to take on the work that genuinely matters.

Capability got us to the demo. Reliability is what gets us to production. The reckoning is just the industry, a little belatedly, agreeing on which one was ever the hard part.


We're building Lodol for the production side of that line.
If you want automation you can put your name on, deterministic where it counts and intelligent where it matters, we'd love to show you what that looks like.

Share this post

Ship your first automation in days, not months.

See how Lodol keeps your workflows reliable as your team grows. Create your first workflow today.