Kamran Akbar
Back to blog

99 Percent of Companies Are Building AI Agents. Only 1 in 10 Ever Ships.

By Kamran Akbar · August 26, 2026 · 4 min read

Key takeaways

  • 99% of companies plan agentic AI deployment, but only 9 to 14% have gotten a system into production, according to new research from Ness Digital Engineering.
  • The two biggest reasons pilots stall are a loss of trust in probabilistic outputs and a lack of visible change in daily operations.
  • Info-Tech Research Group found that pilot-era AI stacks, built fast to prove a concept, often can't survive real operational demands like integration load and cost governance.
  • Companies that reach production tend to assess data and infrastructure readiness first, measure success at the user experience level, and track AI costs against real ROI.
  • A good demo is not proof of a good deployment. The real test is whether the system can be trusted unsupervised and whether anyone would notice if it disappeared.

Nearly every company in the country now says it's building AI agents. Almost none of them have actually gotten one into production.

New research from Ness Digital Engineering, published this week, puts numbers on a pattern most operators have already felt in their bones. Ninety nine percent of companies plan to deploy agentic AI. Only 9 to 14 percent have actually moved a system into production. The report calls the space between those two numbers "Death Valley," the stretch where a promising pilot quietly stalls and never becomes a working part of the business.

A separate report from Info-Tech Research Group, also published this week, backs this up from a different angle. It found that the agentic AI stacks companies build during the pilot phase, thrown together fast to prove a concept, tend to break down the moment they're asked to run for real. Integrations turn out to be brittle. Costs run past budget with no governance in place to catch it. Data pulled from stale or untrusted sources quietly poisons the output. As the researchers put it, "agent demonstrations look alike, but operational realities do not."

Two reasons pilots stall

Strip away the jargon and the Ness report points to two root causes.

The first is trust. Leaders greenlight a pilot, watch it perform well in a demo, then hesitate to hand it real authority over a customer interaction, a financial decision, or an operational process. That hesitation is rational. A probabilistic system that's right 95 percent of the time is still wrong one time in twenty, and most businesses have not built the guardrails to catch that failure before a customer does.

The second is that the pilot doesn't change anything anyone can see. Employees don't notice a difference in their daily work. Customers don't notice a difference in their experience. When a project produces no visible change, it's the easiest line item to quietly deprioritize at budget season, whatever the demo looked like six months earlier.

Why the architecture matters as much as the use case

The Info-Tech report adds a piece that's easy to miss when you're focused on picking the right use case. A pilot built on a stack that was never meant to scale will hit a ceiling no matter how good the underlying idea is. Their recommendation is to think in six layers, from the application a person actually touches, down through the orchestration engine that governs what an agent is allowed to do, down to the data platform and infrastructure underneath. Vendors and internal builds should be judged on whether an agent running on top of them is easy to observe, explain, debug, and shut off safely, not just on how impressive the pilot demo looks.

What the companies in the 9 to 14 percent are doing differently

A few patterns separate the businesses that make it across Death Valley from the ones stuck on the near side of it.

They assess before they build, looking honestly at whether their data is clean enough, their APIs are exposed enough, and the business domain is well defined enough for an agent to operate in it without constant human rescue.

They measure success at the level of the actual user experience, not the level of an isolated task. A support agent that resolves tickets faster but leaves customers navigating five different tools to get there hasn't actually improved anything worth keeping.

They consolidate rather than sprawl, pulling agent interactions into a single interface instead of bolting one more tool onto an already crowded workflow.

And they track cost the way they'd track any other line of the business, watching token consumption, cloud spend, and subscription fees against a real return, not a hoped for one.

What this means if you're running the business, not just watching the trend

None of this is an argument against agentic AI. It's an argument against treating a good demo as proof of a good deployment. The gap between the two is where most of this year's AI budget is quietly disappearing.

If you're evaluating a pilot right now, the useful question isn't "does this work in the demo." It's "would I trust this to run unsupervised on a bad day, and would my team or my customers actually notice if it disappeared tomorrow." If the honest answer to either question is no, you're still standing at the edge of Death Valley, and that's a much better place to know you're standing than to find out six months and one budget cycle later.

My take on the agentic AI production gap

I've sat through more agentic AI demos this year than I can count, and almost every one of them looked great. That's exactly the trap this research is pointing at. A demo tells you almost nothing about whether a system can survive contact with your actual business.

When I evaluate whether to move something past a pilot, I stop asking whether it works and start asking two blunter questions. Would I trust this running unsupervised on our worst day, not our best one. And would my team or my customers actually notice if I turned it off tomorrow. If either answer is no, I'm not looking at a production system. I'm looking at an expensive proof of concept that's going to quietly die at the next budget review.

The trust problem is real and I don't think it gets solved by a better model. It gets solved by building the guardrails first, clear logging, clear boundaries on what the agent can touch, and a fast way to catch it when it's wrong. The visibility problem matters just as much. If nobody outside the project team can feel the difference an agent makes, it was never going to survive contact with a spreadsheet at renewal time.

My advice to any operator right now is to spend less energy chasing the next flashy use case and more energy on the boring parts, data readiness, a single interface, honest cost tracking. That's what actually gets you out of Death Valley.

Questions people ask

What is the agentic AI Death Valley

It's the gap between companies piloting agentic AI and those actually running it in production. New research found 99% of companies plan deployment, but only 9 to 14% have gotten a system live.

Why do most AI agent pilots fail to reach production

Two main reasons stand out in the Ness Digital Engineering report, a loss of trust in probabilistic outputs and a lack of visible change in daily operations for employees or customers.

What should a business do before scaling an AI agent pilot

Assess data and API readiness first, measure success at the level of the user experience rather than a single task, consolidate agent interactions into one interface, and track cost against real ROI.

Does a good AI agent demo mean it's ready for production

Not necessarily. Info-Tech Research Group found that stacks built fast for a pilot often can't handle real operational demands like integration load, cost governance, or data quality at scale.

Want help putting this into practice?

Let's talk