AI agents are becoming remarkably capable. The harder question for businesses is whether they remain reliable when they leave the test environment and meet the realities of everyday work.
The AI agent passed its evaluation.
It followed the instructions, used the right tools and completed the task. Everything suggested it was ready.
Then it went to work.
An API timed out. Someone phrased an instruction differently. A permission had changed. The information it needed had moved.
Nothing particularly dramatic happened. In fact, these are exactly the kinds of small problems that happen every day in a working business.
But suddenly, an AI system that had performed extremely well during testing wasn’t quite as dependable.
And that raises a much more useful question for businesses considering AI agents.
It isn’t simply can the agent do the job?
It is can it still do the job when things don’t go according to plan?
A benchmark isn’t a workplace
Traditional software testing is relatively predictable. Give the software a particular input and you generally expect a particular result.
AI agents don’t always work that way.
An agent might need to interpret an instruction, decide which tool to use, complete several steps, respond to whatever happens along the way and change its approach if something unexpected occurs.
That flexibility is part of what makes AI agents so interesting.
It is also what makes evaluating them difficult.
Anthropic has pointed to this problem in its own work on agent evaluation. Unlike a straightforward prompt-and-response system, an agent may operate over several turns, use different tools and make changes to its environment before it reaches the final result.
Researchers are now looking more closely at what happens when those agents encounter conditions that look more like the real world.
ReliabilityBench, for example, tested agents against problems such as timeouts, rate limits, incomplete responses and changing schemas. The research found that relatively small changes in the environment could have a noticeable effect on whether the agent successfully completed its task.
There is another complication.
Even the infrastructure running the evaluation can influence the result. Anthropic found that infrastructure configuration alone could shift results on an agentic coding benchmark by six percentage points. That’s potentially a bigger difference than the gap between some of the leading models on a leaderboard.
For a business trying to decide whether an AI system is ready to use, that’s important.
A benchmark tells us something about what an agent can do.
It doesn’t necessarily tell us what will happen at 9:15 on Monday morning when a system is slow, a password has expired and somebody hasn’t filled in the customer record properly.
Being capable isn’t the same as being reliable
Imagine an AI agent handling routine customer requests.
During testing, it completes nine out of ten successfully.
At first glance, that’s impressive.
But the tenth case may tell us far more about whether the system is actually ready for the business.
What happens if authentication fails halfway through?
Does the agent stop and explain the problem? Does it try again? Does it recognise that only part of the process was completed?
And does a person know that something went wrong?
That last question may be the most important one.
A failed task that is clearly flagged can be dealt with. A failed task that everybody assumes was completed can become a much bigger problem.
This is why newer approaches to enterprise-agent evaluation are beginning to look beyond raw capability. One recently published framework proposes asking not only how well an agent performs a task, but under what conditions, and at what cost, it can be deployed reliably.
That’s a much more practical question.
A system that achieves 90% accuracy but needs someone checking almost everything it does may ultimately be less useful than a slightly less capable system whose limitations are understood and easy to manage.
The last 10% may be where the real work begins
AI demonstrations are naturally designed to show us what the technology can do.
Can it analyse this document?
Can it update the CRM?
Can it investigate a support ticket?
Can it complete an entire workflow?
Those demonstrations are useful. But once AI moves from a demonstration into a working business, the questions need to change.
How often does it work?
What causes it to fail?
Does it know when it has failed?
When does somebody need to check its work?
And what happens to the business when it gets something wrong?
The answer won’t be the same for every task.
If an AI agent makes a mistake while drafting an internal meeting summary, somebody may need to spend five minutes correcting it.
If it makes a mistake while changing a customer’s account, handling financial information or interacting with production infrastructure, the consequences can be very different.
So reliability isn’t simply a technical measure.
It depends on what we are asking the AI to do and what happens if it gets it wrong.
Humans may not be leaving the loop just yet
There is understandably a great deal of excitement around autonomous AI agents.
But successful business adoption may turn out to be less about removing people from the process and more about working out where people are actually needed.
One enterprise deployment study provides an interesting example.
Two AI systems achieved almost identical autonomous accuracy: 72.8% and 72.5%.
On a benchmark, there isn’t much separating them.
But when researchers looked at the human involvement needed to reach the same reliability target, the difference became much more significant. One required human review in 39.2% of cases. The other required it in 29.6%.
That difference matters to a business.
Someone has to perform those reviews. They take time. They cost money. And they affect how quickly the overall process can run.
Suddenly, two systems that appeared almost identical from an accuracy score can look quite different when you consider how they would actually operate inside a company.
Perhaps we need to test the mess
Real workplaces are messy.
People phrase things differently. Information goes missing. APIs fail. Permissions change. Systems slow down. Someone enters the wrong information into a field.
So perhaps our testing needs to become a little messier too.
Test expired credentials.
Test unclear instructions.
Test missing information.
Test API failures and unexpected responses.
Run the same task more than once.
And perhaps most importantly, test what happens when the agent reaches a point where it simply doesn’t know what to do.
The aim isn’t to build an AI agent that never fails. That isn’t a realistic standard for technology, or for people.
What a business needs to understand is how the agent fails, how often it happens, what those failures could cost and when somebody needs to step in.
Benchmarks remain useful. They help us understand what these increasingly capable systems can achieve.
But businesses eventually have to answer a more ordinary, and arguably more important, question.
The test is finished. The AI agent is now doing real work.
Can we trust it?