As AI agents take on more responsibility in customer service, containment alone can’t tell you whether they’re delivering the right outcomes. Joseph Kim, CEO of Druid AI, explains why customer service leaders need a broader definition of AI performance.
For years, customer service automation has been judged on one metric: how many conversations get resolved without involving a human?
But AI agents are changing jobs. They’re moving beyond answering questions to completing tasks, touching enterprise systems, applying rules and deciding what happens next. Containment doesn’t measure any of that.
An agent resolving 96% of conversations looks more effective than one resolving 70% — until you ask why. The second agent might simply be better at recognising when human judgement is required, when policy demands intervention, or when a specialist would get the customer a better outcome.
Customer service leaders need a broader definition of AI performance.
Instead of judging AI primarily on containment, organisations should ask whether the agent understands its role, completes the right tasks, stays within its boundaries and hands off well when a human is the right outcome.
The containment trap
Containment made sense when automation had one job: answer repetitive questions and keep easy contacts out of the queue. Agentic AI does something else entirely — it retrieves customer information, works across enterprise systems, executes workflows and takes action on the customer’s behalf.
That changes what success looks like. Say one agent contains 96% of conversations and another contains 80%.
The first looks better on a dashboard. But the second recognised that a banking customer had reached a point where only a specialist would do. It handled what was in its remit, followed the organisation’s rules, and made a warm handoff to the right person.
That interaction failed containment. The agent still did exactly what it was supposed to do.
This is why you measure an AI agent against its actual role — what it’s authorised to do, the rules it must follow, and the point where responsibility passes to someone else. Holding onto a conversation isn’t success if the agent should have let go.
A good AI agent needs to know when to hand over
Escalation therefore needs a rethink.
A customer transferred because an AI agent got confused or misread the request — that’s a failure. An agent recognising that the next step needs human judgement, approval or specialist expertise is something else entirely.
Treat both as failures and you create the wrong incentive. Optimise relentlessly for containment, and you train your AI agents to hang onto conversations they should be escalating.
The better measure is whether the handoff happened for the right reason, and whether it worked. Nobody wants to explain their problem to an AI agent, get transferred, and start over from zero.
A good handoff preserves context. The human agent should already know who the customer is, what they need, what’s already happened, what’s been tried and why the AI escalated.
This distinction matters as organisations blend AI-powered self-service with human agent assist. The goal was never to remove humans from every interaction — it’s to bring them in exactly where their judgement, authority or expertise adds value, and to make that handover invisible to the customer.
How to measure whether the job actually got done
As AI agents move from answering questions to doing the work, task completion is what matters.
Take a patient asking an AI agent to reschedule a hospital appointment. A fluent reply isn’t enough. Did the agent understand the request? Did it authenticate the patient, apply the right rules and use the right system? And the question that actually matters: was the appointment changed?
The same principle holds across customer service. A retail customer requests a return — was it created? Someone reports a lost bank card — was the right action taken? A student asks about their application — was the correct information retrieved?
These are the moments where a successful conversation and a completed task turn out to be two very different things.
So organisations need to define what “done” means, task by task: did the agent understand what was required, follow the right rules, use the right tools and systems, and reach the correct end state?
Sometimes the correct end state is escalation to a human. What matters is that the agent recognised it and handed over well.
More autonomy demands a broader definition of success
The more responsibility organisations give AI agents, the less useful any single metric becomes.
An agent answering FAQs can be judged on the accuracy and usefulness of its answers, full stop. An agent authorised to access records, change bookings or execute transactions needs far more scrutiny than that.
Accuracy still matters. So do governance, policy adherence, system access and operating boundaries. Organisations need an auditable record of every action an agent took, every tool it used, whether the task reached its intended outcome, and when responsibility passed to a human.
Put together, that’s a far more useful way to measure performance. Did the agent understand the task, follow the rules, use the right systems and tools, and complete it correctly? And where it shouldn’t complete the task itself, did it hand over appropriately and preserve the context the human needed to pick up where it left off?
Containment still has value — it tells you something real about capacity and efficiency. What it can’t tell you, on its own, is whether your AI agent is doing a good job.
As AI agents take on more consequential work, the question changes. It’s no longer “Did the AI avoid involving a human?” It’s the one that actually matters: “Did the AI do right by your business and your customers?”
About the Author

Joseph Kim is the President & CEO of Druid AI. He has over 25 years of experience as an operating executive in Application, Infrastructure, and Security industries. He has seen success in both start-up companies as well as the largest enterprises – such as General Electric, Hewlett-Packard and Citrix.
Joseph is passionate about helping customers address complex challenges through the delivery of powerful innovations and technologies. He is also an expert in technology and process standards, as well as various industry verticals, bringing a unique 360-degree perspective that aligns development and engineering vision and strategies to advance corporate business objectives.
Joseph also serves on the Board of Directors for SmartBear and Andela.
Druid AI is an end-to-end, enterprise AI platform that enables lightning-fast development and deployment of AI Agents for teams looking to automate business processes and improve access to knowledge bases. Druid AI Agents enable personalised, omnichannel and secure interactions while seamlessly integrating with existing business systems and RPA technologies.
Since 2018, Druid AI has been actively pursuing its vision of providing each employee with an intelligent virtual assistant, establishing an extensive 200+ partner network, and servicing 300+ clients worldwide.