Are Vendor AI Benchmarks Trustworthy? How to Read CX Agent Numbers in 2026

Are Vendor AI Benchmarks Trustworthy? How to Read CX Agent Numbers in 2026
AI CX benchmark comparing vendor claims with live results: 15 deployments answered queries, but only 4 completed tasks.
AI CX Benchmarks: How to Evaluate Vendor Claims | Alhena AI

AI CX Benchmarks: How to Separate Evidence From Vendor Claims

A benchmark with no failure analysis is not a benchmark. It is a brochure with a methodology section.

What should a credible AI customer service study tell you?

A credible AI customer service study identifies who tested the system, where it ran, what counted as success, and what failed. In Alhena’s 2026 Agentic CX Stress Test, all 15 live deployments could answer questions, but only 4 demonstrated a real action. That finding describes observed capability across deployments, not a universal task-resolution rate. 1

The question is not whether AI can produce a convincing response. It is whether the customer got the answer, action or assistance they actually needed.

What four questions separate evidence from marketing?

  1. Who ran the test? Identify the author, funder and evaluator. First-hand testing is not necessarily independent testing. A study published by an AI vendor should disclose that interest, including a study published by Alhena. Independence and methodological quality are separate questions.
  2. Where was it run? Distinguish a sandbox, a pilot and a production deployment. Alhena’s methodology used each provider’s largest verifiable live deployment on a real brand’s public storefront, with the tester behaving as an ordinary shopper. That tests an implemented experience, not everything a platform might support. 2
  3. Which metric was measured? Ask for the numerator, denominator, exclusions and observation period. An AI conversation that ends without escalation is not automatically evidence of a solved problem. Establish whether success means a correct answer, a completed action or a successful handoff.
  4. What failed? A useful audit includes unsuccessful attempts, their frequency and the point where execution stopped. Alhena reported seven failure modes, including answer-only fallback in 10 of 15 deployments and unsafe confidence in 3 of 15. Failure counts make the findings more useful than a headline alone. 3

Which metric definitions can hide the most?

Before comparing percentages, compare definitions. These labels can describe different events depending on the measurement method.

Reported metric What it may count What to ask instead
Deflection rate Contacts that did not become human-handled tickets Was the need resolved, abandoned or moved elsewhere?
Containment rate Conversations that stayed in the automated channel Did the customer get what they needed without escalation?
Resolution rate Outcomes classified as solved by a model, customer or reviewer Who verified success, and what evidence did they use?
“Handles X% of queries” Requests that received a response How many eligible requests reached their intended outcome?
CSAT Customer satisfaction among survey respondents Who received the survey, who responded, and was the issue resolved?

Automated classification is not inherently useless. For example, Ada’s documentation distinguishes resolution from containment and describes assessing the customer’s inquiry and the system’s responses. The important question is how that assessment is validated against outcomes. 4

For an informational request, a correct reply may be the complete solution. A shopper asking about a return deadline may need nothing more.

A transactional request is different. When the shopper asks an AI agent to issue an eligible refund, explaining the policy does not complete the requested action. The same principle applies when an ecommerce customer asks to cancel an order rather than learn how cancellations work.

For chatbots and other conversational AI systems, score against the request—not how fluent the response sounds.

Answer quality and task completion belong in separate columns. Neither should stand in for the other.

What does an honest testing methodology look like?

Alhena’s approach combined one continuous conversation across discovery, a post-purchase issue, an agentic action and an emotional moment. It scored eight dimensions on a 1–3 scale and inferred no capability from marketing. The purpose was to observe behaviour across a customer journey rather than collect isolated answers. 2

Memory needs its own test. Each memory tier answers a different question:

Memory tier What to test
Tier 1: Within the conversation Does the agent retain a budget, preference or constraint introduced earlier?
Tier 2: Logged-in context Does it appropriately use authorised account information and purchase history?
Tier 3: Return visit Does relevant context persist when the customer returns hours or a day later?

Passing one tier does not prove the others work. Record what the AI recalled, whether that information was appropriate to use, and what the shopper had to repeat.

For your own evaluation, define the expected result before testing. Keep the prompts, timestamps, conversation record and supporting evidence. For an action, include confirmation from the system that owns the order or account—not only the final chat message.

A live test also has limits. Alhena’s framework does not establish performance at scale, comprehensive security, multilingual coverage or behaviour during outages. Those require additional testing rather than assumptions based on a successful demonstration. 2

Which field scores provide useful calibration?

These are selected dimension averages from Alhena’s 15-deployment study, using its 1–3 scoring scale. 5

Dimension Average score
Context 3.00
Accuracy 2.47
Agentic Capabilities 1.60

Treat these as results for the tested sample, not an independently established industry average. A high Context score does not mean understanding is solved for every language, customer or deployment. An average of 2.47 does not make accuracy failures unimportant.

The practical implication is narrower: do not let strong answers substitute for evidence that an AI system can finish the work. A claim to outperform the field on execution should be demonstrated against the same scoring criteria.

How should you audit AI customer service in your own support stack?

A useful comparison starts with your customer support environment, not a feature checklist. Whether you are evaluating Zendesk, Ada, Freshworks or Alhena, ask each provider to demonstrate the same requests under comparable conditions. These are procurement examples, not identities assigned to the study’s anonymised deployments.

Start with the customer’s intended outcome

Give an AI assistant two different jobs. First, ask it to retrieve information from your knowledge base, such as a returns deadline. Then ask it to check order status and complete an authorised change.

Score information retrieval separately from the transactional workflow. Can the system identify the correct order, check permissions, perform the requested operation and confirm the result?

An autonomous AI claim should come with evidence of what the system can do autonomously and where approval is required. Ask for an audit trail showing the backend request, result and customer confirmation. A confident message is not proof that the action succeeded.

Read the resolution definition and the bill together

For any SaaS platform, request the measurement rules alongside the commercial terms. Ask which events count toward usage, how repeat contacts are treated, and which charges apply when a human finishes the work.

Zendesk’s documentation distinguishes account configurations and pricing models for automated resolutions. Check the rules that apply to your account rather than assuming an older case study uses the same definitions. 6

Then open the dashboard and inspect the underlying conversations. Can your team trace a reported success to its evidence? A procurement review should reconcile the displayed outcome with the customer’s result and the bill.

Check CSAT against completed work

CSAT measures customer satisfaction; it does not independently verify that a requested change happened. For a positive-response CSAT score, calculate positive valid responses divided by all valid responses, multiplied by 100. Specify the survey scale and which responses qualify as positive. 7

Report survey invitations and response rates alongside the result. Separate AI-only interactions from those completed by a human agent, and compare similar request types. Freshdesk’s survey reporting, for example, includes surveys sent and responses received—not just a headline rating. 8

A pleasant conversation with an unresolved issue should remain visible in your reporting.

Test the handoff, including the queue

When an agent needs help, check what reaches the human agent in your helpdesk. The handoff should include the customer’s goal, relevant account context, actions attempted and the reason for escalation.

Check that the ticket enters the correct queue with an owner and a clear next step. To reduce agent rework, test whether the person taking over can continue without asking the customer to start again.

Judge the AI on appropriate escalation, not on avoiding every transfer. A well-handled exception can be a better customer experience than another automated response.

Review autonomous actions under real operating conditions

Before you automate more work, test the workflow with missing information, an unavailable integration and an action that needs approval. Check what the AI does when it cannot safely proceed.

For high volume operations, request separate evidence on latency, failed actions and repeat contacts. Do not extrapolate capacity from one successful chat. Your operational review should also establish who investigates failures and which changes trigger retesting.

Build the next audit around unresolved cases, not only successful ones. Keep enough evidence to reproduce the failure and confirm that the fix holds.

What are the five red flags in an AI benchmark?

  • No failure section. Without unsuccessful cases, you cannot see where the system stops being reliable.
  • No reproducible methodology. Missing dates, prompts, environments or scoring rules prevent meaningful scrutiny.
  • An unexplained resolution label. Model-based assessment can be useful, but the evaluator should state how it works and how classifications are checked.
  • Aggregate results without relevant breakdowns. Ask for differences by request type, channel and vertical. A single percentage can conceal the work your business actually needs done.
  • No long-conversation or return-visit testing. Short exchanges do not test whether the AI retains context across a longer journey. Alhena’s framework explicitly tests both conversational continuity and return visits. 2

Key takeaways

  • Define success before reading the score. A correct answer, a completed task and a successful handoff are different outcomes. Your reporting should preserve those differences.
  • Evaluate the evidence, including its limits. Live observations, failure analysis and reproducible methods are useful. Neither a polished dashboard nor a named provider makes a study independent.
  • Use Alhena’s findings as a prompt to test your own deployment. All 15 systems in its 2026 study could answer; only 4 demonstrated an action. The question for your business is which customer requests your implementation can actually finish. 1

FAQ: Evaluating AI customer service claims

What should an AI customer service benchmark include?

It should state who ran the study, the test dates, deployment environment, request types, scoring definitions and exclusions. Look for observed outcomes, failure frequencies and limitations. Alhena’s 2026 Agentic CX Stress Test offers a live-testing framework; apply the same scrutiny to its findings as to any other vendor’s research. 2

How do I compare Zendesk and Ada with Alhena?

Use the same customer requests, authorisation requirements and definition of success. Evaluate the specific product configuration connected to your helpdesk, not the brand name alone. Request conversation evidence, completed-action records, escalation outcomes and applicable usage terms. This is a method for testing providers, not a claim that their capabilities are identical.

What is the difference between deflection and resolution?

Deflection concerns avoided human contact; resolution concerns the customer’s need being met. An AI agent may retrieve information and fully resolve a policy question. But describing how to request a refund does not resolve a request to process one. Define the expected outcome from the customer’s intent before assigning credit.

What is a useful CSAT target for AI support?

Start with your own baseline for comparable requests, channels and survey methods rather than borrowing an unrelated headline. Compare AI-only and human-assisted results, disclose response rates, and review whether issues were completed. Survey reporting should show whose feedback the number represents, not imply that every customer responded. 8

What is a realistic containment rate for ecommerce support?

Containment alone is not a reliable buying target. Pair it with verified outcomes and repeat contacts over a stated period, such as 24 hours. An AI agent can keep a conversation in channel while leaving the work unfinished. Choose targets based on your request mix and acceptable escalation behaviour. 4

How do I sanity-check an AI case study before sharing it internally?

Ask which deployment produced the result, when it was measured, how many requests were included and what the baseline was. Confirm whether AI handled comparable work before and after the change. Include exclusions and unsuccessful cases in your review instead of repeating the most favourable number without context.

Should I trust research published by an AI vendor, including Alhena?

Treat it as an interested source, not automatically as independent research. Assess its methodology, disclosures, evidence and reproducibility. A commercial author can publish useful findings, but the company’s own account of its study is not independent verification. Apply that distinction consistently, including to Alhena.

What does measurement technique mean, and why does it matter?

It means how the evidence was collected: scripted tests, public-storefront conversations, production records or customer surveys. Each answers a different question. A knowledge base test can assess answer retrieval; it cannot by itself establish whether an agent completes account changes. Ask how the technique matches the claim being made.

Why are competing deployments anonymised?

Alhena’s stated rationale is that naming a storefront would identify its technology provider. Anonymisation keeps attention on category-level behaviour, but it also limits outside verification of individual results. A credible report should acknowledge that trade-off and still publish enough detail about selection, testing and scoring to make its method assessable.

How old can an AI benchmark be before it becomes less useful?

There is no universal expiry date. Check whether the model, integrations, permissions or evaluation method have changed since testing. A recent study can already describe an outdated configuration. When teams use AI for new tasks, retest those tasks rather than treating the publication year as proof of current performance.

How do I verify an autonomous agent’s capabilities?

Ask the AI assistant to complete an authorised task in an environment that represents your intended deployment. Inspect the resulting system record, not only the response. Include a case that requires escalation. A successful demonstration establishes that specific observed capability; repeated testing is needed to assess reliability.

What is the most misleading way to report AI performance?

Treating requests answered as tasks completed. In Alhena’s 15-deployment study, every system could answer, but only 4 demonstrated an action. That is a count of observed capabilities across deployments—not the percentage of all customer tasks those systems resolved. Mixing the two makes a finding sound broader than it is. 1

See a benchmark built on behaviour, not claims

Explore Alhena’s full methodology, eight-dimension scorecards, seven failure modes and findings across eleven verticals.

Power Up Your Store with Revenue-Driven AI