Score each question yes or no based on behaviour you have personally observed in production. Not on a demo. Not on a capability page. Observed.
Ten questions, drawn from Alhena's live stress test of 15 AI shopping and CX agents, separate a working AI agent from an answer engine. Evaluate yours against all ten. Most deployments answer yes to two or three. If yours does, it is not acting. It is explaining.
What are the ten diagnostic questions?
QUESTION 1 · 4 OF 15Can it complete a real action end to end?
A return, a cancellation, an address change: performed, not described. Task completion is the whole test here, and every other question is downstream of it. The hardest bar in the study, cleared by 4 of 15, and Agentic Capabilities scored 1.6 of 3.0.
QUESTION 2 · 6 OF 15Can it guide a shopper to a specific product page?
Not a category. Not a search result. The specific product page, opened from the conversation. Only 6 of 15 could; dead-end recommendation appeared in 6 of 15. Evaluate this one end to end, from the recommendation to the page actually opening.
QUESTION 3 · 6 OF 15Does it preserve context as the shopper moves across pages?
If the conversation resets on navigation, chat and storefront are two separate surfaces and no amount of tool use will bridge them. UI disconnect: 6 of 15.
QUESTION 4Can it recommend a bundle under multiple constraints?
Not one product: a coordinated set, which is a multi-step reasoning problem rather than a lookup. A full fragrance-free AM/PM routine under budget. An outfit. A necklace and matching earrings, both nickel-free.
QUESTION 5 · THE FASTESTCan it explain why a product fits?
Ask “why this one?” Evaluate the answer, not the confidence: a reasoning system explains the causal link. A ranking system restates the description or cites a star rating.
This is the fastest single diagnostic on the list, and it detects catalog dumping (5 of 15).
QUESTION 6Does it use reviews at the decision stage?
Not a wall of stars up front, but the one review that addresses the specific doubt the shopper just voiced. Most agents surface neither.
QUESTION 7Can it handle an emotional moment with context intact?
Frustration, urgency, a gift with a deadline. Watch for coherence: does the tone shift without the constraints falling out? Every conversation in the study included an emotional moment for exactly this reason.
QUESTION 8 · 1 OF 15Does it remember responsibly across sessions?
Close the tab. Return tomorrow. This is the one question where a multi-day gap is the whole point. Only 1 of 15 deployments cleared this: the single most discriminating question on the list.
QUESTION 9 · 3 OF 15 FAILEDDoes it avoid overclaiming in safety-sensitive categories?
Ask something it should refuse. The correct answer is a graceful decline with a named referral, while keeping the rest of the conversation useful. Unsafe confidence: 3 of 15.
QUESTION 10Can it show measurable impact?
Conversion, AOV, deflection and ticket quality, measured in production rather than modelled in a deck. An AI agent that explains a return without completing it looks excellent on containment and resolves nothing.
How should you score the result?
| Yes answers | What you have |
|---|---|
| 0–2 | An answer engine with a chat interface. Fine for FAQs. It will not move conversion or resolution. |
| 3–5 | A capable assistant with a structural ceiling, usually a search or personalisation engine. It recommends; it does not act. |
| 6–8 | A genuine agentic assistant. Ahead of most of the market. |
| 9–10 | Elite. In a field of 15 live deployments, almost nobody cleared this. |
For calibration against the wider benchmark, field-wide dimension averages ran from 3.0 on Context down to 1.6 on Agentic Capabilities. Understanding is solved. Doing is not.
What does a “no” answer actually tell you?
A score is a diagnosis, not a verdict. Each no on this list points at a specific missing layer, and knowing which one turns an evaluation into a conversation you can actually have with a vendor.
| Question | A “no” usually means | Ask the vendor |
|---|---|---|
| 1. Complete an action | No write access, or no identity model to authorise it | Which systems can the AI agent write to, and on whose authorisation? |
| 2. Route to a product | No UI control from inside the conversation | Open a specific product page from the thread, live |
| 3. Preserve context | State is scoped to the widget, not the session | Where does conversation state live between page loads? |
| 4. Bundle under constraints | Ranking rather than multi-step reasoning | Build a coordinated set under three constraints at once |
| 5. Explain why | A ranking model with no causal chain to show | Show the reasoning, not the product description |
| 6. Reviews at decision | Reviews sit outside what the agent can retrieve | Quote the review that answers this specific doubt |
| 7. Emotional moment | Context drops when the topic or tone shifts | Re-run the same test immediately after a complaint |
| 8. Remember across sessions | No memory store keyed to an identity | What is memory keyed on, and how long does it persist? |
| 9. Restraint | No policy layer and no defined refusal boundary | Show the documented list of what it will never answer |
| 10. Measurable impact | Nothing measured in production beyond containment | Show resolution and 24-hour repeat contacts by intent |
Read the pattern rather than the individual rows. Failures clustered in questions 1 to 3 are an execution problem: the AI agent understands and cannot act. Failures in 4 to 6 are a reasoning problem, usually a ranking model wearing a conversational interface. Failures in 7 and 8 are a state problem. Failures in 9 and 10 are governance, and they tend to be the cheapest to fix.
This is also where a trace earns its place. When an AI agent fails task completion on question 1, the output the shopper sees looks identical whether an API call returned an error, the tool was never called at all, or the write was refused by your own permissions. Only a trace of the execution separates those three, and only your vendor can produce that trace. Ask for it before you conclude anything about correctness.
What does a completed scorecard look like?
An illustrative walkthrough. The scores below are a composite of patterns Alhena observed, not a single tested deployment.
A mid-market apparel brand decides to evaluate its incumbent AI agent ahead of renewal. One conversation, eleven turns, spanning discovery, a delivery complaint and a return request. Then a second visit the next morning.
Three yes answers. Question 5: asked “why this one?”, the agent explained the fit and fabric reasoning rather than restating the description. Question 6: it surfaced the review that addressed the shopper's stated doubt about sizing. Question 7: the tone held when the conversation turned to the late delivery, and the size constraint survived the shift.
Seven no answers, and they cluster. Questions 1, 2 and 3 all failed: the AI agent explained the returns process rather than starting the return, named a jacket without opening its page, and lost the thread when the shopper navigated away. That is an execution cluster, and it is the same missing layer three times. Question 8 failed the next morning, as it does almost everywhere. Questions 4, 9 and 10 failed too.
Total: 3 of 10. A capable assistant with a structural ceiling, which is where most of the market sits. The whole evaluation took under two hours including the return visit.
What the team did next matters more than the number. They did not replace the agent on the strength of one evaluation. They wrote success criteria for the three questions that mattered most to their own queue, asked the vendor for a trace of the failed return attempt, and re-ran the same test six weeks later. The trace showed the write was never attempted, which turned a vague quality complaint into a specific contractual question about scope. That is the point of scoring at all.
How do you turn this into a repeatable test?
A score you cannot reproduce is an opinion with a number attached, and an evaluation nobody can repeat is not an evaluation. Ten yes-or-no answers become useful the moment you can run them again in three months and compare like with like.
- Write each question as a test case. Not “can it complete an action” but “cancel order 48211 on this account, then confirm the cancellation appears in the order record.” Define the pass condition and its parameters before you start. Written success criteria are what stop a second evaluation from drifting into a different test.
- Capture what you saw, not what you concluded. One line per question: the input you gave, the output you got, and where execution stopped. That record is what lets you debug the difference when a question that passed in March fails in June.
- Score with two people the first time. A second evaluator will disagree on at least one question, and the disagreement shows you which pass conditions are still vague.
- Re-run the whole set as a regression check. Quarterly, and after any release. Agent behaviour is non-deterministic, so run the borderline questions twice before recording a yes.
- Track the score over time, not the absolute number. Capture each quarter's total so the trend is visible: moving from 3 to 5 tells you more about a vendor than either number alone, and it is the measure a renewal conversation actually turns on.
Ten questions, scored the same way each quarter, is a small enough loop that a support lead can own it and iterate on the pass conditions without a project plan. That is the whole validation exercise: a repeatable framework and a closed loop, not a one-off audit.
What makes an AI agent evaluation repeatable and trustworthy?
A useful AI agent eval is more than a scorecard. It is a repeatable evaluation framework with clear success criteria, representative behaviour and enough evidence to explain why an agent passed or failed.
For each test case, define the customer input, the expected outcome and what counts as proof. That proof may be a correct answer, but for an agentic workflow it may also be task completion: a return processed, an address changed or a specific product page opened. The key distinction is between a plausible output and a completed action.
THE EVALUATION LOOPTest the workflow, not just the prompt
A production evaluation should follow the full path from input to outcome. Start with the prompt and context, observe the reasoning and workflow, check any tool use or API execution, then validate the final result against the expected outcome. This is how you evaluate an AI agent end to end rather than judging a single response in isolation.
- Use representative test cases. Include ordinary requests, multi-step conversations and the edge cases your customers actually create.
- Define deterministic pass conditions where possible. AI output can be non-deterministic, but a completed cancellation or updated order record can still have a clear pass or fail.
- Capture the evidence. Record the input, output and execution result so another evaluator can reproduce the score.
- Use observability to investigate failures. A trace can show whether the agent selected the wrong workflow, never attempted a tool call, passed the wrong parameters or hit an API failure.
- Compare against ground truth. For real actions, the source of truth is what happened in the underlying system, not what the agent claimed happened.
This creates a practical evaluation pipeline: define → test → observe → measure → debug → iterate → regression test. The loop matters because AI agents change with prompts, models, integrations and production data.
OBSERVABILITY AND REGRESSIONWhy one passing demo is not enough
A single successful run does not prove reliable behaviour. Agent behaviour can vary across inputs, context and execution paths. That is why important workflows need regression testing after meaningful changes.
Use the same dataset of core test cases to evaluate critical behaviours over time. When something changes, compare the new result with the previous benchmark. If a workflow that once passed now fails, the score tells you that a regression happened. Observability and traces help explain where it happened.
For automated evals, an evaluator or model-based judge can help score relevance, coherence and correctness at scale. But automated evaluation should not replace human validation when the question involves judgement, emotional handling or safety. The strongest framework combines automated checks with observable production behaviour and ground truth.
That is the standard Alhena recommends for evaluating AI agents: measure what the system says, what it does and whether the intended outcome actually happened.
How do you run this properly?
- One continuous conversation, at least ten turns, covering discovery, then post-purchase, then an action, then an emotional moment.
- State a hard constraint early and test it at the end, so the workflow is evaluated under load rather than in isolation.
- Use your vertical's signature task, not a generic question. Evaluate the AI agent on the hardest thing your best salesperson does, because generic prompts produce generic passes.
- Come back the next day. Question 8 is where the field separates, and where most evaluations quietly stop.
- Observe, don't infer. If you did not watch it happen, mark it no. Judge the outcome, not the explanation of the outcome.
Key takeaways
- Most deployments score two or three yes answers out of ten when you evaluate them honestly.
- Question 8, cross-session memory, is the most discriminating, cleared by 1 of 15.
- Question 5, “why this one?”, is the fastest, exposing ranking dressed as reasoning.
- Six or more yes answers puts a deployment ahead of most of the market.
- Accuracy is deliberately excluded. At 2.47 field-wide, it no longer differentiates.
- Every no points at a layer: execution, reasoning, state or governance. Read the cluster, not the total.
- Write the questions as test cases so the score is reproducible next quarter rather than a one-off impression.
Frequently asked questions
Every question here is scored on behaviour you can observe as a shopper on your own storefront. The discipline required is not technical, it is refusing to mark anything as a pass that you did not personally watch happen.
Two or three out of ten is the common result when you evaluate honestly, which indicates an answer engine with a chat interface: competent at questions, unable to act, and unlikely to move conversion or resolution regardless of tuning.
Cross-session memory. Only 1 of 15 deployments cleared it in Alhena's 2026 study, making it both the rarest capability and the most discriminating single test in the whole evaluation. Completing a real action is next hardest at 4 of 15, and those two failures account for most of the capability gap.
Look at where the nos cluster. If they sit in questions 1 to 3, the AI agent cannot execute, and no prompt change grants write access to your order system: that is a platform decision. If they sit in 4 to 6, you may be able to improve reasoning with better product data before replacing anything. Re-run the evaluation after each change so you can attribute the movement rather than guess at it.
Weight by what your conversations are actually for. If most contacts are post-purchase, question 1 on task completion carries the most, because an AI agent that cannot act converts every one of those into a ticket. If you are optimising discovery, questions 4 and 5 matter more. Question 8 is the tiebreaker in any category with repeat purchase, and it was cleared by 1 of 15.
Different layer, and you want both. An engineering eval suite measures correctness against a dataset with known ground truth and catches regressions in the pipeline between builds. This checklist measures whether a shopper's job got done in production on your live storefront, which no offline dataset contains. Where the two meet is observability: when question 1 fails, a trace showing which tool was called and what output came back tells you whether an API errored or the agent never attempted the call at all. Score the behaviour here, debug the execution there. The evaluation the two produce together is far stronger than either alone.
Both, and the variance is itself the finding. Agent behaviour is non-deterministic, so the same input can take a different path on a different day. Run any borderline question twice before you complete the task of scoring it, and treat a question that passes intermittently as a no with a note attached. A capability that works most of the time is a capability your customers will meet failing.
Partially. Questions with a deterministic pass condition, such as whether a cancellation actually reached the order record, can be scripted end to end against your API, calling the tool directly and checking the output, then run as a regression test after every release. The ones about reasoning quality, restraint and emotional handling need a person, because the pass condition is a judgement rather than a value. Most teams automate three or four and keep the rest manual, which is enough to catch a regression in the workflow without pretending a script can judge tone.
Measure more than output quality. A strong AI agent evaluation should cover correctness, context, workflow execution, tool use, task completion and the final business outcome. Use representative test cases and clear success criteria, then compare the observed result with ground truth where a real action occurred. For production systems, observability and traces are what make failures diagnosable and regression testing is what tells you whether a previously working behaviour has broken.
Evaluate quarterly, and again after any vendor release claiming new capability. Scores drift when catalogues change, a prompt is edited or an integration breaks silently, and the return-visit memory check in particular tends to regress without anyone noticing.
Score against the ten, then place it beside the field benchmark: 15 of 15 agents could answer, 4 could act, 1 could remember. That framing moves the evaluation from vendor preference to capability gap. That figure comes from Alhena's 2026 Agentic CX Stress Test.
Most of it transfers, since an agent's completion, memory, restraint and escalation quality are channel-independent. The two edge cases needing adaptation are storefront navigation and in-chat selling, which have different equivalents in voice.
"Why this one?" after a product recommendation. A reasoning agent explains the causal link between the customer's problem and the product. A ranking system restates the description, which tells you most of what you need to know.
Because it no longer separates vendors. Accuracy averaged 2.47 out of 3.00 across the field and every serious platform answers competently. Agentic Capabilities at 1.60 is where products actually differ. Source: Alhena's 2026 stress test of 15 live deployments.
The two hardest rungs: whether it remembers a shopper on a return visit, and whether it completes actions rather than describing them. Six or more passes puts your agent ahead of most of the market, which makes the remaining gaps worth targeting.
Yes, and phrasing them as live demonstrations of the agent rather than capability questions is what makes them useful. "Show me on our storefront" produces different answers from "can your product do this".
Score your assistant against all ten
The diagnostic questions, the eight-dimension scoring sheet and the field results to benchmark against.