Every AI agent on every storefront can hold a conversation now. Almost none of them can finish a job. That gap is the whole story of 2026.
Agentic CX is customer experience delivered by AI that takes real actions (processing a return, changing an order, routing a shopper to the right product) rather than only answering questions. In Alhena's 2026 stress test of 15 live AI agents, all 15 could answer. Only 4 could act. Only 1 remembered the shopper on a return visit.
What did the 2026 stress test actually measure?
Alhena's research team opened 15 live storefronts belonging to real brands and behaved like ordinary shoppers. No sandboxes, no vendor demos, no scripted flows.
Each assistant was pushed through one continuous conversation covering product discovery, a post-purchase problem, an agentic action, and an emotional moment. Every verdict reflects observed behaviour. No capability was inferred from a marketing page.
| Capability | What it means | Deployments that demonstrated it |
|---|---|---|
| Answer | Understands the question and responds accurately | 15 of 15 |
| Sell | Sells inside the conversation | 9 of 15 |
| Act · in UI | Navigates the shopper to the right product | 6 of 15 |
| Act | Completes a return, cancellation or address change | 4 of 15 |
| Remember | Recalls the shopper on a return visit | 1 of 15 |
Why can every AI agent answer but almost none act?
Answering is a retrieval problem. Point a model at a help centre and a product catalogue and you get a competent answer engine, which is why the technology is now commoditised.
Acting is a systems problem. It needs write access to order systems, control of the storefront UI, an identity model, and a safe rollback path when something fails.
Vendors who bolted a language model onto a knowledge base got the first rung free. The fourth and fifth rungs have to be designed for.
What do the eight scoring dimensions reveal?
Field-wide averages across all 15 deployments, scored 1–3.
Notice the shape. The industry spent three years perfecting understanding and comparatively little time on execution.
What is the Agentic CX Maturity Model?
The report organises capability into six rungs. Each one depends on the rung beneath it: you cannot sell what you cannot recommend, and you cannot act on what you never understood.
The important finding is not that most vendors sit low on the ladder. It is that many have a structural ceiling they cannot climb past.
Which platform types hit a hard ceiling?
Personalisation engines
Plateau at Recommend. Excellent at matching, no concept of a task to complete.
AI search & discovery
Reach Sell, then route the action to a help centre or form.
Agentic assistants
The only archetype that reaches Act and Remember.
Why it matters
If your chat box is a ranking model in a conversational wrapper, no prompt tuning will make it process a return.
How did the 15 deployments score?
| Deployment | Vertical | Score (of 3.00) | Tier |
|---|---|---|---|
| Alhena | All 11 verticals | 3.00 | Elite |
| Platform A | Beauty / haircare | 2.75 | Upper-middle |
| Platform B | Jewellery | 2.63 | Upper-middle |
| Platform C | Fashion | 2.50 | Upper-middle |
| Platform D | Fashion-rental & mattress | 2.25 | Upper-middle |
| Platform E | Hair colour | 2.25 | Upper-middle |
| Platform F | Apparel | 2.25 | Upper-middle |
| Platform G | Bridesmaid | 1.88 | Lower-middle |
| Platform H | Mattress | 1.63 | Lower-middle |
| Platform I | Beauty & skincare | 1.50 | Lower-middle |
| Platform J | Jewellery | 1.25 | Lower-middle |
Competitors are anonymised because naming a brand identifies its vendor. Three further platforms were assessed by archetype rather than ranked, since their limits are structural.
The most instructive number is the jewellery spread: 2.63 versus 1.25. Same category, same shopper problem, more than double the performance.
How is agentic CX different from conversational AI?
Conversational AI describes the interface: natural language in, natural language out. Agentic CX describes the outcome: whether the conversation changes anything.
A conversational AI can be excellent and still resolve nothing. Every one of the 15 deployments tested was conversational. Four were agentic.
| Conversational AI | Agentic CX | |
|---|---|---|
| What it optimises | Understanding and response quality | Task completion and outcome |
| Typical metric | Containment, CSAT, deflection | Resolution, conversion, AOV |
| System access | Read-only retrieval | Authorised write access plus UI control |
| State | Session-scoped | Persistent across sessions |
| Field result 2026 | 15 of 15 deployments | 4 of 15 deployments |
What changed between 2025 and 2026?
In 2025 the live question was whether AI could understand a shopper well enough to be trusted in front of customers. That question is now closed, Context scored a perfect 3.00 across every deployment tested.
The 2026 question is different: can it finish the job? Selling, acting and remembering are what separate deployments now, and they are the three dimensions where the field scores worst.
This is why buying criteria written in 2025 mislead in 2026. A shortlist built around accuracy and tone will rank fifteen platforms that are all effectively tied.
What does “good” look like in 2026?
Six criteria that work directly as RFP language.
- Completes the task: not “here’s how to start a return,” but the return, started.
- Sells inside the chat: product cards, comparisons and add-to-cart in-thread.
- Moves the storefront: conversation and commerce are one surface.
- Remembers responsibly: recalls the constraint that matters, transparently.
- Has vertical expertise: reasons about undertones, arch support, motion transfer, ingredient conflicts.
- Knows its limits: declines what it should and escalates with full context.
Phase one of AI in customer experience was about comprehension, and it is over. Every vendor passed. Phase two is action, selling, memory and vertical depth, and four out of fifteen platforms have started it.
Key takeaways
- Answering is no longer a differentiator. All 15 live agents tested could answer; Context scored a perfect 3.00 field-wide.
- Only 4 of 15 could complete a real action such as a return, cancellation or address change.
- Only 1 of 15 remembered a shopper across sessions. Memory is the genuine frontier.
- Ceilings are structural, not roadmap items. Personalisation engines stop at Recommend; search engines stop at Sell.
- Alhena scored 3.00 across all 11 verticals, the only deployment in the elite tier.
Frequently asked questions
A chatbot follows scripted decision trees. An AI agent uses an LLM to interpret intent in context, then acts on connected systems: processing the return rather than describing it. The difference is execution, not language quality. Every one of the 15 deployments in Alhena's 2026 stress test was conversational. Four could take action.
Generative AI delivered the language layer, and that layer is now commoditised. Every platform tested in 2026 scored a perfect 3.00 on context. Agentic CX is the shift that follows it: intelligence wired into your order systems and storefront, so customer support becomes something the agent executes rather than explains.
Agentic commerce usually describes AI agents transacting across the wider buying journey, including agent-to-agent purchasing. Agentic CX is the layer on your own storefront, where the agent sells, resolves and remembers for your shoppers. Most brands need the second before the first matters.
It applies directly, and smaller catalogues often make it easier to get right, since vertical reasoning is simpler to tune. The requirements are identical at any size and independent of enterprise scale: the agent needs write access to your order system and control of the storefront.
Yes, within the categories it has authorised access to. Returns, cancellations, address changes and order status can be resolved end to end without a live agent. Alhena's 2026 testing found only 4 of 15 deployments could do this. The other 11 explained the process and handed the work back to the shopper. Resolution rate, not containment, is the metric that reflects this.
It should recognise the limit early and escalate with the full conversation attached, so the human agent opens with context instead of a transcript. A clean handoff removes the friction of the shopper repeating themselves. Weak escalation was one of the seven failure modes Alhena documented in 2026, and it is usually a symptom of an agent that could not act in the first place.
Enough to complete bounded, reversible tasks autonomously, with guardrails on anything financial or irreversible. An intelligent agent declines confidently and escalates cleanly. Genuinely autonomous AI is defined as much by what it refuses as by what it completes, and over-claiming scope does more reputational damage than a narrow scope honestly described.
Yes: flagging a delayed shipment before the shopper asks, or surfacing a fit constraint mid-browse. Proactive behaviour rests on the same foundations as action, namely system access and persistent context. Almost no deployment tested in 2026 attempted it, which makes it a realistic differentiator rather than a solved feature.
Ask it to cancel a real order on your live storefront, then say "take me to it" after a recommendation, then return the next day and see whether anything is remembered. Alhena's 2026 methodology rested on one rule: infer no capability from marketing. Every verdict in the report came from observed behaviour on a live deployment.
Alhena scored 3.00 out of 3.00 across all 11 verticals in its 2026 Agentic CX Stress Test, the only deployment in the elite tier. The strongest anonymised competitor scored 2.75 and the weakest scored 1.25.
Roughly a quarter. Four of the 15 live deployments Alhena tested in 2026 could complete a return, cancellation or address change end to end. The other 11 explained the process and handed the work back to the customer.
Four things: authorised write access to the order system, an identity model, control of the storefront UI, and a safe rollback path. Without them the agent can orchestrate nothing. It can only describe a workflow and hand it back to the customer. This is the structural reason personalisation engines and search tools plateau below Act.
It is rarely a quarter, because it is not a tuning exercise. Acting requires authorised write access, an identity model, storefront UI control and a safe rollback path. If that AI architecture was not designed in, the change is closer to a rebuild than a release.
Start by automating the high-volume post-purchase workflows: returns, exchanges, cancellations and address changes. They are repetitive, rule-bound and safe to hand to automation, and they are the interactions where returning work to the customer costs you the most. Discovery and recommendation automation compounds faster once execution is already proven.
Track resolution rate, conversion and AOV on assisted sessions rather than containment or deflection. Containment counts conversations that ended. Resolution counts problems that were solved. A self service portal can post excellent containment while every shopper leaves the task unfinished, which is precisely the gap the 2026 data exposes.
Only if it holds state. Real time context within a single session is now standard across the field. Carrying a shopper's constraint across sessions and touchpoints is not: one deployment of 15 remembered anything on a return visit. Memory is what turns a good interaction into a relationship.
Execution. Three of the seven failure modes Alhena documented in 2026 share one root cause: the agent cannot act on your systems or move your storefront. Fixing that also improves escalation quality, because fewer conversations fail in the first place.
Get the full 2026 scorecard
Eleven verticals, fifteen live deployments, eight dimensions each, plus the seven failure modes and the ten questions to ask your vendor.