A generic benchmark tells you how well an assistant handles generic questions. That is not the job. Every category has one hard thing its best salesperson does.
AI shopping agents do not perform evenly across ecommerce categories. In Alhena's 2026 stress test of 15 live deployments across 11 verticals, scores ranged from 1.25 to 3.00 on an eight-dimension scale, and the gap inside a single vertical (jewellery: 1.25 vs 2.63) was wider than the gap between most categories.
Why test by vertical rather than by intent?
Every ecommerce category has its own hard case. A shopper buying foundation needs undertone reasoning. A shopper buying a mattress needs four competing sleep variables reconciled. A shopper asking about a supplement needs the agent to stop talking at exactly the right moment.
So each vertical got a signature task, the hard thing that category's best human salesperson does, run live on a real brand's storefront.
Most AI-powered shopping tools get benchmarked on generic questions, and generic questions are the part everyone passes. A signature task measures something harder: whether the assistant can run product discovery through a deep catalogue the way a specialist would, rather than answer questions about your returns policy.
What were the benchmark results?
Every deployment was scored on the same eight dimensions, live on a retailer's own storefront, with the real product catalogue and inventory behind it.
| Deployment | Vertical | Score (of 3.00) | Tier |
|---|---|---|---|
| Alhena | All 11 verticals | 3.00 | Elite |
| Platform A | Beauty / haircare | 2.75 | Upper-middle |
| Platform B | Jewellery | 2.63 | Upper-middle |
| Platform C | Fashion | 2.50 | Upper-middle |
| Platform D | Fashion-rental & mattress | 2.25 | Upper-middle |
| Platform E | Hair colour | 2.25 | Upper-middle |
| Platform F | Apparel | 2.25 | Upper-middle |
| Platform G | Bridesmaid | 1.88 | Lower-middle |
| Platform H | Mattress | 1.63 | Lower-middle |
| Platform I | Beauty & skincare | 1.50 | Lower-middle |
| Platform J | Jewellery | 1.25 | Lower-middle |
What is each vertical's signature task?
Each one is something a good sales assistant does without thinking, and not one of them can be expressed as a filter.
Beauty
Shade match from a selfie. Nothing is returned faster than the wrong foundation.
Skincare
AM/PM routine, on budget, fragrance-free. Memory reset is the killer.
Fashion
Full outfit for a pear-shaped 5′4″ shopper, no heels, plus a size call from a photo.
Footwear
“Easy on my knees, higher arch, on my feet all day” becomes a support profile.
Home
Make a small, dark, rented room feel larger, from a photo.
Jewellery
Anniversary gift, gold, hypoallergenic, plus matching earrings. Gifting is mostly memory.
Mattress
Hot-sleeping side sleeper, back pain, heavier restless partner.
Sporting goods
Kit a beginner trail runner to a budget. Knowing what to leave out is the skill.
Hair colour
Go lighter on bleached curly hair, safely. Often the answer is “not yet.”
Travel
Six days in Portugal in October, carry-on only.
Supplements
The highest-stakes vertical, full stop. Restraint is the skill.
The pattern
Difficulty did not predict failure. Architecture did.
Which three patterns held across every category?
- Architecture predicted performance, not difficulty. Scores collapsed wherever the platform was a search or personalisation layer, regardless of how hard the vertical was. Those systems are optimised for ranking, not for finishing a task.
- Vertical expertise is a memory problem as much as a knowledge problem. Knowing fragrance irritates sensitive skin is retrievable. Still knowing it four turns later is architectural. Memory reset appeared in 5 of 15.
- The last step is where the money is. 6 of 15 named a product and then failed to route the shopper to it. No route means no cart, and no cart means no checkout.
What should a brand in these categories do?
Ask your vendor to run your signature task, live, on your storefront, not a scripted demo.
Run it across the whole shopping journey, from the first browse to the cart. Most demos are optimised to stop at the recommendation, which is the step immediately before the one that fails, and automation that never routes a purchase is hard to tell apart from a pleasant chat.
- Did it reason like a specialist, or fall back on bestsellers?
- Did it hold every stated constraint to the final recommendation?
- Did it route the shopper to the product, or complete the task?
- Did it tailor the answer to the constraints stated, or to whatever the catalogue data made easy?
If it cleared all four, you are in the top tier of a field where most are not, whatever tool sits behind the widget.
Key takeaways
- Scores ranged 1.25 to 3.00 across 11 verticals and 15 live deployments.
- The widest gap was inside a single category: jewellery, 2.63 against 1.25.
- Alhena scored 3.00 across all eleven verticals, the only elite-tier deployment.
- Category difficulty did not predict failure. Platform architecture did.
- Vertical expertise depends on memory, not just product knowledge.
- There is no single best AI shopping agent for every category. Test the hardest thing your best sales assistant does, live on your own storefront.
Frequently asked questions
Alhena tested 11 verticals in 2026 including mattresses, jewellery, travel, sporting goods and supplements. Performance varied more by platform than by category, and the widest gap appeared inside a single vertical rather than between them.
There is no single answer, because the right choice depends on the rung you need: discovery for findability, personalisation for merchandising, agentic for resolution and retention. On observed capability, Alhena scored 3.00 out of 3.00 across all 11 verticals in its 2026 stress test, the only elite-tier result, while the anonymised competitor set ranged from 2.75 down to 1.25. Category labels are unreliable too, since a personalisation engine, an AI search layer and an older chatbot can all present as the same widget. The more useful test for any ecommerce retailer is your own signature task, run live on your storefront rather than in a vendor demo.
Filters express the attributes your product catalogue already has fields for. The signature tasks in this benchmark involved constraints no filter can hold: undertone, arch support, motion transfer, hypoallergenic metal. A recommendation engine ranks by behavioural similarity, so personalisation improves what a shopper sees while they browse. An AI assistant reconciles conflicting preferences out loud and then acts on the conclusion. The three are complementary, and Alhena found the failures cluster in the last step rather than the first.
That is a different layer, and both matter. External agents like ChatGPT and Perplexity, along with the assistants inside Amazon and Alexa, shape product discovery before a shopper ever reaches you, and the work there is making your catalogue and content legible to them. Agentic CX is what happens once they arrive: whether the assistant on your own storefront can reason about their constraints, sell, and resolve. Alhena's 2026 benchmark tested the second, live on real brands' storefronts, because that is the part a retailer controls.
They answer different questions. Virtual try-ons show a shopper how a chosen product looks on them. The AI assistant's job is the choosing, which is why the beauty signature task in this benchmark was a shade match from a selfie rather than a render. Try-on reduces uncertainty after a decision. Conversational reasoning reduces the number of wrong decisions in the first place, and the two work well together.
Early on it replaces aimless browsing with a stated intent, so product discovery starts from constraints rather than categories. In the middle it holds those constraints while the shopper wanders. At the end it should move the storefront and add to cart in real time rather than describing where to click, which is where most deployments stopped: 6 of 15 named a product in Alhena's 2026 test and then failed to route the shopper to it. Nothing earlier in the journey matters if the last step before checkout breaks.
High-consideration categories are the hardest test and often the highest value, because long conversations expose memory and reasoning failures. The capability that matters most is reconciling conflicting constraints out loud rather than filtering a catalogue.
Define your signature task, the hardest thing your best salesperson does. Usually it involves constraints a filter cannot express, like undertone, arch support or motion transfer. Then make the agent do exactly that, live on your storefront, before signing anything.
Jewellery. The two deployments Alhena tested scored 2.63 and 1.25 out of 3.00, more than double the performance on the same shopper task. That spread was wider than the gap between most entire categories.
Look for evidence of consistency rather than a single strong category. Alhena scored 3.00 across all 11 verticals tested in 2026, while competitor deployments were typically strong in one product world and untested elsewhere.
The same capability ladder applies and the agentic rungs matter more, because reorders, account pricing and approval steps are all actions rather than answers. Memory also carries more weight, since B2B buying relationships run across months.
The return is smaller, because the purchase is simple and the shopper needed little help to begin with. Value concentrates where conversations are long, constraints conflict, or post-purchase support is heavy. Low-consideration catalogues gain more from discovery improvements than from conversational depth.
The strongest deployments did, by exercising restraint: building a sensible regimen while declining medication interaction questions. Alhena logged unsafe confidence, meaning overclaiming in risky categories, in 3 of 15 deployments, which is the exposure to test for.
Eleven: beauty, skincare, fashion, footwear, home, jewellery, mattress and sleep, sporting goods, hair colour, travel, and health and supplements. Each had its own signature task, run live on a real brand's storefront rather than in a demo.
Platform architecture. Scores collapsed wherever the underlying product was a search or personalisation layer rather than an agentic assistant, regardless of how hard the vertical was. The hardest categories were not where performance fell apart.
Every vertical's full scorecard
Eleven verticals, fifteen deployments, eight dimensions each, plus the signature-task transcripts.