Most benchmark tasks ask an agent to find something. These four ask it to work something out first. That is a different capability entirely.
Four verticals in Alhena's 2026 stress test required reasoning rather than retrieval: translating a comfort complaint into a support profile, a room photo into a spatial plan, a trip into a packing list, and a beginner's ambition into a budget kit. UI disconnect (6 of 15) and catalog dumping (5 of 15) did the most damage.
Can AI recommend shoes for knee pain and high arches?
The task requires translating a symptom into a spec. The shopper has not described a product. They have described a body and a job.
The agent must get from “sore knees, high arch, twelve-hour shifts” to cushioning depth, arch support structure, heel drop, midsole density and probably a wider toe box.
Where it broke: the reasoning was often fine and the last step failed. UI disconnect appeared in 6 of 15 deployments, dead-end recommendation in another 6 of 15. The shopper lands on a category page with sixty models and no memory of which was recommended.
Can AI design a room from a photo?
This requires spatial reasoning under a hard constraint. “Rented” eliminates most standard advice about wall colour, built-ins and light fixtures.
The agent must reason about scale, light, sightlines and permanence, mirrors and layered lamps, because the ceiling fixture is fixed.
Where it broke: catalog dumping, 5 of 15. A shopper uploads their bedroom and receives the bestselling floor lamp. The photo was accepted and ignored.
Can AI plan a trip and recommend the products?
Six days in Portugal in October, carry-on only. That is a chain of inferences before any product appears: mild days, cool evenings, real rain risk, a laundry assumption, volume and weight limits, the liquids rule.
Only then does a packable rain layer, versatile footwear and a compliant bag become relevant.
Where it broke: agents that treated it as a product query produced a generic list. The ones that reasoned first produced a packing plan with products attached: more useful, and a much larger basket.
Can AI kit a beginner without over-selling?
The expert behaviour here is subtraction. Trail shoes and decent socks matter; the hydration vest, poles, GPS watch and gaiters can wait.
A beginner over-kitted at the start spends their whole budget on things they will not use and will not come back.
What pattern connects all four?
| Vertical | The reasoning step | Dominant failure |
|---|---|---|
| Footwear | Symptom → support profile | UI disconnect / dead-end recommendation (6 of 15 each) |
| Home | Photo → spatial plan under a rental constraint | Catalog dumping (5 of 15) |
| Travel | Trip → packing logic → products | Product listing without the reasoning chain |
| Sporting goods | Ambition + budget → minimum viable kit | Over-kitting the novice |
- Retrieval is not reasoning. Every one of these can be answered badly by a very good search engine. The gap only shows when you ask “why this one?”
- The last mile is where value leaks. Only 6 of 15 could navigate a shopper to the right product, and Agentic Capabilities scored 1.6 of 3.0: the lowest dimension in the study.
Excellent reasoning that ends in a link is a conversation, not a sale.
Key takeaways
- Four verticals demanded an inference chain before any product became relevant.
- Footwear failed at the last step: good reasoning, then no route to the product.
- Home failed at the first: the uploaded photo was accepted and ignored.
- Sporting goods rewards subtraction. Knowing what to leave out is the expert move.
- “Why this one?” is the diagnostic. Ranking models cannot answer it.
Frequently asked questions
That is exactly the capability gap. Getting from "sore knees, twelve-hour shifts" to a support profile of cushioning, arch structure and heel drop is reasoning rather than retrieval. A very good search engine answers these badly, which is why the difference only shows under pressure.
The image is being accepted and ignored, then popularity ranking fills the gap. Alhena logged this as catalog dumping in 5 of 15 deployments. A ranking model has nowhere to put a photo or a constraint like renting.
The best deployments reason through season, duration and luggage limits first, then attach products to the resulting plan. Agents that skip the inference chain produce a generic list, which is both less useful and a much smaller basket.
The expert move is subtraction: name the few items that matter and explain what to skip and when it will be worth buying. Recommending everything is what a catalogue does, and it spends a beginner's budget on things they will not use.
Usually the last step. Alhena found 6 of 15 agents could navigate a shopper to the right product, so reasoning was often sound and the handoff to purchase was not. Excellent advice that ends in a link is a conversation, not a sale.
Stronger agents translate the complaint into a support profile and explain which feature addresses which symptom. The common failure is naming a model and leaving the shopper to find it among sixty options on a category page.
Give it a constraint that is not a product facet, like renting or a twelve-hour shift, then ask "why this one?" A reasoning system explains the causal link. A filtering system restates the product description or cites a star rating.
Not different technology, but the same architectural requirements show up faster. Alhena tested footwear, home, travel and sporting goods alongside eight other verticals, and performance tracked platform architecture rather than category difficulty.
That is a reasoning chain rather than a filter: destination and month imply temperature range and rain risk, which imply layers, which imply products. Agents that jump straight to products skip the constraints that make the recommendation usable.
"Why this one?" It is the fastest diagnostic in the whole benchmark, because ranking models cannot answer it. Everything else about reasoning quality follows from whether the agent can explain the link between the problem and the product.
All eleven verticals, one report
Full signature tasks, transcripts and eight-dimension scorecards for every category tested.