Can AI Recommend Shoes for Knee Pain or Design a Room From a Photo?

Can AI Recommend Shoes for Knee Pain or Design a Room From a Photo?
AI reasoning connects footwear, home, travel, and outdoor product recommendations.

Most benchmark tasks ask an agent to find something. These four ask it to work something out first. That is a different capability entirely.

Short answer

Four verticals in Alhena's 2026 stress test required reasoning rather than retrieval: translating a comfort complaint into a support profile, a room photo into a spatial plan, a trip into a packing list, and a beginner's ambition into a budget kit. UI disconnect (6 of 15) and catalog dumping (5 of 15) did the most damage.

Can AI recommend shoes for knee pain and high arches?

The task requires translating a symptom into a spec. The shopper has not described a product. They have described a body and a job.

The agent must get from “sore knees, high arch, twelve-hour shifts” to cushioning depth, arch support structure, heel drop, midsole density and probably a wider toe box.

Where it broke: the reasoning was often fine and the last step failed. UI disconnect appeared in 6 of 15 deployments, dead-end recommendation in another 6 of 15. The shopper lands on a category page with sixty models and no memory of which was recommended.

Can AI design a room from a photo?

This requires spatial reasoning under a hard constraint. “Rented” eliminates most standard advice about wall colour, built-ins and light fixtures.

The agent must reason about scale, light, sightlines and permanence, mirrors and layered lamps, because the ceiling fixture is fixed.

Where it broke: catalog dumping, 5 of 15. A shopper uploads their bedroom and receives the bestselling floor lamp. The photo was accepted and ignored.

Can AI plan a trip and recommend the products?

Six days in Portugal in October, carry-on only. That is a chain of inferences before any product appears: mild days, cool evenings, real rain risk, a laundry assumption, volume and weight limits, the liquids rule.

Only then does a packable rain layer, versatile footwear and a compliant bag become relevant.

Where it broke: agents that treated it as a product query produced a generic list. The ones that reasoned first produced a packing plan with products attached: more useful, and a much larger basket.

Can AI kit a beginner without over-selling?

The expert behaviour here is subtraction. Trail shoes and decent socks matter; the hydration vest, poles, GPS watch and gaiters can wait.

A beginner over-kitted at the start spends their whole budget on things they will not use and will not come back.

Recommending everything is what a catalogue does. Recommending three things and explaining why not the other six is what a specialist does.

What pattern connects all four?

VerticalThe reasoning stepDominant failure
FootwearSymptom → support profileUI disconnect / dead-end recommendation (6 of 15 each)
HomePhoto → spatial plan under a rental constraintCatalog dumping (5 of 15)
TravelTrip → packing logic → productsProduct listing without the reasoning chain
Sporting goodsAmbition + budget → minimum viable kitOver-kitting the novice
  • Retrieval is not reasoning. Every one of these can be answered badly by a very good search engine. The gap only shows when you ask “why this one?”
  • The last mile is where value leaks. Only 6 of 15 could navigate a shopper to the right product, and Agentic Capabilities scored 1.6 of 3.0: the lowest dimension in the study.

Excellent reasoning that ends in a link is a conversation, not a sale.

Key takeaways

  • Four verticals demanded an inference chain before any product became relevant.
  • Footwear failed at the last step: good reasoning, then no route to the product.
  • Home failed at the first: the uploaded photo was accepted and ignored.
  • Sporting goods rewards subtraction. Knowing what to leave out is the expert move.
  • “Why this one?” is the diagnostic. Ranking models cannot answer it.

Frequently asked questions

Our customers describe problems, not products. Can AI translate that into a recommendation?

That is exactly the capability gap. Getting from "sore knees, twelve-hour shifts" to a support profile of cushioning, arch structure and heel drop is reasoning rather than retrieval. A very good search engine answers these badly, which is why the difference only shows under pressure.

Customers upload photos of their room and get generic furniture back. What's happening?

The image is being accepted and ignored, then popularity ranking fills the gap. Alhena logged this as catalog dumping in 5 of 15 deployments. A ranking model has nowhere to put a photo or a constraint like renting.

Can an AI plan a trip and attach the right products, or does it just list gear?

The best deployments reason through season, duration and luggage limits first, then attach products to the resulting plan. Agents that skip the inference chain produce a generic list, which is both less useful and a much smaller basket.

Our AI over-recommends to beginners and they bounce. How do we fix that?

The expert move is subtraction: name the few items that matter and explain what to skip and when it will be worth buying. Recommending everything is what a catalogue does, and it spends a beginner's budget on things they will not use.

Why does good AI reasoning still fail to convert on our site?

Usually the last step. Alhena found 6 of 15 agents could navigate a shopper to the right product, so reasoning was often sound and the handoff to purchase was not. Excellent advice that ends in a link is a conversation, not a sale.

Can AI recommend footwear for a specific physical complaint like overpronation?

Stronger agents translate the complaint into a support profile and explain which feature addresses which symptom. The common failure is naming a model and leaving the shopper to find it among sixty options on a category page.

How do I test whether an agent reasons or just filters our catalogue?

Give it a constraint that is not a product facet, like renting or a twelve-hour shift, then ask "why this one?" A reasoning system explains the causal link. A filtering system restates the product description or cites a star rating.

Do these reasoning-heavy categories need different AI than fashion or beauty?

Not different technology, but the same architectural requirements show up faster. Alhena tested footwear, home, travel and sporting goods alongside eight other verticals, and performance tracked platform architecture rather than category difficulty.

Can AI account for climate and season when recommending clothing?

That is a reasoning chain rather than a filter: destination and month imply temperature range and rain risk, which imply layers, which imply products. Agents that jump straight to products skip the constraints that make the recommendation usable.

What single question exposes a weak AI shopping agent fastest?

"Why this one?" It is the fastest diagnostic in the whole benchmark, because ranking models cannot answer it. Everything else about reasoning quality follows from whether the agent can explain the link between the problem and the product.

All eleven verticals, one report

Full signature tasks, transcripts and eight-dimension scorecards for every category tested.

Power Up Your Store with Revenue-Driven AI