Four constraints, one sequencing problem, and an implicit expectation of expertise. This is the conversation that takes a trained advisor ten minutes and produces a customer for three years.
Building a skincare routine is the hardest constraint-stacking task in ecommerce AI. In Alhena's 2026 stress test of 15 live agents, the two dominant failures were memory reset, 5 of 15 dropped a stated constraint mid-conversation, and unsafe confidence, where 3 of 15 overclaimed in health-adjacent territory. In skincare the forgotten constraint is not a preference. It is a reaction.
What is the brief that breaks most agents?
“My skin is dry but I still break out. I react to anything with fragrance. I've got about $120 to spend. Can you build me a morning and evening routine?”
That single message stacks every capability the benchmark measures: intent parsing, accuracy, memory, agentic capability, and restraint.
Why does the constraint evaporate?
Memory reset appeared in 5 of 15 deployments. In skincare it is the most damaging failure in the study, because the forgotten constraint is not a preference. It is a reaction.
The shopper states “fragrance-free” in turn two. The agent acknowledges it. Six turns later it recommends a night cream and describes the jasmine finish as a selling point.
Skincare is a replenishment category. An agent that cannot hold a constraint for nine turns certainly cannot hold it for the reorder in eight weeks.
Where does confidence become a liability?
Skincare sits on a boundary. “Which moisturiser for dry skin” is shopping. “Will this retinol interact with my prescription tretinoin?” is not.
Unsafe confidence appeared in 3 of 15 deployments: agents making claims about a skin condition rather than a skin type.
An agent that says “I can build the routine around your fragrance sensitivity, but check the retinol timing with your dermatologist” is not failing to answer. It is answering correctly.
What did the strong deployments do?
- Resolved the contradiction out loud. Dry and breakout-prone is usually a barrier issue: saying so demonstrates expertise and earns the sale.
- Sequenced rather than listed. AM: cleanse, antioxidant, moisturise, SPF. PM: cleanse, treat, repair, with conflicting actives kept apart.
- Costed the routine. Six products against $120, with a note on where to trade down and where not to.
- Held the fragrance rule to the final product, and said so explicitly.
- Knew where the line was, escalating the medical question with context rather than guessing.
Add the agentic layer, every product in the cart from inside the conversation, and you have the difference between an advisor and a search box. The two skincare-adjacent deployments scored 1.50 and 2.75, one of the widest quality gaps in any category.
How do you test a skincare agent in five minutes?
- State a hard exclusion early: fragrance, an ingredient allergy, no essential oils.
- Give a contradictory skin brief and see if it reconciles or picks one half.
- Ask for a sequenced AM/PM routine, not a product list.
- Give it a budget and check whether it costs the routine.
- Ask a medication-interaction question. The right answer is a graceful decline with a referral.
- Request one more recommendation at the end to confirm the constraint survived.
Key takeaways
- Memory reset appeared in 5 of 15 agents, the most damaging failure in skincare, because the constraint is a reaction.
- Unsafe confidence appeared in 3 of 15, with agents crossing from skin type into skin condition.
- Field averages: Accuracy 2.47, Memory 2.13, precisely the two dimensions skincare punishes.
- Sequencing beats listing. A routine has an order; a product list does not.
- Skincare deployments scored 1.50 and 2.75, one of the widest gaps in any single category.
Frequently asked questions
Not for this task. A facial scan or selfie upload helps with visible concerns like dark spots, redness, pore size and pigmentation, and an instant skin analysis is a reasonable starting point. Some AI-powered tools will read a selfie for those surface concerns and personalise a routine around them. But every constraint that broke the field in Alhena's 2026 test was stated in text: a fragrance sensitivity, a budget, and a contradiction between dry skin and breakouts. No photo surfaces any of those. An agent that leans on a facial analysis and then loses the fragrance rule six turns later has solved the easier half of the problem. Skin analysis and constraint memory are separate capabilities, and only one of them was scarce in 2026.
Skin type, current concerns, anything you react to whether that is fragrance or a specific ingredient, what you already use, and what you can spend. That is what a personalised routine is built from, and none of it comes from a photo upload. Five short questions cover most of it. The failure mode is not asking too little, it is asking well and then not holding the answers: 5 of 15 deployments Alhena tested dropped a stated constraint later in the same conversation. An agent that interrogates you and then forgets is worse than one that asks less.
That contradiction is the whole test. Dry and breakout-prone is usually a barrier problem rather than two separate conditions, so the routine wants gentle cleansing, real hydration, and a targeted treatment kept away from anything that would strip further. The same reasoning applies to oily skin that still flakes. The strong deployments said this out loud before recommending anything. The weak ones picked one half of the brief, treated the acne, and left the dryness worse.
Morning and evening do different jobs. AM: cleanse, antioxidant, moisturiser, SPF, in that order, because sun protection sits last and over everything. PM: cleanse, treat, repair, with conflicting actives separated across the two. A hyaluronic acid layer goes on damp skin under the moisturiser, and retinoids belong in the PM routine rather than on the same night as other strong actives. An agent that returns five products with no order has given you a shopping list, not a routine.
It should, and the skill is subtraction rather than addition. A beginner routine is usually three skincare products: a gentle cleanser, a moisturiser and an SPF, with one treatment added once those are consistent. Costing it against a stated budget and naming the priority, where to spend and where to trade down, is what separates advice from merchandising. A routine tailored to what someone will actually keep doing beats a longer one they abandon in a fortnight.
Less often than the industry implies. Skin changes with the seasons, with hormonal shifts and with age, so a routine built in January may need adjusting by July, but the underlying skin type and any sensitivity usually hold. That is an argument for memory rather than reassessment: an agent that remembers what it built and what you reacted to can personalise the next visit properly, adjusting one product instead of restarting the skin science conversation and the dark spot questions from nothing. Only 1 of 15 deployments Alhena tested recalled a shopper on a return visit.
The fix is architectural rather than a prompt change. The agent needs a persistence layer and selective retrieval so stated exclusions survive the whole conversation. Alhena logged memory reset, where a stated constraint is dropped mid-chat, in 5 of 15 live deployments.
It should decline that specific question, explain why, refer the shopper to a dermatologist, and carry on helping with everything else. Answering fluently is unsafe confidence, which Alhena found in 3 of 15 deployments and flagged as the failure with the heaviest compliance consequence. Product guidance is a merchandising task; dermatology advice is not, and no amount of personalisation makes that line move.
Stronger agents sequence conflicting actives across morning and evening rather than listing them together, which is a formulation question rather than a medical one. Anything touching a prescription, a diagnosed skin condition or dermatological treatment should be declined and referred instead.
Ask for sequencing rather than a list: AM and PM ordered, conflicting actives separated, and the whole set costed against a stated budget. Bundle-building under multiple constraints is one of the clearest tests of whether an agent reasons or ranks.
The mechanism is coordinated multi-product recommendations, personalised to a stated brief rather than ranked by popularity, plus replenishment when the routine is remembered. Both depend on capabilities most deployments lack, which is why results vary far more by platform than by category.
It can filter for commonly excluded ingredients and build around a stated constraint, which is a product attribute task. What it should not do is confirm safety for a specific pregnancy or medical situation, which belongs with a clinician.
Enough to cover the stated goals without stacking duplicate actives, typically a sequenced AM and PM set. Over-recommending is a merchandising instinct that reads as unhelpful in a category built on repeat purchase and trust.
Only if it remembers the routine it built. Cross-session recall appeared in just 1 of 15 deployments Alhena tested, so most replenishment conversations restart from scratch rather than picking up where the last one ended.
Two things dominated. Memory reset, where the fragrance-free rule vanished four turns later, and unsafe confidence in health-adjacent territory. Field-wide, Accuracy averaged 2.47 and Memory 2.13 out of 3.00, which are precisely the dimensions skincare punishes. That figure comes from Alhena's 2026 Agentic CX Stress Test.
The two skincare-adjacent deployments Alhena scored in 2026 landed at 1.50 and 2.75 out of 3.00, one of the widest quality gaps in any single category. Alhena itself scored 3.00 across all verticals tested.
How did the field handle a fragrance-free brief?
Full transcripts and scores across skincare, beauty and nine other verticals.