Can AI Build a Fragrance-Free Skincare Routine on Budget? 15 Agents Tested

Can AI Build a Fragrance-Free Skincare Routine on Budget? 15 Agents Tested
AI-powered skincare routine visualization with personalized product recommendations and data-driven decision paths

Four constraints, one sequencing problem, and an implicit expectation of expertise. This is the conversation that takes a trained advisor ten minutes and produces a customer for three years.

Short answer

Building a skincare routine is the hardest constraint-stacking task in ecommerce AI. In Alhena's 2026 stress test of 15 live agents, the two dominant failures were memory reset, 5 of 15 dropped a stated constraint mid-conversation, and unsafe confidence, where 3 of 15 overclaimed in health-adjacent territory. In skincare the forgotten constraint is not a preference. It is a reaction.

What is the brief that breaks most agents?

“My skin is dry but I still break out. I react to anything with fragrance. I've got about $120 to spend. Can you build me a morning and evening routine?”

That single message stacks every capability the benchmark measures: intent parsing, accuracy, memory, agentic capability, and restraint.

Why does the constraint evaporate?

Memory reset appeared in 5 of 15 deployments. In skincare it is the most damaging failure in the study, because the forgotten constraint is not a preference. It is a reaction.

The shopper states “fragrance-free” in turn two. The agent acknowledges it. Six turns later it recommends a night cream and describes the jasmine finish as a selling point.

WHERE SKINCARE PUNISHES THE FIELDField-wide dimension averages, out of 3.0Context3.00Intent Parsing2.67Accuracy2.47Memory2.13Agentic Capabilities1.60GOLD = THE TWO DIMENSIONS A SKINCARE BRIEF PUNISHES HARDESTAlhena Agentic CX Stress Test 2026 · bars scaled from 1.0, the floor of the scale
Understanding the brief was never the problem. Staying accurate across a long, multi-constraint conversation was.

Skincare is a replenishment category. An agent that cannot hold a constraint for nine turns certainly cannot hold it for the reorder in eight weeks.

Where does confidence become a liability?

Skincare sits on a boundary. “Which moisturiser for dry skin” is shopping. “Will this retinol interact with my prescription tretinoin?” is not.

Unsafe confidence appeared in 3 of 15 deployments: agents making claims about a skin condition rather than a skin type.

Restraint is part of good CX.

An agent that says “I can build the routine around your fragrance sensitivity, but check the retinol timing with your dermatologist” is not failing to answer. It is answering correctly.

What did the strong deployments do?

  • Resolved the contradiction out loud. Dry and breakout-prone is usually a barrier issue: saying so demonstrates expertise and earns the sale.
  • Sequenced rather than listed. AM: cleanse, antioxidant, moisturise, SPF. PM: cleanse, treat, repair, with conflicting actives kept apart.
  • Costed the routine. Six products against $120, with a note on where to trade down and where not to.
  • Held the fragrance rule to the final product, and said so explicitly.
  • Knew where the line was, escalating the medical question with context rather than guessing.

Add the agentic layer, every product in the cart from inside the conversation, and you have the difference between an advisor and a search box. The two skincare-adjacent deployments scored 1.50 and 2.75, one of the widest quality gaps in any category.

How do you test a skincare agent in five minutes?

  1. State a hard exclusion early: fragrance, an ingredient allergy, no essential oils.
  2. Give a contradictory skin brief and see if it reconciles or picks one half.
  3. Ask for a sequenced AM/PM routine, not a product list.
  4. Give it a budget and check whether it costs the routine.
  5. Ask a medication-interaction question. The right answer is a graceful decline with a referral.
  6. Request one more recommendation at the end to confirm the constraint survived.

Key takeaways

  • Memory reset appeared in 5 of 15 agents, the most damaging failure in skincare, because the constraint is a reaction.
  • Unsafe confidence appeared in 3 of 15, with agents crossing from skin type into skin condition.
  • Field averages: Accuracy 2.47, Memory 2.13, precisely the two dimensions skincare punishes.
  • Sequencing beats listing. A routine has an order; a product list does not.
  • Skincare deployments scored 1.50 and 2.75, one of the widest gaps in any single category.

Frequently asked questions

Does an AI need a selfie or skin analysis to build a skincare routine?

Not for this task. A facial scan or selfie upload helps with visible concerns like dark spots, redness, pore size and pigmentation, and an instant skin analysis is a reasonable starting point. Some AI-powered tools will read a selfie for those surface concerns and personalise a routine around them. But every constraint that broke the field in Alhena's 2026 test was stated in text: a fragrance sensitivity, a budget, and a contradiction between dry skin and breakouts. No photo surfaces any of those. An agent that leans on a facial analysis and then loses the fragrance rule six turns later has solved the easier half of the problem. Skin analysis and constraint memory are separate capabilities, and only one of them was scarce in 2026.

What should an AI ask before recommending a skincare routine?

Skin type, current concerns, anything you react to whether that is fragrance or a specific ingredient, what you already use, and what you can spend. That is what a personalised routine is built from, and none of it comes from a photo upload. Five short questions cover most of it. The failure mode is not asking too little, it is asking well and then not holding the answers: 5 of 15 deployments Alhena tested dropped a stated constraint later in the same conversation. An agent that interrogates you and then forgets is worse than one that asks less.

Can AI build a skincare routine for dry but acne-prone skin?

That contradiction is the whole test. Dry and breakout-prone is usually a barrier problem rather than two separate conditions, so the routine wants gentle cleansing, real hydration, and a targeted treatment kept away from anything that would strip further. The same reasoning applies to oily skin that still flakes. The strong deployments said this out loud before recommending anything. The weak ones picked one half of the brief, treated the acne, and left the dryness worse.

What does a good AM and PM skincare routine order actually look like?

Morning and evening do different jobs. AM: cleanse, antioxidant, moisturiser, SPF, in that order, because sun protection sits last and over everything. PM: cleanse, treat, repair, with conflicting actives separated across the two. A hyaluronic acid layer goes on damp skin under the moisturiser, and retinoids belong in the PM routine rather than on the same night as other strong actives. An agent that returns five products with no order has given you a shopping list, not a routine.

Can an AI recommend a routine for a beginner on a small budget?

It should, and the skill is subtraction rather than addition. A beginner routine is usually three skincare products: a gentle cleanser, a moisturiser and an SPF, with one treatment added once those are consistent. Costing it against a stated budget and naming the priority, where to spend and where to trade down, is what separates advice from merchandising. A routine tailored to what someone will actually keep doing beats a longer one they abandon in a fortnight.

How often should a skincare routine change?

Less often than the industry implies. Skin changes with the seasons, with hormonal shifts and with age, so a routine built in January may need adjusting by July, but the underlying skin type and any sensitivity usually hold. That is an argument for memory rather than reassessment: an agent that remembers what it built and what you reacted to can personalise the next visit properly, adjusting one product instead of restarting the skin science conversation and the dark spot questions from nothing. Only 1 of 15 deployments Alhena tested recalled a shopper on a return visit.

Customers tell our chatbot about their allergies and it forgets. How do we stop that?

The fix is architectural rather than a prompt change. The agent needs a persistence layer and selective retrieval so stated exclusions survive the whole conversation. Alhena logged memory reset, where a stated constraint is dropped mid-chat, in 5 of 15 live deployments.

Can an AI safely recommend skincare when customers mention prescription treatments?

It should decline that specific question, explain why, refer the shopper to a dermatologist, and carry on helping with everything else. Answering fluently is unsafe confidence, which Alhena found in 3 of 15 deployments and flagged as the failure with the heaviest compliance consequence. Product guidance is a merchandising task; dermatology advice is not, and no amount of personalisation makes that line move.

Can AI work out which active ingredients shouldn't be used together?

Stronger agents sequence conflicting actives across morning and evening rather than listing them together, which is a formulation question rather than a medical one. Anything touching a prescription, a diagnosed skin condition or dermatological treatment should be declined and referred instead.

How do we get an AI to build a full routine instead of recommending one product at a time?

Ask for sequencing rather than a list: AM and PM ordered, conflicting actives separated, and the whole set costed against a stated budget. Bundle-building under multiple constraints is one of the clearest tests of whether an agent reasons or ranks.

Will an AI routine builder actually increase our average order value?

The mechanism is coordinated multi-product recommendations, personalised to a stated brief rather than ranked by popularity, plus replenishment when the routine is remembered. Both depend on capabilities most deployments lack, which is why results vary far more by platform than by category.

Our customers ask about pregnancy-safe and teen skincare. Can AI handle that?

It can filter for commonly excluded ingredients and build around a stated constraint, which is a product attribute task. What it should not do is confirm safety for a specific pregnancy or medical situation, which belongs with a clinician.

How many products should an AI recommend in a routine before it feels like upselling?

Enough to cover the stated goals without stacking duplicate actives, typically a sequenced AM and PM set. Over-recommending is a merchandising instinct that reads as unhelpful in a category built on repeat purchase and trust.

Can the agent prompt customers to reorder their routine at the right interval?

Only if it remembers the routine it built. Cross-session recall appeared in just 1 of 15 deployments Alhena tested, so most replenishment conversations restart from scratch rather than picking up where the last one ended.

What went wrong most often when AI agents were given a skincare brief?

Two things dominated. Memory reset, where the fragrance-free rule vanished four turns later, and unsafe confidence in health-adjacent territory. Field-wide, Accuracy averaged 2.47 and Memory 2.13 out of 3.00, which are precisely the dimensions skincare punishes. That figure comes from Alhena's 2026 Agentic CX Stress Test.

How do skincare agents compare with the rest of the market?

The two skincare-adjacent deployments Alhena scored in 2026 landed at 1.50 and 2.75 out of 3.00, one of the widest quality gaps in any single category. Alhena itself scored 3.00 across all verticals tested.

How did the field handle a fragrance-free brief?

Full transcripts and scores across skincare, beauty and nine other verticals.

Power Up Your Store with Revenue-Driven AI