Measuring AI Agents for Wellness Brands: Benchmarks and an Honest Attribution Model

Illustration of a three-tier measurement ladder for AI commerce benchmarks
Person on a ladder measuring rising bar chart columns with a caliper the three tiers of AI agent measurement.

Measure an AI shopping agent on three tiers: engagement (what it did), attribution (what it was present for), and incrementality (what it caused). Wellness brands see 4.68% conversion from LLM-referred traffic — but that number measures traffic quality, not your agent. Only a holdout test proves cause.

The short version

  • Sort every AI number into a tier before you trust it. Engagement is auditable, attribution is correlational, incrementality is causal.
  • The 4.68% wellness benchmark is a traffic-source figure, not evidence that any on-site agent caused a sale.
  • No brand or vendor has published a completed holdout result for an on-site AI shopping agent as of July 2026 — including Alhena.
  • In supplements the unit of measurement is subscriber lifetime value, not order value. Health and wellness is the largest subscription category on Recharge.
  • One holdout — 2 to 4 weeks, roughly 15,000–20,000 visitors per cohort — is the cheapest way to turn a benchmark into a fact.

What is the published benchmark for AI agents in wellness?

Short answer: LLM-referred shoppers convert at 4.68% on health and supplement stores, against a 2.47% cross-vertical average, across Alhena's 329-brand study of US and EU merchants.

Read that precisely. It is the conversion rate of visitors who arrive from an assistant like ChatGPT, aggregated across hundreds of brands. It tells you who is knocking on the door, not what happens after they walk in.

The distinction matters more than the number. A traffic-source benchmark cannot forecast what your on-site agent will do, because it measures the shopper's intent on arrival — and people who research magnesium dosing inside an AI assistant arrive unusually far down the funnel. For the vertical breakdown and a census of what is actually live in the category, see Alhena's 2026 operator's guide to AI agents in health and wellness.

The 4.68% measures the quality of the traffic, not the effect of the agent.

What are the three tiers of AI agent measurement?

Short answer: Engagement counts what the agent did. Attribution links it to a commercial outcome. Incrementality isolates what it caused. Most arguments about AI ROI are really arguments about which tier a number belongs to.

THE MEASUREMENT LADDER Three tiers. Three different questions. Only one of them says "caused." HARDER · MORE CAUSAL 1 Engagement What did the agent do? Chats handled · resolution · CSAT · deflection AUDITABLE, NOT CAUSAL 2 Attribution What was the agent present for? Assisted conversion · AOV · revenue share CORRELATIONAL 3 Incrementality What did the agent cause? Holdout · randomized exposure · revenue per visitor ALMOST NEVER PUBLISHED Most vendor numbers live on Tier 2 — and get read as Tier 3. That gap is the whole problem.
Each tier answers a different question. The ladder runs from most verifiable to most causal.

Tier 1 — Engagement: what the agent did

Tier 1 counts what happened inside the conversation: chats handled, questions answered, resolution rate, CSAT. These are the most trustworthy numbers you will ever have, because you can audit them line by line.

What Tier 1 cannot do is show the agent made money. Deflection is cost avoidance, and well-tuned agents tend to land somewhere in the 60–86% range before the remaining tickets get genuinely hard — so a brand that measures only deflection will conclude the agent is finished improving long before it has touched revenue. The metric worth reporting is resolved deflection rather than raw containment.

Tier 2 — Attribution: what the agent was present for

Tier 2 links the agent, or the channel feeding it, to a commercial outcome: assisted conversion, engaged-versus-unengaged AOV, attributed revenue share. Alhena credits a conversation only when it produced a helpful, product-grounded response, inside a fixed 24-hour attribution window.

Every number in this tier is correlational. "Shoppers who engaged converted higher" is not "engaging caused them to convert," because the shoppers who open a chat are self-selected. Tier 2 is where honest brands and overreaching brands use identical words and mean different things.

Tier 3 — Incrementality: what the agent caused

Tier 3 isolates cause. You withhold the agent from a random slice of traffic, or randomize exposure, and compare revenue per visitor between the groups.

Only Tier 3 answers the question your CFO actually asked: how much of this revenue would we have earned anyway? Practitioners who run these tests consistently report that true incremental lift lands well below the headline attributed figure — often in the single digits once self-selection is stripped out.

MetricTierWhat it provesSelf-selection risk
Resolution rate, CSAT, deflectionEngagementThe agent handled real volume at qualityNone — directly auditable
LLM-referred conversion (4.68%)AttributionThe traffic source arrives with high intentChannel-level; says nothing about the agent
Engaged vs. unengaged conversion and AOVAttributionEngaged shoppers behave differentlyHigh — chat-openers are already leaning in
Attributed revenue shareAttributionYour attribution rule credited the agentDepends entirely on an often-undisclosed window
Revenue per visitor, exposed vs. held outIncrementalityCausal, incremental revenueNone — randomization removes it

Has any brand published a controlled test of an AI shopping agent?

Short answer: no. As of July 2026, no brand and no vendor — Alhena included — has published a completed holdout or randomized test reporting the incremental revenue lift of an on-site AI shopping agent.

That absence is worth sitting with. Every headline figure in this category, in wellness and everywhere else, is Tier 1 or Tier 2. The methodology exists and is well documented; the published results do not.

It is also not a knock on the category. It is the starting condition every honest measurement plan has to design around, and it is why the recommendation at the end of this guide is a test rather than a benchmark.

If a vendor — including us — shows you attribution numbers and calls them incremental revenue, push back on the word.

That line is from Alhena's own guide to AI search revenue attribution, and it is the standard this post applies to Alhena's published numbers as readily as to anyone else's.

How can you tell what a vendor's AI number actually means?

Short answer: sort it into one of four buckets. The wording almost never tells you which one you are in, so you have to ask.

Before and after

Weakest form. The season, ad budget, catalog and economy all changed too.

Engaged vs. unengaged

Better, and what most credible case studies report. Cannot escape selection bias.

Attributed share

A rule credited the agent. Ask what the window is before you read the number.

Controlled

The only bucket that supports the word "caused." Nearly empty in commerce.

Why a single before-and-after snapshot is never a fact

Adobe Analytics reported that AI-referred shoppers converted 54% better than other traffic in May 2026, with revenue per visit up 53%. A year earlier, that same traffic converted at roughly half the rate of other channels.

Same metric, opposite direction, twelve months apart. If a channel-level reading can flip that hard, no single snapshot is a fact about your agent.

Why a big multiple often means a small denominator

Engagement lift figures are frequently quoted as ranges spanning an order of magnitude. The top of those ranges usually comes from a channel whose unassisted baseline is close to zero — short-form video, for instance — so almost any engaged cohort looks enormous against it.

When you see an eye-catching multiple, ask what the denominator was before you ask what the agent did.

Does a conservative "attribution factor" fix the problem?

No. Applying a 10–25% discount to chatbot-influenced revenue, as most vendor business cases recommend, is sensible spreadsheet hygiene — and it is still Tier 2. A discount applied to an attributed number produces a smaller attributed number, not evidence of causation.

Use the factor to keep the CFO conversation honest. Run the holdout to end the argument.

Why do subscriptions change the math for supplement brands?

Short answer: because a single transaction is the wrong unit. Health and wellness is the largest subscription category on Recharge, with 11.2 million active subscribers — about 1.9 million ahead of beauty at 9.4 million.

When most of your revenue recurs, first-order conversion and AOV describe the smallest slice of what the agent affects. The value of a converted shopper is a lifetime value, and lifetime value is governed by churn.

11.2M
active health & wellness subscribers — the #1 subscription category
Recharge, 2026
~8.8%
monthly churn in health & wellness subscriptions, up from ~4.2% in 2023
Recharge, 2026
19.3%
of online sales returned overall — a lever that barely applies to consumables
NRF & Happy Returns, 2025

Why churn, not AOV, is the number that compounds

An agent that lifts first-order AOV by 12% but does nothing for retention is worth far less than one that lifts month-two survival by three points, because the second effect repeats on every future shipment.

Subscriber lifespan is roughly the inverse of monthly churn, so small movements in the churn rate swing lifetime value hard. Churn in this category is also front-loaded — a large share of cancellations happen before the second shipment ever ships.

WHERE WELLNESS REVENUE ACTUALLY ACCRUES A three-point churn improvement outruns any first-order AOV lift CUMULATIVE REVENUE PER SUBSCRIBER first order ≈ 13% of year-one revenue +16% by month 12 5.8% monthly churn (3 points better) 8.8% monthly churn (category benchmark) 136912 SHIPMENT MONTH Illustrative arithmetic, not measured results. Assumes a fixed order value and constant monthly retention.
Worked example only. The point is the shape: retention effects compound, first-order effects do not.

Which conversations should count as agent revenue?

In wellness, the agent's most valuable conversations are often the ones that keep revenue rather than start it. A skip taken instead of a cancellation, a reorder retimed, a stack adjusted, a "where is my shipment" answered before frustration sets in — each is a revenue event.

All of them are invisible on a dashboard that only counts pre-sale chats. Retention outcomes belong in Tier 2 alongside conversion, with their own attribution windows defined in advance.

Should returns be part of a wellness AI ROI model?

Barely. Returns are a headline cost in apparel and auto parts, where roughly 19.3% of online sales come back. Consumable supplements return at a small fraction of that rate.

The equivalent lever in wellness is churn, and that is where the measurement effort should go.

What belongs in a wellness AI measurement stack?

Short answer: five metrics, each labeled by tier, with subscription outcomes sitting alongside conversion rather than beneath it.

What to trackTierWhy it matters in supplements
Resolved deflection + CSAT1Cost avoidance, and an early warning system for answer quality
Assisted conversion and AOV
fixed 24-hour window, set before launch
2The standard commerce read — necessary, and the least important number here
Subscription starts and month-2 survival2Where most category revenue lives; churn is front-loaded at the first reorder
Skip-instead-of-cancel and cancellation reversals2Retention events the agent creates that no pre-sale dashboard captures
Answer accuracy and escalation rate1Ingredients, interactions and dosing make this a liability metric, not a nice-to-have
Revenue per visitor, exposed vs. held out3The only line on this table that supports a causal claim

Define your attribution windows before you launch, not after you see the numbers. The exact windows matter less than fixing them in advance so the figure cannot be retrofitted to look good. Alhena documents its own convention in the Shopify AI chatbot ROI guide.

How do you run a holdout test on an AI shopping agent?

Short answer: hide the agent from a random slice of visitors for two to four weeks, then compare revenue per visitor between the two groups. The gap is the incremental lift.

HOW A HOLDOUT TEST WORKS Randomize exposure, then compare revenue per visitor Incoming traffic randomly assigned Exposed · agent on 90% of visitors Held out · no agent 10% of visitors Compare revenue per visitor the gap is the incremental lift Run 2–4 weeks · roughly 15,000–20,000 visitors per cohort to detect a 5% lift at 95% confidence
The cheapest route to a causal answer. Sizing guidance from Alhena's holdout testing methodology.
  1. Pick the split. 90/10 or 80/20. A 10% holdout costs very little revenue and is usually enough.
  2. Fix the primary metric. Revenue per visitor, not conversion rate — it captures AOV and conversion in one number.
  3. Size the cohorts. Roughly 15,000–20,000 visitors per group to detect a 5% lift at 95% confidence.
  4. Run for 2–4 weeks. Long enough to cover a full weekly cycle and any promotional rhythm.
  5. Add the wellness layer. Compare subscription starts and month-two survival across both groups, not just first-order revenue.
  6. Declare a verdict. Won, lost, or inconclusive. "Inconclusive" is a legitimate and common result.

Alhena's holdout testing guide and Experiments feature exist to run exactly this. One clean holdout is worth more than a year of engaged-versus-unengaged dashboards.

How should you measure AI answer accuracy for supplements?

Short answer: score it continuously against verified label and policy data, and treat it as a revenue metric rather than a compliance afterthought.

This matters more in wellness than anywhere else, because the questions are about ingredients, interactions, dosing and pregnancy safety. A 2026 audit published in BMJ Open found that 49.6% of responses from five general-purpose consumer chatbots to medical questions were problematic.

Read that finding precisely, in the spirit of this guide. It tested free-tier general assistants under adversarial prompting, which overstates real-world rates — and a grounded commerce agent restricted to your catalog and policy data is a different system entirely. That difference is exactly what an accuracy metric is for: it is the only way to know yours is the safe kind.

Four things to score

  • Grounding rate. What share of answers trace to live product, label or policy data rather than free generation.
  • Disclaimer coverage. What share of claim-bearing answers carry a compliant structure/function disclaimer.
  • Disease claims. The target is zero. The FTC applies the same substantiation standard regardless of how a claim is framed, and a disclaimer does not cure a deceptive one.
  • Escalation reliability. How consistently the agent routes medical questions to a human instead of answering them.

Consumer scepticism gives this urgency. In a Pew survey, 48% of adults who use AI chatbots for health information rated them highly convenient — but far fewer rated them highly accurate, and KFF found 32% of US adults had used a chatbot for health information in the past year. An agent that converts well and advises badly is a liability, not an asset.

How does agentic checkout change what you can measure?

Short answer: it fragments attribution, which makes your own on-site conversation the one measurement surface you fully control.

The durable 2026 pattern is "discover in AI, buy on your own site." Walmart reported that in-chat ChatGPT checkout converted roughly three times worse than sending shoppers through to its own site — while driving about twice the new-customer rate.

The practical implication for supplement brands is that upstream visibility and downstream measurement are now one loop. Whether your brand is present and correctly described inside an AI answer is a measurable input to that 4.68%, and it belongs in the same review as on-site conversion — Alhena covers that side in its guide to AI visibility for health and wellness brands.

Key takeaways

What to remember

  • Label every number by tier. Never let an engaged-versus-unengaged figure wear the clothes of a causal one.
  • The 4.68% is a traffic-quality benchmark. It is a reference point, not a target your agent will hit.
  • Attribution windows go in before launch. A window chosen after seeing results is not a measurement, it is a decoration.
  • Measure the agent as a retention engine. In a category where revenue recurs, lifetime value and churn are the real scoreboard.
  • Accuracy and revenue are the same asset when the questions are about what people put in their bodies.
  • One holdout ends the argument. Nobody in wellness has published one yet — which means the first brand to run one will know something its competitors do not.

What should you do next?

  1. Audit your current dashboard this week. Put every AI metric you report into one of the three tiers. Most brands find they have no Tier 3 column at all.
  2. Write down your attribution window. If you cannot state it in one sentence, your attributed revenue number is not yet meaningful.
  3. Add two subscription metrics. Month-two survival and cancellation reversals, split by whether the shopper engaged the agent.
  4. Stand up an accuracy score. Sample 100 answers against verified label and policy data, and set a monthly cadence.
  5. Schedule one holdout. Ten percent of traffic, four weeks, revenue per visitor as the primary metric.

A measurement stance is a product requirement, not a slide. It is also a fair test of any vendor, Alhena included: can the platform separate assisted revenue from support deflection, expose the attribution window, and give you the data to run a holdout?

Run the test that settles the argument

See how Alhena separates assisted revenue from deflection, exposes the attribution window, and holds out traffic to measure what the agent actually caused.

Frequently asked questions

What is a good conversion rate for a wellness AI shopping agent?

The published benchmark is 4.68% for LLM-referred traffic in health and supplements, against a 2.47% cross-vertical average across Alhena's 329-brand dataset. Treat it as a reference point rather than a target. It measures the quality of LLM-referred traffic across many brands, not any single agent's effect, and individual brands vary widely by catalog, price point and how much demand already arrives with intent.

Does an AI agent cause the conversion lift it reports?

Usually you cannot tell from the reported number. Most published lifts, including Alhena's, are engaged-versus-unengaged comparisons, which cannot separate the agent's effect from the fact that shoppers who open a chat are already higher-intent. The only way to establish cause is a controlled test.

What is the difference between attribution and incrementality?

Attribution credits the agent for outcomes it was present for, using a defined rule such as a 24-hour click-through window. Incrementality measures what would not have happened without the agent, using a holdout or randomized experiment. Attribution answers “what did the agent touch”; incrementality answers “what did the agent add.” Only incrementality supports the word “caused.”

What attribution window should a wellness brand use?

Any defensible window works as long as you fix it before launch and disclose it. Alhena's own convention is a 24-hour window, counting only conversations that produced a helpful, product-grounded response. Because most wellness revenue recurs, add a separate, longer window for retention events such as skips and cancellation reversals — those sit on a slower pathway than a first purchase and a same-day window will miss them entirely.

How big does a holdout test need to be?

Roughly 15,000 to 20,000 visitors per cohort to detect a 5% lift in revenue per visitor at 95% confidence, run over two to four weeks. A 90/10 or 80/20 split is standard, which keeps the revenue cost of the test small. Lower-traffic brands should either extend the window or accept that they can only detect a larger effect.

Why is deflection rate a weak measure of AI ROI?

Deflection counts cost avoidance, not revenue, and it plateaus. Once containment reaches the top of its practical range, the remaining tickets are complex cases that need a human, so pushing higher tends to hurt satisfaction. Track resolved deflection alongside CSAT, and treat both as Tier 1 evidence about the workflow rather than proof of commercial impact.

How do subscriptions change AI measurement for supplement brands?

They move the important number from order value to lifetime value. Health and wellness is the largest subscription category on Recharge at 11.2 million active subscribers, so most revenue recurs. Count every time the agent changes a delivery schedule instead of ending it, retimes a replenishment, or reverses a cancellation. Those retention effects compound across every future shipment and usually outweigh anything the agent does to the first order.

What should an AI assistant be measured on during subscriber onboarding?

Month-two survival, split by whether the shopper interacted with the assistant in their first week. Churn in this category is front-loaded, so onboarding is where an AI assistant earns or loses most of the lifetime value it will ever influence. A useful secondary read: how many new subscribers ask a dosage or timing question, and whether that single interaction changes their odds of reaching a second shipment.

Can you measure personalization separately from the agent itself?

Yes, and you should. Run the holdout on the agent first to establish that it moves revenue at all, then test personalization variants inside the exposed group. How well the agent adapts to a returning shopper — recalling a previous stack, or a stated sensitivity — is a different experiment from whether the agent exists. Collapsing the two produces a number nobody can act on.

Should an AI agent use wearable or app data to personalize supplement recommendations?

Only if you can measure the accuracy cost. Wearable inputs make personalization feel sharper, but they widen the range of claims the agent might make and raise the stakes on every answer. If you connect that data, treat it as its own holdout variant and watch escalation rate as closely as conversion rate. A wearable-informed recommendation drifting toward diagnosis is the failure mode to instrument for.

How should a wellness brand measure AI answer accuracy?

Score four things continuously: grounding rate against live label and policy data, disclaimer coverage on claim-bearing answers, disease claims with a target of zero, and escalation reliability on medical questions. A 2026 BMJ Open audit found 49.6% of general chatbot responses to medical questions were problematic, which is the benchmark a grounded commerce agent should be measured against — not a rate to accept.

What compliance metrics belong on an AI agent dashboard?

The same ones your legal team would ask for in an ad review, reported weekly rather than annually. Disclaimer coverage, disease-claim count, and how often the agent overstates clinical evidence are all countable. The FTC applies the same substantiation standard however a claim is framed, so an agent that improvises around evidence creates regulatory risk no conversion lift offsets. Compliance and revenue belong on one dashboard, not two.

What data does an AI agent need to collect for measurement to work?

Session-level assignment, conversation quality flags, and checkout events joined to one visitor identity. That is enough to compute revenue per visitor and run a holdout. Data minimization is the safer default for an agent fielding health questions — collect what measurement requires and nothing more. EU shoppers bring GDPR obligations regardless of where your platform sits, and a wellness brand earns more trust by storing less than by centralizing everything.

How do you measure an AI agent during a new product launch?

Compare the agent's assisted conversion on the new SKU against your catalog baseline for the same period, not against the product's own pre-launch numbers, which do not exist. A product launch is also the cleanest use case for testing whether the agent can sell something it has no conversation history on. If it can, your grounding is working. If it defaults to bestsellers, your catalog sync is stale.

How should brands measure agentic checkout and AI-referred demand?

Separately from on-site conversion, because agentic surfaces fragment attribution. The durable 2026 pattern is discovery inside an assistant followed by purchase on your own storefront, so classify AI-referred sessions as their own channel and hold them to the same window rules as everything else. Your on-site conversation remains the one surface where you control the full measurement workflow end to end.

Should returns be part of a wellness AI ROI model?

Less than in most verticals. Returns are a major cost in apparel and auto parts, where roughly 19.3% of online orders come back, but consumable supplements return at a small fraction of that rate. The wellness equivalent of the returns lever is churn, so put the measurement effort into retention instead.

What is the minimum honest AI ROI test a wellness brand can run?

One holdout. Withhold the agent from 10% of traffic for a few weeks and compare revenue per visitor, plus subscription starts and month-two survival, between the exposed and held-out groups. Pair it with attribution windows fixed before launch and you have a measurement plan more rigorous than most of the category.

Power Up Your Store with Revenue-Driven AI