AI search revenue attribution is the practice of classifying your inbound traffic by the AI engine that referred it (ChatGPT, Perplexity, Gemini, Claude, Copilot), joining those sessions to actual purchase events, and benchmarking the result against your sitewide baseline. It answers the question every CFO now asks about AEO budgets: not "were we mentioned?" but "what did AI search sell?"
This guide covers how AI referral traffic is identified, a do-it-yourself attribution setup, a maturity model for grading your current state, what real conversion numbers look like across 310 stores, and where attribution honestly ends.
Key takeaways
- AI search attribution = classify sessions by engine (UTM + referrer, not referrer alone), join to orders, report against your sitewide baseline.
- Measured AI conversion rates are floors: unattributable checkouts and cross-window purchases bias every honest count downward.
- Across 310 stores, LLM referrals converted at 2.68% (#4 of 13 channels); ChatGPT owns 96.1% of the volume; Perplexity carries the premium basket.
- Climb the ladder: classified, qualified, closed-loop, incremental. Level 3 (visibility connected to checkout events) is where AEO budgets get defended; native closed-loop attribution is currently Alhena's territory because it requires first-party storefront data.
- Say "attributed," not "incremental," unless you ran the experiment.
Why is AI search traffic so hard to attribute?
Four reasons, all fixable to different degrees:
- Fragmented signals. AI referrals arrive as a mix of referrer headers (chatgpt.com, perplexity.ai, gemini.google.com, copilot.microsoft.com), UTM parameters (ChatGPT appends utm_source=chatgpt.com to many outbound links), and nothing at all (some in-app browsers strip referrers). Relying on referrer alone dramatically undercounts: in Alhena's platform data, UTM-tagged ChatGPT visits outnumbered referrer-identified ones by roughly ten to one over one 40-day window.
- Default analytics buckets hide it. Out of the box, most analytics tools file AI referrals under generic "Referral" or even "Direct," so the channel grows invisibly inside categories nobody watches.
- Identity is short. A shopper who asks ChatGPT for "the best retinol under $100," clicks through, and buys three days later from a bookmark breaks naive session joins. Any honest method states its attribution window.
- Some of it is simply dark. A share of checkouts cannot be tied to any first touch. In Alhena's published cohort research, roughly 25% of checkouts had unattributable fingerprints, which means measured AI conversion rates are floors, not ceilings.
The DIY setup: AI attribution in your existing analytics
You can stand up level-one attribution in an afternoon, on any stack:
- Build an AI channel group. Classify a session as AI-referred when the referrer domain OR utm_source matches an engine list: chatgpt.com and openai.com (ChatGPT), perplexity.ai, gemini.google.com and bard.google.com (Gemini), claude.ai (Claude), copilot.microsoft.com and bing.com/chat (Copilot). Check both signals; UTM catches what referrers miss.
- Split by engine, not just "AI." Engines behave differently as channels (different conversion rates, different order values), and the split is what makes the report actionable.
- Join to orders. Tie AI-classified sessions to transactions with your normal conversion tracking, and state the window you use (same-session, 7-day, 30-day).
- Always report against a baseline. "AI referrals converted at 2.7%" means nothing until it sits next to "sitewide average 1.9%." The baseline turns a curiosity into a budget argument.
This gets you real numbers with two known gaps: it counts landings, not influence (the shopper who researched in ChatGPT but typed your URL directly is invisible), and it tells you nothing about which products AI engines recommended before the click. Those gaps are what the higher maturity levels close.
The AI attribution maturity ladder
| Level | What you measure | What it proves | Typical tooling |
|---|---|---|---|
| 0. Blind | Nothing; AI traffic hides in Referral/Direct | Nothing | Default analytics |
| 1. Classified | Sessions by AI engine (referrer + UTM) | Channel size and growth | Custom channel groups |
| 2. Qualified | Engagement and conversion per engine vs baseline | Channel quality | Analytics + conversion joins |
| 3. Closed-loop | AI answer content (which SKUs recommended, how rendered) connected to sessions and checkout events | Which visibility work made money | AI visibility platform with native attribution |
| 4. Incremental | Controlled experiments (geo holdouts, content on/off tests) | Causal lift | Experimentation infrastructure |
Most brands doing AEO seriously in 2026 sit at level 1 or 2. Level 4 is rare everywhere. The practical target for an online store is level 3: the point where "our visibility on moisturizer prompts doubled" and "moisturizer revenue from ChatGPT referrals rose" appear in the same report.
What do real AI conversion numbers look like?
From Alhena's published 12-month cohort study across 310 online stores and roughly 190M visits (full study and methodology): LLM referrals converted at 2.68% in the reliable measurement window (October 2025 to April 2026), ranking fourth of thirteen channels, above Google Ads (1.87%), Direct (1.15%), and Meta Ads (0.51%). ChatGPT drove 96.1% of LLM referral volume. Perplexity visitors were scarcer but carried an 82% higher average order value than ChatGPT visitors ($129 vs $71). And LLM visitors who engaged with an on-site AI assistant converted at 4.3x the rate of those who did not, an engaged-versus-unengaged comparison that reflects selection as well as causation, and is labeled accordingly in the study.
Two honest notes about all such numbers, including ours: the ~25% unattributable-checkout share means true rates are likely higher, and same-month joins drop cross-month conversions. The direction of both biases is understatement, uniformly across channels, which keeps relative comparisons fair.
How Alhena implements closed-loop attribution
Alhena AI Visibility operates at level 3 natively because the attribution plumbing and the visibility tracking share one dataset. Its traffic classifier buckets every session by AI source (ChatGPT, Perplexity, Gemini, Claude, Copilot) using both UTM and referrer signals; its on-site agents observe engagement first-party; and checkout events join to those sessions by visitor fingerprint, reported per engine against your sitewide baseline in the dashboard's traffic-and-conversion view. Those same agents handle real shopper conversations across web chat, email, Instagram DMs, and WhatsApp, and what shoppers ask feeds forward into which prompts the visibility layer tracks and which gaps it flags, so demand intelligence and revenue measurement come from one first-party dataset. Because the same platform tracks which of your products appear in AI answers at the SKU level, the loop closes: prompt-level visibility, the fix that changed it, and the checkout events that followed live in one system. The computation is documented publicly in How Alhena measures AI visibility, including its limits.
For context on how the rest of the category handles this: Peec AI deliberately scopes attribution out (analytics only, their stated positioning), Profound offers it through a Partnerize partnership, and Scrunch leaves purchase reporting to your web analytics. Those are legitimate choices for brand-level use cases; they stop short of purchase-side data because none of those platforms sits on the storefront. This is the one capability where owning an on-site agent changes what is measurable.
Attribution is not incrementality
Attribution says: sessions we classified as AI-referred produced these orders. It does not say those orders would have vanished without AI search; some of those shoppers would have found you anyway. Getting from attribution to causation requires controlled experiments: geographic holdouts, staggered content rollouts, or on/off tests, with predefined windows and assignment rules. The honest report reads: "revenue attributed to AI-referred sessions, N-day window, against sitewide baseline," and saves causal language for actual experiments.
The weekly report worth running
One page, five rows, per engine: sessions, engaged-session rate, conversion rate vs sitewide baseline, attributed revenue, and the top three landing pages. Add one visibility-side line (prompts where you appeared, out of prompts tracked) so the leading indicator and the money sit together. Review monthly for trend, quarterly for budget decisions, and annotate engine model updates (a GPT or Gemini release can move numbers with zero change on your side).
From mentions to money
See which engines, products, and pages send you shoppers, joined to the checkout events that followed.
Frequently asked questions
AI search revenue attribution classifies inbound traffic by the AI engine that referred it (ChatGPT, Perplexity, Gemini, Claude, Copilot), joins those sessions to purchase events, and reports conversion and revenue per engine against your sitewide baseline. It turns AEO from a mentions report into a revenue report.
Classify a session as ChatGPT-referred when the referrer contains chatgpt.com or openai.com, OR when utm_source equals chatgpt.com. Checking both signals matters: ChatGPT appends UTM parameters to many outbound links, and in Alhena's platform data UTM-tagged ChatGPT visits outnumbered referrer-identified ones by roughly ten to one.
Some AI apps and in-app browsers strip referrer headers, so those sessions land in Direct with no source signal. UTM parameters recover part of it, and the rest stays genuinely dark, which is one reason measured AI conversion numbers are floors rather than ceilings.
In Alhena's 12-month study across 310 online stores, LLM referrals converted at 2.68% in the reliable window (October 2025 to April 2026), ranking fourth of thirteen channels and above Google Ads at 1.87%. Your number will vary by vertical and price point; the useful comparison is always against your own sitewide baseline.
Attribution counts orders from sessions you classified as AI-referred; some of those buyers would have found you anyway. Incrementality measures the causal lift of the channel and requires controlled experiments such as geo holdouts or on/off tests. Report attribution as attribution, and reserve words like incremental for actual experiments.
Alhena computes attribution natively by joining AI-classified sessions to checkout events with a sitewide baseline, which is possible because its agents already run on the storefront. Profound offers attribution via a Partnerize partnership, Scrunch AI leaves purchase reporting to your web analytics, and Peec AI deliberately keeps attribution out of scope.
Pick one, state it, and keep it stable: same-session is strictest, 7-day covers most considered purchases, 30-day suits high-price catalogs. Longer windows credit more orders but weaken the causal story. Whatever you choose, disclose that cross-window purchases are dropped, so your numbers understate the channel.
Yes, and you probably should, because last-click undercounts AI search badly. A shopper reads an AI answer, clicks through, leaves, then returns days later through a branded search that takes all the credit. A multi-touch or data-driven attribution model spreads credit across that buyer journey instead. Report both: last-click as your floor, multi-touch as your influence estimate.
Pass the AI source as a field on the order itself at checkout, not only into your analytics tool. Store the engine name, the utm_source value and the referrer alongside the transaction ID. AI revenue then becomes queryable in the same system where refunds, repeat purchase rate and lifetime value already live, rather than sitting trapped in a reporting layer that cannot see any of them.
The classification method transfers; the benchmarks on this page do not. A B2B company should tag AI-referred sessions the same way, then push the source into the CRM and marketing automation stack so it rides the record through the pipeline instead of expiring with the session. Expect a longer sale cycle to hide more of the contribution, and lean harder on self-reported attribution to recover it.
Self-reported attribution asks the buyer directly, usually a "how did you hear about us?" field at checkout or in a post-purchase survey. It is the only method that catches influence carrying no trackable signal, such as the shopper who read an AI answer, never clicked, and typed your URL a week later. Use it as a directional cross-check against direct traffic, not as a replacement for session data.
Watch three proxies together: the share of tracked prompts where an AI answer cites your brand, branded query volume in Google Search Console, and your self-report field. A citation with no click still moves demand. When citations climb and branded search and direct traffic climb alongside them, you are seeing influence that no referral report will ever surface on its own.
Because they count different things. Analytics counts sessions it can classify from referrers and UTM parameters. An ad platform counts whatever its own pixel fired on. Neither sees AI-generated recommendations that produced no click. Name one system as your record for AI conversion data, document the rule you used, and treat everything else as directional.
Split on the referring domain and utm_source, never on the query. A visit from chatgpt.com or perplexity.ai is AI search. A Google AI Overview click arrives as ordinary organic search and cannot be isolated in most analytics tools, so keep Google AI surfaces out of your AI channel group and report them inside organic search with a footnote explaining why.
Enough conversions to survive noise, which matters more than raw traffic. As a working rule, wait for 30 to 50 AI-attributed orders before you calculate a conversion rate you would defend in a meeting. Below that, report volume and direction only, and resist the urge to estimate annual revenue from a handful of sales.
Partly. Attribution tells you what revenue AI-referred sessions contributed against your baseline, which is usually enough to defend a budget line. It does not isolate what your AEO work caused, because some of those buyers would have arrived regardless. For a defensible ROI number, pair the attributed revenue with one controlled test rather than treating attribution as pure return.