Your product feed is not a data export. It is the only version of your catalog that ChatGPT, Perplexity, and Google AI will ever read, and most brands are handing them a title, a price, an image, and nothing else.
For a decade, a product feed was plumbing. You mapped it once, pointed it at a shopping channel, and forgot about it. The people making buying decisions were on your site, looking at your photography, reading your copy.
That is no longer where the decision starts. It starts in a conversation with an assistant that never sees your product page, only your fields. And it does not read them the way a shopper does. It parses them, scores its own confidence, and quietly drops anything it cannot describe with certainty. There is no penalty screen and no error log. Your product simply does not come up.
The feed became the storefront
AI shopping assistants do not browse. They convert your product attributes into vector embeddings and match them against a shopper's actual sentence, something like "a lightweight sunscreen I can wear running that won't sting my eyes," using semantic similarity. Every product gets a confidence score. Low confidence products are not ranked lower. They are excluded.
This is why a description like "lightweight daily sunscreen for outdoor runners" outperforms "SPF 50 sunscreen" in a conversational context. One is a spec. The other tells the model who this is for, when they use it, and what problem it solves. Keyword density is irrelevant here. Interpretability is everything.
Six fields decide whether you appear
Not sixty. Six, and most catalogs are missing three of them.
01GTIN
The identifier that lets a platform cross-reference you against global databases and aggregate reviews. No GTIN, no verification, no confidence.
02Intended purpose
The strongest single predictor of surfacing. It maps your product to the shape of the question, "best moisturizer for sensitive skin," rather than to a keyword.
03Material and composition
"Organic cotton bedding." "Sulfate-free shampoo." These are filters applied before ranking. Miss the field and you are absent from the whole query class.
04Use-case copy
Free text is where semantic matching gets its strongest signal: the buyer, the scenario, the benefit. It is also the field most brands leave to marketing and never fill.
05Category taxonomy
"Athletic Shoes > Trail Running" gets matched to trail queries. "Shoes" gets matched to nothing in particular. Precision here costs you nothing.
06Availability and fulfillment
Out-of-stock and slow-shipping items get deprioritized in real time, and repeat stockouts leave a mark. "Need it by Friday" is now a ranking input.
None of these are exotic. Every one is a column your PIM already has and your feed probably drops.
What the completeness gap is worth
Analysis from eFulfillment Service's Complete Product Data Optimization Guide for Google's AI Shopping (January 2026) puts stores with near-complete attribute coverage at 3 to 4 times the visibility in AI recommendations compared with sparse feeds. Treat the multiple as directional rather than precise. The direction, though, is not in dispute, and it matches what happens when brands close the gaps on their own catalogs.
Tatcha ran Alhena's shopping assistant against a fully mapped catalog on Salesforce Commerce Cloud, and the assistant became a measurable sales channel rather than a support cost.
All of it at 81% CSAT, while deflecting 82% of incoming chats. Read the full Tatcha story.
Where recommendation engines fit in
AI recommendation systems blend user signals, meaning browsing and purchase history, with product signals, which are the six fields above. Collaborative filtering alone cannot do this work because it has no product-level context to reason over, which is also why it stumbles on anything new enough to lack interaction data. For the full mechanics of collaborative, content-based, and hybrid systems, see how AI recommendation engines power upselling and cross-selling.
Fixing this is configuration, not a project
Alhena's shopping assistant reads these exact fields and grounds every recommendation in them, which is why it answers "I don't have that detail" instead of inventing a fabric blend. Native connectors for Shopify, WooCommerce, Magento, and Salesforce Commerce Cloud pull your feed in without manual mapping, and most brands are live inside two days with no engineering tickets.
Getting your fields right is also what makes you legible to platforms you do not control. That is the job Alhena AI Visibility does, tracking how your products actually surface in AI answers so feed work stops being guesswork. Two adjacent pieces of the same problem are worth reading next: the JSON-LD implementation, and the on-page PDP checklist.
Audit your feed against the six. Close the gaps in the order above, GTIN and intended purpose first, because they gate the most queries. The brands doing this now are not buying an advantage so much as avoiding a slow, invisible subtraction.
See how your feed actually reads
We'll run your catalog through an AI shopping context and show you which fields are costing you queries.
Frequently asked questions
Platforms convert structured attributes into vector embeddings and score each product by confidence. Products missing GTIN, intended purpose, or material score low and get excluded before ranking begins, and there is no notification when it happens.
Intended purpose, consistently. It maps a product to the shape of a shopper's question rather than to a keyword, which is what conversational retrieval matches on. Use-case copy is a close second, and the two reinforce each other.
Related, but not the same job. This post covers being surfaced by external platforms you don't control. For what your own assistant needs to answer shoppers correctly, and the failure modes when it can't, see catalog data quality and the accuracy ceiling.
Yes, and they are different surfaces. The feed is what shopping platforms ingest. Schema is what crawlers parse on the page itself. The schema markup guide covers the JSON-LD implementation.
Connecting Alhena to an existing store is a configuration task rather than a dev project, and most brands are answering shoppers within two days. Headless stacks take a little more setup alongside our solutions team.