How structured retailer data supports discovery in AI search | simlir blog
How structured retailer data supports discovery in AI search
What AI search systems can read from public retailer pages, and how crawlability, structured data, internal links and stable identity support accurate discovery.
Structured retailer data gives AI search systems clearer paths to discover and understand products. Credit: Generated for Simlir
The short answer
Retailers improve their chance of accurate discovery when product information is crawlable, clearly structured, internally linked and consistent with what a shopper actually sees. None of this guarantees inclusion or ranking in any search engine or assistant. Google states that structured data makes a page eligible for rich results rather than entitled to them, and no generative surface offers guaranteed placement. What good data reliably does is remove the reasons your product cannot be read, matched or described correctly.
The catalogue fields shopping assistants rely on: stable identifiers, comparable specs, price and availability context, images and canonical links. With a readiness checklist.
What AI search systems can retrieve from public pages
Set aside the branding for a moment. Whatever an AI search feature is doing at the end, the input is largely familiar: it reads public web pages that a crawler was allowed to fetch, and it works with what those pages made explicit.
That means the practical inputs are the ones you already influence:
Visible page text — the product name, the description, the specification table, the price, the availability statement.
Structured markup — schema.org Product data, where it matches the visible content.
Page structure — the title element, headings, tables and lists, which signal what is a fact and what is decoration.
Link context — how you link to the page, and with what anchor text.
Identifiers — GTIN and MPN, which allow your listing to be recognised as the same product a shopper saw elsewhere.
The important consequence: anything only implied is likely to be lost. A capacity that exists only inside a variant dropdown, a noise level that appears only in a product photograph, a stock status conveyed only by a greyed-out button — these are legible to a human looking at the page and absent from the extracted record.
Crawlability and index eligibility
This is the unglamorous foundation, and it is where more retail catalogues fail than anyone expects — usually as a regression rather than a decision.
Product pages return 200 and are not blockedCheck robots.txt and meta robots tags on the live product template, not on a page you happened to have open.
Key content does not require interactionSpecifications behind a tab that only loads on click may not be seen at all.
Products are reachable through linksNot only through internal search or a filter interface. If the only route to a product is a search box, assume it is undiscoverable.
Faceted URLs do not overwhelm the crawlCanonicalise or exclude filter and sort permutations so crawl budget reaches actual products.
Sitemaps list canonical product URLsCurrent, complete, and reflecting the canonical form.
The crawl posture is reviewed after every releaseAn accidental disallow or a template-level noindex can remove an entire category from every downstream system in one deployment.
Product structured data that matches visible content
Schema.org Product markup remains worth implementing, with one rule above all others: the markup must describe what the page actually shows. Google's guidance is explicit that structured data should reflect visible content, and mismatches risk losing rich-result eligibility entirely — a worse outcome than not marking up at all.
Property
Populate with
Common failure
name
The product name shown on the page
Keyword-stuffed variant that differs from the visible heading
gtin, mpn, sku
Real identifiers
Omitted entirely — the most consequential gap on this list
brand
The manufacturer brand
Set to the retailer name instead
image
Durable HTTPS product images
Expiring or signed CDN URLs
offers.price
The price the shopper sees
Stale value, or the pre-discount price while the page shows the discounted one
offers.availability
Actual stock status
Hardcoded to InStock across the template
aggregateRating
Genuine reviews for this product
Ratings aggregated across a whole range, or invented
Two habits prevent most of these: validate markup after every release, because template changes break it routinely; and treat markup as a projection of your product data rather than as a separate artefact maintained by hand. Hand-maintained markup drifts from the page within a quarter.
Internal links and descriptive page titles
Internal linking does two jobs at once: it makes products reachable, and it tells a reader — human or machine — what the destination is about.
Descriptive anchor text. “Acme SlimWash 45 slimline dishwasher” carries meaning; “view product” carries none. Repeated across a category, the difference is substantial.
Title elements that name the product specifically. Brand, product name and the distinguishing variant attribute. Title templates that produce “Dishwashers | Buy Online | Retailer” on every page waste the single most-read line on your site.
Category and buying-guide pages that link into products. These give a crawler a route and give a reader context about how products relate to each other.
Related-product links. Comparable models, accessories and successor products establish relationships that would otherwise have to be inferred.
Headings that mark real structure. A specification section under a heading called Specifications is easier to extract than the same table under a styled div.
Stable product and variant identity
Identity is what allows your product to be recognised as the same product a shopper is looking at elsewhere. Without it, every other improvement is attached to something that cannot be matched.
Publish GTINs wherever the manufacturer assigns one. It is the only genuinely cross-retailer key.
Publish MPN or model number for categories where GTIN assignment is inconsistent.
Give each purchasable variant its own canonical URL where price or identifier differs.
Keep URLs and identifiers stable across site releases. Where a URL must change, use a permanent single-hop redirect to the equivalent product, not to the category page.
Never reuse a retired identifier. It corrupts every downstream record built from it.
The cost of getting this wrong is quiet. A retailer with strong specifications and no GTINs simply does not appear in comparisons that other retailers appear in, and there is no error message anywhere to tell them why.
Freshness, availability and price context
Downstream systems treat retailer prices as observations with a timestamp — simlir returns a price snapshot with the retailer it was seen at and an as_of date, for exactly this reason. That convention works in your favour, but only if your page is accurate when it is read.
Push price changes and stock-outs to the live page quickly. The gap between your system and your page becomes a wrong number in someone else's answer.
State availability explicitly rather than implying it through a button state.
Distinguish discontinued from temporarily unavailable. They deserve different treatment everywhere.
Keep discontinued product pages resolving sensibly — with a clear status, or redirected to a successor. Silent 404s remove accumulated context.
Show unit pricing where the category calls for it. It supports comparison across pack sizes and is a legal requirement in many UK categories.
Measuring referral and search visibility
Measurement is harder here than in classic search, but it is not the black box it is often described as. Two concrete instruments exist.
Signal
What it tells you
Limitation
Search Console generative AI performance report
How your content performs in Google's generative AI features on Search and Discover
Covers Google surfaces only; a site must also be included in Search generative AI features to be eligible at all
ChatGPT referral traffic
OpenAI appends utm_source=chatgpt.com to referral URLs, so ChatGPT-sourced sessions are directly attributable in your analytics
Requires OAI-SearchBot to be allowed; only captures click-throughs, not mentions without a click
Referrals from other assistant domains
Traffic arriving from named assistant products
Referrer data is inconsistent, and many sessions arrive with none
Direct and brand search volume
Whether awareness is rising overall
Correlational, and moves for many other reasons
Manual spot checks
Whether your products are described accurately right now
Not a sample, and results vary by user and session
Data completeness rates
Whether the inputs you control are improving
An input measure, not an outcome
The most useful metric is the least glamorous one: the percentage of products in a category carrying a GTIN, a durable hero image and the deciding specifications. It is entirely within your control, it moves when you do the work, and it is the input every downstream system depends on. Track it alongside the visibility signals rather than instead of them.
What cannot be guaranteed
Anyone promising you placement in AI search results is selling something they cannot deliver. Here is the honest boundary:
Inclusion is not guaranteed. No amount of structured data entitles a page to appear in any search feature or assistant answer.
Ranking is not purchasable or configurable. There is no submission process or setting that determines position.
Rich results are eligibility, not entitlement. Google states this directly about structured data.
Behaviour changes. Surfaces, formats and citation conventions are moving quickly, and a tactic tuned to today's behaviour may not survive the year.
Attribution will stay imperfect. Plan for directional measurement rather than precise channel accounting.
Things you can safely ignore
Google has been unusually direct about tactics that do not help on its surfaces, which is useful because several of them are actively sold as services:
llms.txt and similar files. Google Search does not use them. Maintaining one for other systems is fine and will neither help nor harm your visibility in Google Search.
“Chunking” content into small pieces. Not required. There is no ideal page length.
Rewriting pages specifically for AI systems. They understand synonyms and intent, so you do not need to capture every phrasing a shopper might use.
Overfocusing on structured data. It is not required for generative AI features. Keep using it because it supports rich-result eligibility in ordinary Search — not because it is an AI tactic.
Chasing inauthentic mentions. Spam systems and quality ranking both apply to generative features.
Set against that list, the recommendations in this article are deliberately boring: crawlable pages, explicit facts, stable identity, accurate markup, prompt freshness. That is the work.
What you can control is whether your products are readable, matchable and accurately described when they are read. That is a smaller claim than most of what is written about this subject, and it is the part that actually compounds — because the same work improves your own site search, your feeds, your comparison partners and your conventional SEO at the same time.
Frequently asked questions
Is there a way to submit my catalogue to AI search products?+−
Not in the sense of a submission that guarantees inclusion. Some platforms accept merchant feeds for specific shopping features, and those are worth using where they exist. But for generative answers drawn from the open web, the mechanism is the same as for search: your pages have to be crawlable, readable and accurate. There is no queue to join.
Should I block AI crawlers?+−
It is a commercial judgement rather than a technical default. Blocking gives you more control over how your content is used and reduces the chance of your products being described in assistant surfaces at all. Allowing improves the chance of accurate representation without guaranteeing inclusion. The failure mode is neither choice — it is inheriting a decision nobody made, from a robots.txt written years ago for a different web.
Does this replace conventional SEO?+−
No, and the overlap is larger than the difference. Crawlability, canonical URLs, descriptive titles, internal linking, accurate structured data and content that matches what shoppers ask about serve both. The additional emphasis for AI surfaces is on explicit facts over implied ones, and on stable identity that lets your product be matched to itself across retailers.
Do we need an llms.txt file?+−
Not for Google. Google Search states that it does not use llms.txt or similar machine-readable files, and that maintaining one will neither help nor harm your visibility there. If another system you care about consumes such a file, publishing one is harmless — just do not treat it as a substitute for crawlable pages, explicit facts and stable identifiers, which is what actually determines whether your products can be read.
How do I know if my products are being described accurately?+−
Spot-check them. Ask the assistants and search features your customers use about products in one of your categories, and compare what comes back against your live pages. It is not a statistically valid sample — results vary by user and session — but it surfaces systematic problems quickly, and misdescriptions almost always trace back to a specific field your pages left implicit.
We have thousands of products. Where do we start?+−
One category, chosen for commercial importance, and the identity fields first. Publish GTINs, give variants stable URLs, then add the two or three deciding specifications for that category with consistent names and units. Measure completeness before and after. A single category done properly gives you both a template and an evidence base for the next one.
A price without its currency, retailer and checked time cannot be explained. How to model price snapshots honestly, and what to do with stale or missing values.