How to Check What AI Assistants Say About Your Products (Without Trusting a Single Screenshot)
Someone on your team asks ChatGPT for the best merino base layer, your brand is missing, and a screenshot lands in Slack. Ten minutes later a colleague asks the same question and your brand is second on the list. Both screenshots are true and neither is a measurement. AI answers vary from one run to the next, so the only honest way to know how assistants treat your products is to ask the same buying questions many times and count.
It is the same logic behind the dated baseline we capture at the start of a client engagement, adapted so an established Shopify brand can run it in-house: which questions to test, how to run them, what to record, where Google now gives you free first-party numbers, and how to read the results without overreacting.
Why is one AI answer not a measurement?
Because the list of recommended brands changes almost every time the question is asked. In a study published in January 2026, SparkToro and Gumshoe.ai had 600 volunteers run 2,961 prompts through ChatGPT, Claude and Google's AI Overviews and AI Mode, and found under a 1 in 100 chance of getting the same list of brands twice, and roughly 1 in 1,000 of getting it in the same order.
The study's categories included shopping questions such as chef's knives under $300 and headphones for travelling families, so it describes exactly the kind of answer your customers see. Its practical conclusion is the one to adopt: track how often you appear across many runs of many questions, not where you ranked in one answer.
Which questions should you test?
Test the questions a real buyer in a specific market would ask before choosing a product like yours, written in their words and with their constraints. Start with 15 to 30 questions for one product family and one market, spread across the four types below, and keep the list fixed so next month's results are comparable.
| Question type | Example | What it tells you |
|---|---|---|
| Category, with constraints | Best merino base layer for women under $120 that won't itch | Whether you are found when nobody names you |
| Comparison | [Your brand] vs [competitor] for winter hiking | How you are positioned against the brand shoppers compare you with |
| Brand fact | Does [your brand] run true to size? Where is it made? | Whether the facts an assistant repeats about you are correct and current |
| Policy and market | Does [your brand] ship to the UK, and are duties included? | Whether your shipping and returns terms reach the answer, per country |
Name the country in the question when the market matters, or run the checks from that market. An answer about US delivery tells you nothing about how a UK shopper is served.
How do you run the check?
Run every question several times in each assistant you care about, in fresh sessions, and record the same fields every time. Five runs per question per assistant is a reasonable first pass: enough to separate "always there" from "sometimes" from "never" without turning it into a research project.
- Fix the question list. Same wording every time, saved in a sheet.
- Use fresh sessions. A new chat each run, and a logged-out or clean profile where you can, so earlier conversations do not steer the answer.
- Repeat. At least five runs per question per assistant, on the same day.
- Record everything. Save the full answer text, not just whether you were named.
- Date it. Every run gets a date, so next month's comparison is fair.
What should you record, and how do you read it?
Record whether you were named, how you were described, which sources were cited and which competitors appeared. Then read the results as rates and patterns, not as individual answers. A description that is wrong three times in five is a real problem; a single odd answer is not.
| Field | Why it matters |
|---|---|
| Named or not | Your appearance rate per question and per assistant |
| Position | Useful only as an average; single positions are noise |
| Description accuracy | Wrong materials, discontinued products, outdated prices or policies |
| Cited URLs | Which of your pages, and which outside sources, the answer drew on |
| Competitors named | Who you are compared with, which is not always who you expect |
Patterns point to causes. Named when asked directly but never in category questions often points to weak facts or weak outside mentions for the category.
Described with last year's product line points to outdated pages or listings somewhere. Cited from a retailer instead of your own site suggests your own pages are missing the answer. The fixes are covered in why ranking first no longer means AI will recommend you.
Where can you get free first-party numbers?
For Google's AI features, Search Console now reports them. Google introduced generative AI performance reports in June 2026 and, as of 31 August 2026, rolled them out to all sites worldwide. They show impressions of your URLs in AI Overviews and AI Mode, with the pages, countries, devices and dates involved.
That covers Google's surfaces only, and it counts impressions, not whether you were recommended or described correctly. For other assistants, your analytics can show visits referred from chat assistants, and the manual method above covers the rest. When we run these checks for clients, our automated checks cover Claude and Google AI Overviews; other assistants are not part of that automated coverage.
How often should you repeat it?
Monthly for the agreed questions, and again after any change that could alter an answer: a new product line, a price change, a new returns policy or a move into a new market. Keep the question list stable so the months compare, and add questions rather than rewriting old ones.
Resist the urge to react to a single month. Answers move, models are updated and competitors change their pages too. A pattern that holds for two or three checks is worth acting on; a one-month dip usually is not.
FAQ
How many runs does a question need?
Five runs per question per assistant is a practical minimum for a first baseline. The SparkToro study found identical lists repeat under 1 time in 100, so a single run cannot tell you whether an appearance was typical or luck.
Should I test while logged in to my own accounts?
Preferably not. Assistants can take earlier conversations and personal context into account, and someone who works on your brand has asked about it many times. Fresh sessions and clean profiles give you something closer to what a new customer sees.
An assistant says something wrong about my product. Can I correct it?
Not directly in the answer. You correct the sources it draws on: your own product and policy pages, your Google & YouTube channel and catalog data, and outside listings such as retailers and marketplaces. Then check the same question again next month.
Which assistants should an established brand check?
The ones your customers use in the market you care about. For most brands that includes Google's AI Overviews and AI Mode, because they sit inside Google Search, plus ChatGPT and whichever others show up as referrers in your analytics.



