How to Measure Your Brand’s Visibility in ChatGPT and Perplexity Without Any Special Tools
AI visibility auditing is the process of checking whether answer engines like ChatGPT and Perplexity mention your brand, cite it with a live link, or surface competitors instead. You can audit your brand’s AI visibility this week with nothing more than a spreadsheet and two browser tabs. By the end, you will know which buyer questions mention your brand, which ones cite you with a live link, where competitors are outranking you, and what public evidence they have that you do not.
This matters because the major answer engines still give founders almost no native reporting. Ahrefs reported in 2026 that ChatGPT provides no impressions data, no analytics dashboard, and no built-in record of what it says about you or competitors; the same analysis said only about 20% of ChatGPT mentions include a clickable citation link that would ever appear in GA4. Ahrefs: How to Track ChatGPT Traffic
The gap is large enough to justify manual checks. Demand Gen Report found that 89% of B2B buyers use generative AI during the buying process, and Ahrefs reported that ChatGPT handles more than 2.5 billion prompts a day. One free AI visibility grader from Profound published benchmark snapshots showing an average score of 31 out of 100, while stronger brands scored above 80. Demand Gen Report: 2024 B2B Buyers Survey Report Ahrefs Profound GEO Grader
What will this audit show you?
A simple manual audit reveals four things: whether your brand is mentioned, whether it is cited with a clickable source, which competitors appear instead of you, and which public proof signals those competitors have. That distinction matters because a mention shows awareness, while a citation shows the model found a source strong enough to attach to a claim.
You do not need paid software to get a useful first read. You need a spreadsheet, access to ChatGPT and Perplexity, and 30 to 50 realistic buyer questions that reflect how your customers actually research.
Step 1: How do you build a prompt set of 30 to 50 real buyer questions?
Your prompt set should cover branded, category, and comparison searches because buyers do not ask only one kind of question. A balanced set gives you a clearer picture than testing only your company name.
- Create a spreadsheet with these columns: Query, Query Type, Platform, Mentioned?, Cited with Live Link?, First Brand Named, Competitors Mentioned, Notes.
- Add 10 to 15 branded queries using your company name, product names, and founder names if buyers use them publicly.
- Add 10 to 15 category queries describing the problem you solve without naming your company.
- Add 10 to 20 comparison queries that include your brand versus a competitor, or “top tools” style questions buyers ask before purchase.
Use plain buyer language, not SEO language. Good prompts sound like “What are the best tools for measuring AI visibility in B2B SaaS?” or “How does [brand] compare with [competitor] for tracking AI citations?” rather than keyword fragments.
A worked prompt set for a two-person SaaS company
Take a small B2B software company with one core product and three visible competitors. A practical 30-query set would include 10 branded queries such as “Is [brand] good for AI search tracking?”, 10 category queries such as “How do startups measure visibility in ChatGPT?”, and 10 comparison queries such as “[brand] vs [competitor] for GEO reporting.”
Success at this checkpoint looks like a list that reflects the questions a buyer would ask before a demo, not the terms your team wants to rank for. If your list feels like marketing copy, rewrite it.
A concrete example using Profound’s public benchmark
Profound’s public GEO Grader gives one grounded way to sanity-check your manual results. In the benchmark examples shown on its free grader, the average AI visibility score sits at 31 out of 100, while stronger brands score above 80, which is a spread of at least 49 points. That kind of gap is large enough that a 30 to 50 prompt manual audit usually shows the underlying pattern fast: weak category presence, weak citations, or both. Profound GEO Grader
Step 2: How do you run each prompt in ChatGPT and Perplexity?
Testing the same prompt across both platforms shows where your visibility is platform-specific. That matters because Perplexity and ChatGPT reward different kinds of evidence and do not cite at the same rate.
Run every prompt once in ChatGPT and once in Perplexity. Use a clean browser state if possible: sign out of other business accounts, avoid custom GPTs, and do not use a conversation thread that already contains brand context.
- Open a fresh spreadsheet row for each prompt-platform pair.
- Paste the exact prompt into ChatGPT.
- Record whether your brand is mentioned at all.
- Record whether your brand is cited with a clickable source link.
- Record the first brand named in the answer.
- List any competitors mentioned.
- Repeat the same process in Perplexity with the exact same wording.
Do not rewrite prompts mid-test. If you change wording, you are no longer comparing platforms fairly.
Step 3: How do you score mention, citation, and absence correctly?
Use three labels only: Mention, Citation, and Absent. A mention means the model names your brand in the answer; a citation means it attaches a clickable source to your site or another source clearly about you; absent means your brand does not appear at all.
Keep mention and citation separate in your sheet. A brand can be mentioned without a citation, and that weaker signal is common in ChatGPT, where Ahrefs reported that only about 20% of mentions include a clickable citation link. Ahrefs
- Mention: Your brand name appears in the answer text.
- Citation: The answer includes a live, clickable source tied to your brand or a page about your brand.
- Absent: Your brand is not named anywhere in the answer.
If the engine cites a review site, analyst page, or news article about your company, count it as a citation. If it only says your name with no source, count mention only.
Step 4: Track who gets named first
First mention matters because answer engines show position bias similar to traditional search. If a competitor appears first across many category and comparison prompts, that brand is getting the attention advantage even when your company is still listed later.
Add a “First Brand Named” column and fill it for every result. This takes seconds per query and gives you a powerful pattern to review later.
Research from SE Ranking, published in 2026, found that Perplexity attributed claims to a specific source in 78% of tested complex research questions versus 62% for ChatGPT. Those figures came from a side-by-side prompt comparison on research-style queries rather than consumer shopping prompts, so use them as a platform behavior signal, not as your own benchmark. SE Ranking: ChatGPT vs Perplexity
Step 5: Calculate your baseline visibility
Your baseline is the percentage of prompts where your brand is mentioned, cited, and named first. Those three numbers together are more useful than a single “visibility score” because they separate awareness from source-backed authority.
Create three simple formulas for each platform:
- Mention rate = prompts where your brand was mentioned divided by total prompts.
- Citation rate = prompts where your brand had a live citation divided by total prompts.
- First-name rate = prompts where your brand was named first divided by total prompts.
Add one more view: category-level performance. You may find that branded visibility is strong, while category visibility is weak and comparison visibility is dominated by one rival.
Success here looks like a table you can explain in one sentence, such as: “We are mentioned in branded prompts, mostly absent in category prompts, and cited rarely in ChatGPT.” That tells you what to fix first.
If you want to compare your manual baseline with a tool-generated diagnostic later, use the same prompt categories and definitions so the results stay comparable. For example, a GEO scoring tool can be useful for trend tracking, but the manual audit still helps you see the exact prompts, citations, and competitor evidence behind the score.
Step 6: Why should you inspect competitor evidence to diagnose bias?
The practical question is not “What trick did they use?” but “What public evidence did the answer engine find for them that it could not find for us?” Competitor bias usually reflects evidence gaps, not hidden prompt magic.
Check the pages, sources, and proof signals attached to competitors that keep appearing. In many cases, the answer is visible on the page: original research, comparison pages, third-party reviews, pricing transparency, strong product definitions, structured data, or repeated citations from authoritative sites.
- Open the cited competitor pages for the prompts where you were absent.
- List what type of evidence each page provides: comparison content, customer reviews, research data, product specs, pricing, schema markup, press coverage, or glossary-style definitions.
- Compare that list with your own public footprint.
- Mark the missing items in your spreadsheet.
A study discussed by Search Engine Journal in 2026 reported that adding authoritative citations to web content improved LLM extraction likelihood by roughly 30% to 40% in the tested setup. That figure came from observed extraction-rate differences between pages with and without explicit authoritative references, not from a universal ranking factor. Search Engine Journal: LLM Citation Study
Step 7: Turn the findings into a fix list
Your next actions should come directly from the evidence gaps you found. The goal is to publish or improve the pages that answer engines were already looking for, using facts, sources, and comparisons a model can extract cleanly.
- Create or improve one comparison page for each rival that repeatedly outranked you.
- Publish one category-defining page that answers the broad buyer question where you were absent most often.
- Add structured data where it fits your content model.
- Strengthen pages with named sources, real statistics, and outbound citations to primary references.
- Make key facts easy to extract: product definition, use case, target customer, pricing model, integrations, and implementation details.
If your team already uses tools for schema deployment, semantic content creation, or perception analysis, use the audit findings to prioritize where those tools should be applied first. The manual review should drive the fix list, not the other way around.
A good fix list is short and specific. “Write better content” is not a task; “publish a [your brand] vs [competitor] page with pricing model, setup time, export options, and sourced comparisons” is.
What if your brand is never cited even when it is mentioned?
If your brand is mentioned but rarely cited, the engines likely know your name but do not find enough source-worthy evidence to attach to it. That usually points to weak public proof, not a discovery problem.
Start with the pages closest to purchase: homepage, product page, comparison pages, documentation, and review profiles. Tighten factual claims, add named sources where appropriate, and make sure important pages are public, crawlable, and specific enough to quote.
If the issue appears to be structural rather than editorial, review whether your important pages expose clear Organization, Product, and FAQ signals. In some cases, schema improvements can make citation-ready facts easier for AI systems to interpret.
Troubleshooting common problems
The answers keep changing between runs
That is normal. Models are probabilistic and retrieval layers change, so use one run per platform for your baseline, then repeat the full test monthly instead of trying to force identical outputs in one session.
I only tested my company name and the results look good
That can produce a false sense of security. Buyers spend much of their journey in category and comparison research, and Demand Gen Report found that 89% of B2B buyers now use generative AI during purchasing, so branded prompts alone miss the higher-risk part of visibility. Demand Gen Report
A competitor is cited from review sites, not from their own website
Count it anyway. If third-party pages are supplying the evidence layer for that competitor, they are part of why the answer engine trusts and surfaces that brand.
What the finished outcome looks like
At the end of this process, you should have a spreadsheet that shows where your brand is visible, where it is invisible, who gets named first, and what evidence gap explains the difference. That gives you a practical roadmap for content, comparison pages, reviews, and citation-ready facts.
The useful outcome is not a vanity score. It is a ranked list of missing public proof signals you can fix in order, starting with the buyer questions that matter most to revenue.
FAQ
How often should you repeat this AI visibility audit?
Run the full audit once a month and again after any major content, pricing, or review change. A 30 to 50 prompt set across 2 platforms creates 60 to 100 observations per cycle, which is enough to spot directional movement without overreading day-to-day model variation.
How many prompts are enough for a reliable first audit?
Thirty prompts is enough for a first baseline if they are split across branded, category, and comparison intent. Fifty prompts gives a stronger read because it produces 100 platform-level observations when you test each query in both ChatGPT and Perplexity.
Does this method work for local or niche businesses?
Yes, but the prompt set needs tighter scope. For a local business, a 30-query set might include 10 location prompts, 10 service prompts, and 10 comparison prompts so you can see whether the engine cites you for the exact metro area and offer that drive revenue.
What is a good citation rate in ChatGPT or Perplexity?
A good citation rate is one that rises from your own baseline, because platform behavior varies by query type. As a reference point, Ahrefs reported that only about 20% of ChatGPT mentions include a clickable citation link, while SE Ranking reported source attribution in 62% of tested complex questions for ChatGPT and 78% for Perplexity, so expectations should differ by platform and prompt type. Ahrefs SE Ranking



