Your buyers ask ChatGPT or Perplexity for a shortlist before they search, and your analytics only see the clicks, never the answers that left you out. You can measure it, but not with a screenshot. This is the method: which questions to ask, how many times, which numbers to keep, how to tell a real change from noise, and what to do with the result. It works with a spreadsheet and patience, or with an agent doing the asking for you.
Ask the AI engines your buyers use a fixed set of buyer questions, several times each, on a regular schedule, and count. Report the share of answers that mention you (mention rate), your share of all brand mentions against named competitors (share of voice), and the domains cited in answers that left you out. Call a change real only when it is larger than normal run-to-run variation.
Unit of measure
A mention rate over many answers, never one answer
Prompt set
10 to 30 buyer questions, each with a permanent id
Samples
At least 3 per prompt, per engine, per run
Cadence
Weekly runs, a monthly review of models and prompts
The numbers worth keeping
Six numbers per engine, per run. Every one is computed from the same stored answers, so you can recompute history when you fix an alias.
Number
How to compute it
What it tells you
Mention rate
Answers that name you, divided by all answers
How often you are in the conversation at all
Share of voice
Your mentions, divided by mentions of every tracked brand
Your slice of the shortlist against named competitors
Rank
Your position among tracked brands, by first mention
Whether you are the first name or an afterthought
Citation rate
Answers that link to one of your domains
Whether the engine uses your pages as a source
Cited where you are absent
Domains cited in answers that did not mention you
Where the engines get their shortlist, and where you are missing
Prompts gained and lost
Prompts where you went from zero mentions to some, or back
Which buyer questions moved, so you know where to read
Why one spot check tells you nothing
The usual first attempt is to type your category into ChatGPT, see your name, and feel fine. Then a colleague asks the same thing an hour later and gets a different list. Generated answers are sampled, so the same question can produce different brands on different runs, and each engine retrieves and ranks sources its own way. The consumer apps can also personalize: ChatGPT, for example, can use saved memories and past chats when you turn that on. A single answer is an anecdote about one run on one account.
Measure rates, not yes or no: "mentioned in 7 of 9 answers" survives a rerun, "it mentioned us" does not.
Measure every engine your buyers use separately. Visibility in one says little about another.
Use a clean, repeatable setup rather than your own logged-in account, so last week and this week are comparable.
The prompt set is your keyword list and your time series at once, so it deserves more care than anything else here. Write 10 to 30 questions the way a real buyer types them before choosing, not the way your marketing team would phrase them. Give each one a permanent id. When you reword a prompt, it becomes a new prompt with a new id, because a reworded question is a different measurement and quietly swapping it breaks every trend line that ran through it.
Category questions, where you have to earn the mention: "best invoicing tool for freelancers".
Comparison and alternatives questions: "X vs Y", "alternatives to a competitor".
Problem questions that lead to your category: "how do I stop chasing late invoices".
A few branded questions, where the test is accuracy rather than presence: "does X have an API".
Weight the set toward category questions. A prompt that names you will mention you, which proves nothing.
Step 2: name the field and choose the sampling
List your own brand with every alias and domain it goes by, then the three to eight competitors a buyer would actually weigh you against. Count mentions for all of them, because your number alone has no context: a mention rate of 30 percent is weak if two rivals sit at 80 and strong if nobody else clears 20. Then pick the sampling. Ask each prompt at least three times per engine per run, with web search turned on so the engine answers from what it retrieves today rather than only from training data. If you query through an API rather than the consumer app, say so in every report: the API gives you a consistent panel for trends and sources, not a copy of a buyer's screen, and API and app answers can name quite different brands.
Calls per run are prompts times engines times samples: 20 prompts, 5 engines and 3 samples is 300 answers.
At 3 samples, 20 prompts gives you 60 answers per engine, enough to compute a rate you can compare.
Aliases matter. A brand that is also a common word needs a stricter match or it will be over-counted.
Store every raw answer and every cited link, not only the tally, so you can recount after fixing an alias.
This is the step most reports skip, and it is where most false wins come from. With 60 answers per engine, a move from 12 mentions to 18 looks like 50 percent growth, but a standard two-proportion test puts it at about 1.3 standard errors: well within the variation you get by rerunning the same week. A move from 15 to 27 is about 2.3 standard errors, which is worth calling a trend. A simple rule holds up: treat a change of two standard errors or more as likely real, and label everything smaller as noise, however good the story sounds.
Keep the series comparable. Changing the model behind an engine, the sample count or a prompt breaks comparison; note the date and mark the break in the next report.
Read the answers behind every change before reporting it. Counting misses paraphrases and catches false positives.
Check how you were described, not only whether you were named. A mention with the wrong price or a dead feature is worse than silence on a money prompt.
Keep your own change log: every page you shipped or listing you claimed, with its date, so you can line it up against the numbers.
Step 4: act on the citations
The most useful output is not your mention rate. It is the list of domains the engines cite in answers that left you out. Those pages are where the shortlist comes from: a review site category, a comparison article, a community thread, a documentation page the engine keeps misquoting. Each one points at a concrete move. Keep recommendations few and specific, tie each one to the prompt, the engine and the evidence, and label a guess as a guess. The research paper that coined the term generative engine optimization found that some content changes raised visibility in its benchmark, which is encouraging, but there is no switch that guarantees a mention.
A review or directory site keeps getting cited: get listed there, with a complete and accurate profile.
A comparison page wins a prompt you lose: publish your own plain, honest comparison for that question.
An engine states a wrong fact about you: answer the question directly on your own site, where it can be retrieved.
No more than three recommendations a week, then watch the prompts they targeted and report what happened, including when nothing moved.
None of the method is hard. The problem is the repetition: five engines, twenty prompts, three times each, every week, with the tally, the noise test and the reading done before Monday. That is the homework, and it is exactly the kind people drop after the third week. Qoren's AI visibility tracker template runs this method as an always-on agent: it collects on a schedule, stores every answer and citation, flags which changes are likely real, and sends a one-page brief with at most three recommendations. The fixing is still yours.
Between 10 and 30 is the useful range. Fewer and one odd prompt swings the whole number; more and the cost and the reading grow faster than the insight. Pick the questions that carry commercial weight, weighted toward category questions that do not name you.
Run OpenClaw or Hermes without managing infrastructure.
Deploy a managed agent environment, configure the runtime, and keep the agent online without Docker, VPS setup, or server maintenance.