New template

6 min read

Track Brand Mentions in ChatGPT: A New Qoren Template

Our new agent template asks five AI engines your buyers' questions every week, counts who gets mentioned, and flags which changes are real. How it works.

Published by David Silva

A buyer asks ChatGPT for the three best tools in your category. Your name is not on the list. They never visit your site, so nothing in your analytics records that it happened. Today we are shipping a template for exactly this: the AI Visibility Tracker, an agent that asks ChatGPT, Claude, Perplexity, Grok and Gemini the questions your buyers ask, every week, and tells you where you stand and what moved.

Why can't you just check by hand?

You can, once. You type your category into ChatGPT, see your name, and feel fine. A colleague asks the same question an hour later and gets a different list. Generated answers vary from one run to the next, differ from one engine to another, and shift as the engines' models and search results change.

So a spot check is an anecdote. The real measurement is a rate: ask each engine the same set of questions several times and count. Twenty prompts, five engines, three samples each is 300 answers a week. Nobody keeps that tally by hand for long, and the week you stop is the week something changes. That tally is the homework. The agent does it.

We wrote the full method up separately, in a guide on how to measure AI search visibility, so you can run it yourself with a spreadsheet if you prefer. This post is about the agent that runs it for you.

What does the tracker actually do?

It works in four scheduled jobs and one conversation.

  • Setup, in your first chat. The agent drafts 10 to 30 buyer prompts from what your business sells and you cut and edit them. It records your brand with its aliases and domains, plus three to eight competitors. Then it shows you the number of calls per run and an estimated cost, and waits for your yes before the first collection.
  • Weekly collection, every Monday. It sends every prompt to every engine, three times each by default, with web search on. Every raw answer and every cited link is stored in a database on the agent's own machine, so history can be recounted later.
  • A catch-up run. If the collection did not finish, a second job finishes it rather than skipping the week.
  • The Monday brief. One page, deltas only, delivered to the messaging channel you connected and saved as a dated report.
  • A monthly model review. It checks whether a model it uses has been retired or superseded, and proposes up to three prompts to add or retire. It never switches a model without your explicit yes, because a switch breaks the week-over-week comparison.

Counting is done by a script, not by the model: plain alias matching against your brand and competitor list. The agent's job is the part a counter cannot do. Before the brief claims you won or lost a prompt, it reads the answers behind the change, checks how you were described, and fixes its own alias list when it finds a miss.

What does the Monday brief look like?

Here is an example with made-up numbers and a made-up rival called Rival Books:

  • Perplexity mentioned you in 27 of 60 answers, up from 15. Likely real: your new comparison page is cited in 8 of them.
  • ChatGPT: 19 of 60, was 21, within noise. Claude, Grok and Gemini: steady.
  • Lost: "best invoicing tool for freelancers" on Gemini, where Rival Books took the first slot in all 3 samples.
  • Factual error: Grok says your starter plan has no API access. It does. Quote saved.
  • Cited where you are absent: a G2 category page (14 answers), a Reddit thread (6).
  • Do next: claim and complete the G2 listing; publish a plain answer to the API question on your pricing page.

Three things in that brief are deliberate.

First, every change is labeled. The tracker runs a two-proportion test on each engine's mention rate against the previous week and calls a change "likely real" only when it clears two standard errors. With 60 answers, a move from 15 to 27 clears that bar; a move from 21 to 19 does not, so it is reported as noise rather than a trend.

Second, it names the cited domains in answers that left you out. That list is usually more useful than your mention rate, because it tells you where the engines get their shortlist: a review site, a comparison article, a community thread. Each one is a concrete move.

Third, it stops at three recommendations, each tied to a prompt, an engine and the evidence. Then it watches those prompts and tells you what happened, including when nothing moved.

What are the honest limits?

  • It measures the API, not the app. The agent queries each engine through its API with web search on. That gives you a consistent, repeatable panel, not a copy of what a buyer sees: the consumer apps can add memory, personalization and their own instructions. ChatGPT, for example, can use saved memories and past chats when a user turns that on. The gap is not small. A published comparison of scraped and API answers found the two methods named the same brands only about 16 to 24 percent of the time, though part of that is answers varying from run to run, and the study comes from a vendor that sells app scraping. So treat the numbers as a trend line on a fixed panel, not a screenshot of anyone's screen. The brief says so, once.
  • It does not change what the engines say. It tells you where you are missing and what to fix. The fixing, the pages and listings and corrections, is still work you or your team do.
  • Counting has edges. Alias matching can miss a paraphrase and can catch a brand name that is also a common word. That is why the agent reads the answers and re-scores history when it fixes an alias.
  • Model changes break the series. When a provider retires a model and you approve the switch, the next run is marked as a break so nobody reads the jump as a win.
  • It costs model usage. Cost scales with prompts times engines times samples. A per-run call cap enforces the budget you approved, and anything that would raise spend needs your yes first: more prompts, engines or samples, or a pricier model.
  • It is not for everyone. If your buyers never ask an AI assistant before choosing, classic search tracking is the better use of the money.

It is also quiet by design. A collection that worked ends without a message; it interrupts only when collection is broken or when an engine states a factual error about you on a prompt that matters. And it treats every AI answer and every cited page as data to report on, never as instructions to follow.

How do you deploy it?

Pick AI Visibility Tracker from the agent templates, choose an environment, name the agent, and deploy. Model access, including the calls to each engine, runs on your plan's managed key or your own OpenRouter key, injected automatically. Two keys are optional: Firecrawl, so the agent can read the pages the engines cite, and a Slack bot token if you want the brief in Slack.

From the CLI it is one command:

qoren agent create --env <environment-id> --template ai-visibility-tracker --name "AI visibility"

Then say hello. The agent asks what to call it and your timezone, builds the prompt list with you, shows you the calls and the estimated cost per run, and runs the first collection as your baseline once you approve. Your first week-over-week comparison comes with the next Monday brief.

If you want the longer version of the job first, read the AI search visibility tracking use case. For the vocabulary, we added glossary entries for generative engine optimization and AI share of voice.

Keep reading

  • AI search visibilityTrack whether ChatGPT, Claude, Perplexity, Grok and Gemini mention your brand when buyers ask, how that changes each week, and which sources they cite.
  • How to Measure AI Search Visibility: A Repeatable MethodMeasure how often ChatGPT, Perplexity and other AI engines mention your brand: a fixed prompt set, repeated samples, a noise test, and citations to act on.
  • Generative engine optimization (GEO)Generative engine optimization (GEO) is the practice of making a brand more likely to be mentioned and cited in answers from AI engines such as ChatGPT.
  • AI share of voiceAI share of voice is the portion of all brand mentions in AI answers to a fixed set of buyer prompts that go to your brand rather than to your competitors.
  • Agent templatesReady-to-deploy agents with a job and a clock.

Get the next one by email

Always On goes out every Friday: one idea worth keeping, one agent recipe you can copy, and a note from the field. Three minutes to read.

Subscribe