How to see if ChatGPT and Gemini cite your brand
by Crownler6 min read
Asking once is not measuring. The fixed question, the fixed model, the weekly run and the stored answer, and what to record from each one.

Open ChatGPT, type your category question, read the answer, feel good or feel bad. Almost every team has done this. It is not a measurement, and the reason matters more than the verdict.
Ask the same question twice and you get two different paragraphs. Ask it in your own account and the model may be carrying your history. Ask it tomorrow and the web search behind it hit a different set of pages. There is nothing to compare, because nothing was held still.
Here is how to turn that into something you can put in a report.
What a measurement needs
Four things, and all four are boring.
A fixed question. Written once, changed never, or changed with a note about the day it changed. The moment you tweak the wording you have started a new series.
A fixed model. "ChatGPT" is not a model. The answer from the free tier and the answer from the top tier differ, and so do two versions of the same family. Write down which one you asked.
A fixed cadence. Weekly is enough for most categories. Daily costs more and moves less than people expect.
The whole answer, stored. This is the one that gets skipped, and it is the one that makes the rest defensible. A stored answer lets you show a client the sentence that named a competitor. A score alone lets you show them a number they have no reason to trust.
The method, step by step
Pick between five and fifteen questions. Mix the intents: something informational ("what is X"), something comparative ("X versus Y"), something transactional ("best X for a small team"), something navigational (your brand name plus a qualifier). This is the same idea as treating the intention, and not the keyword, as the unit. Different intents retrieve different pages, and a set made only of comparisons will tell you your category is a review-site category when it is not.
Decide who counts as you. Your brand has spellings. The legal name, the short name, the product name, the way people misspell it. All of them count as one mention, otherwise the brand with more nicknames wins on nothing.
Decide who counts as a rival. Name them in advance. Discovering competitors from the answers is a separate, later exercise; if you leave the rival list open, every run produces a different denominator.
Run, store, and count three things. Whether the brand appeared. Where the first mention falls in the text, measured from the start. Which domains the answer cited. That third one is the finding that surprises teams most often.
Repeat the same run against a couple of assistants. They do not agree. We regularly see a brand named by one and absent from another for the same question in the same week, which is the whole argument for measuring more than one surface. Holding Google and losing ChatGPT for the same intention is an ordinary situation, not an anomaly.
The traps
Asking your own brand name. "What do you know about Acme" will produce a paragraph about Acme. Of course it will. That question measures the model's memory, not your visibility in a decision. The useful questions are the ones where your name is not in the input.
Believing the politeness. Assistants soften. An answer that says "Acme is one option among several" and an answer that opens with Acme are different results, and a yes-or-no mention count flattens them into the same thing. Recording where the first mention lands is a cheap way to keep that difference.
Measuring on a model your buyers do not use. If the person researching your category is on the free tier, that is where the measurement belongs. Choosing the expensive model gives you a friendlier chart about a population that does not exist.
Turning one run into a trend. Two data points are not a line. Give it a month before you tell anyone the number is moving.
Confusing mentions with traffic. A mention is presence in an answer. It is not a visit, and there is no report anywhere today that tells you how many people saw it. Anyone selling you an AI impressions volume is modelling. Say so out loud before someone else does.
What we record, per run
In Crownler this is a scheduled job, and it is worth describing precisely because the shape of what gets stored is the whole argument.
Once a week, each active question goes to ChatGPT, Claude and Gemini, with web search on, and the complete answer is written to the database. Alongside it we store the model that answered, the tokens, the cost of that call, and the exact time.
From the stored text, a deterministic pass extracts the mentions. Deterministic, meaning no second model is asked to judge: it is string matching with a tolerant word boundary, so "acme" does not match inside "acmetools" and a two-word nickname still matches across a line break. Overlapping hits at the same position count once, so a brand with a long name and a short nickname does not double its own score.
Each mention carries the brand, whether it was flagged as a rival, the snippet around it, the character offset of the first appearance, how many times it appeared, a sentiment read, and the URL attached to it when the answer attached one. The sentiment is a heuristic over the words in a hundred and twenty character window, and we describe it as a heuristic on the screen too, because it is good enough to separate praise from a complaint and not good enough to judge a brand.
Separately we store every source the answer cited, with its domain and its position in the list. Over a few weeks that turns into the most useful chart of the set: which handful of domains your category's answers keep coming back to.
Perplexity is in the data model and is not in the weekly round yet, because its key is not configured on our side. We would rather say that than run a call we know will fail and show you an empty column.
What to do with the result
Three moves, in order of how often they work.
Go after the cited sources, not the answer. If the same three domains feed every answer in your category, being absent from them is the finding. That is a diplomacy and press problem before it is a content problem.
Fix the page that should have been cited. Take the question, find the page on your site that answers it, and read the first two hundred words as a stranger. Most of the time the page answers the question in paragraph six, after the positioning.
Watch the intention across surfaces, not the keyword. The same intention has a front on Google and a front inside each assistant, and they move independently. Treating them as one number hides the half that is losing.
Every step above works with a spreadsheet and a calendar reminder. The tool exists because doing it by hand for forty questions across three assistants every week is where teams quietly stop.
Crownler asks the questions on a schedule, keeps every answer, and shows who was named beside you and which sources fed the answer. What the tool does · Plans · What generative engine optimization is, and what it is not · AI Overviews, measured