"Record a baseline" has become standard advice, which is progress. What has not become standard is any description of what a baseline contains, and the result is a lot of screenshots in a folder that nobody can do anything with six months later.
This is the structure we use, and the reasoning behind each part of it.
A baseline is an artefact, not a metric
The most common mistake is treating the baseline as a score. Someone runs a few prompts, concludes "we appear about 20% of the time", writes that down and moves on.
That number is nearly useless. It cannot tell you why, it cannot be audited, and when it changes you have no way to distinguish real movement from sampling variance.
A baseline is a saved set of full response transcripts, dated, alongside a structured reading of each one. The transcripts are the evidence. The structured reading is what makes it actionable.
The four columns
For every prompt, on every assistant, record these four things.
What to record per prompt
Columns three and four are the ones teams skip and later wish they had.
- 1
Appeared
Named at all, yes or no. The blunt binary.
- 2
Position
First, second, third, or mentioned in passing.
- 3
Reason given
The justification the model attached to each name.
- 4
Competitor set
Every other business the answer named.
Why the reason column is the most valuable
When a model names a competitor and explains why, it has handed you a specification.
"They publish transparent pricing." "They have detailed case studies." "They are certified by the relevant industry body." Each of those is a concrete, addressable gap. You do not have to theorise about what would improve your position — the answer is written in the transcript.
This is also the column that reveals when a model is working from stale or wrong information about you, which is a different and more urgent problem than absence.
If a model cites a reason for choosing a competitor that you also satisfy, the problem is not capability — it is that your capability is not legible. That distinction changes what you should fix.
Why the competitor column is a free market map
The set of businesses returned alongside you is a live statement of who a model considers comparable to you.
That set is frequently not the competitor list your sales team maintains. Sometimes it includes a company you have never heard of. Sometimes it excludes your biggest rival entirely, which tells you something about how the market is being categorised.
Recording it monthly also shows movement — a name appearing consistently that was absent a quarter ago is worth a look before it becomes entrenched.
Clean sessions, and why this is non-negotiable
Assistants personalise. Conversation memory, account history and location all shift responses toward things you have engaged with — including your own business.
Run a baseline while logged into your normal account and you will see yourself far more often than a stranger would. The resulting baseline is flattering, wrong, and worse than having none, because it produces confidence rather than uncertainty.
- Use a clean browser profile or a private window.
- Disable memory and personalisation features in the assistant settings.
- Do not prime the session with earlier questions about your industry.
- Where you sell into several regions, run from each region or use a location setting.
- Never edit the prompt after seeing a bad result — that is how a prompt set becomes a flattery device.
Sampling beats precision
Responses vary between runs. The same prompt asked twice within a minute can return different names.
This is not a defect in your method and it does not make measurement impossible. It means you report proportions across a set rather than a value for a single query.
Two runs per prompt is the practical minimum. Report as "named in seven of forty prompt-runs" rather than "we rank second". The first statement survives scrutiny; the second does not.
Building the prompt set honestly
The temptation is to write prompts you would win. It is a strong temptation because a good-looking baseline feels like a good start.
Write the questions your hardest prospect asks instead. Include the comparison question that names your strongest competitor. Include the price question. Include the one where a buyer asks what the downsides are.
Twenty to forty prompts across four intent bands — category, comparison, problem and local — is enough to see a pattern without becoming unsustainable to re-run every month.
Common mistakes
- Saving screenshots instead of full text, then being unable to search or compare them.
- Recording only whether you appeared, and losing the reason and competitor context.
- Running while logged in, producing a flattering and invalid picture.
- Changing the prompt set between measurements, which resets the comparison.
- Running each prompt once and treating the result as fact.
- Never dating the file, making a before-and-after claim unsupportable.
What good looks like after three months
Three dated folders, each containing the full transcripts for the same forty prompts across four assistants, plus a four-column summary sheet.
From that you can state, with evidence anyone can audit, how often you were named, whether the reasons cited changed, and which competitors entered or left the consideration set.
That is a genuinely defensible measurement practice, and building it costs an afternoon a month.