There is a great deal of confident writing about ChatGPT optimisation and not much of it separates what is known from what is inferred. This guide tries to hold that line, because the difference matters commercially: work built on speculation is expensive and unfalsifiable.
What is actually known
ChatGPT answers questions from two sources: parametric knowledge acquired during training, and live retrieval when the question needs current information or the model judges retrieval useful.
That second path is the one you can influence in a reasonable timeframe. Content published this month cannot be in a training set that closed earlier, but it can absolutely be retrieved.
OpenAI publishes and documents its crawlers, which means access is a decision you control explicitly rather than something you have to guess at.
Two paths to being mentioned
They operate on completely different timescales.
Training data
- Fixed at training time
- Cannot be influenced quickly
- Governed by GPTBot access
- Favours long-established sources
- No way to correct errors directly
Live retrieval
- Reflects current web content
- Influenceable within months
- Governed by OAI-SearchBot access
- Favours clear, corroborated sources
- Corrections propagate as sources update
Crawler access — check this first
Three user agents are relevant and they do different things. GPTBot collects data that may be used for training. OAI-SearchBot powers ChatGPT's search results. ChatGPT-User handles fetches triggered by a user in conversation.
They are controlled independently in robots.txt, which means you can allow search visibility while declining training use, if that is your position.
The common failure is inheriting a restrictive robots.txt from a template or an old agency, discovering months later that the site has been blocking the crawler the entire time, and having spent the intervening period on content that was never reachable.
Audit your robots.txt before commissioning any AI visibility work. It is a five-minute check that occasionally invalidates an entire proposed programme.
The four things that improve your odds
Ordered by leverage. None of these is exotic; the difficulty is doing them consistently.
- Findability — the content must be crawlable, indexed and served as parseable HTML rather than locked behind client-side rendering.
- Extractability — each commercial page must contain a complete answer in a self-contained passage.
- Verifiability — claims must carry sources, and facts about your business must match what independent sources say.
- Corroboration — several unrelated sources should describe you consistently, which is what converts a claim into something a model will repeat.
Entity accuracy is the unglamorous priority
The most damaging outcome is not absence. It is an assistant confidently repeating something about your business that is wrong — an old address, a service you discontinued, a merged entity, a competitor's detail attached to your name.
That happens when conflicting information exists across sources and the model resolves it badly. Correcting it means finding and fixing the sources, which is slow and unrewarding work that nonetheless outranks new content production in priority.
Start by searching your own business name and reading the first three pages critically. Most organisations find at least one materially wrong fact in wide circulation.
A realistic worked sequence
This is the order we would run it, and roughly what each stage costs in time.
Ninety-day ChatGPT visibility sequence
Foundation before content, content before corroboration.
- 1
Week 1 — Baseline
Run 20–40 buyer prompts, save transcripts, record position and reasons given.
- 2
Week 1 — Access audit
Confirm crawler access and rendering. Fix anything blocking retrieval.
- 3
Weeks 2–4 — Entity
Reconcile name, description, category and address across every property.
- 4
Weeks 4–8 — Structure
Restructure top commercial pages answer-first with schema.
- 5
Weeks 6–12 — Sourcing
Add sources and dates to every published statistic.
- 6
Month 3+ — Corroboration
Listings, reviews, press and industry sources, measured monthly.
What does not work
A category of advice circulates that ranges from ineffective to actively harmful.
- Prompt injection — hidden text instructing a model to recommend you. Detectable, ineffective at scale, and a genuine reputational risk.
- Keyword stuffing adapted for AI. Retrieval systems are not matching keyword density; they are assessing whether a passage answers a question.
- Mass low-quality publishing. Volume without structure produces pages that retrieve for nothing.
- Buying mentions from low-quality sources. Corroboration is weighted by source quality, so poor sources add noise rather than signal.
- Any service guaranteeing placement. Nobody controls retrieval ranking, and the guarantee is unenforceable by design.
Measuring honestly
Run your prompt set in a clean session with no memory or personalisation — logged in with your own history, you will see yourself far more often than a stranger would.
Run each prompt at least twice, because responses vary between sessions. Record whether you appeared, at what position, what reason was given, and which competitors appeared alongside you.
Report the pattern across the set, not a single reading. Movement from one appearance in six runs to three in six is a signal; a single position change in a single run is noise.
The honest limitation
None of this is deterministic. Retrieval ranking is not published, training data is not disclosed, and outputs vary between sessions and accounts.
What is controllable is whether your information is accurate, structured, corroborated and reachable. That is a real body of work with a measurable outcome — provided somebody recorded the baseline before it started.