AI responses vary by design. Model temperature introduces deliberate randomness, personalisation adjusts for account history and location, and models update continuously. Two people asking an identical question minutes apart can receive materially different answers with different sources.
This is why single spot-checks are worthless as measurement — one query is one sample of a distribution.
Reliable measurement requires repeated sampling of the same prompts over time.
It also means a competitor 'appearing above you' in one screenshot proves nothing.
