Research · Published September 2026
By Marin Tintchev, Founder, GeoSynapse & China GEO EU
Why AI Answers Change Between Runs — And What That Means for Measuring Visibility
A single test of a single AI system tells you very little. Here is the research behind why, and what it means for how visibility should actually be measured.
Why do AI answers change between runs?
Independent research on AI search visibility has found that results vary across repeated runs of the same question, even when nothing about the question or the underlying model changes[1]. This is not a claim about any one platform being unreliable — it is a property of how large language models are served at scale in production.
FACT
Is this a documented, verified phenomenon?
Yes. Analysis of production LLM inference has identified batch-level non-determinism in serving infrastructure as a documented source of output variation — distinct from, and larger in effect than, ordinary floating-point rounding differences. This means identical queries, submitted through the same interface, can return meaningfully different completions even when the system's temperature setting is set to zero[1].
INFERENCE
What does this mean for measuring AI visibility?
If a single query can return different results on different runs, then a single test of "does this AI system mention my company" is not a reliable measurement — it is one sample from a distribution of possible answers. A visibility measurement built on one run per platform, per question, risks reporting noise as signal in either direction: a false negative (missed on this particular run) or a false positive (mentioned only on this particular run, not representative of typical behavior).
How should a credible AI-visibility methodology handle this?
A credible AI visibility methodology has to run each question multiple times, across multiple platforms, before drawing a conclusion about whether — and how — a company is being surfaced. This is the basis for the repeated-run protocol described on our Methodology page, and it is the same discipline applied in our own AI Visibility Audit.
It also means we're cautious about any AI-visibility claim — ours or a competitor's — that is based on a single query result. A single favorable mention is not evidence of a stable pattern, in either direction.
References
- Schulte, Bleeker & Kaufmann, "Don't Measure Once: Measuring Visibility in AI Search (GEO)," arXiv:2604.07585 (2026).
Frequently asked questions
Does this mean AI-visibility testing is inherently unreliable?
No — it means single-run testing is unreliable. A methodology that runs each question multiple times across multiple platforms, as described in our Methodology, produces a stable enough signal to act on despite per-run variation.
Does answer variation affect all AI platforms equally?
The underlying serving-infrastructure cause (batch-level non-determinism) is a general property of how large language models are deployed at scale, not specific to one vendor — but this piece doesn't claim to have measured whether the magnitude of variation differs meaningfully between platforms; that would require platform-specific empirical testing we haven't published.
Should a single AI query result ever be trusted as a measurement?
Not on its own. Treat a single response as one data point, not a verdict — the same caution that applies to a single customer review or a single analyst quote.
About the author
Marin Tintchev is the founder of GeoSynapse and China GEO EU. He has 30+ years in ICT and international business across China and Europe, was educated at Tsinghua University, and is fluent in Mandarin.