Why AI Answers Change: Variance Between Runs, Models and Locations

Updated · Citepoint · 6 min read

Short answer

AI answers change because the model samples its words at random, the engine runs fresh web searches that it writes itself, different products use different models, and location, context and the web all shift over time. In SparkToro's study of 2,961 runs, AI tools gave the same list of brands twice less than once in 100 times. So never judge a brand on one answer. Ask repeatedly and measure how often it appears.

Key takeaways

  • Even with randomness set to zero, Anthropic says results aren't fully deterministic, and OpenAI says determinism isn't guaranteed even with a fixed seed.
  • Engines often search the web as they answer, with queries they write themselves, so the pages they read differ from run to run.
  • Lists and order are unstable, but how often a brand appears holds up well enough to measure: one brand appeared in 85 of 95 Google AI responses in SparkToro's data.
  • Over a month, 40% to 60% of the domains cited for the same prompts were different in Profound's data, so compare periods, not single answers.

Why does the same question get different answers?

Four things move an AI answer between runs, and a fifth moves it over time.

Cause What happens Evidence
Sampling The model picks each next word from a probability distribution, so wording and list order vary Thinking Machines Lab, Anthropic and OpenAI documentation
Search The engine writes its own searches, often several, and reads different pages each time Google's documentation on query fan-out
Product and model Google AI Overviews and AI Mode may use different models and techniques Google's documentation, and an Ahrefs study of 540,000 query pairs
Location and context Results can be refined by location, and Google announced personal context for AI Mode OpenAI's web search guide, Google's I/O 2025 announcement
Time The web, the index and the models keep changing Ahrefs and Profound studies

Is the randomness built in?

Yes. Thinking Machines Lab, an AI research company, explains that getting a result from a language model "involves 'sampling', a process that converts the language model's output into a probability distribution and probabilistically selects a token."

Turning the randomness down doesn't remove it. Anthropic's API documentation says that "even with temperature of 0.0, the results will not be fully deterministic." OpenAI says of its seed parameter that "Determinism is not guaranteed." In Thinking Machines' test, 1,000 completions of the same prompt at temperature 0 produced 80 unique answers, and the most common one appeared 78 times. They traced the main cause to server load, which changes how requests are batched. You can't tune any of this from inside the apps your buyers use.

Why do the searches differ from run to run?

Answers about products often draw on live search. Google's documentation says AI Overviews and AI Mode "may use a 'query fan-out' technique" that issues multiple related searches across subtopics and data sources to build a response. OpenAI's documentation describes reasoning models that can run web searches as part of their chain of thought, "analyze results, and decide whether to keep searching."

So each run can take a different path: different searches, different pages, different sources. Ahrefs found that between consecutive observations of a Google AI Overview, only 54.5% of cited URLs overlapped on average.

Do different products and models give different answers?

Yes, even inside Google. Its documentation says "AI Mode and AI Overviews may use different models and techniques, so the set of responses and links they show will vary." An Ahrefs study of September 2025 US data found the two cited the same URLs only 13.7% of the time across 540,000 query pairs, yet reached similar conclusions, with 86% semantic similarity across 730,000 pairs. If a brand was mentioned in an AI Overview, there was a 61% chance it also appeared in AI Mode.

Google also says AI Overviews "often don't trigger", so sometimes there is no answer at all. The practical rule: measure each engine separately, and count a missing answer as neither a win nor a loss.

Do location and context matter?

They can. OpenAI's web search tool accepts an approximate user location, by country, city, region or timezone, to refine results. Google announced that AI Mode would offer "personalized suggestions based on your past searches" and let people opt in to connect Gmail for more personal context. Two buyers asking the same question can therefore see different answers.

That is why Citepoint asks every question from your country and in your language, and says openly that it doesn't model logged-in personalization, memory or follow-up questions.

How fast do answers change over time?

Quickly. Ahrefs tracked more than 43,000 keywords, each with at least 16 recorded AI Overviews, over a month. An AI Overview had a 70% chance of changing from one observation to the next, and its content lasted 2.15 days on average. The meaning stayed close, with an average cosine similarity of 0.95, even as the wording and sources changed.

Profound compared the same prompts between June 11 to 13 and July 11 to 13, 2025, with about 80,000 prompts per platform. It found that 59.3% of the domains cited by Google AI Overviews in July weren't cited in June for the same prompt. The figure was 54.1% for ChatGPT, 53.4% for Microsoft Copilot and 40.5% for Perplexity. Comparing January with July pushed the figures to 70% to 90%. Both Ahrefs and Profound sell AI visibility tracking, so treat their studies as vendor research.

How different are two answers, really?

SparkToro and Gumshoe had 600 volunteers run 12 prompts through ChatGPT, Claude and Google's AI Overviews or AI Mode, 2,961 runs in all, with each prompt run 60 to 100 times, in November and December 2025. Their findings:

  • There was a "<1 in 100 chance" that ChatGPT or Google's AI would give the same list of brands in any two responses.
  • Order was even less stable, closer to 1 in 1,000.
  • The number of items on a list ranged from 2 or 3 to 10 or more.

But the same data shows what holds steady. A digital marketing agency appeared in 85 of 95 Google AI responses. A Los Angeles hospital showed up in 69 of 71 ChatGPT answers, a 97% visibility rate, yet was the top mention in only 25 of them. Rand Fishkin concluded that visibility percentage across dozens to hundreds of prompts, each run multiple times, is a reasonable metric. Rank position is not. One of SparkToro's co-researchers works at Gumshoe, an AI tracking vendor, which the post discloses.

How many times should you ask?

Enough that chance stops dominating. For a brand named in 30% of answers, the 95% margin of error is about 20 points with 20 runs (anywhere from 10% to 50%), about 9 points with 100 runs (21% to 39%) and about 4.5 points with 400 runs. The formula comes from standard statistics: z × √(p(1 - p) ÷ n), valid when you have more than five answers that name you and more than five that don't. SparkToro suggests at least 60 to 100 runs per prompt.

How should you sample?

  1. Fix the question set. Changing questions changes the number for reasons unrelated to your brand.
  2. Fix the market and language, and keep each engine separate.
  3. Ask on a schedule, daily for the engines your buyers use most and weekly for the rest.
  4. Pool answers over rolling periods, such as the last 7 days against the 7 before.
  5. Keep the answer text, so you can see which sources changed when a result moves.

What should you do when an answer changes?

Don't react to one answer. Look at the trend over weeks, then check which sources the engine cited when your result moved. If a competitor appeared, find the page or review that brought it in. If an answer says something wrong about you, see how to fix what AI says about your brand.

Frequently asked questions

Is the variation a bug the AI companies will fix?

Not in any way that helps you. Sampling randomness is part of how the models work, and answers also change because the web, the search results and the models themselves change. Plan to measure rates across many answers.

Does using incognito or clearing history make answers consistent?

No. Ahrefs captured two AI Overviews for the same query two minutes apart in incognito mode and found the phrasing and content differed. Clearing history removes personal context but not the randomness or the fresh searches.

How many times should I ask each question?

SparkToro suggests at least 60 to 100 runs per prompt to see an AI's set of recommendations. With 100 runs, a 30% rate carries a margin of error of about 9 points. More questions across more days also add samples.

Why does my screenshot show me first when tools say I'm rarely named?

A screenshot is one sample. SparkToro found that even a brand named in 97% of answers was the top mention in only 25 of 71. Position moves a lot between runs, so count how often you appear instead.

Do all AI engines vary equally?

No. In Profound's data, the domains cited by Perplexity drifted less over a month (40.5%) than those cited by Google AI Overviews (59.3%). Each engine searches differently, which is why they should be measured separately.

Sources

  1. SparkToro: AIs are highly inconsistent when recommending brands or products (Jan 2026)
  2. Ahrefs: AI Overviews Change Every 2 Days (But Never Change Their Mind)
  3. Ahrefs: Are AI Mode and AI Overviews Just Different Versions of the Same Answer? (730K responses)
  4. Profound: AI Search Volatility, why AI search results keep changing
  5. Thinking Machines Lab: Defeating Nondeterminism in LLM Inference
  6. Anthropic API reference: Create a Message (temperature)
  7. OpenAI API OpenAPI specification, openapi.yaml (seed parameter)
  8. OpenAI API docs: Web search
  9. Google Search Central: AI features and your website
  10. Google: AI Mode in Google Search, updates from Google I/O 2025
  11. OpenStax, Introductory Statistics 2e: A Population Proportion
  12. Citepoint: How we measure AI visibility (methodology)