AI Visibility

You Asked ChatGPT About Your Brand and It Looked Fine — But That’s Not What Everyone Sees

Metricus Research · April 10, 2026 · 8 min read

AI recommendation lists repeat less than 1% of the time. SparkToro research across 2,961 queries found that ChatGPT, Claude, and Gemini almost never return the same brand list twice for the same prompt. A single spot-check of what AI says about your brand captures only a small part of the pattern.

The spot-check illusion

You typed your company name into ChatGPT. The answer looked reasonable — maybe it got your product category right, mentioned a couple of features, listed you alongside competitors. You closed the tab feeling fine about it. Then your colleague ran the same query from their laptop. Different answer. Different competitors listed. Your brand mentioned third instead of first — or not mentioned at all.

This is not a bug. This is how large language models work. Every response is generated fresh, and the output varies based on built-in randomness, context, timing, and infrastructure factors that no user controls. What you saw in your one query was one sample from a distribution of possible answers — and that distribution is far wider than most people realize.

The core problem: A single ChatGPT query about your brand is like checking the weather by glancing out the window once. You might catch sunshine, you might catch rain. Neither tells you the climate.

Why ChatGPT gives different answers every time

ChatGPT does not retrieve pre-written answers from a database. It generates every response from scratch by predicting what words should come next, based on patterns learned during training. This prediction process is inherently probabilistic — the model assigns probability weights to thousands of possible next tokens and samples from that distribution. The result is that two identical prompts can produce meaningfully different outputs.

Multiple technical factors compound this variability:

The net effect is that even when OpenAI sets temperature to zero internally, outputs are not guaranteed to be identical. A 2026 study documented that even under the most deterministic settings possible, LLM outputs still varied across runs. For consumer-facing products like ChatGPT, where temperature is set higher to make conversations feel natural, the variation is substantially greater.

What the research shows

The scale of this variability is not theoretical. Multiple studies have now quantified it.

SparkToro: 2,961 queries, less than 1% list repetition

SparkToro recruited 600 volunteers who ran 2,961 individual AI queries across 12 different prompts on three major AI platforms. Each prompt was run 60 to 100 times per platform to generate statistically meaningful samples. The finding: AI tools returned the same brand recommendation list less than 1% of the time. When it came to ordering, researchers estimated you would need roughly 1,000 runs before seeing two identical ordered lists.

Despite this extreme per-query variability, the research did find that AI platforms drew from a relatively consistent consideration set. For headphone recommendations, brands like Bose, Sony, Sennheiser, and Apple appeared in 55% to 77% of the 994 responses — but their position, the surrounding recommendations, and the framing changed nearly every time.

Washington State University: 73% consistency rate

A March 2026 WSU study tested ChatGPT by feeding it the same prompts 10 times each. The model produced consistent answers only about 73% of the time. It frequently contradicted itself, sometimes flipping answers on the same question across runs. When adjusted for random chance, the accuracy was just 60% better than guessing.

What this means for your brand

The implications for brand visibility in AI are significant. When you spot-checked ChatGPT and saw your brand mentioned, you were seeing one outcome from a slot machine that spins differently every time. Your colleague who got a different answer was seeing another equally valid pull.

This creates several concrete problems:

The real question is not “what did ChatGPT say?” It is “across 100 buyer queries on multiple AI platforms, how often does my brand appear, where is it positioned, and is the information accurate?”

How many queries it takes to see the real picture

SparkToro’s research suggests a minimum of 60 to 100 queries per prompt per platform to establish a statistically meaningful pattern.

But the number of queries per prompt is only one dimension. To understand your actual AI visibility, you also need:

The math adds up quickly. If you need 60+ queries across 20+ prompt variations across 3+ platforms, you are looking at 3,600+ individual AI queries to build a reliable picture of your brand’s AI visibility. This is not something you can do by hand over lunch.

What to do instead of spot-checking

The first step is accepting that your single query told you almost nothing. That is not a comfortable realization, but it is the necessary starting point.

From there, you need systematic measurement. Not one query, not five, not even twenty — but a statistically valid sample across the prompts real buyers use, on the platforms they actually use, through the interfaces they actually interact with.

What systematic measurement reveals that spot-checking cannot:

Metricus shows the answer AI gives when a customer asks for what you sell, with every business it names. Its AI Visibility Score shows how often AI mentions and recommends you across your customers' questions, today and with new text for your website.

Your one ChatGPT query was not wrong. It was just one data point out of thousands. The question is what the other thousands look like.

Frequently asked questions

Does ChatGPT give the same answer to everyone?

No. ChatGPT is probabilistic, not deterministic. Two people asking the exact same question will often receive different answers with different brand recommendations, different ordering, and different details. Research from SparkToro found that AI recommendation lists repeat less than 1% of the time.

Why does ChatGPT give different answers about my brand each time?

LLMs like ChatGPT use probabilistic token prediction, meaning each response is generated fresh. Built-in randomness (temperature), floating-point arithmetic, hardware parallelism, and batching effects all contribute to variation. Even at the lowest randomness settings, outputs are not fully deterministic.

How many times do I need to ask ChatGPT to get an accurate picture of my brand visibility?

Research suggests 60 to 100 queries per prompt per platform to establish a statistically meaningful pattern. A single spot-check captures only a small part of the pattern. What matters is the aggregate pattern across many queries, not any single answer.

Can I trust what ChatGPT says about my company?

Any single response from ChatGPT is unreliable as a measure of how your brand appears to AI users overall. The answer your colleague sees may be completely different from yours. What matters is the aggregate pattern across many queries, which requires systematic measurement rather than spot-checking.

Last updated: September 2026

Cite as: Metricus — You Asked ChatGPT About Your Brand and It Looked Fine — But That’s Not What Everyone Sees

What do your customers ask AI?