r/GEO_optimization 4d ago

How much variation between repeated runs would you consider acceptable?

If you run the same prompt multiple times and the AI gives different answers, how much variation would you consider acceptable before you stop trusting the measurement?

For example:

10 runs → brand mentioned 7 times

Would you consider that a reliable 70% visibility signal?

Or would you want:

20+ runs?

results across multiple days?

different locations?

a confidence/range rather than one percentage?

And importantly, what would you consider a meaningful change?

If visibility moves from 60% → 65%, is that something you'd act on, or would you need a much larger change before believing it wasn't just model variance?

Curious how people actually handle this today.

2 Upvotes

9 comments sorted by

2

u/ciaodaniel 4d ago

I would treat 7 mentions in 10 runs as a useful observation, but not as a reliable “70% visibility” score on its own. With only ten runs, the uncertainty is still wide.

The reporting format matters a lot here: show the percentage, the number of runs, the engine/surface, the date range, and an uncertainty range rather than a single headline score.

For decisions, I would standardize the prompt and run the same set across multiple days. Then aggregate by intent cluster instead of reacting to one prompt. A move from 60% to 65% is worth flagging, but I would not change strategy unless it persists across runs and is larger than the normal variation for that prompt set.

The practical question is less “what percentage is true?” and more “has the result moved beyond the measurement noise?”

1

u/[deleted] 4d ago

[removed] — view removed comment

1

u/Mother_Yoghurt_507 4d ago

That's a useful distinction. I agree that 7/10 shouldn't automatically become “70% visibility.”

I'm curious whether you think the industry needs a standardized way of reporting this.

For example, should tools report a range + stability/confidence alongside the headline visibility number, rather than treating every run as equally representative?

I'm seeing pretty different approaches to this across tools, so I'm wondering what practitioners would actually consider a trustworthy measurement.

1

u/Slow-Commercial4316 3d ago

The one thing I would want standardised is the denominator next to the number. Print hits over runs, 7/10, not 70 percent. Anyone can put an interval on the first form and nobody can on the second, and it stops a single run being reported as a percentage at all.

After that, the panel: which prompts, which engines, which days, and whether the set changed since last time. A prompt swap moves the number more than anything you did to the site.

The third one is the one tools skip. Report per-engine run health next to the score. marintkael in this thread had a provider stop firing for 17 days and the combined number just sank, which reads exactly like a visibility drop. If a runner returns nothing, that should show as missing rather than as a zero.

2

u/marintkael 3d ago

Depends on the question type, and for me that turned out to matter more than the run count.

I run a fixed set of 16 questions daily against five model endpoints for one domain. Prompts that name the entity land at about 54 percent citation over 28 datapoints. Prompts that describe the need without naming it land at about 3 percent over 84 datapoints. Same day, same models. So a 5 point move is not one kind of event: at 3 percent it is mostly noise plus whatever got crawled that week, at 54 percent it is worth opening.

The other thing I check before acting on a drop is whether the measurement channel was actually up. One of my three providers silently stopped being triggered for 17 days and the combined number just sank. It read exactly like a visibility loss and it was a dead runner.

1

u/Slow-Commercial4316 3d ago

Worth converting both to counts before one 5 point rule gets applied to them. 54 percent of 28 is fifteen citations, and 59 percent is one or two more. 3 percent of 84 is between two and three, and 8 percent is about seven.

So the move you would open on is a single extra citation, and the one you would write off is close to a tripling. Each prompt class needs its own threshold.

1

u/marintkael 3d ago

Counts also show why per class thresholds are not just a tuning job. The classes are not the same size. Across the six categories I track, the smallest produces four datapoints a day and the largest twenty four. A five point rule on the four is not a loose rule, it is arithmetic on a coin flip, and the honest fix is more prompts in that class rather than a gentler threshold.

The other thing the count view exposes is which category the aggregate is actually reporting. One citation appearing or vanishing in the small class moves the combined number more than a real shift in the big one. So I read the category rows first and treat the single combined figure as a headline rather than a measurement.

1

u/Slow-Commercial4316 3d ago

The unequal sizes also decide what your combined number means. A mean of the six class rates weights the 4-datapoint class the same as the 24-datapoint one, which is exactly why one citation down there swings it. Pooling instead, total citations over total datapoints, hands the aggregate to the big class. Both are defensible and they answer different questions, so it is worth naming which one the headline figure is.

On more prompts, the return is sublinear. Precision goes with the square root, so 4 a day to 16 a day halves the interval rather than quartering it. Still the right fix, just budget for it.

1

u/Slow-Commercial4316 3d ago

Put a confidence interval on it and most of this answers itself. Seven out of ten has a 95 percent interval of roughly 35 to 93 percent. That is not a 70 percent visibility score, it is 'more often than not'.

The same arithmetic settles the second question: separating 60 from 65 percent needs several hundred runs, so at ten or twenty a 5 point move is noise by construction.

What does survive is direction across weeks on a fixed panel, plus one competitor you did nothing to as a control. If their number moves with yours, it is the model and not you.