What a consideration set is, and why a position in an AI answer is not a real number
Ask an assistant the same question five times and you get five slightly different lists. That is not a flaw in your measurement. It is the thing being measured, and it determines the only honest way to report it.
The experiment worth running yourself
Take a question a buyer in your category would actually ask. Ask it five times in fresh sessions. Write down which companies are named each time.
Almost every category produces the same shape of result: a small group named nearly every time, a larger group named sometimes, and a long tail named once. The lists overlap without matching.
This takes twenty minutes and it changes how you read every visibility claim you will ever be shown, including ours.
What a consideration set is
The consideration set is the group of companies a model is willing to name for a question. Membership is probabilistic: a company appears in some proportion of runs rather than occupying a fixed place.
So the meaningful measurement is a share — appeared in eleven of twenty runs — reported with the number of runs it is based on. That single figure carries both the estimate and its reliability, which is why it is the form worth insisting on.
Two companies at eleven of twenty and nine of twenty are not meaningfully apart. Two at nineteen and three are. The share makes that visible; a position in a list hides it.
Why an ordinal position misleads
A position implies a stable ordering that survives asking again. In search results that assumption roughly holds. In model answers it does not, and the same question minutes apart reorders the list.
Reporting a position from one reading takes a sample and presents it as a fact. Worse, it invites a comparison over time — we moved from fourth to second — where both readings are single samples and the difference is inside the noise.
This is the reason our own product refuses to produce that number, and refusing it costs us something. A single position is easier to put in a report and easier to celebrate. It is also the figure most likely to collapse under a follow-up question.
How many runs are enough
Five runs tells you whether you are in the set at all, which is often the only question that matters.
Twenty runs gives a share stable enough to compare between periods, provided the question and the sampling stay identical.
More than twenty buys precision you probably cannot act on. The interesting movements in this measurement are large — from absent to present, or present to absent — and detecting a five point change in share is rarely worth what it costs to detect.
Reading anyone's visibility claim
Three questions dispose of most claims in this category. How many runs is this based on? Which model and which version? Was the question asked with browsing enabled or not?
A vendor that cannot answer all three is reporting a sample as a fact. That is not necessarily dishonesty — it is often just how the tool was built — but it means the number carries less than it appears to.
Apply the same three questions to us. A measurement claim that cannot survive them does not deserve to be believed, whoever is making it.
Designing a question set worth measuring
The questions you measure determine everything downstream, and most sets are chosen badly. A question containing your brand name tests whether the model knows you exist, which is rarely the question anyone needs answered.
Category questions — the ones a buyer would ask before knowing any vendor — are what reveal whether you are in the set at all. Those are harder to write because they require describing your category the way a buyer would rather than the way you would.
Aim for questions specific enough to have a real answer. Best software is too broad to produce a stable set from any model; software for a stated task in a stated situation produces something you can measure and act on.
Fix the set and leave it alone. Changing questions between periods produces movement that reflects your editing rather than your visibility, and it is the easiest way to accidentally manufacture a result.
Why this shapes what you should buy
Once you accept that presence is a distribution rather than a position, most of the category's reporting starts looking overconfident, and that is a useful lens when evaluating any vendor.
The questions to ask are the same three every time: how many runs, which model version, and browsing enabled or not. A tool that cannot answer them is not measuring badly so much as measuring something less specific than it appears to.
Apply the test to us as readily as to anyone else. A measurement claim that cannot survive those three questions does not deserve to be believed regardless of who makes it.
There is a practical consequence worth stating for anyone about to buy something in this category. If a tool reports a position and cannot tell you how many runs it is based on, you are being sold a single sample presented as a fact, and the price is the same as for a tool that samples properly.
The distinction is invisible in a demo, because a single reading and a well-sampled share look identical on a dashboard. It becomes visible the first time a figure moves sharply for no reason anyone can explain, which is usually a quarter or two after purchase.
The practical defence is to ask the three questions during evaluation rather than after purchase. A vendor that answers them clearly is telling you something real about how the product was built, and a vendor that deflects is telling you something too. Neither answer takes long to obtain and both are more informative than any feature list.
Questions
- Why does the same question give different answers?
- Because model outputs are probabilistic and the retrieval behind them varies between runs. Two identical questions minutes apart can name different companies, which is a property of the system rather than an error in your measurement.
- How many runs should a visibility figure be based on?
- Five to establish presence, around twenty for a share stable enough to compare across periods. Any figure quoted without a run count should be treated as a single sample.
- Is a position in an AI answer ever meaningful?
- Only within a single reading, and it does not survive to the next one. A share of runs with the count attached carries the same information without implying a stability that does not exist.