How to compare AI traffic month over month without fooling yourself
Small numbers move a lot for reasons that have nothing to do with your work. Here is how to tell a real change from noise, and which comparisons are worth making at all.
Why small channels mislead
A channel that sent 120 sessions last month and 150 this month is up twenty-five per cent, and that headline is almost meaningless. Thirty sessions is within the range a small channel moves for no reason at all.
Percentage change on a small base is the single most misleading figure in analytics, and AI channels are small on most sites. The percentage looks dramatic precisely because the base is thin.
Report absolute numbers alongside any percentage. It is a small discipline that prevents a large class of confident wrong conclusions.
Hold the definition still
A comparison only means something if both months were measured the same way. Adding a host to your AI channel mid-quarter produces an increase that is entirely definitional.
Keep a written record of the definition — which hosts, whether bing.com is included, which source dimension, which channel group. When it changes, note the date and treat the series as broken at that point rather than pretending it continues.
This sounds bureaucratic until the first time someone asks why a number moved and the honest answer turns out to be that you changed the filter.
Compare the things that move for real reasons
Landing page mix is the most informative comparison available. If assistants started citing a different set of your pages, that is a genuine change in how you are being read, and it shows up long before totals move.
Conversion rate within the channel is next, provided you have the volume to support it. A stable rate with rising sessions is growth; a falling rate with rising sessions is usually a content or intent mismatch.
Crawler activity by page is the leading indicator. Assistants read before they cite, so a shift in what gets crawled tends to precede a shift in what gets cited by weeks.
Watch the calendar, not just the chart
Months have different numbers of working days, and B2B traffic follows the working week closely. February against January is a comparison of twenty working days against twenty-two before anything else is considered.
Holidays move between months across years, which makes year-on-year comparisons for the same month unreliable in exactly the periods people most want to compare.
Where the volume allows, compare rolling four-week windows instead of calendar months. It removes most of this and costs nothing but a slightly less familiar chart.
Before you attribute a change to something you did
Ask what else changed. A model version update, a shift in how an assistant formats citations, or a competitor publishing something can move your figures more than your own work does, and none of it appears in your analytics.
Ask whether the change is larger than the channel's normal variation. If you have not looked at week-to-week variation within a stable period, you do not know what normal looks like, and you cannot say whether this month is unusual.
And be willing to conclude that no effect was detectable. That is a legitimate result, it is far more common than published case studies suggest, and reporting it honestly is what makes the times you do detect something believable.
Establishing what normal variation looks like
You cannot judge whether a month is unusual without knowing how much the channel moves when nothing happens. That baseline is cheap to build and almost nobody builds it.
Take a stable period of eight to twelve weeks with no known changes, and look at week-to-week movement in sessions. The range you see is your noise floor, and any single month inside that range should be described as flat regardless of what the percentage says.
Rebuild the baseline after any definitional change, and after any large shift in the channel's size. A noise floor computed when the channel sent forty sessions a week does not apply once it sends four hundred.
Once it exists, the baseline changes how the conversation goes. Movement inside the range gets reported as flat, movement outside it gets investigated, and nobody spends a week explaining a fluctuation that was never a signal.
The comparison worth making instead
If month-over-month totals are too noisy to read, compare composition instead of volume. Which pages received the traffic, and how did that set change?
Composition is more stable than volume on small channels, and it is more actionable: a shift in which pages get cited tells you something you can respond to, while a shift in total tells you only that a small number moved.
Pair it with the crawl distribution from your logs. When both the crawled set and the landing page set move the same way, you are looking at a real change in how assistants read your site rather than at a fluctuation.
One last discipline, and it is the one most likely to be skipped: decide what would count as a real change before you look at the numbers. A threshold chosen in advance is a threshold; a threshold chosen after seeing the data is a justification.
This matters more here than in larger channels, because small numbers offer enough movement in either direction to support whatever conclusion somebody arrives wanting. Writing the threshold down first is the cheapest protection available against that, and it takes a sentence.
The same discipline applies to the comparison window itself. Choosing between month-over-month and rolling four weeks after seeing which one shows the result you hoped for is the same error wearing different clothes, and it is easier to spot in someone else's analysis than in your own. Fix the window when you set the threshold, and let the data say whatever it says.
And be sceptical of your own explanations for movement you did like. A month that went up gets attributed to the work; a month that went down gets attributed to seasonality. Applying the same standard of evidence in both directions is the whole discipline, and it is uncomfortable in exactly one of them.
Questions
- How much month-to-month movement is normal?
- It depends on your volume, and you can only find out by looking at variation within a period when nothing changed. Without that baseline you cannot tell an unusual month from an ordinary one.
- Should I compare calendar months or rolling windows?
- Rolling four-week windows where volume allows. Calendar months differ in working days and move holidays between years, both of which distort exactly the comparisons people most want to make.
- What if I cannot tell whether my change worked?
- Then say so. No detectable effect is a real result, it is common, and reporting it plainly is what makes your positive findings worth believing.