Data Quality

A Privacy Statistic Without Its Sample Size Is Decoration

By Rawsoft Team | September 2026 | 8 min read

The omission

Most privacy and tracking statistics arrive as a percentage and nothing else. Some share of websites do this. Some proportion of consumers feel that. No sample size, no collection window, no description of the population it was drawn from.

The figure gets quoted, then quoted from the quote, and within a year it is an established fact with no way back to whatever produced it.

The four-second test

Two questions, asked of any statistic, before anything else. What was the n, and when was it collected?

Without an n you cannot judge precision. A difference of four points means one thing across 50,000 sites and nothing at all across 200.

Without a window you cannot judge relevance. Tracking behaviour is not a constant. A figure collected before a major browser change, a platform policy shift or a regulation taking effect is describing a web that no longer exists, and nothing about the number itself will tell you that.

The third question nobody asks

Population. "10,000 websites" means nothing until you know how those ten thousand were selected and from what frame.

Ten thousand sites drawn from a traffic ranking is a study of large sites. Ten thousand drawn from a vendor's own customer base is a study of that vendor's customers. Ten thousand drawn from one industry in one country is a study of that industry in that country. All three can honestly be described as ten thousand websites, and all three would produce wildly different answers to the same question.

Selection is where most of the distortion actually lives, and it is almost never described.

Ours, in the format we are arguing for

So that this is not a lecture delivered from behind a hedge, here are our own figures with their denominators. From the Website Privacy Index, read on 20 September 2026:

Note how much work the word detected is doing in the fourth line. It is not a synonym for absent, and we keep it there deliberately.

The part that cost us something

Here is why we think this format is worth the trouble.

We rebuilt this index. The crawl ran on, the sample grew from roughly eighteen thousand scored sites to more than forty-seven thousand, and every headline figure moved. Some moved by several points. Posts we had already published carried the sample they were written from, and those numbers no longer describe the index.

We want to be careful about what that does and does not mean. The earlier figures were not wrong when they were collected. They were correct measurements of a smaller sample. What changed is the sample, and therefore what the figures describe.

The point is that the drift was visible at all. It was visible because we had published denominators, so a reader could line the old n against the new one and see exactly what had changed. Had we shipped bare percentages, nobody could have caught that, including us. This is the same class of problem as a historical number that moves when you re-run the query, which we wrote about in state is not history.

Why the incentive runs the wrong way

A checkable statistic can be falsified. An uncheckable one cannot.

So the less rigorous format is also the safer one commercially, which is worth naming without moralising about it. Publishing an n is choosing to be catchable. Most organizations, reasonably enough, would rather not be, and there is rarely any consequence for that choice.

The rarest disclosure is the most valuable one

Re-running the same method on a later sample, and saying what moved.

Almost nobody does this, because a figure that moves looks like an admission. It is the opposite: a figure that has never been re-measured is a figure nobody knows the shelf life of. If a number moved on a re-run, say so, and say by how much. If it has never been re-run, that is also worth saying.

Three questions for your next meeting

  1. What was the sample, and from what population was it drawn?
  2. When was it collected?
  3. Has the method been re-run since, and did the figure move?

None of them are hostile, all three are answerable in a sentence by anyone who has the answer, and the silence when nobody does is informative in itself.

The ask

This is not a request to distrust numbers. It is a request to be able to check them. A statistic that cannot be checked is not a weaker claim than one that can. It is a different kind of object, and it belongs on a slide rather than in a decision.

Our free privacy scan runs the same method on any domain in about a minute, with no account. Run it on your own site, or book a data and tracking audit if you want the whole tag layer reviewed rather than the front page.

Not legal advice. Rawsoft determines what a system technically does: which tags fire, when, under what consent state, and what data is transmitted. We do not determine which laws apply to your organization, how a regulator would read them, or whether your organization is in compliance. Those are determinations for your counsel. Everything above describes behavior an automated scan observed on public pages at the time of the scan, and nothing in it is a legal conclusion about any site.

About Rawsoft

Rawsoft is an Atlanta-based digital data agency specializing in analytics implementation, privacy compliance, and media tracking for enterprise brands.

More from the blog

Data Quality
The Date Column Is Lying To You

A record stamped with today's date can be mostly months old. Partial updates rewrite the timestamp but not the fields, and the row that results passes every data quality check while quietly destroying period-over-period comparison.

August 2026 Read →
Data Quality
State Is Not History

Our scan chart showed fewer domains for August 23 than we recorded on August 23. Nothing was deleted. A store keyed by identity rewrites its own past every time a row is updated, and most dashboards never notice.

August 2026 Read →