Coverage and completeness are two different questions that buyers routinely merge into one, and merging them is how a dataset gets bought that answers the wrong question.
Coverage is presence. Which platforms, categories, geographies and time periods does the dataset contain at all? It is a boundary, and a boundary can be written down and checked.
Completeness is depth inside the boundary. Of everything that happened within a covered slice, how much does the dataset actually hold? It is a proportion, and proportions are much harder to establish than boundaries.
A dataset can be broad and shallow, narrow and deep, or any combination. None of those is inherently wrong. What is wrong is not knowing which one you have bought.
The four axes of coverage#
| Axis | The question | Where it usually breaks |
|---|---|---|
| Platform | Which marketplaces and social platforms are collected? | An emerging platform added years after the others, with no restated history |
| Category | Which product categories, at what depth of the classification tree? | Categories that do not map cleanly to a platform's own taxonomy |
| Geography | Which countries and, within them, which channels? | Cross-border and offline sales sitting outside the collection boundary |
| Time | How far back, at what consistency? | History assembled after the fact, or a methodology change mid-series |
The time axis deserves particular attention because it is the one buyers check last and regret first. Multi-year analysis assumes a consistent series. If a platform entered the dataset partway through, or a category definition changed, or a source was added and the history was not restated, then the series contains a step that has nothing to do with the market. Anyone computing growth across that step measures the vendor's operations rather than consumer behaviour. Ask directly: has this series been restated, and if so, when and why? Data Provenance and Due Diligence treats restatement policy as a first-class procurement question for exactly this reason.
Why completeness resists a single number#
Buyers frequently ask for one figure — "what percentage of the category do you cover?" — and vendors frequently supply one. The figure is rarely defensible, because computing it honestly requires knowing the true size of the category, which is the quantity the dataset exists to estimate. Either the vendor is citing an external benchmark, in which case the benchmark should be named and dated, or it is dividing by its own estimate, in which case the number is circular.
The defensible form of a completeness claim is structural rather than numeric:
- Named inclusions and exclusions. Which platforms are collected; which are not; which channels — offline, private, cross-border — sit outside the boundary entirely.
- A stated resolution floor. Below what level of sales, review volume or listing activity does a product stop being reliably represented? Every observed dataset has such a floor, and a vendor that claims none has not looked.
- Per-slice statements rather than aggregates. Coverage varies by platform, by category and by price tier. An aggregate figure conceals precisely the variation a buyer needs, and the variation is often large.
A vendor willing to state exclusions plainly is giving you more usable information than one quoting a high aggregate percentage, even though the first sounds more modest. Modesty that can be checked beats confidence that cannot.
Gaps are systematic, and that is the problem#
If coverage gaps fell randomly, they would be a nuisance: they would add noise, and noise averages out over enough observations. Real gaps do not fall randomly. They cluster where collection is hardest, which means they cluster by channel and by product type:
- Offline retail, which for many categories is still the majority of sales
- Private and social commerce — group buying, direct messaging, closed communities
- Cross-border purchases, recorded in neither the origin nor the destination market's retail data
- Unbranded and white-label goods, which resist brand attribution
- Bundles and multipacks whose contents are not itemised, so unit volumes are understated
- Products below a platform's listing or reporting thresholds
Each of these is a consistent, directional shortfall. The consequence is that errors in the dataset lean the same way every time, and collecting more data from the same source does not correct the lean. A category dominated by offline sales will look small in any online-observed dataset, no matter how much of the online portion is captured. That is not an argument against the dataset; it is an argument for knowing its shape before drawing a conclusion from it.
Testing coverage before you buy#
Four tests, all runnable on a sample dataset in a single working day. They pair with the broader set in Sample Data Verification.
1. The known-item test. List ten products you already know perform well in your category, deliberately including two or three challenger brands rather than only category leaders. Ask to see each in the sample, with history. Missing leaders reveal a category gap. Present leaders with missing challengers reveal a resolution floor — the more consequential finding for anyone whose actual question is about emerging competition.
2. The boundary test. Ask for the list of platforms and channels collected, then ask what proportion of your category's sales you believe happens outside that list. The vendor cannot answer the second half; you can, roughly, from your own commercial knowledge. If your estimate of the outside portion is large, the dataset is a partial view and should be used as one.
3. The back-period test. Repeat the known-item test for the earliest period you intend to analyse. Coverage today says nothing about coverage three years ago, and multi-year analysis depends entirely on the older end.
4. The classification test. Take a product whose categorisation is genuinely ambiguous — a hybrid product, a gift set, a subscription refill — and check where the dataset puts it. Then check whether it is put there consistently. Classification instability is invisible in aggregate figures and destroys any analysis that depends on a category boundary.
Coverage is a use-case question, not a quality score#
There is no coverage level that is correct in the abstract. A buyer sizing a national category needs breadth and can tolerate a resolution floor. A buyer benchmarking their own SKUs against three named competitors needs depth in one narrow slice and does not care what happens elsewhere. A buyer researching emerging entrants needs specifically the thing most datasets are worst at — resolution below the level of established brands.
So the right sequence is: write down the questions you actually need answered, derive the coverage those questions require, and evaluate vendors against that. Evaluating against a generic notion of "good coverage" produces a shortlist optimised for breadth, which is the dimension vendors market on and often not the dimension that decides whether the dataset answers your question.
Where to look next#
For what generates coverage in the first place, see Market Intelligence Data Sources. For how often a covered slice is rebuilt, see Data Refresh Frequency. For the full side-by-side framework, see Market Intelligence Platform Comparison Criteria and How to Evaluate Market Intelligence Providers.
Common questions#
What is the difference between coverage and completeness?#
Coverage is a question about presence: does the dataset contain this platform, this category, this country, this period at all? Completeness is a question about depth within something that is covered: of everything that happened in that covered slice, what proportion does the dataset actually hold? A vendor can have excellent coverage and poor completeness — the full platform list present, but only the best-selling products on each. It can also have narrow coverage and near-perfect completeness, which is often the more useful dataset, because a buyer can reason about a known boundary but cannot reason about an unknown shortfall inside one.
How can a buyer verify a coverage claim before purchasing?#
The cheapest reliable test is the known-item check, and it takes about an hour. Pick ten products you already know sell well in your category — ideally including two or three challenger brands rather than only market leaders — and ask the vendor to show each one in the sample data, with its history. Missing leaders indicate a category gap. Present leaders but missing challengers indicate a resolution floor, which matters most to anyone researching emerging competition. Then repeat the exercise for the earliest period you intend to use, because coverage in the current period says nothing about coverage three years ago.
Why is a single coverage percentage usually not meaningful?#
Because computing one requires a denominator the vendor does not have. To state that a dataset covers a given share of a category, someone must know the size of the whole category — which is the very thing the dataset was bought to estimate. A vendor quoting a single figure is either citing an external benchmark, which should be named, or estimating its own denominator, which makes the figure circular. The useful form of the claim is structural and checkable instead: which platforms are in, which are out, which channels are outside the collection boundary entirely, and where the resolution floor sits.
Are coverage gaps random or predictable?#
Predictable, almost always, and that is what makes them dangerous. Gaps cluster where collection is hard: offline retail, private and social commerce channels, cross-border purchases, unbranded and white-label goods, bundles and multipacks whose contents are not itemised, and products sold below the platform's listing thresholds. A gap that clustered randomly would add noise and average out across a large enough sample. A gap that clusters by channel or by product type introduces a consistent directional error, so every figure drawn from the dataset leans the same way — and no amount of additional data from the same source corrects it.