Market intelligence data sources are the raw inputs a dataset is built from. Six families dominate commercial practice, and each observes a different slice of a market with a different structural blind spot. The family a number came from determines what that number can honestly be used for, which is why "where did this come from?" is a first-order procurement question rather than a technical footnote.
Vendors rarely use one family alone. Most commercial datasets are composites, and a composite is only as sound as its weakest constituent and the rules used to join them together.
The six source families#
| Family | What it observes | Typical resolution | Structural blind spot |
|---|---|---|---|
| Marketplace observation | Public listing, price, position and review signals on e-commerce platforms | SKU, daily to monthly | Anything sold outside the observed platforms |
| Retailer and brand feeds | Point-of-sale or first-party sales records supplied under agreement | SKU, weekly | Only the retailers and brands that agreed to supply |
| Consumer panels | Purchases reported by a recruited, weighted sample of shoppers | Basket, weekly to monthly | Small-share brands sit below the sample's resolution |
| Survey research | Stated attitudes, intent and recalled behaviour | Respondent, per wave | Stated behaviour is not observed behaviour |
| Social and review corpora | Public posts, comments and product reviews | Document, daily to monthly | Only consumers who post, skewed by platform demographics |
| Trade and regulatory records | Customs declarations, filings, product registrations | Shipment or filing, monthly to annual | Says nothing about sell-through to consumers |
Marketplace observation#
The dominant source for online-heavy markets. A vendor observes what a platform publishes — listings, prices, positions, review counts, promotional state — and derives sales estimates from those signals. Its great advantage is resolution: it reaches individual SKUs and individual sellers, which no sampled method does. Its great limitation is boundary. It sees the platforms it observes and nothing else, so a brand's offline business, its private-channel business and its cross-border business are all outside the frame unless a separate source supplies them.
Marketplace observation is also the family where estimation matters most, because most platforms publish demand signals rather than sales figures. The conversion from signal to sales volume is a model, and models have assumptions. See Data Provenance and Due Diligence for how to interrogate one.
Retailer and brand feeds#
Sales records supplied directly, under a commercial agreement, by the party that made the sale. This is the highest-fidelity family available — it is the actual transaction record rather than an inference about it. Its limitation is the agreement: the dataset covers exactly the retailers and brands that consented to supply, and consent is not randomly distributed. Retailers that supply data tend to be larger, more organised and more modern than those that do not, so the resulting picture is systematically skewed toward the organised trade.
Consumer panels#
A recruited sample of households or shoppers reports its purchases, and the sample is weighted up to represent a population. Panels answer questions no observed source can: who bought, what else was in the basket, whether this purchase replaced another brand, and how often the same shopper returns. That is genuinely different information, not merely another route to the same sales figure.
The cost is sample resolution. A brand with a small share of a category may be represented by a handful of panellists, and an estimate built on a handful of panellists carries an error band wide enough to swallow the entire figure. Panels are dependable at category and major-brand level and become unreliable at the granularity where challenger brands live.
Survey research#
Respondents are asked what they think, intend or remember doing. Surveys are the only practical source for attitudes, unmet needs, brand perception and purchase intent — none of which leave an observable trace anywhere else. The standing caveat is that stated behaviour and observed behaviour diverge, sometimes sharply, and the divergence is not random: respondents over-report socially approved behaviour and under-report the rest. Treat survey output as evidence about perception, and use an observed source when the question is about what actually happened.
Social and review corpora#
Public posts, comments and product reviews, collected as text. This family answers why — the language consumers use, the attributes they raise unprompted, the problems they report after purchase. It is the natural input for Sentiment Analysis and Social Listening, and it is the only source that surfaces a concern before it shows up in sales.
Its skew is severe and should be stated plainly: the corpus represents consumers who post, on the platforms observed, in the languages processed. That population is younger, more urban and more opinionated than the buying population. Volume is not representativeness, and a corpus of millions of documents can still be a biased sample of a market.
Trade and regulatory records#
Customs declarations, company filings and product registrations. Slow, coarse and unusually reliable, because the records exist for legal reasons rather than commercial ones. Useful for import and export flows, for confirming that a company operates where it claims to, and for identifying products entering a market before they appear at retail. Useless for anything about consumer demand, because a shipment clearing customs is not a sale.
Composing sources: why one family is never enough#
A serious market question almost always needs at least two families. Sizing a category needs an observed source for volume and often a panel or survey to account for the part of the market the observed source cannot see. Explaining a change in volume needs a social or review corpus alongside the sales figure. Assessing a competitor's position needs both what they sell and how consumers describe them.
The composition itself is where quality is won or lost. Three questions decide whether a composite dataset is trustworthy:
- Can each number be attributed to its source? If a figure cannot be traced back to the family it came from, its error characteristics are unknown and it cannot be defended when challenged.
- How are incompatible units reconciled? A panel produces a weighted population estimate; marketplace observation produces an SKU-level count. Combining them requires an explicit rule, and the rule should be written down.
- What happens where sources disagree? Reasonable answers include preferring one family for one class of question, publishing both figures with their provenance, or reconciling through a documented model. An unreasonable answer is silence.
Where observation stops and modelling starts#
Every commercial dataset contains a boundary between what was observed and what was inferred, and the honest ones mark it. Common inference points:
- Converting a demand or position signal into a sales-volume estimate
- Grossing a sample up to a population
- Attributing an unlabelled product to a brand or a category
- Filling a period where collection failed
- Reconciling one platform's reporting conventions with another's
None of these are illegitimate. All of them are places where a number becomes an estimate, and a buyer is entitled to know which of the figures in front of them are which. A vendor that cannot draw the line between observation and inference has usually not drawn it internally either.
Questions to ask a vendor about its sources#
- For each metric we plan to use, which source family produces it?
- Where a figure is estimated rather than observed, what is the estimation method, and what is its expected error?
- What is the collection boundary — which platforms, which geographies, which channels are outside it?
- When two of your sources disagree, what is the published rule?
- Has the source mix for this metric changed during the history we would be buying? If so, when, and was the history restated?
The last question catches a specific and common problem. A vendor that added a source three years ago and did not restate the earlier history has a series with a discontinuity in it, and a buyer computing multi-year growth across that discontinuity will measure the source change rather than the market. Data Coverage and Completeness covers how to test for that, and Data Refresh Frequency covers the related question of how often each source is rebuilt.
Where to look next#
For the full comparison framework, see Market Intelligence Platform Comparison Criteria. For the procurement process built on top of it, see How to Evaluate Market Intelligence Providers. For testing a source before you buy it, see Sample Data Verification.
Common questions#
What is the difference between observed data and panel data?#
Observed data records events the vendor can see directly — a listing's price, a review being posted, a shipment clearing customs. Panel data records what a recruited sample of shoppers reports buying, then weights that sample up to a population estimate. The practical difference is where the error lives. Observed data is precise about what it sees and silent about what it cannot see, so its error is a coverage gap. Panel data speaks about the whole market but through a sample, so its error is sampling error and it grows as you drill into smaller brands and narrower categories. Neither is better in the abstract; they fail in opposite directions, which is why serious buyers ask which one a given number came from.
Why do two vendors report different numbers for the same category?#
Usually for one of four reasons, and it is worth establishing which before treating either number as wrong. First, different source families — one may be observing marketplace activity while the other weights a consumer panel. Second, different category definitions, which is the most common and least visible cause: "wellness beverages" is not a standard classification and each vendor assembles it from a different set of sub-categories. Third, different treatment of returns, cancellations and bundled items. Fourth, different estimation models where sales are not directly observable. Ask each vendor to state its category definition and its treatment of returns before comparing any figure.
Should a buyer prefer a single-source or a multi-source dataset?#
Multi-source, in almost every case, but with one condition: the vendor must be able to say which source each number came from. A composite dataset that cannot be decomposed is worse than a single-source one, because its errors are invisible and its blind spots are inherited from constituents the buyer never sees. The value of multiple sources is triangulation — sales figures say what happened, review and social corpora say why, and the two together answer questions neither answers alone. The risk of multiple sources is silent joining, where figures from incompatible collection methods are averaged into one number that means nothing precise.
How much does source choice actually change a conclusion?#
Enough to reverse it, routinely. A category that looks flat in a consumer panel can look like it is growing quickly in marketplace observation, because online-first challenger brands are often too small to register in a panel sample but are fully visible in listing-level data. The reverse also happens: a category with heavy offline sales looks smaller than it is in any online-observed dataset. This is not a flaw in either source — it is the source describing the part of the market it can see. The failure is using one source to answer a question the other source was needed for.