Skip to main content
Market Intelligence Wiki

E-Commerce Transaction Data

Last updated August 2026

Definition

E-commerce transaction data records what was bought online — product, price, quantity, time. It is observed rather than reported, which makes it precise about what it sees and silent about the rest.

E-commerce transaction data is the record of what was actually bought through online channels: product, price, quantity, time. It is the substrate under most modern category analysis, and its usefulness and its limits come from the same property — it is observed, not reported.

Nobody was asked to remember a purchase or estimate a share. Within its boundary the data is precise in a way survey and panel methods cannot match. Outside that boundary it is silent, and it will not tell you that it is silent.

What it contains#

Field Typically observable Notes
Product identity Yes At variant level — see SPU vs SKU
Price Yes Selling price, not list price, is the useful one
Time Yes Down to day, sometimes finer
Quantity / volume Often derived Most platforms do not publish it
Buyer identity No Which is why repeat behaviour is hard
Channel Yes But only the channels observed

The third row is the one that decides how a dataset should be read.

Where observation stops and modelling starts#

Most platforms publish demand signals — listing position, review accumulation, promotional state — rather than sales figures. Sales volume is therefore commonly derived from those signals through a model.

This is legitimate and unavoidable. What is not legitimate is presenting a modelled figure in the same visual register as an observed one, with nothing marking the difference. A buyer who cannot tell which numbers on a dashboard were counted and which were inferred has no basis for deciding how much weight any of them can carry.

So the question to ask of any transaction dataset is not "is this real data?" but "which of these figures are observed, which are estimated, and by what method?" Data Provenance and Due Diligence treats that as a first-order procurement question for exactly this reason.

Its structural blind spots#

Because the data is generated by observing specific channels, everything outside them is absent — and the absences cluster:

  • Offline retail, still the majority of sales in many categories
  • Private and social commerce — group buying, direct messaging, closed communities
  • Cross-border purchases, recorded in neither the origin nor the destination market
  • Unbranded and white-label goods, which resist brand attribution
  • Bundles and multipacks not itemised, understating unit volume
  • Products below listing or reporting thresholds

These are not random gaps that average out. They cluster by channel and by product type, which means the errors lean the same way every time. A category dominated by offline sales will look small in any online-observed dataset regardless of how completely the online portion is captured. See Data Coverage and Completeness.

What it is unusually good for#

Set against those limits, transaction data does several things nothing else does:

Granularity. It reaches individual products and variants. Sampled methods cannot, because a small brand may be represented by a handful of panellists or none.

Challenger visibility. Because there is no sample floor, brands too small to register in a panel are fully visible. For anyone researching emerging competition this is the decisive advantage, and it is why a category can look stable in panel data while restructuring underneath.

Price reality. Actual selling prices, which is what Price Band Analysis needs and what list-price sources cannot supply.

Speed and consistency. Observed continuously and on one definition, so periods are comparable within a platform.

What it is bad for#

Anything about people. It rarely links purchases to a buyer, so penetration, repeat rate and basket composition are either unavailable or estimated. See Category Penetration and Repurchase and Panel vs Census Data.

Total market size. It measures a portion of a market and cannot establish what portion without an external denominator.

Why anything happened. It records the outcome and contains no account of the reason. That is what review and social text is for — see Review Mining.

Using it well#

  1. Establish the boundary first. Which channels, which geographies, from when.
  2. Ask which figures are observed and which are modelled, per metric.
  3. Do not read it as total market unless an external denominator has been supplied and named.
  4. Pair it with a source that covers people or reasons — a panel for who, review text for why.
  5. Check when each platform entered the series. A platform added mid-history creates a step that is not a market movement; see Tmall vs JD vs Douyin.

Where to look next#

For the full source taxonomy, see Market Intelligence Data Sources. For the sampled alternative and how the two fail differently, see Panel vs Census Data. For the headline metric built on this data, see GMV. For coverage limits, see Data Coverage and Completeness.

Common questions#

What exactly is e-commerce transaction data?#

It is the record of purchases made through online channels, at the level of the individual transaction or the individual product: what was bought, at what price, in what quantity, and when. Its defining property is that it is observed rather than reported — nobody was asked to recall or estimate anything. That gives it a precision no survey or panel can match within its boundary, and it is also the source of its central limitation: it contains only what happened inside the channels being observed, and says nothing at all about the rest of the market.

How is it different from panel data?#

Panel data follows a recruited sample of shoppers and weights them up to represent a population, so it can describe the whole market and can link purchases to the same buyer over time. Transaction data observes activity directly within specific channels, so it is precise about those channels and blind beyond them, and it usually cannot tell you that two purchases came from the same person. The practical division: transaction data answers what sold, at what price, at what granularity; panel data answers who bought, how often, and what else was in the basket. Neither substitutes for the other.

Is transaction data actual sales or an estimate?#

Both, depending on the figure, which is why the distinction has to be marked. Some values are directly observable — a listed price, a review count, a product's presence and position. Sales volume frequently is not, because most platforms do not publish it, so it is derived from observable demand signals through a model. That derivation is legitimate and unavoidable, but it means a category total is usually part observation and part estimate. A buyer is entitled to know which figures are which, and a dataset that presents both in the same visual register without distinguishing them is withholding something material.

What are the main blind spots?#

Everything outside the observed channels, and the gaps are systematic rather than random. Offline retail, which in many categories is still the majority of sales. Private and social commerce — group buying, direct messaging, closed communities. Cross-border purchases, recorded in neither market's retail data. Unbranded goods that resist brand attribution. Bundles whose contents are not itemised, which understates unit volume. And products below a platform's listing thresholds. Because these cluster by channel and product type rather than falling randomly, they produce a consistent directional error that more data from the same source does not correct.

Talk to a Market Intelligence Specialist

Get a sample of our market intelligence data — covering your category, your platforms, your markets.

/ 50 characters minimum

For a faster response, use your work email. We never share it — by submitting, you agree to our Privacy Policy .