Retail Intelligence

Syndicated Data Providers: What Each Method Can Actually See

14 min read

A proposal lands with “syndicated data” in the first line, a six-figure annual figure near the bottom, and almost nothing in between about how the numbers were produced. You are supposed to know already. Most people don’t, because the term describes how the data is sold rather than how it was gathered.

So, the short version. Syndicated data is research a provider funds and collects once, standardizes into a single format, and licenses to many subscribers. That shared cost is the whole point of the model. It is also why the design is fixed before you arrive: nobody asked which stores, which weeks, or which definition of price mattered to you.

Which means “syndicated” tells you the business arrangement and says nothing about the instrument. Two datasets can both be syndicated, both cover beverage alcohol, both arrive weekly, and be built from measurements that have almost nothing in common. One counts what a recruited household says it bought. Another counts what a register rang up. A third counts cases leaving a warehouse. They disagree, and they are supposed to, because they are watching different events.

We’re gmware, a software and data engineering firm in Austin, TX, with delivery centers in Bangalore and Mohali, India. We build the warehouses brands land these feeds into, and we run Shield Suite, our own retail-intelligence platform covering 60,000+ storefronts. So we sell one of the methods below. We’ll tell you which one and where it stops.

MethodWhat it countsBlind where
Household panelPurchases a recruited shopper reportsThe store itself
Retailer censusRegister transactions from participantsNon-participating stores
Distributor depletionCases shipped to an accountPast the back door
Invoice reconciliationWhat a retailer was billedConsumer price
Storefront observationShelf state at one momentUnits sold
Field-rep auditA completed store visitStores off the route

What “syndicated” actually means

The textbook definition is narrower than the way vendors use the word. Syndicated data is “external and secondary data supplied in a standardized format,” made available “to multiple subscribers known as a syndicate,” and held in a common database, so “the information is not tailor made to meet the needs of any particular company” (MBA Knowledge Base). Market-research firms describe the funding structure the same way: the study is paid for by the research firm, which owns the data and sells it, and because “the same research findings are bought by multiple buyers, this spreads out the cost” (Drive Research).

Two consequences follow, and buyers underrate both.

The first is what the research literature calls the data-fit problem. Measurement units, category definitions, geography and recency were all decided by someone solving a general problem, not yours. If your brand lives in a state where the class of store that may sell your product is unusual, a standard geography can flatten exactly the thing you needed to see.

The second is that exclusivity is gone by design. Your competitor can buy the same file, read the same cut, and reach the same conclusion in the same week. That is not a defect. It does mean a syndicated subscription is table stakes rather than advantage, and any edge you get has to come from something you did to the data or something you observed that the syndicate didn’t.

Where each instrument sits on the chain

Product travels from a distributor warehouse, through an invoice, onto a shelf, across a register, and into a household. Every method below is a sensor bolted to one of those five points. Once you see it that way, the disagreements stop being mysterious.

Household panel

A provider recruits a sample of households, keeps them enrolled over long periods, and has them record what they bought. NIQ describes Homescan as a “static longitudinal consumer panel” spanning “more than 250,000 households across 25 countries,” with UPC-level product granularity (NIQ). Collection mechanics vary by country and by vintage: a scanner in the kitchen drawer, an app, sometimes a survey layered on top.

The unit is a purchase occasion as the shopper reported it, weighted and projected up to a population. Collection runs continuously; the release you read is periodic.

The blind spot is the store. A panel is the only method here that can tell you who your buyer is, what else was in the basket, and whether they came back eight weeks later. It cannot tell you what the shelf looked like when they decided, because the shelf was never the thing being measured. Panels also thin out fast under fine cuts. One SKU, one metro, one week, and you’re reading a handful of households with a projection weight on top. The number stays arithmetically correct and stops meaning much.

So a panel answers who is buying me and why they come back. Nothing else below can answer that at all.

Retailer census and scan

Participating retailers send register data, which the provider standardizes and projects into a measured universe. This is what most people picture when they hear syndicated data, and for category share it is the right buy. We laid out the vendor-level version of this in our Circana vs NielsenIQ vs store-level buyer’s guide, so we won’t repeat it here.

Units and dollars scanned at a register, rolled up to store, chain or market depending on the grain you bought. Weekly periods, plus a release lag.

The first blind spot is the obvious one. It is a census of participants. Stores outside the arrangement get modelled, weighted around, or simply left out, and which of those three happened is a methodology question with real consequences in a fragmented category. Beverage alcohol is fragmented, which is why we wrote a whole post on where independent liquor store coverage goes blind.

The subtler blind spot is price. Scan data gives you dollars and units, and dividing one by the other gives an average, not a tag. Assume a bottle listed at $24.99, on promotion at $19.99 for two weeks of a four-week period, and assume units split evenly across the four weeks. Dollars over units comes out at $22.49. That is a genuine average of two genuine prices, and it was never printed on a shelf anywhere. Take it into a pricing conversation and you are arguing about a number no shopper was asked to pay. That gap is exactly why Price to Consumer is built on the observed tag rather than a derived one.

Ask census scan about category share and rate of sale. That is what it is for and it is genuinely good at it. Ask it what the tag said on a Tuesday and it has no idea.

One variant deserves naming. In control jurisdictions the state sits in the supply chain, so the transaction records are government records rather than a commercial panel, and the coverage arithmetic works differently. That’s a category of its own and we treat it separately in our control-state data guide.

Distributor depletion

Distributors record every case that leaves the dock for a licensed account, and aggregators standardize those records across contributors. WSWA describes SipSource as built from “single-product transactions between wholesalers and retailers,” reported as 9-litre case equivalent volume on a rolling twelve-month basis, with quarterly reporting (WSWA).

The unit is a case, or a 9-litre case equivalent, shipped to an account. Distributor systems record it close to real time and the syndicated aggregation lands quarterly, which is the widest gap on this whole list between how fast a source captures and how fast the product built on it ships.

The blind spot is the back door. Depletion is a shipment, not a sale, so it leads consumer demand and overstates it whenever an account buys ahead. We covered the mechanics and the reconciliation traps in the depletion data field guide. Coverage is also contributor-gated, the same shape of hole as the census above with a different party holding the gate. Volume moving through a distributor that doesn’t contribute stays invisible no matter how large it is.

What depletion answers well is where am I distributed and how fast is inventory clearing the middle tier.

Invoice reconciliation

A payments and invoice network sits between retailers and their distributors, and the transaction records that flow through it get packaged back to suppliers. Fintech describes itself as an “Alcohol Invoice Data & Payments Platform” and offers distributors and supply-chain partners “SKU- and store-level insights” out of its scan-based trading service (Fintech).

The unit is an invoice line. Specifically, what the retailer was billed. It lands as invoices clear, so close to continuous for accounts on the network.

The blind spot is the retailer’s margin decision. Invoice data gives you cost in, scan gives you realized price out, and the shelf tag sits between them where neither one records it. Note that the gate has moved again. A store is in this dataset because of a payment relationship rather than a data-sharing agreement, which changes who is missing without shrinking the problem.

Ask it what this account actually paid, and whether your costs and chargebacks hold up across markets. It answers both of those better than scan data does.

Storefront observation

Somebody or something looks at the store and records its state. Presence or absence of an authorized SKU, facings, the printed price, whether the case stack exists. This is our method, so read the next paragraph with that in mind.

One SKU, at one store, at one moment. As often as you look. Shelf state changes daily, so daily is the cadence worth designing for.

The blind spot here is bigger than the others and we’d rather say it plainly: observation counts nothing. It records a state, not a flow. It cannot tell you units sold, and we are not going to pretend it hands you a projectable market share. A store with two facings and no gap looks identical whether it sold forty bottles last week or four. Anyone selling you observation as a replacement for measurement is selling you a problem.

What it does answer is whether the product is physically there, at what price, and whether the program you paid for got built. That last one has no counterpart in any register or invoice feed above, which is the whole reason Marketing Monitoring exists.

Field-rep audit

Reps or contracted merchandisers visit stores on a schedule and record what they find. The data is a by-product of the visit. Retail-execution platforms describe covering “planogram compliance, on-shelf availability, competitor tracking, and more” from the rep’s phone (Repsly).

A completed store visit is the unit, and the route schedule is the cadence. Weekly for key accounts, quarterly or never for the tail.

Coverage equals wherever a human physically went, and cost per visit sets a hard ceiling on that. The stores most likely to be broken are usually the ones the route skips, which is the panel’s thin-cut trap arriving by a different road. There is also an evidence-class difference between a rep ticking a box and a photograph of the shelf, and that gap shows up worst in the accounts nobody wants to admit are neglected.

The thing an audit does that nothing above does: the person who finds the problem can fix it before leaving the building.

The absence rule

There’s one question we’d put on the first page of any data evaluation, and almost nobody asks it. What does this dataset do when the event happens somewhere the instrument isn’t installed?

Every method has an answer, and the answers are not interchangeable.

Projection is not dishonest. It is the correct statistical response to a designed sample, and panel providers document it. The danger is what happens downstream, when a projected figure and a directly recorded figure land in the same table with the same number of decimal places and nothing in the column header distinguishes them. Six months later someone builds a forecast on both.

Our practical rule: carry the provenance through your warehouse as a column, not as tribal knowledge. Which instrument produced this row, and was the store in-universe or projected into it. This is unglamorous work and it is most of what our data analytics and business intelligence team actually does on these projects.

Which question goes to which method

Two rows on that grid are worth a second look because brands routinely send them to the wrong place.

“Where am I losing distribution” gets sent to census data, which reports a share decline, and share declines are ambiguous. A market softening five percent could be four hundred stores each slipping a little or thirty stores that stopped carrying you outright. Those need opposite responses. Only store-grain records separate them.

“Is my pricing holding” gets sent to whatever price field is nearest, which is almost always derived. See the $22.49 above.

Platform is not method

The syndicated data platforms question deserves its own answer, because platform and instrument get conflated constantly and they price separately.

A platform is delivery. Dashboard, API, scheduled flat file, or a governed semantic layer that other people’s data flows into. Dimensional Insight, for instance, describes its Diver Platform as “a governed analytics platform that unifies data, business rules, and analytics into a single, trusted foundation” (Dimensional Insight). That is a real and useful job. It is also not a sensor, so it reports precisely what you load into it and nothing you didn’t.

Practically, this means two things. A dataset you cannot export is worth much less than the same dataset as a file, because the value shows up when you blend it with your own depletions and your own account master. And a beautiful dashboard over a thin instrument is still a thin instrument. Judge the sensor and the plumbing on separate lines.

One scope note while we’re being precise: everything above is off-premise. Bars and restaurants run on different sources with a different set of holes, and we mapped that in on-premise versus off-premise data.

What we’d recommend

Write down the event you need to observe before you shortlist anything. Not the metric, the event. “A shopper paid money” is a different event from “a case left a warehouse,” which is a different event from “a bottle occupied a shelf facing.” Once the event is written down, the method falls out almost automatically, and the vendor conversation becomes short.

Then ask every provider the absence question, including us. What does your dataset do when this happens in a store you don’t cover. A provider who answers that crisply is telling you they understand their own instrument. A provider who deflects is telling you something too.

If you’re buying your first syndicated subscription, buy the one that answers the decision already on your calendar this quarter, and buy nothing else. If you already own census and depletion feeds and keep hitting questions neither can answer, the gap is usually the shelf, and it’s usually the cheapest layer to add because you don’t need national scope to start.

Tell us which feeds you already pay for and which question keeps going unanswered, and we’ll give you a straight answer on whether the missing piece is a method you don’t have or a reconciliation problem in what you already own. We do this work with beverage-alcohol brands and distributors most weeks, and often enough the answer is that your existing contracts already cover it.

  • syndicated data
  • syndicated data providers
  • retail data methods
  • beverage alcohol data
FAQ

Common questions, answered

What is syndicated data?
Syndicated data is information a research provider collects on its own initiative, standardizes into one format, and then licenses to many subscribers. Because the same dataset is sold repeatedly, the cost per buyer drops well below what the same study would cost commissioned. The trade-off is that the design is fixed. Nobody asked you which stores, which weeks, or which price definition mattered before the collection started.
What are the main types of syndicated data providers?
Group them by instrument rather than by brand. Household panels recruit shoppers who report their own purchases. Retailer census providers ingest register data from participating chains. Depletion aggregators collect shipment records from distributors. Invoice networks read what retailers were billed. Storefront observation records the shelf itself. Field-rep platforms produce data as a by-product of scheduled store visits. Two providers can both call their product syndicated data and share no method at all.
What are syndicated data platforms?
A platform is the delivery layer, not the sensor. It is the dashboard, API, flat-file drop or BI semantic layer through which a dataset reaches you. Some of the best-known names in retail analytics are platforms in this sense and originate no observation of their own, so they report exactly what you feed them and nothing more. Judge the instrument and the delivery separately, because a strong instrument behind a login you cannot export from is close to unusable in a warehouse.
Does syndicated data show what is on the shelf?
Almost none of it does, and this is the most common misread we see. Register and invoice data both describe money changing hands, so they can tell you a bottle sold or was billed without telling you whether a second bottle was ever stocked behind it. Shelf presence, facings, the printed tag price and whether a display was physically built are states of a physical location, and the only methods that record them involve someone or something looking at the store.
Is syndicated data worth it for a small emerging brand?
Usually not as a first purchase. Syndicated pricing is built around a shared cost across many subscribers, which makes it cheap relative to custom research and still expensive relative to a small brand's whole analytics budget. Below the point where you cannot personally call your accounts, distributor reports you already receive plus a spreadsheet will outperform a subscription. Buy syndicated data when the decision it feeds is worth more than the licence.

See it on your own data.

Book a 30-minute discovery call and we'll walk through your use case.