Scan data is a record from a store’s register. A UPC scanned at a price, in a store, in a week. The retailer’s point-of-sale system produces it as a by-product of ringing up sales, a measurement provider licenses it, standardizes it, and sells it back to the brands whose products appear in it. That’s the entire mechanism. Almost everything suppliers get wrong about scan data happens in that licensing step.
Two other things carry the same name, so let’s clear them out first. If you operate convenience stores, “scan data” most likely means the manufacturer incentive programs where you transmit your own tobacco sales files upstream and get paid for doing it. And if you’re looking for software that turns paper documents into text, that’s document scanning, and this is the wrong page entirely. This post is about the third meaning: retail scan data as a measurement product that a beverage-alcohol supplier buys in order to see what sold.
We’re gmware, a software and data engineering firm in Austin, TX, with delivery centers in Bangalore and Mohali, India. We also run Shield Suite, our retail-intelligence platform for beverage-alcohol brands across 60,000+ storefronts. We spend most of our time on the plumbing underneath these files, which means we’ve read more scan extracts than is healthy and we know which columns brand teams misread.
This is for the analyst who just inherited a scan subscription, or the brand manager about to sign one, and wants to know what the number on the dashboard is actually a measurement of.
| Term | What it is in one line |
|---|---|
| Scan data | A register record: UPC, units, money, store, period |
| Retail scan data | Same thing, worded to separate it from document scanning |
| Store level data | Reported per store instead of rolled up to a market |
| Depletion data | Cases shipped distributor to account, not a register record |
| Scan-based trading | A payment arrangement between retailer and vendor, not a dataset |
Three unrelated things get called scan data
The confusion is real and it wastes people’s time in meetings, so here they are, sorted.
The first is retail measurement, the subject of this post. A provider licenses register files from retailers, harmonizes them into a single product hierarchy, and sells brands a view of their own category. That’s the file your category manager lives in.
The second runs in the opposite direction. In convenience retail, a scan data program is an agreement where the retailer sends item-level tobacco sales files to a manufacturer and receives money back. Cenex’s retailer guidance describes scan data as “the information captured at the point of sale (POS) every time a tobacco product is purchased,” passed to manufacturers “to demonstrate compliance, validate promotions, and access incentive programs” (Cenex, tobacco scan data 101). Point-of-sale vendors describe the same loop from the store’s side: the register logs the sale, a certified provider formats and transmits the file, the manufacturer validates it against the program agreement, and payment follows, usually monthly (KORONA POS). Altria’s is the program both of those sources name. So when a c-store operator and a spirits brand manager each say “our scan data,” they mean opposite things, and the money moves the opposite way.
The third is scan-based trading, sometimes called pay-on-scan, where the retailer doesn’t own the inventory until it scans at checkout. Fintech markets separate retailer and vendor offerings under exactly that name (Fintech). It’s a settlement arrangement, not a dataset, though it produces data as a side effect.
Only the first one answers “how much of my tequila sold last month.” Keep them separate in conversation and you’ll save yourself a meeting.
What’s actually inside a scan record
Commercial providers treat their methodology as proprietary, which makes the research versions of these files unusually useful for understanding the shape. The Kilts Center at Chicago Booth documents the NielsenIQ Retail Scanner dataset in public, and it’s the clearest description of the grain we’ve found anywhere.
The data is weekly, by store, by UPC. For each of those cells, participating stores report “units, price, price multiplier, baseline units, baseline price, feature indicator, and display indicator,” alongside store attributes covering “store chain code, channel type, and area location” (Kilts Center, NielsenIQ Retail Scanner data). Retailer names are masked. Private-label items get a masked UPC so the retailer can’t be reverse-identified from its own products.
Read that field list slowly, because it’s the honest boundary of the whole category. There is a price, and it’s period-level. There is a display indicator, and the same documentation notes the feature and display variables come only “from a subset of stores,” not the full sample. There is no shelf position, no facing count, no photograph, no inventory on hand, and nothing whatsoever about the other brands sharing that bay.
Lane to dashboard, five hops
The word doing all the work is “participating”
The sentence worth memorising is this one. Scan data exists for a store only if two conditions are both true: the store’s point-of-sale system can produce a clean, consistently structured export on a schedule, and somebody at that retailer has agreed to send it.
Neither condition is about how much your brand sells there. A store can be your best account in the state and still contribute nothing, because nobody there ever signed. And the absence doesn’t show up as bad news in your file. It shows up as nothing at all. The store isn’t a zero in the store level data, it’s outside the frame the file was built from.
The published figures make the shape of that frame concrete. The Kilts documentation describes the research file as generated by point-of-sale systems at “more than 90 participating retail chains across all US markets,” covering 35,000 to 50,000 participating stores across food, drug, mass merchandise, convenience and liquor channels, and representing “more than half the total sales volume of US grocery and drug stores.”
The published shape of one scanner file
Note the unit on the first number. Chains, not stores. Licensing agreements get signed by organizations that employ somebody whose job is signing them, which is exactly why the frame has the edges it has. In beverage alcohol the stores sitting outside those edges cluster hard in single-store independents, and sizing your own exposure to that is a whole exercise on its own. We wrote it up separately in where your independent-store coverage goes blind, including the reconciliation that turns the argument into one defensible percentage.
The last thing to say here is what nobody can tell you. What share of US beverage-alcohol category volume sits outside a given provider’s universe is not a published number, because universe definitions are commercial property. Anyone who quotes you one is guessing.
Scan data and depletion data are not two views of one thing
Suppliers blur these constantly, usually because both arrive as spreadsheets and both have a case count in them.
A depletion is a case leaving a distributor’s warehouse for a retail account. A scan is a bottle leaving a store in a shopper’s hand. Depletions arrive earlier, cover every account a distributor ships to including ones no panel measures, and say nothing about whether a human bought anything. Scan arrives later, only from participating stores, and is the only one of the two that reflects an actual purchase. When they disagree, the gap is usually inventory sitting somewhere between the two events, and it resolves eventually. That’s not a data quality problem, it’s the two instruments doing their jobs. The mechanics of the depletion side, plus how VIP, iDIG and SipSource relate to each other, are covered in our depletion data explainer.
Five conclusions people draw from a scan number
One of these five holds up. The other four get said out loud in real meetings by people who are good at their jobs.
“We grew 6% last quarter.” That one holds, with the fine print attached. You grew 6% in measured stores, in that reporting period, measured in money collected. Whether the business as a whole grew depends on what the unmeasured portion did, and the file has nothing to say about it.
“Our shelf price is $24.99.” No. A scan-derived price is money collected divided by units moved across a period, blended across promoted and unpromoted transactions, net of whatever discounted at the register. The tag on the shelf right now is a different quantity. During a funded promotion the two can sit a long way apart, and the retailer’s real margin decision happens in the space between them. That gap is why price policy enforcement needs the observed tag instead of a derived average, and it’s what our price to consumer module exists to capture.
“That store is out of stock.” No. Zero units in a store-week reads identically whether the shelf is empty, the shelf is full and nothing is moving, or the store never carried the item at all. Three different problems, one row. Separating them is the job of void and out-of-stock reporting, and it needs an instrument the register doesn’t have.
“The display we paid for went up.” No, and this is the one that costs the most money. The display indicator is a flag, reported from a subset of stores, and a flag is not evidence you can put in front of a retailer’s trade team. Settling a spend dispute takes dated proof from the floor, which is the entire premise of retail execution tracking.
“A competitor took our shelf space.” No. Scan tells you their units moved and yours didn’t, which is a result, not a cause. Space is physical: how many facings you hold, and where those facings sit. Cold box and back wall are not the same business. A register has no opinion about either one.
What the number does and doesn't support
When scan data is the right buy
When you need a defensible share number and your volume runs through measured chains. Nothing else in this category produces a figure that survives a boardroom the way projected retail measurement does, and we say that as people who sell a different kind of data. If a national account is about to review your brand and you need to argue category performance, this is the file, and choosing between providers is its own decision that we walked through in our bev-alc data buyer’s guide.
When it isn’t the right buy: you’re in two markets, your distribution skews to independents, and your sales lead can still name every account from memory. A subscription then buys you a partial view of a business you already understand better than the file does. Spend the money on getting product placed instead.
The genuinely awkward case is the brand in between. Real volume, many markets, distribution split between chains you’re measured in and a long independent tail you aren’t, running national decisions off a file that describes the measured half. That brand doesn’t need to cancel anything. It needs a second instrument covering the stores the first one structurally can’t reach, and the two answer different questions on purpose.
What we’d recommend
Keep the scan subscription and stop asking it questions it was never built to answer. Use it for share and for category trend inside the measured universe, and treat those answers as strong, because they are. Then write down the four claims from the list above that it can’t support, and decide separately how you’re covering each one. That exercise takes an afternoon and it’s the cheapest data work you’ll do this year.
All four of the unsupported claims share one property. Each is a physical fact about one store on one day: the number on the tag, the bottle on the shelf, the display on the floor, the facings a competitor holds. A register can’t know any of them, because a register only ever sees what crossed the scanner. That’s what we built Shield Suite competitive intelligence to do, and we’ll be equally clear about its boundary: we observe what’s on the shelf and what it costs across 60,000+ storefronts, and we do not count units sold. A brand expecting us to replace scan will be disappointed. A brand using us to cover the stores and the facts scan can’t reach gets exactly what it came for.
Send us the store list from your current scan file and your distributor’s authorized-account list, and we’ll show you which stores appear in one and not the other. That comparison is usually more persuasive than anything in a deck. Tell us what you’re working with and we’ll give you a straight answer on whether closing the gap is worth the money. We do this with beverage-alcohol brands and distributors constantly, and the data engineering underneath it is the part most teams underestimate.