Retail Intelligence

Scan Data Explained: What It Captures and What It Misses

11 min read

Scan data is a record from a store’s register. A UPC scanned at a price, in a store, in a week. The retailer’s point-of-sale system produces it as a by-product of ringing up sales, a measurement provider licenses it, standardizes it, and sells it back to the brands whose products appear in it. That’s the entire mechanism. Almost everything suppliers get wrong about scan data happens in that licensing step.

Two other things carry the same name, so let’s clear them out first. If you operate convenience stores, “scan data” most likely means the manufacturer incentive programs where you transmit your own tobacco sales files upstream and get paid for doing it. And if you’re looking for software that turns paper documents into text, that’s document scanning, and this is the wrong page entirely. This post is about the third meaning: retail scan data as a measurement product that a beverage-alcohol supplier buys in order to see what sold.

We’re gmware, a software and data engineering firm in Austin, TX, with delivery centers in Bangalore and Mohali, India. We also run Shield Suite, our retail-intelligence platform for beverage-alcohol brands across 60,000+ storefronts. We spend most of our time on the plumbing underneath these files, which means we’ve read more scan extracts than is healthy and we know which columns brand teams misread.

This is for the analyst who just inherited a scan subscription, or the brand manager about to sign one, and wants to know what the number on the dashboard is actually a measurement of.

TermWhat it is in one line
Scan dataA register record: UPC, units, money, store, period
Retail scan dataSame thing, worded to separate it from document scanning
Store level dataReported per store instead of rolled up to a market
Depletion dataCases shipped distributor to account, not a register record
Scan-based tradingA payment arrangement between retailer and vendor, not a dataset

Three unrelated things get called scan data

The confusion is real and it wastes people’s time in meetings, so here they are, sorted.

The first is retail measurement, the subject of this post. A provider licenses register files from retailers, harmonizes them into a single product hierarchy, and sells brands a view of their own category. That’s the file your category manager lives in.

The second runs in the opposite direction. In convenience retail, a scan data program is an agreement where the retailer sends item-level tobacco sales files to a manufacturer and receives money back. Cenex’s retailer guidance describes scan data as “the information captured at the point of sale (POS) every time a tobacco product is purchased,” passed to manufacturers “to demonstrate compliance, validate promotions, and access incentive programs” (Cenex, tobacco scan data 101). Point-of-sale vendors describe the same loop from the store’s side: the register logs the sale, a certified provider formats and transmits the file, the manufacturer validates it against the program agreement, and payment follows, usually monthly (KORONA POS). Altria’s is the program both of those sources name. So when a c-store operator and a spirits brand manager each say “our scan data,” they mean opposite things, and the money moves the opposite way.

The third is scan-based trading, sometimes called pay-on-scan, where the retailer doesn’t own the inventory until it scans at checkout. Fintech markets separate retailer and vendor offerings under exactly that name (Fintech). It’s a settlement arrangement, not a dataset, though it produces data as a side effect.

Only the first one answers “how much of my tequila sold last month.” Keep them separate in conversation and you’ll save yourself a meeting.

What’s actually inside a scan record

Commercial providers treat their methodology as proprietary, which makes the research versions of these files unusually useful for understanding the shape. The Kilts Center at Chicago Booth documents the NielsenIQ Retail Scanner dataset in public, and it’s the clearest description of the grain we’ve found anywhere.

The data is weekly, by store, by UPC. For each of those cells, participating stores report “units, price, price multiplier, baseline units, baseline price, feature indicator, and display indicator,” alongside store attributes covering “store chain code, channel type, and area location” (Kilts Center, NielsenIQ Retail Scanner data). Retailer names are masked. Private-label items get a masked UPC so the retailer can’t be reverse-identified from its own products.

Read that field list slowly, because it’s the honest boundary of the whole category. There is a price, and it’s period-level. There is a display indicator, and the same documentation notes the feature and display variables come only “from a subset of stores,” not the full sample. There is no shelf position, no facing count, no photograph, no inventory on hand, and nothing whatsoever about the other brands sharing that bay.

The word doing all the work is “participating”

The sentence worth memorising is this one. Scan data exists for a store only if two conditions are both true: the store’s point-of-sale system can produce a clean, consistently structured export on a schedule, and somebody at that retailer has agreed to send it.

Neither condition is about how much your brand sells there. A store can be your best account in the state and still contribute nothing, because nobody there ever signed. And the absence doesn’t show up as bad news in your file. It shows up as nothing at all. The store isn’t a zero in the store level data, it’s outside the frame the file was built from.

The published figures make the shape of that frame concrete. The Kilts documentation describes the research file as generated by point-of-sale systems at “more than 90 participating retail chains across all US markets,” covering 35,000 to 50,000 participating stores across food, drug, mass merchandise, convenience and liquor channels, and representing “more than half the total sales volume of US grocery and drug stores.”

Note the unit on the first number. Chains, not stores. Licensing agreements get signed by organizations that employ somebody whose job is signing them, which is exactly why the frame has the edges it has. In beverage alcohol the stores sitting outside those edges cluster hard in single-store independents, and sizing your own exposure to that is a whole exercise on its own. We wrote it up separately in where your independent-store coverage goes blind, including the reconciliation that turns the argument into one defensible percentage.

The last thing to say here is what nobody can tell you. What share of US beverage-alcohol category volume sits outside a given provider’s universe is not a published number, because universe definitions are commercial property. Anyone who quotes you one is guessing.

Scan data and depletion data are not two views of one thing

Suppliers blur these constantly, usually because both arrive as spreadsheets and both have a case count in them.

A depletion is a case leaving a distributor’s warehouse for a retail account. A scan is a bottle leaving a store in a shopper’s hand. Depletions arrive earlier, cover every account a distributor ships to including ones no panel measures, and say nothing about whether a human bought anything. Scan arrives later, only from participating stores, and is the only one of the two that reflects an actual purchase. When they disagree, the gap is usually inventory sitting somewhere between the two events, and it resolves eventually. That’s not a data quality problem, it’s the two instruments doing their jobs. The mechanics of the depletion side, plus how VIP, iDIG and SipSource relate to each other, are covered in our depletion data explainer.

Five conclusions people draw from a scan number

One of these five holds up. The other four get said out loud in real meetings by people who are good at their jobs.

“We grew 6% last quarter.” That one holds, with the fine print attached. You grew 6% in measured stores, in that reporting period, measured in money collected. Whether the business as a whole grew depends on what the unmeasured portion did, and the file has nothing to say about it.

“Our shelf price is $24.99.” No. A scan-derived price is money collected divided by units moved across a period, blended across promoted and unpromoted transactions, net of whatever discounted at the register. The tag on the shelf right now is a different quantity. During a funded promotion the two can sit a long way apart, and the retailer’s real margin decision happens in the space between them. That gap is why price policy enforcement needs the observed tag instead of a derived average, and it’s what our price to consumer module exists to capture.

“That store is out of stock.” No. Zero units in a store-week reads identically whether the shelf is empty, the shelf is full and nothing is moving, or the store never carried the item at all. Three different problems, one row. Separating them is the job of void and out-of-stock reporting, and it needs an instrument the register doesn’t have.

“The display we paid for went up.” No, and this is the one that costs the most money. The display indicator is a flag, reported from a subset of stores, and a flag is not evidence you can put in front of a retailer’s trade team. Settling a spend dispute takes dated proof from the floor, which is the entire premise of retail execution tracking.

“A competitor took our shelf space.” No. Scan tells you their units moved and yours didn’t, which is a result, not a cause. Space is physical: how many facings you hold, and where those facings sit. Cold box and back wall are not the same business. A register has no opinion about either one.

When scan data is the right buy

When you need a defensible share number and your volume runs through measured chains. Nothing else in this category produces a figure that survives a boardroom the way projected retail measurement does, and we say that as people who sell a different kind of data. If a national account is about to review your brand and you need to argue category performance, this is the file, and choosing between providers is its own decision that we walked through in our bev-alc data buyer’s guide.

When it isn’t the right buy: you’re in two markets, your distribution skews to independents, and your sales lead can still name every account from memory. A subscription then buys you a partial view of a business you already understand better than the file does. Spend the money on getting product placed instead.

The genuinely awkward case is the brand in between. Real volume, many markets, distribution split between chains you’re measured in and a long independent tail you aren’t, running national decisions off a file that describes the measured half. That brand doesn’t need to cancel anything. It needs a second instrument covering the stores the first one structurally can’t reach, and the two answer different questions on purpose.

What we’d recommend

Keep the scan subscription and stop asking it questions it was never built to answer. Use it for share and for category trend inside the measured universe, and treat those answers as strong, because they are. Then write down the four claims from the list above that it can’t support, and decide separately how you’re covering each one. That exercise takes an afternoon and it’s the cheapest data work you’ll do this year.

All four of the unsupported claims share one property. Each is a physical fact about one store on one day: the number on the tag, the bottle on the shelf, the display on the floor, the facings a competitor holds. A register can’t know any of them, because a register only ever sees what crossed the scanner. That’s what we built Shield Suite competitive intelligence to do, and we’ll be equally clear about its boundary: we observe what’s on the shelf and what it costs across 60,000+ storefronts, and we do not count units sold. A brand expecting us to replace scan will be disappointed. A brand using us to cover the stores and the facts scan can’t reach gets exactly what it came for.

Send us the store list from your current scan file and your distributor’s authorized-account list, and we’ll show you which stores appear in one and not the other. That comparison is usually more persuasive than anything in a deck. Tell us what you’re working with and we’ll give you a straight answer on whether closing the gap is worth the money. We do this with beverage-alcohol brands and distributors constantly, and the data engineering underneath it is the part most teams underestimate.

  • scan data
  • retail scan data
  • store level data
  • off-premise data
FAQ

Common questions, answered

What is scan data?
Scan data is the record a retailer's point-of-sale system creates when an item is scanned at checkout: the UPC, the units, the money collected, the store, and the time period. Measurement providers license those records from retailers, clean and standardize them, and sell the result back to the brands whose products appear in it. The academic version of one such file reports weekly units and price for every UPC in every participating store, which is the clearest public picture of what the grain actually looks like.
Where does retail scan data come from?
From retailers who have agreed to hand their register files over, under a licensing arrangement, and who have a point-of-sale system capable of producing a clean export on a schedule. The Kilts Center documentation for the NielsenIQ research file describes it as generated by point-of-sale systems at more than 90 participating retail chains. Both conditions matter. A store that fails either one contributes nothing, and it does not appear as a zero either.
How is scan data different from depletion data?
They measure two different events. A depletion is a case leaving a distributor's warehouse for a retail account. A scan is a bottle leaving a store in a shopper's hand. Depletions arrive earlier and cover every account a distributor ships to, including ones no panel measures. Scan arrives later, only from participating stores, and is the only one of the two that reflects a consumer actually buying something.
Does scan data show the price on the shelf?
No. A scan-derived price is money collected divided by units moved over a reporting period, blended across promoted and unpromoted transactions and net of whatever discounts applied at the register. The number printed on the shelf tag today is a different quantity, and the two can diverge by a wide margin during a promotion. If you are enforcing a price policy, you need the tag, observed.
Can scan data tell me if a store is out of stock?
Not reliably. Zero units in a store-week is consistent with an empty shelf, with a shelf that is full but not selling, and with a store that never carried the item. Those are three different problems with three different fixes, and the scan record renders them identically. Distinguishing them takes either an inventory feed from that retailer or a direct observation of the shelf.

See it on your own data.

Book a 30-minute discovery call and we'll walk through your use case.