The Say/Do Gap in CPG and Retail Research (August 2026)
Aug 3, 2026 by Ethan Pidgeon
On this page▼
Roughly 38% of online shoppers don't follow through on what they tell you they'll do, per Horizon's consumer research. That's not a rounding error in your tracker. It's a structural feature of how surveys work, and once you know the four mechanisms behind it, you start reading your data very differently.
TLDR:
- The say/do gap is structural: roughly 38% of US online shoppers do not follow through on stated behavior, per consumer research.
- Four mechanisms drive the gap: social desirability bias, hypothetical framing, recall decay, and situational override at purchase.
- A consumer who says one thing and buys another is often blocked by friction, not lying. That friction is a distribution or pricing problem you can fix.
- DTC repeat rate versus retail sell-through divergence is a live say/do signal you can read weeks before syndicated data catches it.
- Merciv joins cross-retailer reviews, social conversation, internal POS, and licensed syndicated research in a single query with confidence-scored, cited outputs.
What the Say/Do Gap Actually Is
The say/do gap is the distance between what a consumer tells you in a survey and what they do at the shelf, on the site, or in the reorder cycle. She says she buys clean beauty. Her cart says drugstore staples. He tells your tracker sustainability comes first. His purchases skew to whichever brand ran the sharpest promo that week.
For a CPG consumer insights lead, this is why a concept test greenlights a launch that misses velocity, or a tracker shows rising affinity while sell-through softens.
One anchor: roughly 38% of US online shoppers do not follow through on stated behavior, per Horizon's consumer research. Stated preference is a signal. Revealed preference is the decision.
The Psychology Behind Why the Gap Exists
The gap is structural. Four mechanisms do most of the work.
- Social desirability bias. Answering a tracker is a social act. Saying you buy refill pouches, choose the fragrance-free SKU, or read the ingredient panel makes you sound like the person you want to be. The checkout does not care.
- Hypothetical framing. Surveys strip out price, promo, out-of-stock, the coupon that hit her inbox that morning, and the competitor at eye level. Intent forms in a room where trade-offs do not exist.
- Recall decay. Memory is reconstructive. Asking what she bought last month returns a tidied narrative, not the actual basket.
- Situational override. The moment of purchase belongs to context, not conviction.
Why Traditional Survey Research Amplifies the Problem
Survey instruments were built to measure attitudes, not to predict a Tuesday-night basket. That design choice, which is also why traditional consumer research has limits, widens the gap in three ways worth naming.
- Articulate answers score highest. Likert scales and forced-choice trade-offs reward the respondent who can construct a coherent self-image on the spot, which is not the same person standing in the aisle.
- Hypothetical purchases carry no cost. No wallet opens, no substitute tempts, no shelf tag interrupts. Willingness to pay inflates because paying is not happening.
- Aspirational categories distort most. Health, sustainability, and clean ingredients pull answers toward the person the respondent wants to be, a core reason CPG brands misread their shoppers, so intent readings run hot by design.
This is a method boundary. Researchers are not the failure point.
The Sustainability Version: Where the Gap Is Most Visible
Sustainability is where the gap is loudest, best documented, and easiest to measure against actual purchase.
A Global Sustainability Study found 50% of consumers rank sustainability among their top five purchase drivers, and 66% say they will pay more for it. The same Forbes analysis reports 66% of retail executives believe the opposite about their shoppers. In the UK, 20% of consumers expressed intent to buy a battery electric vehicle. BEVs came in at 7.5% of new car sales in 2021.
The read most teams reach for is that consumers are lying. A cleaner read, per the Say Do Company, is that stated intent runs into friction the survey never modeled: sticker price, charging access, refill availability, the coupon in her inbox. A lying consumer is a research problem. A blocked consumer is a merchandising, pricing, or distribution problem, and consumer behavior analysis for CPG is where the fix lives inside the business.
What Behavioral Signals Reveal That Surveys Miss
Behavioral data is the counterweight. Reviews post within days of purchase, so complaint clusters and repeat-language surface on their own timeline. Repeat-buy rates tell you whether trial converted or died on the second bottle. Social listening vs consumer intelligence clarifies why unprompted social language surfaces what she actually talks about when nobody is holding a survey in front of her.
| Signal Type | What It Captures | What It Misses | Say/Do Gap Role |
|---|---|---|---|
| Survey / tracker | Stated intent, attitudes, aspirational self-image | Price, promo timing, out-of-stock, shelf context: the friction that kills intent | Measures the say; inflates by design in aspirational categories |
| Cross-retailer reviews | Post-purchase verbatims within days of the transaction; complaint clusters; repeat language | Pre-purchase intent; why a shopper never tried in the first place | Surfaces the do, including the gap between trial and the second bottle |
| DTC repeat-buy rate | Whether highest-intent buyers reordered; true loyalty signal | Shelf execution, distribution gaps, in-store competitive pressure | Leading indicator: DTC repeat climbing while retail softens = distribution problem, not demand |
| Retail sell-through (POS) | Actual units scanned; habitual buyers and promo lift | Why velocity changed; early perception movement at top of funnel | Lagging indicator: confirms the gap weeks after behavioral signals have already flagged it |
| Unprompted social language | What consumers say about a brand when no survey prompt is present | Purchase frequency, basket size, channel preference | Bridges stated and revealed: shows awareness vs. routine adoption split |
The PwC 2025 holiday read is the cleanest live example. Gen Z respondents said they would pull back, then spent nearly 21% more than the prior year. The survey caught the sentiment. The behavioral layer caught the wallet.
Keep the tracker. Add the layer that shows what happened after the answer was given.
The DTC vs. Retail Divergence: A Live Say/Do Signal
The gap gets concrete at the SKU level. Two divergence patterns show up before any tracker or syndicated read catches them.
- DTC repeat climbing, retail sell-through softening. Highest-intent buyers are reordering; the shelf is not moving. Stated demand is real. The friction is distribution or in-store execution: out-of-stocks, shelf position, a planogram reset that buried the hero SKU, a competitor who bought endcap. A concept test cannot see any of this.
- DTC conversion dropping, retail velocity holding. Retail is coasting on habitual buyers and promo lift while direct comparers quietly walk. Perception erodes at the top of the funnel before scan data catches it.
Both are diagnosable weeks before syndicated reports ratify the shift, if you apply CPG and retail shopper insights by joining DTC analytics, retailer POS, and review verbatims on the same timeline.
The Beauty and Wellness Version: Ingredient Claims vs. Actual Repeat Purchase
Beauty and wellness are where this gap gets expensive fast. A concept test says "clean," "fragrance-free," or "no seed oils" drives purchase intent. The claim trends on TikTok. Trial lifts. Then the second bottle never sells through, and the tracker calls it a conversion problem without naming what actually happened.
Cross-retailer reviews across Sephora, Ulta, Target, and Amazon read the second half of the story. In our work with beauty and personal care brands, trial-driving claims and repeat-driving claims are often different claims, a pattern covered in depth in beauty brand research driving and hurting growth, and the survey cannot see the split because it never asks about the second purchase.
"The say-do gap is as real as it ever has been." Tucker Mitchell, Pernod Ricard
A claim trending on social confirms awareness, but multi-source consumer intelligence confirms whether it survived the routine.
Now What: 3 Actions to Close the Gap This Week
Three moves you can run this week without a research budget or a new vendor.
- Audit one live tracker for aspirational framing. Pull the instrument for whichever U&A or brand tracker is closest to fielding. Flag questions where the socially attractive answer is obvious: sustainability priority ranking, ingredient panel reading frequency, willingness to pay a premium for clean or refillable formats. Find three. Rewrite them as forced trade-offs with real cost attached, or move them to a behavioral proxy from POS or review data.
- Read the hero SKU against itself. Pull cross-retailer reviews on your top SKU across Sephora, Ulta, Target, or Amazon for the last 90 days (the kind of read covered in the Beauty Consumer Intelligence Report 2026). Cluster the top three complaint themes. Put them next to the top three stated purchase drivers from your most recent U&A. Where the lists do not overlap, the survey is measuring the wrong thing.
- Line up DTC repeat against retail sell-through. Same SKU, last two quarters, one chart. If DTC repeat is climbing while retail softens, the problem is distribution or shelf, not demand. If DTC conversion is dropping while retail holds, perception is eroding at the top of the funnel and scan data has not caught it yet.
How Merciv Connects Stated and Revealed Preference in One Place
Most brands never close the gap because the two halves of the answer live in different systems. The tracker sits in one vendor, reviews in another, POS in a third, syndicated behind a fourth login. Nobody joins them against the same question on the same day.
Merciv runs that join. Triangulating syndicated, qual, quant, and reviews into one story is exactly how cross-retailer reviews, social conversation, internal POS, and licensed syndicated research get pulled in a single query, with every claim cited and confidence-scored on a three-tier scale (High, Directional, Exploratory) and a clickable audit trail back to the underlying feed. Multi-week synthesis compresses to minutes. For a team of one, that is the readout landing on time. For a VP presenting upward, it is a finding the CMO can pressure-test at the source.
Final Thoughts on Using Behavioral Data to Bridge the Say/Do Gap
The say/do gap is structural, and no survey redesign fully closes it. What you can do is stop treating stated preference as the whole answer and start reading it alongside what buyers actually did. Audit one tracker question, pull cross-retailer reviews against your top SKU, and line up DTC repeat against retail sell-through. Those three moves cost nothing and start changing how the team reads signal. When that work outgrows a spreadsheet, Merciv joins the behavioral and stated layers in a single query with every source cited and confidence-scored.
FAQ
What is the say/do gap and why does it keep showing up in sustainability and clean beauty research?
The say/do gap is the distance between what a consumer reports in a survey and what they actually do at the shelf or in the reorder cycle, and it runs hottest in categories where the socially attractive answer is obvious. In sustainability and clean beauty, stated intent readings inflate by design because the survey never models the real friction: sticker price, promo timing, out-of-stock, or a competitor at eye level. The gap is not a sign that consumers are lying; it is a sign that the survey instrument was never built to capture a Tuesday-night basket.
How can you reduce the value-action gap in a CPG tracker without rebuilding the entire research instrument?
Start with one audit before the next fielding cycle: pull the live tracker instrument and flag every question where the socially desirable answer is obvious, then rewrite those items as forced trade-offs with real cost attached or replace them with a behavioral proxy from POS or cross-retailer review data. That single pass narrows the gap more reliably than adding survey questions, because it removes the design conditions that inflate aspirational answers in the first place.
What behavioral signals actually close the say-do gap that surveys miss?
Repeat-buy rates, cross-retailer review verbatims, and unprompted social language each capture revealed preference that stated-intent surveys structurally cannot. Reviews post within days of purchase, so complaint clusters surface on their own timeline, typically three to six weeks before a syndicated velocity read catches the same signal. DTC repeat rate versus retail sell-through on the same SKU over the same two quarters is the fastest live diagnostic: if DTC repeat is climbing while retail softens, the problem is distribution or shelf execution, not demand. No concept test will surface that split.
Is Merciv useful for diagnosing the say/do gap, or does closing the value-action gap still require separate survey and behavioral data tools?
The gap persists at most brands because the two halves of the answer (tracker data in one vendor, reviews in another, POS in a third, syndicated behind a fourth login) are never joined against the same question on the same day. Merciv runs that join: cross-retailer reviews, social conversation, internal POS, and licensed syndicated research pulled in a single cited query with a three-tier confidence score (High, Directional, Exploratory) and a clickable audit trail on every finding. The tracker still has a role; what changes is that revealed preference is no longer a separate workstream assembled by hand a week after the readout.
Why is there a gap between what consumers say and what they do, even when surveys use rigorous methodology?
Four mechanisms do most of the work regardless of how the survey is designed: social desirability bias (answering a tracker is a social act), hypothetical framing (surveys strip out price, promo, and shelf context), recall decay (memory reconstructs a tidier basket than the actual one), and situational override at the moment of purchase. These are structural properties of survey research, not execution failures. The instrument was built to measure attitudes, not to predict a specific purchase occasion. Recognizing that boundary is what separates a fixable design problem from a method limitation that requires a behavioral data layer alongside the tracker.