Mystery shopping advantages and disadvantages

Understanding the mystery shopping advantages and disadvantages matters when you use the method to make decisions about stores, service standards or frontline performance. Mystery shopping can show whether a defined process happened. But a scheduled visit also gives you a limited number of observations, which affects what you can confidently infer about everyday customer experience.

Mystery shopping gives structured, step-by-step evaluation of whether defined service standards are being followed, which customer surveys cannot always do. Its limitations include cost per visit, small sample sizes, differences between evaluators, reporting delay and the fact that shoppers are observing an experience rather than simply living it as normal customers.

That does not make mystery shopping a bad method. It means the method needs to be used for the jobs it is good at.

What is mystery shopping?

If you are asking what is mystery shopping or how does mystery shopping work, the simplest definition is a structured observation of a customer experience by someone following a predefined brief. Academic researchers describe mystery shopping as a “participant observation method”. The shopper acts like a customer, follows a scenario, observes specific parts of the service process and records the results.

A typical mystery shop report might record whether an associate:

  • greeted the shopper
  • asked a discovery question
  • recommended a product
  • offered an additional item
  • explained a promotion correctly
  • followed a required process
  • completed a compliance step

That is different from asking a customer whether the experience was good, easy or helpful. A real customer can tell you how an interaction felt. But they may not remember whether every step in a seven-stage service process took place. A mystery shopper is specifically briefed to look for those details.

Mystery shopping is also an established research and operational discipline. In their 2019 Journal of Retailing paper, Gerald Blessing and Martin Natter cite historical MSPA estimates of $2 billion in worldwide mystery shopping spend in 2016 and around 1.5 million mystery shoppers worldwide.

Those are historical figures rather than current market estimates, but they show the scale the method had already reached. MSPA Global represents the industry internationally and sets professional and ethical standards for participating mystery shopping companies.

How much does mystery shopping cost?

How much mystery shopping costs depends on the scope of the program. The number of locations, visit frequency, shopper requirements, reimbursements, reporting depth and complexity of the scenario can all affect total cost.

But cost per visit is not necessarily the most useful number. If the purpose of a mystery shopper program is to support decisions at store level, the more important question is how much it costs to generate enough usable observations at each store.

Increase the number of stores, shifts, dayparts or service scenarios you want to compare and the amount of fieldwork required rises with it. One of the practical tensions in mystery shopping is that more observations can strengthen the evidence behind a decision, but every additional observation requires another visit.

The advantages of mystery shopping

The advantages of mystery shopping and many of the practical benefits of mystery shopping come from structure. A trained evaluator can enter every location with the same shopper brief, follow the same scenario and assess the same predefined standards. That gives the method several genuine strengths.

It can check specific operational standards

One of the main mystery shopping benefits is its ability to answer objective operational questions.

  • Was identification requested?
  • Was the required disclosure given?
  • Was a particular product mentioned?
  • Was the promotion explained correctly?
  • Was a specified process followed?

Those are different questions from asking, “How satisfied were you?” Mystery shopping is therefore well suited to compliance checks, regulated processes, brand standards and operational audits where the business needs evidence that a specific action occurred.

It can evaluate a process step by step

A key advantage of mystery shopping is that the shopper knows exactly what to look for. Customers experience service naturally, so they are unlikely to notice whether every step of a service model was followed or whether a specific phrase was used.

Blessing and Natter identify this as an important strength of the method. Mystery shoppers can notice detailed aspects of personal selling that normal customers may not observe or remember later. That makes mystery shopping useful when a business wants to understand whether a detailed service process is being delivered as designed.

It provides a consistent instrument across locations

Another benefit of mystery shopping is the ability to use the same scorecard across multiple stores, formats or regions. The same questions can be asked in Atlanta, Chicago and Los Angeles. The same scoring framework can be applied across the estate. That can help operations teams identify obvious execution gaps and give field leaders evidence for coaching conversations.

Professional standards also matter. MSPA Global maintains common professional and ethical standards for the industry, helping establish expectations around how mystery shopping research should be carried out.

It can inspect things customers would never report

Another advantage of mystery shopping is that trained evaluators can inspect things normal customers would never be expected to notice. A customer is unlikely to tell you whether every sign was placed correctly, whether a detailed merchandising standard was followed or whether every step of a regulated interaction happened. A mystery shopper can be briefed specifically to look for those things.

It can be used for competitor visits

A mystery shopper program can also be used to observe competitor experiences using a consistent framework. That can make it useful for benchmarking service processes, merchandising, promotions or other observable parts of the customer journey. This distinction matters as continuous customer feedback cannot simply replace every use of mystery shopping.

The disadvantages of mystery shopping

The disadvantages of mystery shopping become more important when businesses use the method as a broad measure of everyday customer experience or store performance. The biggest limitations of mystery shopping are often not about the quality of the checklist. They are about how much evidence sits behind the score and how reliably different evaluators produce that score.

Each additional observation requires another visit

One of the practical limitations of mystery shopping is that every extra observation requires more fieldwork. Every additional data point needs another shopper, another visit and another completed evaluation. That naturally limits how frequently many organizations can measure each store.

The problem becomes more noticeable when a retailer wants to understand execution not only by location, but by weekday, weekend, morning, afternoon, evening or shift. A store-level score may look useful until you start asking what sits underneath it.

The sample can become very small

Small sample sizes are another major disadvantage of mystery shopping when the data is used to make granular operational decisions. One mystery shop is one observation. Four mystery shops are four observations. If a store receives a score of 70%, you need to know whether that represents the way the store normally operates or simply what happened during a handful of interactions.

The issue becomes more pronounced as the business tries to act at a lower level. A brand-level view requires one amount of evidence. A store-level decision requires more. A store-and-shift decision needs more again. We will come back to the sample-size problem below.

Different mystery shoppers can rate the same experience differently

One of the strongest findings in the mystery shopping research is not that individual shoppers are inconsistent. It is that different shoppers may not agree with each other.

Blessing and Natter found a clear difference between the two. When the same mystery shopper assessed a store repeatedly, their ratings were highly consistent. Across 931 repeat evaluations by 43 shoppers, the mean correlation was .81.

But agreement between different mystery shoppers was much lower. Across 18 subjective salesperson attributes in Data Set 2, the mean ICC(1), a measure of inter-rater reliability, was just .03. The researchers compared this with a threshold of .6 for good agreement. A separate reliability cross-check produced an average G-coefficient of .21.

The practical implication is important. A mystery shopper may be consistent with themselves. But if another trained shopper observes a similar interaction and gives it a very different rating, the resulting store score becomes harder to treat as an objective measure of performance.

Blessing and Natter considered this low level of inter-rater reliability a major root cause of the missing relationship between mystery shopper assessments of salespeople and salesperson sales performance. That matters most when scorecards contain subjective judgments such as friendliness, expertise, responsiveness or service quality.

Subjective scorecards can also create a halo effect

Another limitation of mystery shopping appears when evaluators are asked to distinguish between several subjective qualities at once. In Data Set 2, Blessing and Natter found that ratings across 21 salesperson attributes were strongly correlated. Mystery shoppers appeared to have difficulty separating individual qualities, with responses showing evidence of a halo effect.

In simple terms, an overall positive or negative impression of the salesperson could affect ratings across several different attributes. That is very different from an objective behavioral check such as:

“Was an additional product offered?”

Mystery shopping is strongest when the thing being observed is specific and clearly defined. Reliability becomes harder to maintain as the scorecard asks evaluators to make subjective judgments about several overlapping qualities.

The shopper is not experiencing the visit like a normal customer

Another disadvantage of mystery shopping is that a mystery shopper may behave like a customer without experiencing the interaction in exactly the same way.

  • They arrive with a brief.
  • They know what to observe.
  • Their job is to evaluate the interaction.

A normal customer arrives because they want or need something and are deciding whether to spend their own money. Blessing and Natter make this distinction in their research. Even where shoppers resemble the target customer profile, they do not necessarily have the same personal involvement or decision risk as normal customers. That does not stop them observing whether a process happened, but it becomes more important when their assessment is used as a proxy for how real customers feel.

Point-in-time scores can carry too much weight

A further limitation of mystery shopping is that the resulting score represents a snapshot. If visits happen relatively infrequently, an individual interaction can have a large effect on the overall score assigned to a store.

  • One unusually poor visit can drag the result down.
  • One unusually good visit can disguise inconsistent execution.

This becomes a problem if mystery shop scores feed directly into rankings, bonuses, performance management or coaching. The observation may be valid. The question is whether it should represent thousands of other customer interactions that were never observed.

Mystery shopping vs customer surveys vs transaction-linked feedback

Mystery shopping research, customer surveys and transaction-linked feedback answer different questions. The most useful comparison is not which method is universally best, but which one gives you the evidence required for the decision you need to make.

Method Who is measured? How often? Sample per store What can it evaluate? Cost basis Can results link directly to spend?
Mystery shopping Trained or briefed evaluators Scheduled visits or evaluation waves Usually limited by visit frequency Detailed processes, standards, compliance and observable behaviors Fieldwork and visits
Customer surveys Real customers who choose to respond Periodic or triggered Highly dependent on participation Perceptions, satisfaction, attitudes and open-text explanations Platform, research program or survey activity
Transaction-linked feedback Real paying customers during the transaction Continuously during trading Potentially high volume at individual store level Short targeted experience and behavior questions at scale Platform or program basis

Swipe horizontally to view the full comparison.

A compliance audit and a measure of customer reaction are not trying to answer the same question. Problems arise when one measurement method is expected to do both.

For more on the limitations of traditional feedback collection, see our guides to survey fatigue and customer satisfaction survey response rates.

Do mystery shopping scores predict sales?

The strongest mystery shopping research on whether mystery shop scores predict commercial performance comes from Blessing and Natter. Across data from three service retail chains, the researchers found little evidence that overall mystery shopper assessments predicted sales performance.

Customer evaluations, by contrast, were associated with sales in the study. The findings do not mean mystery shopping has no value, but the study focused on particular consultative retail settings, and its results should not be treated as proof that every mystery shopping program fails. The more useful question is what happened when the researchers looked at specific behaviors rather than broad evaluations.

One frontline behavior stood out

In Blessing and Natter’s mystery shopping research, one observable frontline behavior stood apart from the broader scorecard. In Data Set 2, Blessing and Natter tested a range of salesperson attributes and observable behaviors against sales. These included questions asked, use of sales materials, closing behaviors and whether the salesperson offered an additional product.

Of the mystery shopper variables tested, the additional-product offer was the one behavior that showed a significant positive relationship with sales performance. The effect was modest. But average sales volume was slightly higher in stores where mystery shoppers observed salespeople offering an additional product.

That finding is especially interesting because the same type of frontline behavior regularly appears in TruRating data. For example, at Paradies Lagardère, customers were asked:

“Did a team member suggest an item for you today?”

Paradies operates more than 700 retail stores and restaurants across 92 North American airports. Using TruRating, the business collected around 690,000 responses per month to questions including whether customers had received a product suggestion. Customers who reported receiving a suggestion spent 10% more per transaction. The data also showed a major execution gap. Half of stores were delivering the behavior less than 60% of the time.

These are different studies and should not be treated as direct replications. Blessing and Natter examined store-level sales and mystery shopper observations in a different retail environment. Paradies measured customer-reported behavior connected directly to individual transactions.

But independently, both point toward the commercial importance of making an additional product suggestion, which reveals something useful about mystery shopping. Its instinct about what to measure can be exactly right. The constraint is how often it can measure it. A mystery shopper can show that an additional product was offered during a particular visit. Transaction-linked feedback can ask real customers about that same behavior continuously, then connect each response to the basket, turning a useful behavioral check into an ongoing performance signal.

The sample size problem nobody talks about

The sample size problem in mystery shopping is simple… the more granular the decision, the more observations you need.

Imagine a store receives four mystery visits in a quarter. If one is poor, that single visit represents 25% of the quarterly sample. Now imagine the operations team wants to understand three dayparts. Even if those four visits were distributed evenly, that is only about 1.3 observations per daypart. Split the data again between weekdays and weekends, shifts or different service scenarios and the evidence becomes thinner still.

This concern predates the Blessing and Natter paper. In Unmasking a Phantom: A Psychometric Assessment of Mystery Shopping, Finn and Kayandé examined the reliability and validity of mystery shopping data and questioned whether the small number of visits commonly used could produce dependable store-level measures.

Finn developed that argument further in his 2001 paper, Mystery Shopper Benchmarking of Durable-Goods Chains and Stores, arguing that sufficiently generalizable store benchmarking could require at least 20 visits per store, far above the common practice of two to four visits per observational unit.

Blessing and Natter’s own reliability analysis adds another useful piece of evidence. In Data Set 2, stores received an average of 16.69 mystery shopper ratings. Even at that level, the researchers found that the mean store ratings did not reach their .6 threshold for acceptable reliability.

That does not mean 17 visits is a universal cutoff, or that every mystery shopping program needs the same number of observations. It does show how difficult it can be to produce a stable store-level score when different evaluators do not agree strongly with one another.

There is also an important distinction between reliability and sales prediction. Blessing and Natter separately tested whether the number of observations per store changed the relationship between mystery shop assessments and sales performance.

It did not. So the research supports two different conclusions:

  • First, even an average of around 17 observations per store was not enough to produce sufficiently reliable mean ratings in Data Set 2.
  • Second, having more observations did not explain the missing relationship between mystery shop assessments and sales.

Those are not contradictory findings. One is about whether different evaluators produce a stable store score. The other is about whether that score predicts commercial performance. For context, conventional survey sample-size calculations show how quickly evidence requirements rise when you want greater statistical precision. SurveyMonkey’s sample-size guidance gives around 384 responses for a large population when targeting a 95% confidence level and a ±5 percentage-point margin of error.

That does not mean every mystery shopping program needs 384 visits. The methods are different. It simply illustrates why four observations and hundreds of observations give you very different levels of confidence when you begin slicing results by store, shift or trading period.

And volume is only half the issue. Representativeness matters too. As Pew Research Center’s work on survey response has shown in a different research context, a sample can still contain bias depending on who is represented. More data improves precision. It does not automatically eliminate bias. Before acting on a mystery shop score, the useful question is not simply:

“How many visits did we complete?”

It is:

“Are the observations consistent enough, and numerous enough, to support the decision we are about to make?”

When mystery shopping is the right tool

The advantages of mystery shopping are clearest when the job genuinely requires a trained observer. That includes:

  • compliance and regulated checks
  • age-verification processes
  • competitor benchmarking
  • one-off process validation
  • inspection of merchandising or physical standards
  • detailed observation of multi-step interactions
  • checks involving things normal customers would not notice or reliably report

Continuous customer feedback does not make those use cases disappear. If you need to know whether a legally required process occurred exactly as prescribed, asking thousands of customers whether they enjoyed their visit will not answer the question. Use the method that matches the decision.

How to create a continuous pulse

Continuous measurement changes the picture when the thing being measured is simple enough for a real customer to answer reliably. Instead of asking whether a behavior happened during a scheduled evaluation wave, retailers can build a running pulse of how consistently it happens during normal trading.

There is another useful clue in the mystery shopping literature. Blessing and Natter cite earlier research by Alan Wilson in which practitioners reported that mystery shopping could have at least a short-term impact on service standards, but that improvements could eventually reach a “plateau of no further improvement.”

That tension makes sense. A scheduled visit tells you whether a standard was delivered at one point in time, another visit later tells you whether it happened again, but there may be thousands of customer interactions between those observations. Continuous measurement changes the question from:

“Did this happen when we checked?”

to:

“How consistently is this happening across normal trading?”

The additional-product example makes the difference clear. Blessing and Natter found that offering an additional product was the one mystery shopper behavior in Data Set 2 with a positive relationship to sales. Paradies Lagardère then asked essentially the same operational question of real paying customers at scale.

“Did a team member suggest an item for you today?”

Using TruRating, Paradies collected around 690,000 responses per month. Customers who said they received a suggestion spent 10% more per transaction. Because each response was tied to the relevant purchase, Paradies could move beyond knowing whether the behavior happened. They could see:

  • how often it happened
  • where execution was strongest
  • where execution was inconsistent
  • how customers who experienced the behavior spent
  • where coaching could have the greatest potential impact

The data revealed an estimated $1.8 million recoverable revenue opportunity if the bottom half of stores improved suggestion execution by 25 percentage points. That is a modeled opportunity, not booked revenue. But the example shows what changes when a behavioral question moves from an occasional audit to continuous transaction-linked feedback.

The behavioral insight is similar, however, the measurement model is different. Mystery shopping gives a trained observation at a scheduled point in time. Continuous transaction-linked feedback can show whether the behavior is happening across everyday customer interactions, how execution varies by store or shift, and how the answer relates to the transaction itself.

That is the shift from a snapshot to a running pulse. For more on how the model works, see our guide to transaction-linked feedback.

Using mystery shopping and continuous feedback together

A mystery shopper program and continuous customer feedback do not need to compete for the same job.

  • Use mystery shopping where a trained eye is required.
  • Use continuous customer measurement where the thing you need to understand can be answered reliably by real customers and where volume, store-level granularity or transaction context matters.

A retailer could use mystery shoppers to confirm whether an age-verification process is being followed correctly, then use continuous feedback to understand whether everyday service behaviors are being delivered consistently across hundreds of stores. It could use a mystery shop to inspect a detailed ten-step service process, then continuously measure the two or three moments from that process that appear to matter most to customers or commercial performance.

That combination reduces the pressure on either method to do something it was not designed to do. It provides a clearer view of whether service strategies are being delivered consistently, rather than relying on a periodic snapshot. The higher volume also creates a more continuous stream of behavioral and experience data between larger research studies, without asking an occasional mystery shop score to represent every customer interaction. And because the signal can be broken down by store, shift or daypart, it becomes easier to see where execution is slipping and where coaching or operational support should be focused.

Mystery shopping is often right about what matters.

  • Specific behaviors matter.
  • Standards matter.
  • Execution matters.

The next question is whether a handful of scheduled observations gives you enough evidence for the decision you need to make.

If the missing part of your current program is a continuous, store-level view of customer experience connected to actual transaction outcomes, see how TruRating works in store.

FAQ

Frequently asked questions

Answers to common questions about the advantages, disadvantages and practical uses of mystery shopping.

What are the main disadvantages of mystery shopping?
The main disadvantages of mystery shopping are small samples, the cost of adding more observations, low agreement between different evaluators, reporting delay and the fact that a shopper experiences the visit differently from a normal customer. These limitations matter most when a small number of visits are used to make decisions about individual stores, shifts or performance.
What are the advantages of mystery shopping?
The main advantages of mystery shopping are structured, step-by-step evaluation against defined standards, compliance verification, competitor visits and consistent evaluation criteria across locations. A trained shopper can also observe detailed parts of a service process that normal customers may not notice, remember or be able to report reliably.
Do mystery shopping scores predict sales?
Not consistently. A 2019 Journal of Retailing study found little evidence that overall mystery shopper assessments predicted sales across the retail settings studied. The researchers identified low agreement between different mystery shoppers as a major reason. One notable exception was offering an additional product, which showed a small positive relationship with sales.
How much does mystery shopping cost?
Mystery shopping costs depend on factors such as the number of locations, visit frequency, shopper requirements, reimbursements, scenario complexity and reporting needs. Rather than looking only at cost per visit, businesses should consider the cost of generating enough usable observations at each store or operational level where they want to make decisions.
How many mystery shops do you need for reliable results?
There is no fixed number for every mystery shopping program. In one dataset studied by Blessing and Natter, an average of 16.69 mystery shopper ratings per store was still insufficient to raise the reliability of mean store ratings to their .6 threshold. The number needed depends on the decision, the consistency between evaluators and the level at which results will be used.
Can customer feedback replace mystery shopping?
Partly, but not for every job. High-volume customer feedback can provide a more continuous view of real customer experience and operational behaviors, especially when responses are linked to transactions. Mystery shopping remains better suited to trained-observer tasks such as compliance checks, competitor visits and detailed audits that customers cannot reliably perform.
Is mystery shopping still worth it in 2026?
Yes, when it is used for a job that benefits from a trained observer. Mystery shopping remains useful for compliance, brand-standard checks, competitor benchmarking and detailed process audits. Its limitations become more important when a small number of scheduled visits are used as the primary measure of everyday customer experience or store-level performance.
Author

TruRating

Real people, trusted feedback.
At TruRating, we capture real-time, transaction-linked feedback at scale. Integrating with point of sale systems and other touchpoints, we provide retail businesses with reliable customer insights to drive improvements, enhance experiences, and boost performance.

Related content

TruRating for business

Take a more open approach to customer feedback

Share this page

Link copied

Download report

Improving ATV – how to get your frontline to think “sales”, not just “service

atv guide

Request a demo

Connect with a TruRating representative for more information about our solutions.

You’ll be in great company...

Aldi-500x200 1

We use TruRating to confirm what customers actually notice and respond to positively, which allows us to quickly roll out a plan based on real data. Then we can double down on what works across our stores.

ALDI

JD-500x200 1

“With other CX programmes, the stats never matched what we observe in store because of low response rates. With TruRating, the numbers make sense.”

JD Sports

bealls-icon

“TruRating provides a consistent, up-to-date view of performance across every channel, it brings our business together in a way no other tool has.”

bealls