Comparing Location Performance: How to Tell Which Store Actually Has a Problem
The monthly ranking spreadsheet — sort every location by rating, call out the lowest — feels rigorous but fails in specific, predictable ways. Here is what to compare instead: sample size, distribution shape, actionable leading indicators, and grouped comparisons that account for real differences between stores.
- 1. The Ritual That Feels Rigorous but Measures the Wrong Thing
- 2. Problem One: Different Volumes Make Averages Incomparable
- 3. Problem Two: The Average Hides the Shape of the Problem
- 4. Problem Three: Comparing the Result Instead of the Thing That Caused It
- 5. Problem Four: Ranking Creates the Wrong Incentive
- 6. Problem Five: Some Differences Are Not the Location's Fault
- 7. From "This Location Has a Problem" to "This Location Fixed a Specific Thing"
- 8. Where OwnCrew Operate Fits
The Ritual That Feels Rigorous but Measures the Wrong Thing
If you run more than one location, you have probably built some version of this: a spreadsheet with every store's name, its average rating, its review count, maybe a response-rate column, sorted worst to best. Once a month, the bottom name gets a phone call or a meeting invite.
It feels like management. There is a number, it goes into a rank, someone is accountable. But the ritual has a structural problem: it treats a single blended average as if it were a fair, apples-to-apples measurement, when in practice it is sensitive to things that have nothing to do with how well a location is actually run — how many people happened to leave feedback that month, how the underlying experiences are distributed, and whether the number itself is even the right thing to be comparing in the first place.
None of this means comparison is a bad idea. Comparing locations is one of the few ways a multi-location operator can tell where attention is actually needed instead of guessing. The problem is what gets compared and how. The rest of this article walks through five specific ways the naive version breaks, and what a more honest version looks like instead — ending with the part that matters most: how to turn "this location has a problem" into "this location fixed a specific thing," and how to check that it actually did.
Problem One: Different Volumes Make Averages Incomparable
Start with the math of an average. A location that receives a small number of pieces of feedback per month has an average that moves in large steps — one strongly negative piece of feedback pulls the whole number down noticeably. A location that receives a large volume per month has an average that barely moves for the same single incident, because it is diluted across many other data points.
Example (illustrative, not a real store): picture an 8-location chain where Location A gets about 8 pieces of customer feedback in a typical month, and Location B gets about 200. One sharply negative piece of feedback at Location A can move its monthly average noticeably. The same single incident at Location B barely moves its average at all — it is one data point among 200. Ranking these two locations against each other on raw average is not really comparing service quality. It is comparing sample size.
This is not a reason to ignore low-volume locations — a real problem can absolutely show up there. It is a reason not to trust a single month's average from a low-volume location as a verdict. A few things help:
- Look at trend, not a single snapshot. A location's average moving in a consistent direction over several months is a much stronger signal than any one month's number.
- Set a minimum sample size before treating a number as meaningful. Below some threshold, treat the average as "not enough data yet" rather than ranking it alongside high-volume locations.
- Look at the distribution, not just the mean — which is the next problem, and arguably the bigger one.
Problem Two: The Average Hides the Shape of the Problem
Two locations can land on the exact same average rating for completely different reasons, and those reasons call for completely different responses.
One location might land there because most feedback clusters in the middle — customers who found the experience fine, nothing more. The other might land at the same average because the feedback is split: a large majority who found it very good, and a smaller share who had a genuinely bad experience. The average looks identical. The underlying situation is not.
The first pattern usually points to a positioning or expectations question — the experience is consistently mediocre, and no single incident explains it. The second points to an operational inconsistency — something is going wrong for a subset of customers, some of the time, and it is severe when it happens. Sending a "raise your average" directive to both locations treats these as the same problem when they need opposite diagnoses: one needs a broader look at the whole experience, the other needs to find and fix whatever is causing the bad-experience cases.
This is why a comparison built only on the mean is not just incomplete — it can actively point you at the wrong location, or point you at the right location for the wrong reason. Looking at how feedback is distributed, not just where it averages out, is what tells you which situation you are actually in.
Problem Three: Comparing the Result Instead of the Thing That Caused It
A rating is an outcome. It is the sum of everything that happened at a location, filtered through whoever decided to leave feedback. That makes it a reasonable high-level indicator, but a poor tool for management, because you cannot hand a location manager a rating and ask them to go fix it. There is no lever labeled "increase the average by 0.2."
What you can hand someone is a leading indicator tied to something specific:
- The share of feedback mentioning a particular topic — wait time, order accuracy, cleanliness, staff behavior — as a percentage of total feedback, tracked over time.
- Response time — how long it takes the location (or the team responsible for it) to acknowledge and respond to feedback, especially negative feedback.
- Repeat-issue rate — whether the same complaint keeps showing up week after week, which distinguishes a one-off incident from a standing operational gap.
These are things a manager can actually act on, because they point at a cause rather than a summary. "Your rating is 4.1" tells a manager nothing about what to do differently on Monday. "Wait-time complaints made up a third of your feedback this month, and that share has been climbing for six weeks" tells them exactly where to look.
The outcome metric — the rating — still matters as a sanity check and a long-horizon indicator. It is just the wrong thing to lead with when the goal is to change what happens at a location, not just to measure it after the fact.
Problem Four: Ranking Creates the Wrong Incentive
The moment a comparison becomes a public ranking with consequences, people optimize for the ranking, not the underlying thing it was supposed to measure. This is not a flaw specific to any one manager — it is a predictable response to being measured and judged on a single number.
In a review-and-feedback context, that optimization tends to show up in a few recognizable ways: being more selective about who gets asked for feedback (steering requests toward customers who seemed happy, and skipping the ones who seemed less satisfied), timing feedback requests around good moments only, or simply asking less often so there is less data to be judged on. Every one of these moves improves the number without improving the experience — which is the opposite of the goal.
A few design choices reduce how strongly a comparison system pulls in this direction:
- Compare trend and topic mix, not a single ranked score. "Is this location improving" is harder to game than "is this location #1."
- Do not tie a public leaderboard to individual consequences. Visibility for learning is different from visibility for punishment.
- Make the leading indicators (topic share, response time) part of what gets reviewed, not just the outcome. It is much harder to game a response-time metric or a topic-mention rate than it is to game which customers get asked for feedback.
- Normalize for volume before comparing anything, so a manager is not incentivized to suppress collection just to reduce their sample size.
The underlying principle: whatever you put in front of a manager as "the number that matters" is the number they will manage toward. Choose it accordingly.
Problem Five: Some Differences Are Not the Location's Fault
A downtown location with heavy foot traffic and a mall kiosk in a residential suburb are not competing for the same customer, and their feedback will not look the same even if both are run equally well. Catchment area, format, size, price point, and even the local competitive set all shape what feedback tends to look like, independent of operational quality.
Comparing these locations directly against each other on any raw metric conflates "different context" with "different performance." A location in a tougher catchment might be doing an excellent job relative to its circumstances and still show up at the bottom of a flat, ungrouped ranking — while a location in an easy catchment coasts near the top without doing anything particularly well.
The fix is grouping before comparing: put locations into cohorts that share the relevant context — same format, similar size, similar catchment type — and compare within the cohort rather than across the whole chain. A location's most meaningful comparison point is often not another location at all, but its own history: is this location better than it was six months ago, regardless of where it ranks against a very different store across town.
Here is a simple summary of what tends to be a fair comparison and what tends to mislead:
| Compare this | Not this |
|---|---|
| A location's trend over several months | A single month's snapshot |
| Topic-level complaint share (e.g., wait time, accuracy) | Overall star rating alone |
| Locations within the same format/size/catchment cohort | Every location against every other, regardless of context |
| A location against its own past performance | A location against a chain-wide leaderboard rank |
| Response time and repeat-issue rate | Raw review count as a proxy for "caring more" |
| Distribution of feedback (how spread out experiences are) | The mean alone |
From "This Location Has a Problem" to "This Location Fixed a Specific Thing"
A fair comparison — trend-based, distribution-aware, grouped by context, built on leading indicators instead of the raw average — gets you to a much better starting point: instead of "Location C is at the bottom," you get something like "Location C has a rising share of complaints about order accuracy over the past two months, concentrated more than similar-format locations." That is a diagnosis, not just a verdict.
The next step is turning that diagnosis into something a specific person can act on: an owner. Not "do better," but "reduce order-accuracy complaints," tied to whatever operational change makes sense for that specific issue — a process change at the point of order, a training refresh, a staffing adjustment during peak hours. The comparison told you where to look; it does not tell you what to do about it, and that part still requires someone who understands the actual operation.
The part multi-location operators most often skip is verification. A month after a change, the instinct is to check whether the overall rating went up. Resist that instinct — at typical monthly feedback volumes, the overall average is too slow and too noisy to reflect a single operational change within four weeks, for exactly the sample-size reasons covered above. The more honest check is narrower and more direct:
- Compare the same period before and after the change, not "this month vs. an arbitrary earlier baseline."
- Look at the specific topic you targeted, not the overall score. If the issue was order accuracy, look at whether the share of feedback mentioning order accuracy actually dropped — not whether the star average ticked up.
- Give it enough time and volume before concluding either way. A single week of quiet on a topic is not evidence it is fixed; a sustained, multi-week drop in that topic's share is.
This closes the loop that the ranking spreadsheet never could: it does not just point a finger at a location, it produces a specific, checkable claim — "we changed X, and topic-share for the related complaint moved from A to B over the following month" — that can be verified or discarded on its own terms, rather than buried in an average that was never sensitive enough to show it either way.
Where OwnCrew Operate Fits
Everything above is a way of thinking about comparison, not a specific tool. But some of it is much harder to do by hand once you are past a handful of locations and more than one feedback channel.
OwnCrew's Operate plan brings feedback from surveys, QR codes, email, and direct submissions — alongside Google reviews — into a unified feedback inbox, so you are not stitching together numbers from separate systems before you can even start comparing. Feedback is run through topic, sentiment, and severity classification, which is what makes a comparison like "order-accuracy complaint share" possible without reading every piece of feedback by hand. Recurring issue and anomaly detection helps surface exactly the pattern this article argues you should be looking for — a topic that keeps coming back at one location — rather than reacting to a single incident. And location comparison and Google Maps performance views let you look across locations with that context built in, rather than a single flat leaderboard.
Worth being direct about scope: turning a diagnosis into an assigned, tracked fix with a due date, or running a formal before/after scorecard, is operational work that Operate supports the inputs for — the unified inbox and topic classification — but does not yet automate end to end. Use the platform to get to a clear, evidence-based diagnosis; the assignment and follow-through is still a management process you run.
If you are trying to figure out whether consolidating feedback across locations is worth it before you commit, our guide on managing reviews across multiple locations covers the broader operational picture, and the ROI calculator can help estimate the time currently spent stitching together manual comparisons. You can see current plan details on the pricing page.
References
- [1]Online Reviews Statistics and Trends — ReviewTrackers
- [2]Online Review Statistics — Podium
- [3]Global Consumer Insights Survey — PwC
- [4]Consumer Insights — Nielsen
- [5]Google Business Profile Help: Reviews — Google
- [6]Google Business Profile: Edit Your Profile — Google
Frequently Asked Questions
How often should I actually look at location comparisons?+−
How many locations before I need a tool instead of a spreadsheet?+−
Should a location manager see other locations' numbers?+−
What if two locations serve completely different neighborhoods — is comparing them even useful?+−
A low-volume location just got one bad review and its average dropped a lot. Is that a real problem?+−
We changed something at one location last month. How do we know if it worked?+−
Know what needs attention.
OwnCrew Customer Ops brings customer feedback from every location into one place, shows you what keeps coming up, and tracks whether the fix worked.
Start 14-day free trial