Did Last Month's Fix Actually Work? How to Measure a Service Change
You changed something last month — staffing, a process, the menu. Now you're checking your overall rating to see if it worked, and it almost never tells you anything. Here's a method that actually isolates one change from everything else happening at the same time.
- 1. Why the Overall Rating Almost Never Answers This Question
- 2. The Right Metric: Complaint Share, Not the Score
- 3. You Need a Before, Not Just an After
- 4. Choosing the Window: Not Too Short, Not Too Long
- 5. What Will Fool You: Four Common Traps
- 6. "We Can't Tell" Is a Legitimate Result
- 7. A Six-Step Process You Can Run on the Next Fix
Why the Overall Rating Almost Never Answers This Question
You made a change — hired another server for the dinner rush, rewrote the return process, took a slow-moving dish off the menu. A few weeks later, you check your overall rating to see if it moved. It usually hasn't, or it's moved so little you can't tell if it's the fix or noise. This isn't because the fix failed. It's because the overall rating is the wrong instrument for this question.
Three things make it the wrong tool:
It's a cumulative average, not a current reading. Your rating today is built from every review you've ever received, weighted (loosely) by however your platform calculates it. A location with 400 historical reviews and a 4.3 average needs a sustained run of near-perfect reviews to move that number by even a tenth of a point. The change you made last month is one recent data point diluted into a much larger pool.
It reflects everything, and you changed one thing. The overall rating is downstream of food quality, price perception, cleanliness, staff friendliness, wait times, parking, noise level, and a dozen other factors, all mixed into a single number. If you fixed wait times, that's one input among many — and if something else got slightly worse at the same time (a new hire still learning the register, a supplier substitution), it can cancel out entirely in the aggregate.
It moves too slowly to be useful for decisions. By the time an overall rating shifts enough to be meaningfully different, months may have passed. You'll have long since moved on to the next problem, and you won't be able to reconstruct whether the rating change three months ago was due to the fix you made, or something else entirely.
None of this means the overall rating is useless — it's a fine long-horizon health indicator. It's just the wrong tool for the specific question "did the thing I changed actually work." For that, you need a measurement that responds to one variable, not ten.
You Need a Before, Not Just an After
Complaint share only works as a comparison. If you didn't measure it before you made the change, you have nothing to measure it against afterward — you'll be looking at a single number with no reference point, which tells you almost nothing about whether it's good or bad relative to where you started.
This sounds obvious, but it's the step most operators skip, because the moment you decide to fix something, the instinct is to fix it immediately, not to spend ten minutes documenting the current state first. That ten minutes is the difference between having a measurement and having a guess.
Before you make the change, go back through your last few weeks of feedback — reviews, survey responses, direct comments, whatever you collect — and count two things: total pieces of feedback in that window, and how many of them mention the specific issue you're about to address. Write both numbers down along with the exact date range. That's your baseline. It doesn't need to be sophisticated; it needs to exist.
If you're already mid-fix and realize you skipped this step, the honest move is to say so — "we don't have a clean baseline, so we can only speak to the trend going forward from here" — rather than reconstructing a baseline from memory or cherry-picked examples that will bias the comparison toward whatever answer you want to find.
Choosing the Window: Not Too Short, Not Too Long
The measurement window is where most attempts at this fall apart in one of two opposite directions.
Too short, and you're measuring noise, not a trend. A handful of reviews right after a change is a small, unrepresentative sample — a couple of unusually happy or unusually annoyed customers can swing a small sample by a wide margin without meaning anything. It also catches customers who haven't fully experienced the "after" state yet, especially if the change is something like a new menu that takes a few weeks to work through existing customer habits.
Too long, and other things start happening inside your window that have nothing to do with your fix — a seasonal shift, a new competitor opening nearby, a staffing change unrelated to the one you're measuring, a price increase. The longer the window, the more of these accumulate, and the harder it becomes to say the change in complaint share is attributable to your specific fix rather than to everything else that happened during that stretch.
There's no universal number of days that works for every business — it depends entirely on how much feedback you normally get. The practical rule is: make the post-change window long enough to collect roughly as much feedback as your baseline window had, using a comparable period of the business cycle where possible (weekday-heavy vs. weekend-heavy weeks, same rough season, no major promotions running in one window but not the other). A high-volume location might reach that in a week. A low-volume one might need a full quarter. Match the window to your data, not to the calendar.
If, after a reasonable window, you still don't have enough feedback to say anything with confidence, you have three honest options: extend the window further, pool data across comparable locations that got the same change, or conclude — accurately — that this particular change can't be verified with the volume you currently collect. All three are better than forcing a conclusion out of a handful of data points.
What Will Fool You: Four Common Traps
Even with the right metric, the right baseline, and a reasonable window, there are specific ways this measurement can mislead you. Watch for these before drawing a conclusion.
| Trap | How to Spot It | What to Do |
|---|---|---|
| Seasonality | The improvement lines up suspiciously well with a season known for fewer complaints anyway (e.g. comparing a slow summer to a busy holiday season). | Compare against the same period last year if you have that history, or against a comparable-season window rather than the immediately preceding one. |
| Simultaneous changes | You changed staffing and the menu and the layout in the same stretch, and complaint share for one topic dropped. | Stagger changes by at least one full measurement window when possible. If you can't, report the result as attributable to the combined set of changes, not to any single one. |
| Change-induced friction | Complaints about the new thing spike right after rollout, even though the underlying problem is improving — a new process is unfamiliar to both staff and customers at first. | Give a short settling period (days to a couple of weeks depending on complexity) before starting the "after" measurement window, and expect an initial bump that isn't the final state. |
| Regression to the mean | You made the change right after an unusually bad month, and the following month looks better — but an unusually bad month was likely to be followed by a more typical one regardless of what you changed. | Compare against a longer historical baseline (several months, not just the one bad month) so you're not anchoring to an outlier. |
None of these mean you should give up on measuring — they mean you should hold the result loosely until you've checked for them. A clean-looking before/after difference that lines up with one of these four patterns deserves a second look before you announce success or failure.
"We Can't Tell" Is a Legitimate Result
After running this properly — real baseline, reasonable window, checked against the traps above — sometimes the honest answer is that the complaint share didn't move in either direction, or moved by an amount too small to distinguish from noise given how little feedback you collected. That's not a failure of the method. It's information.
The temptation at that point is to reach for a story anyway — to point at a small dip and call it evidence, or to point to a single glowing review and call it proof. Resist that. A small, ambiguous movement in a low-sample metric is not evidence of anything, and treating it as a win sets you up to either repeat a change that didn't actually help, or to stop paying attention to a problem that's still there.
When the result is genuinely inconclusive, you have three honest paths forward, and they're not mutually exclusive:
- Extend the observation. If the sample was too thin, keep collecting and revisit the comparison once you have more data.
- Change the metric. Maybe complaint share was the wrong specific topic to track, or the issue is better captured through a direct question on a survey than through open-ended feedback mentions.
- Accept that the change may not have mattered. Not every operational tweak moves customer perception, even when it feels like it should from the inside. If a fair test says no measurable difference, that's useful — it tells you to stop investing further effort there and look elsewhere.
Any of these is more honest, and more useful to your operation, than declaring victory on a result you can't actually defend.
A Six-Step Process You Can Run on the Next Fix
Putting the above together into something repeatable:
- Name the specific issue and the specific change. Not "improve service" — "reduce complaints about wait time at the register" by adding a second cashier during the 12–2pm rush.
- Record the baseline before you act. Pull the last few weeks of feedback, count total pieces and how many mention the specific issue, and write down the exact date range.
- Make the change, and give it a settling period. A few days to a couple of weeks, depending on how disruptive the change is, before you start counting the "after" period — this avoids mistaking initial friction for the real result.
- Collect the "after" window until it's roughly the same size as the baseline sample, using a comparable period of the business cycle where you can.
- Compare complaint share, not counts — and check the result against the four traps above (seasonality, simultaneous changes, initial friction, regression to the mean) before trusting it.
- Report the honest outcome — improved, unchanged, or inconclusive — and decide the next action based on that, not on what you hoped to find.
This isn't a statistical framework that needs special software — it's a discipline of measuring the same specific thing, the same specific way, before and after, and being honest about what the comparison actually shows.
For step 2 through 5, the hard part in practice is usually operational, not analytical: pulling feedback from multiple channels — reviews, surveys, direct comments — into one place where you can actually count what mentions what. If you're currently doing that by scrolling through review pages manually, that's the part worth fixing first. Turning scattered complaints into a specific operational change goes deeper into that handoff from raw feedback to the issue you decide to act on.
This is where a unified feedback inbox with topic, sentiment and severity classification earns its keep — it's what makes counting "how many pieces of feedback mentioned wait time this month" a query instead of an afternoon of manual reading, and recurring-issue detection helps confirm you're tracking the same underlying problem rather than a string of unrelated one-offs. To be direct about what this doesn't do yet: running the actual before/after comparison in this article — pulling your baseline number and your after-the-fix number and comparing them — is on you today. Automated before/after verification isn't built yet; for now, the six-step process above is something you run using the platform's inbox, classification and location-comparison views as your data source, not a report it generates for you.
References
- [1]Google Business Profile Help: Reviews — Google
- [2]Google Business Profile: Edit Your Profile — Google
- [3]Local Business Structured Data — Google Developers
- [4]Review Snippet Structured Data — Google Developers
- [5]Google Reviews Policy — Google
- [6]Creating Helpful, Reliable, People-First Content — Google Search Central
Frequently Asked Questions
How long do I need to wait before I can tell if a fix worked?+−
What if we barely get any feedback — how do we measure anything?+−
Can we change two things at once and still measure each one?+−
The overall rating didn't move, but complaints about the issue went down. Did the fix work?+−
What's the difference between "the fix didn't work" and "we can't tell yet"?+−
Should we ever go back to just watching the overall rating?+−
Know what needs attention.
OwnCrew Customer Ops brings customer feedback from every location into one place, shows you what keeps coming up, and tracks whether the fix worked.
Start 14-day free trial
