Service Improvement11 min read

Did Last Month's Fix Actually Work? How to Measure a Service Change

You changed something last month — staffing, a process, the menu. Now you're checking your overall rating to see if it worked, and it almost never tells you anything. Here's a method that actually isolates one change from everything else happening at the same time.

OwnCrew Customer Ops Team/
Section 1

Why the Overall Rating Almost Never Answers This Question

You made a change — hired another server for the dinner rush, rewrote the return process, took a slow-moving dish off the menu. A few weeks later, you check your overall rating to see if it moved. It usually hasn't, or it's moved so little you can't tell if it's the fix or noise. This isn't because the fix failed. It's because the overall rating is the wrong instrument for this question.

Three things make it the wrong tool:

It's a cumulative average, not a current reading. Your rating today is built from every review you've ever received, weighted (loosely) by however your platform calculates it. A location with 400 historical reviews and a 4.3 average needs a sustained run of near-perfect reviews to move that number by even a tenth of a point. The change you made last month is one recent data point diluted into a much larger pool.

It reflects everything, and you changed one thing. The overall rating is downstream of food quality, price perception, cleanliness, staff friendliness, wait times, parking, noise level, and a dozen other factors, all mixed into a single number. If you fixed wait times, that's one input among many — and if something else got slightly worse at the same time (a new hire still learning the register, a supplier substitution), it can cancel out entirely in the aggregate.

It moves too slowly to be useful for decisions. By the time an overall rating shifts enough to be meaningfully different, months may have passed. You'll have long since moved on to the next problem, and you won't be able to reconstruct whether the rating change three months ago was due to the fix you made, or something else entirely.

None of this means the overall rating is useless — it's a fine long-horizon health indicator. It's just the wrong tool for the specific question "did the thing I changed actually work." For that, you need a measurement that responds to one variable, not ten.


Section 2

The Right Metric: Complaint Share, Not the Score

If you changed wait times, don't watch the rating — watch the share of feedback that mentions waiting. Specifically: of all the feedback you received in a given window, what percentage raised waiting as an issue? If that percentage goes down after the fix and stays down, you have a real signal. If it doesn't move, you don't — regardless of what the overall rating did in the meantime.

This works because it isolates the one thing you changed from the nine things you didn't. A customer who waits less won't necessarily leave a five-star review — they might still have an opinion about the food, the price, or something unrelated — but they will be less likely to mention waiting specifically. Complaint share tracks that directly instead of hoping it surfaces in an aggregate score.

Why share, not raw count. This is the part people get wrong most often. Take a purely hypothetical example: wait-time complaints drop from 12 to 8 in a month, which looks like progress — until you notice total feedback volume also dropped from 60 to 30 pieces, because it was a slower month. As a share, that's 20% before and roughly 27% after: in this made-up scenario, the problem actually got relatively worse, even though the raw count fell. Foot traffic, seasonality, promotions, and even how many people you're actively inviting to leave feedback all move volume around for reasons that have nothing to do with whether your fix worked. Dividing by total feedback in the same window removes that noise. Always compare a rate to a rate, not a count to a count.

To make this concrete — purely as an illustration, not a claim about any real result — imagine a location that gets around 40 pieces of feedback a month, and roughly a quarter of them mention slow service before a schedule change. After the change, in a comparable month, the location gets around 35 pieces of feedback and mentions of slow service drop to around one in ten. That's a real signal, because the comparison is share-to-share across comparable windows, not count-to-count across windows of different size.


Section 3

You Need a Before, Not Just an After

Complaint share only works as a comparison. If you didn't measure it before you made the change, you have nothing to measure it against afterward — you'll be looking at a single number with no reference point, which tells you almost nothing about whether it's good or bad relative to where you started.

This sounds obvious, but it's the step most operators skip, because the moment you decide to fix something, the instinct is to fix it immediately, not to spend ten minutes documenting the current state first. That ten minutes is the difference between having a measurement and having a guess.

Before you make the change, go back through your last few weeks of feedback — reviews, survey responses, direct comments, whatever you collect — and count two things: total pieces of feedback in that window, and how many of them mention the specific issue you're about to address. Write both numbers down along with the exact date range. That's your baseline. It doesn't need to be sophisticated; it needs to exist.

If you're already mid-fix and realize you skipped this step, the honest move is to say so — "we don't have a clean baseline, so we can only speak to the trend going forward from here" — rather than reconstructing a baseline from memory or cherry-picked examples that will bias the comparison toward whatever answer you want to find.


Section 4

Choosing the Window: Not Too Short, Not Too Long

The measurement window is where most attempts at this fall apart in one of two opposite directions.

Too short, and you're measuring noise, not a trend. A handful of reviews right after a change is a small, unrepresentative sample — a couple of unusually happy or unusually annoyed customers can swing a small sample by a wide margin without meaning anything. It also catches customers who haven't fully experienced the "after" state yet, especially if the change is something like a new menu that takes a few weeks to work through existing customer habits.

Too long, and other things start happening inside your window that have nothing to do with your fix — a seasonal shift, a new competitor opening nearby, a staffing change unrelated to the one you're measuring, a price increase. The longer the window, the more of these accumulate, and the harder it becomes to say the change in complaint share is attributable to your specific fix rather than to everything else that happened during that stretch.

There's no universal number of days that works for every business — it depends entirely on how much feedback you normally get. The practical rule is: make the post-change window long enough to collect roughly as much feedback as your baseline window had, using a comparable period of the business cycle where possible (weekday-heavy vs. weekend-heavy weeks, same rough season, no major promotions running in one window but not the other). A high-volume location might reach that in a week. A low-volume one might need a full quarter. Match the window to your data, not to the calendar.

If, after a reasonable window, you still don't have enough feedback to say anything with confidence, you have three honest options: extend the window further, pool data across comparable locations that got the same change, or conclude — accurately — that this particular change can't be verified with the volume you currently collect. All three are better than forcing a conclusion out of a handful of data points.


Section 5

What Will Fool You: Four Common Traps

Even with the right metric, the right baseline, and a reasonable window, there are specific ways this measurement can mislead you. Watch for these before drawing a conclusion.

| Trap | How to Spot It | What to Do |

|---|---|---|

| Seasonality | The improvement lines up suspiciously well with a season known for fewer complaints anyway (e.g. comparing a slow summer to a busy holiday season). | Compare against the same period last year if you have that history, or against a comparable-season window rather than the immediately preceding one. |

| Simultaneous changes | You changed staffing and the menu and the layout in the same stretch, and complaint share for one topic dropped. | Stagger changes by at least one full measurement window when possible. If you can't, report the result as attributable to the combined set of changes, not to any single one. |

| Change-induced friction | Complaints about the new thing spike right after rollout, even though the underlying problem is improving — a new process is unfamiliar to both staff and customers at first. | Give a short settling period (days to a couple of weeks depending on complexity) before starting the "after" measurement window, and expect an initial bump that isn't the final state. |

| Regression to the mean | You made the change right after an unusually bad month, and the following month looks better — but an unusually bad month was likely to be followed by a more typical one regardless of what you changed. | Compare against a longer historical baseline (several months, not just the one bad month) so you're not anchoring to an outlier. |

None of these mean you should give up on measuring — they mean you should hold the result loosely until you've checked for them. A clean-looking before/after difference that lines up with one of these four patterns deserves a second look before you announce success or failure.


Section 6

"We Can't Tell" Is a Legitimate Result

After running this properly — real baseline, reasonable window, checked against the traps above — sometimes the honest answer is that the complaint share didn't move in either direction, or moved by an amount too small to distinguish from noise given how little feedback you collected. That's not a failure of the method. It's information.

The temptation at that point is to reach for a story anyway — to point at a small dip and call it evidence, or to point to a single glowing review and call it proof. Resist that. A small, ambiguous movement in a low-sample metric is not evidence of anything, and treating it as a win sets you up to either repeat a change that didn't actually help, or to stop paying attention to a problem that's still there.

When the result is genuinely inconclusive, you have three honest paths forward, and they're not mutually exclusive:

  • Extend the observation. If the sample was too thin, keep collecting and revisit the comparison once you have more data.
  • Change the metric. Maybe complaint share was the wrong specific topic to track, or the issue is better captured through a direct question on a survey than through open-ended feedback mentions.
  • Accept that the change may not have mattered. Not every operational tweak moves customer perception, even when it feels like it should from the inside. If a fair test says no measurable difference, that's useful — it tells you to stop investing further effort there and look elsewhere.

Any of these is more honest, and more useful to your operation, than declaring victory on a result you can't actually defend.


Section 7

A Six-Step Process You Can Run on the Next Fix

Putting the above together into something repeatable:

  1. Name the specific issue and the specific change. Not "improve service" — "reduce complaints about wait time at the register" by adding a second cashier during the 12–2pm rush.
  2. Record the baseline before you act. Pull the last few weeks of feedback, count total pieces and how many mention the specific issue, and write down the exact date range.
  3. Make the change, and give it a settling period. A few days to a couple of weeks, depending on how disruptive the change is, before you start counting the "after" period — this avoids mistaking initial friction for the real result.
  4. Collect the "after" window until it's roughly the same size as the baseline sample, using a comparable period of the business cycle where you can.
  5. Compare complaint share, not counts — and check the result against the four traps above (seasonality, simultaneous changes, initial friction, regression to the mean) before trusting it.
  6. Report the honest outcome — improved, unchanged, or inconclusive — and decide the next action based on that, not on what you hoped to find.

This isn't a statistical framework that needs special software — it's a discipline of measuring the same specific thing, the same specific way, before and after, and being honest about what the comparison actually shows.

For step 2 through 5, the hard part in practice is usually operational, not analytical: pulling feedback from multiple channels — reviews, surveys, direct comments — into one place where you can actually count what mentions what. If you're currently doing that by scrolling through review pages manually, that's the part worth fixing first. Turning scattered complaints into a specific operational change goes deeper into that handoff from raw feedback to the issue you decide to act on.

This is where a unified feedback inbox with topic, sentiment and severity classification earns its keep — it's what makes counting "how many pieces of feedback mentioned wait time this month" a query instead of an afternoon of manual reading, and recurring-issue detection helps confirm you're tracking the same underlying problem rather than a string of unrelated one-offs. To be direct about what this doesn't do yet: running the actual before/after comparison in this article — pulling your baseline number and your after-the-fix number and comparing them — is on you today. Automated before/after verification isn't built yet; for now, the six-step process above is something you run using the platform's inbox, classification and location-comparison views as your data source, not a report it generates for you.

References

  1. [1]Google Business Profile Help: Reviews Google
  2. [2]Google Business Profile: Edit Your Profile Google
  3. [3]Local Business Structured Data Google Developers
  4. [4]Review Snippet Structured Data Google Developers
  5. [5]Google Reviews Policy Google
  6. [6]Creating Helpful, Reliable, People-First Content Google Search Central

Frequently Asked Questions

How long do I need to wait before I can tell if a fix worked?+
Long enough to collect a comparable amount of feedback to your pre-change baseline — not a fixed number of days. A location with a handful of reviews a week needs a longer window than one with dozens a day. As a rule of thumb, don't call it before you have roughly as much post-change feedback as you used to establish the baseline. Checking early is fine as a temperature check, but treat anything before that point as "too soon to say," not as a result.
What if we barely get any feedback — how do we measure anything?+
Three options, in order of preference. First, extend the window — a low-volume location just needs more calendar time to accumulate the same sample a busy one gets in a week. Second, if you operate more than one similar location, pool feedback across the ones that got the same change and compare against ones that didn't. Third, accept that some changes at some locations genuinely can't be verified with the feedback volume you have, and say so — that's a legitimate answer, not a failure to analyze.
Can we change two things at once and still measure each one?+
Not cleanly. If you change staffing and the menu in the same week, and complaints about wait times drop, you cannot attribute that to either change alone — both could be responsible, or one could be masking the other making things worse. If you need to move fast on two fronts, stagger them by at least one full measurement window, or accept upfront that you'll only be able to speak to the combined effect, not each change individually.
The overall rating didn't move, but complaints about the issue went down. Did the fix work?+
That can absolutely count as a working fix. The overall rating is a slow-moving cumulative average — a real improvement in one recurring issue can take a long time to show up there, especially if the location has a long review history. If the share of feedback naming that specific issue dropped and stayed down across a full measurement window, that's the more direct signal. The rating catching up later doesn't make the earlier result less real.
What's the difference between "the fix didn't work" and "we can't tell yet"?+
"Didn't work" means you ran a proper before/after comparison over a reasonable window, controlled for the obvious confounders below, and the complaint share for that topic didn't move. "Can't tell yet" means one of the conditions for a clean read wasn't met — not enough feedback, window too short, something else changed at the same time, or you're still in the early period where a new process naturally generates more complaints. Mixing these up leads to either declaring victory too early or abandoning a fix that just hasn't had time to show results.
Should we ever go back to just watching the overall rating?+
It still has a role — as a long-horizon health check, not as a way to evaluate a specific change. Keep watching it for the big picture. But for the question "did the thing I changed three weeks ago work," it's the wrong tool: too slow, too aggregated, and too exposed to everything else happening in the business at the same time.
Tagsservice improvementmeasuring changecustomer feedbackoperationsbefore and afterroot cause

Know what needs attention.

OwnCrew Customer Ops brings customer feedback from every location into one place, shows you what keeps coming up, and tracks whether the fix worked.

Start 14-day free trial