
Your surveys can't reach your competitor's customers, but their public reviews can. Here is the five-step method to benchmark your CX against rivals using the reviews they leave in the open.
To benchmark your CX against a competitor, collect their public reviews from every platform their customers use, score every brand against one shared theme taxonomy with theme-level sentiment, convert to share-of-theme so volumes compare fairly, and rank the gaps into a roadmap. Thematic runs this across all brands at once so the output is a defensible head-to-head, not a word cloud.
Your surveys can tell you how your own customers feel. They cannot tell you how you compare to the competitor your customers almost switched to. You never get to survey that rival's customers. That blind spot is getting more expensive. Survey response rates have fallen 7 to 8 percentage points since 2021, and only 16% of customers strongly believe their feedback changes anything. So fewer people are answering at the same time that leadership wants a sharper read on where you actually stand.
The way to close the gap is to treat your competitors' public reviews as the feedback dataset you were never able to collect. You benchmark your customer experience against a rival by pulling their reviews from the platforms customers already use (app stores, Trustpilot, G2, Yelp, Google, and Reddit), analyzing that text the same way you analyze your own feedback, and comparing the two side by side on the themes that move satisfaction. Thematic does this by applying one shared theme structure and theme-level sentiment scoring across every brand's reviews at once, so "we score 4.4 and they score 2.6" becomes "here is the specific theme where they lose customers and we win." Over 80% of customer feedback is already unstructured, and public reviews are the one at-scale external signal you can legally read, so the raw material is sitting there.
This is competitive benchmarking, which is different from internal benchmarking (your own stores against each other) or industry benchmarking (your NPS against a category average). Below is the five-step method, a worked example with real numbers, and the mistakes that make a competitive benchmark misleading.
Pick the two or three competitors your customers realistically consider, then gather their reviews from the platforms where those customers leave feedback. For consumer apps that means the App Store and Google Play. For software it means G2, Trustpilot, and Capterra. For local and service brands it means Google and Yelp, and for almost every category it now includes Reddit and forum threads, which most benchmarking exercises ignore.
Pull enough volume to be representative and time-stamped so you can track trends. Thematic's own public teardowns used 2,000 iOS reviews per app for a travel comparison and 10,000 Google Play reviews per app for a navigation comparison. Gather your own reviews from the same platforms in the same window so the comparison is fair. Manual reading breaks down past a few hundred reviews, so this step assumes automated collection with scheduled imports.
The single most common mistake is analyzing each competitor separately, which produces categories that do not line up. To benchmark, every brand's reviews have to be scored against the same set of themes: onboarding, pricing, reliability, customer support, search, and so on. Thematic builds this taxonomy bottom up from the review text itself rather than forcing a predefined list, then applies the identical structure to every brand, so a "customer support" complaint means the same thing whether it appears in your reviews or a rival's.
A shared taxonomy is what makes the comparison defensible. Without it you are comparing one company's word cloud to another's, and the two never map cleanly onto each other.
An overall star rating hides the story. Google Maps averaged 3.7 out of 5 and Waze 3.8 in a 10,000-review-per-app comparison, a gap that looks like a tie until you score each theme. At the theme level, Waze drew roughly five times as many mentions of voice prompts and ten times as many of real-time traffic, while other themes dragged it down. The average concealed both the strength and the weakness.
So score sentiment for each theme rather than each review, then express it as a share of that brand's reviews. Share-of-theme is what lets brands with very different review counts compare fairly: "6% of their reviews complain about support" is comparable across a brand with 2,000 reviews and one with 20,000, where a raw count is not.
Now put the brands side by side, theme by theme. In Thematic's Expedia versus Booking.com teardown of 2,000 iOS reviews each, customer support was Expedia's weakest theme at 1.5 out of 5 average, with 80% negative sentiment against Booking.com's 50%, and search functionality ran 56.7% positive for Booking.com against 17.4% for Expedia. Those are the head-to-head gaps a survey of your own customers would never surface.
Four traps make a comparison misleading, so control for each:
A benchmark is only useful if it ends in decisions. Rank the gaps by how often a theme appears, how negative its sentiment is, and how recent the complaints are, then act on the top of that list. A competitor's most frequent, most negative, most recent theme is either your wedge, if you are strong there, or your risk, if you are not.
The direction matters both ways. In a Garmin Connect versus Strava comparison of 4,000 App Store reviews, price appeared in 452 Strava reviews against only 80 for Garmin, a standing wedge Garmin can lead with in positioning. Prioritize where a rival's weakness maps to your strength, and fix the themes where the benchmark shows you are the one behind.
The most complete public example is Thematic's analysis of 7,000 App Store and Google Play reviews across three Southeast Asian banks, presented at CX Asia in Kuala Lumpur. Scoring each review against what customers actually want (ease, reliability, value, fairness, and empathy) and grouping the results into themes surfaced competitive gaps no survey had caught. One bank had 25% of its reviewers complaining about intrusive pop-ups in a single month. For another, the analysis quantified more than SGD 40 million in revenue at churn risk. The study also confirmed, from Google Play data, that simply replying to a negative review lifts that review's rating by 0.7 stars on average, a fix any of the three could act on immediately.
| Theme | Booking.com | Expedia |
|---|---|---|
| Overall star rating | 4.4 / 5 | 2.6 / 5 |
| Customer support (avg score) | Higher; 50% negative | 1.5 / 5; 80% negative |
| Search functionality (positive) | 56.7% | 17.4% |
| Pricing (avg score) | 4.3 | 2.5 |
Acting on public-review signal is not hypothetical. Atom Bank, the UK digital challenger, folds its own app-store reviews and Trustpilot alongside support and call data into one score, and after acting on what that signal showed became the highest-rated UK bank on Trustpilot at 4.6 out of 5. The same review text a competitor's customers leave in public is the text you can benchmark yourself against.
To benchmark your customer experience against a competitor, collect their public reviews from every platform their customers use, score every brand against one shared theme taxonomy with theme-level sentiment, convert to share-of-theme so the comparison is fair, and rank the gaps into a roadmap. Thematic runs this across all brands at once so the output is a defensible head-to-head, not a word cloud. Start with one rival and one platform this week: pull their last 500 reviews, theme them, and see which of their weaknesses is your opening.
Thematic turns fragmented feedback into one consistent source of customer truth — so every team acts on the same customer story. Up and running in days, not quarters.

Transforming customer feedback with AI holds immense potential, but many organizations stumble into unexpected challenges.