
Reading 10,000 open-ended comments isn't an option, and a fluent summary can quietly drop the two complaints that mattered most. The fix isn't a better summarizer, it's a pipeline you can check.
Summarization accuracy has two failure modes, and omission is the more common one. Research shows models drop scope-limiting qualifiers far more often than they invent facts, and the risk grows as the corpus grows. Thematic handles it by constraining generation to an already-validated theme model and keeping every theme summary traceable back to the comments underneath, so a skeptical stakeholder can verify any line in seconds.
Any team with real feedback volume hits the same wall. Ten thousand open-ended survey responses arrive, and someone has to say what's in them by Thursday. Reading all of it isn't an option. Reading a sample means the answer depends on which sample. So a summary gets generated, and it reads well. Nobody in the room can tell whether it quietly dropped the two complaints that mattered most.
Accuracy isn't a property of the summarizer. It's a property of the pipeline around it. Thematic summarizes large volumes of feedback accurately in two ways. First, generation is constrained to an already-validated theme model instead of pointing a language model at raw text. Second, every theme summary stays traceable back to the grouped sentences and the full comments underneath it. That combination is what makes a summary checkable, and a summary you can check is the only kind you should stake a decision on.
Below: what "losing accuracy" means at volume, why the risk grows with the pile, where most approaches break, how Thematic handles it, and what to ask a vendor during a demo.
Summarization accuracy usually gets discussed as if there's one failure mode. There are two, and they behave very differently.
The second one is more common and much harder to notice. A 2025 npj Digital Medicine study had clinicians annotate 12,999 sentences of large language model (LLM) output and found a 3.45% omission rate against a 1.47% hallucination rate. Omissions were twice as frequent. Fabrications were rarer but more serious: 44% were graded major errors, against 16.7% of omissions. Those rates are from clinical notes, so the shape of the finding transfers rather than the numbers.
A 2025 Royal Society Open Science study points the same way. Across 4,900 model-written summaries, it found them roughly five times more likely than human-written ones to contain broad generalizations, with three models over-generalizing in 26 to 73% of cases even when the prompt asked for accuracy. The characteristic error isn't invention. It's a dropped qualifier.
For a CX team, that's concrete. The dangerous summary isn't the one claiming customers complained about a feature that doesn't exist. It's the one that says "customers found onboarding confusing" when the comments said "customers on the legacy plan found the new onboarding confusing after the March change."
A technique that works on 200 comments fails on 20,000 for structural reasons, not model quality.
More sources means more hallucination. A 2025 study in the NAACL Findings track tested five models against two purpose-built multi-document summarization benchmarks. Hallucination increased as the number of input documents increased. It increased further when the sources were less coherent or spanned a wider range of topics. A real feedback corpus is exactly that: thousands of short, unrelated, contradictory texts covering everything from shipping to billing to the mobile app.
Evidence in the middle gets ignored. The "Lost in the Middle" research found performance drops by more than 30% when relevant information sits mid-input rather than at the start or end. That was measured on retrieval and question-answering, not summarization, so treat it as a mechanism rather than a statistic. It explains why pasting 5,000 comments into one prompt is a poor architecture, whatever the context window.
Models answer even when there's nothing to answer. In the same study, models were asked to summarize information absent from the source. GPT-3.5-turbo produced a summary anyway 79.45% of the time, and GPT-4o did so 44% of the time. Ask a general-purpose model what customers said about a theme nobody raised, and you'll usually get a fluent paragraph about it.
Set a realistic bar, though. Human coders aren't perfect either: in peer-reviewed research on coding open-ended survey responses, independent coders agreed roughly 75 to 89% of the time, with Cohen's kappa between 0.46 and 0.75. The goal isn't a summary that's provably perfect. It's a summary whose errors you can find.
| Approach | What it gets right | Where it loses accuracy |
|---|---|---|
| Paste comments into a general chatbot | Fast, no setup, readable output | No coverage guarantee, positional bias on long inputs, produces confident answers about themes that were never mentioned |
| Manual coding by analysts | Judgment, context, defensible reasoning | Does not scale past a few thousand comments; independent coders only agree 75 to 89% of the time |
| Recursive chunk-and-merge summarization | Handles unlimited volume | A detail dropped by an early chunk summary is invisible to every later merge step |
| Summarization with no drill-down | Looks identical to the accurate version | Nobody can check it, so a wrong summary and a right one carry the same weight in the room |
There's also a measurement gap vendors rarely mention. Public hallucination benchmarks, such as Vectara's grounded-summarization leaderboard, only score whether a summary is factually consistent with its source. Leading models sat in the 1.8 to 4.1% range there as of its May 11, 2026 snapshot, across more than 7,700 articles. That says nothing about omission or coverage: a summary that silently drops an entire complaint theme still passes a faithfulness check.
Generation is constrained by a theme model, not pointed at raw text. Thematic doesn't plug a language model into a pile of comments. As Thematic's product team described it, "the existing rules and guardrails within Thematic's model guide the LLMs to generate helpful, accurate summaries." The themes are discovered and validated first, and summaries describe that structure. So the model isn't working out what a corpus means from scratch. It's describing a slice of an analysis you can already inspect.
Every theme summary is traceable back to the comments. For theme and sub-theme summaries, a reader can expand any summarized line to read all the grouped sentences, then click a sentence to see the full comment it came from. That closed loop is the real answer to the accuracy question. It doesn't promise the summary is right. It means a skeptical stakeholder can find out in about fifteen seconds, which is a stronger position than any accuracy percentage.
Summaries run on live data, so they reflect the most recent feedback loaded rather than a cached snapshot from the last analysis run.
The scope is yours to set before summarizing. Filter to Issues, Requests, positive or negative sentiment, or by date, segment, and feedback category, then summarize that slice. Narrowing the input first is the practical countermeasure to volume: a summary of one filtered theme is far easier to verify than a summary of everything.
A human keeps the judgment calls. Community Health System, whose team analyzes open-ended employee feedback across 250 departments, is a documented example. Thematic clustered the comments into themes, generated the summaries, and proposed next steps per department. The Organizational Development team handled the judgment calls, and in at least one case replaced a recommendation suggesting the return of a supervisor role the organization had just phased out. The model wasn't wrong about the feedback. It was wrong about the world.
One limit worth knowing: the documented drill-down path applies to theme and sub-theme summaries. If your verification workflow depends on tracing a dashboard-level or score-change summary back to individual comments, ask about that surface specifically.
DoorDash consolidates and synthesizes feedback from tens of thousands of open-ended net promoter score (NPS) survey responses from Consumers, Dashers, and Merchants. Zach Schendel, DoorDash's Head of Research, described the problem plainly: "It's impossible for me, or anyone else, to find anything useful from that information unless it's tamed in some kind of way." What makes the taming trustworthy is that the summaries don't travel alone. Thematic gives DoorDash AI-generated summaries, thematic insights, and direct verbatims together, so a summary and its evidence arrive in the same place.
Community Health System runs an annual employee engagement survey collecting open-ended feedback from staff across 250 departments. With summarization doing the first pass and the Organizational Development team owning the judgment, they produced 250 standardized one-page department reports in a single three-day sprint. Preparation dropped from about an hour per department to about 20 minutes, saving over 160 hours per survey cycle.
Cross-checking against a tool you already trust works too. A New Zealand hardware retail cooperative, whose three-person insights team handles 20,000 comments a month across 84 stores, reported that "the consistency between our Thematic insights and our Qualtrics analysis gave us confidence in the approach." Any team can run that comparison during an evaluation.
You can't summarize thousands of customer comments accurately by finding a better summarizer. You do it by constraining generation to a validated theme structure, keeping every summary line traceable back to the verbatims underneath it, filtering the input first, and leaving the judgment calls to a person. That's Thematic's approach, and what makes its summaries hold up is the drill-down, not a claimed accuracy score.
One test: pick the most consequential sentence in your current tool's last output and see how long it takes to reach the comments that justify it. If you can't get there, you're guessing with better grammar.
Thematic turns fragmented feedback into one consistent source of customer truth — so every team acts on the same customer story. Up and running in days, not quarters.

Transforming customer feedback with AI holds immense potential, but many organizations stumble into unexpected challenges.