The survey has closed and the export is sitting in a spreadsheet. Six hundred rows, twelve questions, three of them open text, and a meeting on Thursday that expects findings. The analysis usually goes wrong in the first twenty minutes, when the pivot tables start before anyone has decided what the survey was supposed to settle.
What follows is the order that produces conclusions somebody can act on, along with the places where survey data quietly misleads the person reading it.
Decide what the analysis has to settle
Write down the two or three decisions waiting on this survey, in plain sentences, before opening the data. Whether to keep the current onboarding call. Whether the complaints about turnaround time come from one segment or all of them. Whether to hire in support now or in the next quarter.
This matters for a practical reason: survey data will produce an endless supply of true, uninteresting statements. Sixty-two percent of respondents selected email as their preferred channel is true and changes nothing. The decisions list is what separates a finding from a fact.
It also sets the required precision in advance. A decision that turns on a rough direction needs far less certainty than one that turns on a five-point difference between two segments, and knowing which one is in play stops the analysis from either over-claiming or stalling. If the decisions list cannot be written, the survey was probably not ready to send, and the honest outcome of the analysis may be a better second survey.
One more thing belongs on that page: what result would change the decision. Committing to that in advance is the cheapest protection against reading the data to confirm what was already believed.
Clean the data first
Every export contains records that should not be counted. Removing them takes half an hour and changes the numbers more than most of the analysis that follows.
- Test submissions. Entries made during setup, usually with placeholder text. Delete them and note how many.
- Duplicates. The same person submitting twice, common after a reminder or a shared link. Match on email address where one exists and keep the first or last submission consistently.
- Out of scope respondents. Colleagues, partners and internal addresses that came along with a list export.
- Straightliners. Records where every rating is identical down a long grid, particularly where the grid mixes positive and negative statements. These are not opinions.
- Speeders. Submissions completed in a fraction of the median time. If the median is four minutes, a fifty second completion did not involve reading.
- Partials. Decide the rule and apply it to the whole dataset: either include partials question by question, so each question is analysed on the number of people who actually answered it, or exclude them entirely. Including them for one question and excluding them for the next produces a report that contradicts itself.
Keep the removed records rather than deleting them outright, and write down the count and the reason. Somebody will ask why the report says 574 when the tool says 600, and the answer needs to exist.
Counting closed questions without misleading yourself
The three types of answer support different claims, and confusing them is the most common analysis error.
| Answer type | What it produces | What it can support | The usual mistake |
|---|---|---|---|
| Single choice | Counts and shares | A distribution, and comparisons between groups | Reporting shares on a tiny base |
| Multiple choice, select all | Counts per option | Which options matter, in order | Adding the percentages, which will exceed 100 |
| Rating scale | An ordered distribution | Direction and change over time | Averaging as if the gaps between points were equal |
| Open text | Coded themes with counts | What people actually raised, in their words | Reading a memorable quote as a pattern |
Two rules do most of the work. State the base with every percentage. Not the survey base, the question base: the number of people who answered that specific question after cleaning. Percentages quoted without a base are unfalsifiable, and skipped questions mean the base changes from question to question.
Below about thirty answers, report counts and not percentages. With 14 respondents, one person is seven points. A shift from 29 to 36 percent is one person changing their mind, and presenting it as a seven point rise invites a decision nobody would make if the sentence read four of fourteen instead of five of fourteen.
Cross-tabulations are where surveys earn their keep, because a total is rarely actionable and a difference between two groups usually is. Cut by the attributes already known from the list rather than by attributes the survey asked about, where possible: tenure, plan, region, volume, how the person arrived. Those are recorded accurately, while self-reported attributes carry their own error. Before treating a gap between two groups as real, look at the two bases. A 20 point gap between a group of 240 and a group of 19 is mostly a statement about the group of 19.
Rating scales: what they support and what they do not
Scale questions are the most used and most over-read part of a survey. A five point agreement scale is ordinal: the points are in order, but the distance from 2 to 3 is not known to equal the distance from 4 to 5. Averaging them produces a number that feels precise and is not, and a mean of 3.8 hides whether the distribution is a tight cluster in the middle or a split between enthusiasts and detractors.
Two presentations survive scrutiny:
The full distribution. Five bars, with the counts on them. It takes more space and it cannot be misread.
Top box or top two box. The share choosing the highest point, or the top two combined, reported with the base. This compresses the scale into one comparable number without pretending the intervals are equal, and it is stable enough to track across survey rounds.
Means are acceptable for tracking the same question over time, where the interest is the direction of movement rather than the absolute value, as long as the distribution is shown at least once so the reader knows what is moving. What means cannot do is carry a comparison between different questions, or between differently worded scales, or across a change in scale length. A shift from a five point to a seven point scale ends the comparability of the series, and no arithmetic restores it.
Watch for scale effects that are properties of the scale rather than the opinion. Grids with many rows collect less careful answers towards the bottom. Positively worded statements attract more agreement than negatively worded ones asking the same thing. Neither is a finding about the business.
Coding open text so it can be counted
Open answers hold the reasons, which is what the closed questions cannot provide. The problem is that they arrive as prose, and a handful of vivid comments will dominate the reading unless they are counted.
The method is straightforward. Read the first fifty answers without coding anything, just noting recurring subjects. Turn those into a code frame of roughly eight to fifteen labels, kept concrete: slow reply to enquiries, price of the second seat, missing export format. Then code every answer against the frame, allowing more than one code per answer, and add codes as genuinely new subjects appear. Once the frame stabilises, count the mentions.
What comes out is a ranked list of subjects with numbers attached, which can sit beside the closed questions in the same report. Twenty-two of 180 comments raised the delay before a first reply is a sentence that supports a decision. Several customers mentioned delays is not.
Three cautions. Keep the verbatim text next to the counts, because the exact wording is often the most useful output of the entire survey. Do not code sentiment as the primary dimension, since positive and negative counts on the same subject are less useful than knowing the subject came up at all. And if a language model is used to help group the answers, treat its output as a first pass to be checked by reading, not as the count itself, because the counts are what the decision rests on.
Check who is missing before believing any of it
Every survey result is a statement about the people who answered. Whether it extends to the people who did not is a separate question, and it is answerable without extra research.
Compare the respondents with the full invited list on attributes known for both: plan, tenure, region, size, activity level. Where the distributions diverge, the survey over-represents someone. A satisfaction survey answered mostly by customers active in the last month says very little about the ones who quietly stopped using the product, and those are usually the ones the decision is about.
Two reporting habits follow from this. State the known skew in the findings rather than in an appendix, in one sentence: respondents skew towards customers active in the last quarter, who are 71 percent of answers and 44 percent of the list. And where the skew is severe, say what the survey cannot answer. A report that names its own limits gets trusted on the parts it does claim.
Response rate is a rough proxy for this risk, not a measure of it. A high rate with a lopsided respondent profile is worse than a modest rate with a representative one.
Getting from findings to follow-up
Analysis produces two outputs, and most reports only handle one. The aggregate goes into the findings. The individual answers that need a reply go nowhere, and that is the part respondents notice.
Open text frequently contains direct questions, specific complaints and requests that are addressed to a person, not to a chart. If those sit in a spreadsheet column after the presentation, the next survey to the same list gets a lower response rate for reasons no redesign will fix. Handling them means each response keeping an owner and a state, so a comment that needs answering can be assigned, replied to and closed while the analysis proceeds in parallel. That is what response management means in practice, and the uses where it matters most are the ones where answers are individual: applications, support requests, feedback forms with a name attached.
The practical setup is to code the open text for themes and, in the same pass, flag the records that need a human reply. One pass, two outputs, and nothing lost.
What to change first
Before touching the export again, write the two or three decisions the survey is meant to settle and what result would change each one. Then clean the data and record what was removed, because every number afterwards depends on it. If open answers from the last survey were never replied to, assign them now: keeping the analysis and the follow-up in one place, as Halict does per response, is what stops the next round from arriving smaller.
Q1. How many responses are needed before survey data can be analysed?
It depends on how the result will be used and how many subgroups it will be cut into. A single overall question needs far fewer answers than a comparison between four segments, since each cut divides the base. As a working rule, report counts rather than percentages for any group under about thirty, and treat differences between small groups as questions to investigate rather than findings.
Q2. Is it acceptable to average a five point rating scale?
For tracking the same question over time, a mean is workable as long as the distribution is shown at least once. For comparing different questions or different scales it is not, because the intervals between scale points are not known to be equal. Top two box percentages are the safer summary and stay comparable across rounds.
Q3. What is the best way to handle answers that skipped questions?
Analyse each question on the number of people who actually answered it, and state that base beside every figure. Dropping every record with any missing answer usually throws away far more data than the gaps are worth. The important part is that the rule is the same for every question in the report.
Q4. Should a language model be used to analyse open ended answers?
It is useful for a first grouping of a large volume of comments and for drafting a code frame. The counts that feed a decision should come from coding that a person has checked, because an unverified grouping can merge two distinct subjects or invent a theme, and neither is visible in the output. Keep the original text alongside whatever summary is produced.
Q5. How should conflicting results between two questions be reported?
Report both, with their bases, and say plainly which one the decision should rest on and why. Contradictions usually come from question wording, from different bases, or from a filter applied to one question and not the other, so checking those three explains most of them. Hiding the contradiction is the one option that damages the report.