# Likert Scale: Examples, Response Options, and Analysis Errors

Likert scale examples: choose 5 or 7 points, label options clearly, separate uncertainty, and avoid analysis errors.

- Canonical URL: https://www.harmate.com/en/blog/likert-scale-examples-response-options-and-analysis-errors
- Author: Harmate Team
- Published: 2026-09-09
- Updated: 2026-09-09T18:42:12.695318+00:00
- Language: en

## Content

# Likert Scale: Examples, Response Options, and Analysis Errors

A Likert scale turns a perception into an ordered response, for example from “strongly disagree” to “strongly agree.” It makes answers comparable, but it does not prove why an experience occurred or whether a group is representative. Its value depends on the dimension defined, the labels offered, and the decision that follows.

## Define what the item measures

An **item** is the question or statement shown to a respondent. In everyday usage, “Likert scale” often refers to an item’s ordered response options; strictly speaking, it combines several items intended to measure one construct. The technique described by Rensis Likert in 1932 used a series of attitude statements with ordered responses ([the original article](https://legacy.voteview.com/pdf/Likert_1932.pdf)). A single item with five choices is not automatically a validated scale.

Before choosing options, complete this sentence: “I want to compare ______ in order to decide ______.” If an item mixes two subjects—“The training was clear and useful”—split them. If you are still looking for reasons or exceptions, start with an open-ended question. A closed response compares perceptions; it does not explain on its own why they differ.

The [J-PAL survey-design guide](https://www.povertyactionlab.org/resource/survey-design) recommends starting with the concepts to measure and building clear responses suited to the context. A respondent may report feeling able to apply a method without demonstrating that skill in a task. The wording should state what is requested: a perception, agreement, frequency, intention, or reported behaviour.

## Choose the axis and five or seven points

A 5-point format is often easy to read and quick to complete. A 7-point format can add a distinction when respondents genuinely understand the nuance. Neither is universally better. For an agreement item, you might use:

| 5 points | 7 points |
|---|---|
| Strongly disagree | Strongly disagree |
| Disagree | Disagree |
| Neither agree nor disagree | Somewhat disagree |
| Agree | Neither agree nor disagree |
| Strongly agree | Somewhat agree |
| — | Agree |
| — | Strongly agree |

For frequency, use options from “never” to “very often”; for ease, from “very difficult” to “very easy.” The response axis should follow the item, otherwise respondents must translate an unsuitable scale in their heads.

Research on the number of categories does not yield one recipe for every questionnaire. Revilla, Saris, and Krosnick compared agreement scales in four multitrait-multimethod experiments from the European Social Survey. In their design, five categories produced better measurement quality than seven or eleven ([study abstract](https://doi.org/10.1177/0049124113509605)). That result concerns agreement questions and their quality criteria; it does not condemn every seven-point scale.

Preston and Colman compared formats from two to eleven categories in a different service questionnaire and found trade-offs among reliability, validity, and preference ([PubMed abstract](https://pubmed.ncbi.nlm.nih.gov/10769936/)). The practical question is therefore: which additional distinction will change your decision? If respondents cannot distinguish “somewhat” from “rather,” seven points add noise. If a meaningful intermediate distinction is understood, five points may be too coarse. Pilot the words, not only the visual grid.

Labels matter as much as the number of boxes. In a comparison of response formats, Saris and co-authors studied item-specific options alongside agree/disagree options and found a quality advantage for item-specific options in their design ([response-options study](https://ojs.ub.uni-konstanz.de/srm/article/view/2682)). This is not a universal ban on agree/disagree formats; it is a reminder to make the axis describe the judgement being requested.

## Build a training questionnaire

Imagine a workshop on writing neutral questions. The objective is to identify what participants believe they can reuse and what needs more teaching, not to measure satisfaction again.

Instruction: “Choose one response for each statement. If you do not have enough information, choose ‘I don’t know.’ If it does not concern your work, choose ‘Not applicable.’”

Common scale: 1 = strongly disagree, 2 = disagree, 3 = neither agree nor disagree, 4 = agree, 5 = strongly agree. Add “I don’t know” and “Not applicable” only when those outcomes are plausible.

1. “I can distinguish an open-ended question from a question that already suggests an answer.”
2. “I can spot when a question mixes two subjects.”
3. “I can rewrite a question to ask for a concrete example.”
4. “I know when a closed response should be followed by an open probe.”
5. “I can explain which decision each question in my questionnaire will inform.”

Add: “Which question or situation is still difficult to formulate?” The items measure declared ease, not proven skill; a task or observation would be needed to assess actual performance. If the subject does not concern everyone, provide “Not applicable” instead of forcing a judgement.

Keep the time frame visible: “in the past month” and “in the exercise just completed” invite different answers. Avoid asking respondents to rate an abstract “experience” when the decision concerns one identifiable step. A short instruction can protect the interpretation more effectively than adding another response category.

The midpoint is a position on the measured dimension. “I don’t know” indicates missing information. “Not applicable” indicates that the situation does not concern the respondent. “Prefer not to answer” signals refusal. Nonresponse may reflect forgetting, abandonment, or a presentation problem. None should be silently recoded as 3.

## Describe the distribution before the mean

This is an **illustrative example**, not study data. In a questionnaire sent to 84 participants, 80 valid answers are collected for the item “I can apply the method presented in a situation close to my work.” The answers are: 8 strongly disagree, 14 disagree, 18 neither agree nor disagree, 25 agree, and 15 strongly agree. Two people choose “I don’t know” and two do not answer.

A mean coded from 1 to 5 does not tell the whole story. The valid distribution denominator is 80, not 84. The four remaining outcomes describe coverage and should not be silently recoded to the midpoint. The midpoint may contain genuine neutrality, hesitation, or difficulty judging application: inspect counts and related comments before concluding that the training “worked.”

The descriptive calculation is `(8×1 + 14×2 + 18×3 + 25×4 + 15×5) / 80 = 3.3125`, or **3.31/5** under this coding. The two agreement options contain 40 answers: **50% of the 80 scale responses**, but **47.6% of the 84 people invited**. These are not competing estimates: their denominators answer different questions. Always name the denominator. The 22 disagreement responses (27.5% of valid scale answers) remain visible; describing only a slightly positive mean would hide them.

Responses are ordinal: category order is known, but the gap between two categories is not automatically an equal measurable distance. Sullivan and Artino’s academic guide recommends cautious interpretation of Likert-type distributions ([An Introduction to the Analysis of Likert Data](https://pmc.ncbi.nlm.nih.gov/articles/PMC3886444/)). A mean can be a conventional descriptive summary in some settings; it is not causal evidence or an unquestionable continuous measure.

Here is a second **entirely illustrative** comparison of two groups of 20 answers on a 1-to-5 scale. Both groups have exactly the same mean, 3.

| Group | Answers | Calculation | Mean |
|---|---|---:|---:|
| A — polarized | ten answers at 2, ten answers at 4 | (10×2 + 10×4) / 20 = 60 / 20 | 3 |
| B — concentrated at the midpoint | twenty answers at 3 | (20×3) / 20 = 60 / 20 | 3 |

For Group A, nobody selected 3: the mean combines two opposed response positions. Check what distinguishes the subgroups, inspect the context, and read open-ended answers. One intervention may suit one subgroup and worsen the gap for another.

For Group B, answers cluster at the midpoint. That could be a genuinely intermediate position, but it could also indicate a difficult question, insufficient information, an overly convenient midpoint, or reluctance to take a position. The identical mean cannot choose among these interpretations. With 20 answers, show counts, “I don’t know,” nonresponses, and, where possible, a targeted open probe.

Read the table in a fixed order: identify the valid denominator, check how many people selected each category, separate excluded outcomes, then compare the shape of the distributions. Ask whether the groups saw the same wording, period, instructions, and situation. A difference in the proportion of “not applicable” may signal different exposure rather than a weaker opinion. If the decision depends on a threshold, define that threshold before looking at the result; otherwise the summary can drift toward the most convenient interpretation.

For every item, record invited respondents, valid answers, “I don’t know” or “not applicable” outcomes, nonresponses, and the full distribution. A change in the mean may come from the denominator or group composition. Check the same wording, scale, exposure, time frame, and rules for excluded outcomes before comparing groups. If an option is never used, do not delete it after collection: it may have been unnecessary, misunderstood, or inaccessible.

## Avoid analysis errors

“The content was clear and useful” asks for two judgements. “The training met my expectations” lets everyone define expectations differently. Write one item, one time frame, and one dimension.

Reverse-worded items can introduce inattention and coding errors rather than control acquiescence. One study reports these risks ([van Sonderen, Sanderman and Coyne](https://pmc.ncbi.nlm.nih.gov/articles/PMC3729568/)). Keep the same reading direction; if a validated instrument requires reverse wording, document it and check the recoding.

A high mean can hide polarization, nonresponse, or a subgroup not exposed to the situation. Compare only distributions with comparable items, labels, and context. Cronbach’s alpha does not automatically authorize aggregation: justify the construct, item set, and model. High internal consistency can come from redundant items; it does not prove that the intended construct was measured.

The mean, median, and most frequent category answer different descriptive questions. The mean is sensitive to how categories are coded; the median identifies an ordered position; the mode shows the most common answer. None is automatically the “real” result. If the categories are ordinal and the sample is small, show the frequencies first and explain why a numerical summary is being used. When a score is used for a decision, retain the underlying counts so another reviewer can reconstruct the reasoning.

## Explore, stabilize, and pilot

Open-ended questions help at the beginning: they reveal respondents’ words and whether planned categories match their experience. An ordered scale can then stabilize a specific dimension. Keep “other,” exceptions, and contradictions. A targeted open probe is often enough: “What explains your answer?” or “When did this answer not apply?”

The [J-PAL survey-design guide](https://www.povertyactionlab.org/resource/survey-design) stresses that options should be clear, mutually exclusive, and sufficiently exhaustive, while closed questions necessarily restrict the richness of answers. Its [questionnaire-piloting guidance](https://www.povertyactionlab.org/resource/questionnaire-piloting?lang=en) recommends checking missing options, comprehension, translation, and filters with comparable respondents. Ask people to read a few items, ask what each option means, and test mobile display.

See also [Open-Ended Questions: Get Actionable Data Without Bias](/en/blog/open-ended-questions-get-actionable-data-without-bias), [Training Satisfaction Questionnaire](/en/blog/training-satisfaction-questionnaire-why-your-results-are-often-unusable), and [Actionable Questionnaires](/en/blog/actionable-questionnaires-start-with-the-decision-not-the-questions).

A mean, median, or distribution describes the collection that was obtained. It does not turn a self-report into an observation, a perception into competence, or a score into causal evidence. Document the intended construct, item origin, translation, pilot findings, exclusions, and interpretation rule.

In remote collection, Harmate can distribute questions, preserve open answers, and organize a reading by themes and sources. It does not decide what a score means, replace piloting, or decide in place of human review. The useful result is a comparison whose limits, context, and reasons remain visible.