# Word Cloud: How to Create One Without Distorting Your Responses

How to create a word cloud from responses, preserve context, and check frequencies before making a decision.

- Canonical URL: https://www.harmate.com/en/blog/word-cloud-how-to-create-one-without-distorting-your-responses
- Author: Harmate Team
- Published: 2026-08-30
- Updated: 2026-08-30T01:06:26.813918+00:00
- Language: en

## Content

# Word Cloud: How to Create One Without Distorting Your Responses

A word cloud can be produced in minutes. The difficult part starts afterwards: deciding what a word’s size represents, what it hides, and what must be checked before a visual impression becomes a decision.

This guide gives you a six-step method for creating a word cloud from text or open-ended responses, followed by a reading framework that keeps frequency, theme, cause, and evidence distinct.

## The short answer

To create a useful word cloud:

1. define the question the visual should help you examine;
2. set an explicit corpus scope;
3. clean only what you can justify;
4. generate the cloud with reproducible settings;
5. return to the sentences containing the prominent terms;
6. compare subgroups and cases that do not look like the majority.

A cloud is an orientation view. It can help you identify terms worth examining, but it does not prove a theme, a cause, or agreement with a proposal.

## Create a word cloud in six steps

### 1. Start with a question, not a tool

Without a question, a word cloud quickly becomes a decorative image. Before pasting your text into a generator, write down what you are trying to inspect.

Useful questions include:

- Which terms recur in answers about a specific problem?
- Which words distinguish two groups or two periods?
- Which topics deserve a closer reading before a workshop or a report?
- Which terms look important but might only be stock phrases?

The question determines the corpus, the comparisons, and the cleaning you can reasonably justify. A cloud built from answers to “what did you like?” does not answer the same question as one built from “what prevented you from acting?”.

### 2. Define the corpus

Gather texts that actually address the question. Record their source, period, language, and any filters. A very long response should not automatically carry more weight simply because it contains more words: depending on your purpose, you may compare documents, people, groups, or occurrences, but you need to know which unit you are counting.

Keep separate corpora when comparison matters. Do not merge before/after responses if you want to see change over time. Do not mix customers, employees, and managers if their vocabulary differences are part of what you are studying.

A larger corpus is not automatically a more relevant corpus. A heterogeneous collection can produce a clear-looking cloud while combining several different questions.

### 3. Clean without erasing meaning

Text preparation directly affects the result. You may remove technical elements that do not carry meaning for your question—tags, exact duplicates, repeated headers, or stop words—if you record the choice.

Be careful about automatically removing:

- negations such as “not”, “never”, or “without”;
- qualifiers such as “sometimes”, “almost”, or “only”;
- words that mark conditions or exceptions;
- terms used by a minority group;
- multi-word expressions whose meaning disappears when split apart.

Cleaning is not neutral. Removing “not” because it is frequent can turn “not clear” into an artificial presence of “clear”. Reducing every form to a stem without checking the result can merge uses that do not mean the same thing.

Keep the original corpus and a record of transformations. The goal is not the prettiest cloud; it is a result you can explain and reproduce.

### 4. Choose reproducible settings

A generator may offer settings for the maximum number of words, frequency thresholds, ignored terms, case, inflections, colours, orientation, and placement. Do not change several settings at once without recording them.

For a first version, keep the configuration simple:

- state the maximum number of terms;
- document the ignored-word rule;
- document how phrases are handled;
- define what the font size encodes;
- export the cloud with the corpus and settings.

Visual form attracts attention, but it does not replace measurement. An empirical study by [Felix, Franconeri, and Bertini](https://nyuvis.github.io/word-cloud/paper.pdf) shows that the representation affects how easily people extract information depending on the task. Two clouds made from the same text can tell different visual stories when the ignored-word list, scale, or placement changes.

### 5. Return to the context of each term

Take the most visible terms and search for their occurrences in the corpus. Read several surrounding sentences, not only the first one that confirms your impression.

This check may show that:

- a word comes from a formula repeated by one document;
- a term appears mostly inside a negation;
- several words belong to one expression;
- a frequent word describes a cause in some answers and a consequence in others;
- a group uses different vocabulary and disappears in the global cloud.

Corpus-reading tools such as [Corpus Terms](https://docs.voyant-tools.org/docs/tutorial-corpusterms.html) and [Contexts](https://docs.voyant-tools.org/docs/tutorial-contexts.html) illustrate the useful pairing: term counts can guide exploration, while context lets you reread the occurrences.

### 6. Compare before concluding

A global cloud often hides differences. When relevant, create separate views by period, role, question, or situation. Comparison is not about declaring which group is “right”; it is about making differences visible enough to investigate.

Compare the cloud with a frequency list or table as well. A list preserves values and supports checking; a cloud supports a first visual exploration. They do not answer exactly the same need.

## Example: why negation changes the reading

Consider these two responses:

> “The tool is simple, but not clear enough to find an old request.”

> “The tool is clear for a new request, but not simple when an error must be corrected.”

Overly aggressive cleaning may surface `clear`, `simple`, `request`, and `error` without preserving the direction of each judgement. The cloud then shows lexical presence, not a positive or negative assessment.

The check is to return to the sentences, keep negation markers, and record conditions. The useful question is not “do respondents find the tool clear?”, but “in which situations does clarity or simplicity become a problem?”.

## What a cloud shows, hides, and requires you to check

| A cloud may show | A cloud may hide | Check to perform |
|---|---|---|
| The relative presence of terms under a counting rule | The meaning and polarity of a sentence | Read several occurrences in context |
| Words to examine first | Expressions and synonyms carrying the same idea | Search related forms and phrases |
| Vocabulary differences between corpora | Group size and minority responses | Compare subcorpora and retain units |
| A first discussion prompt | Causes, consequences, and conditions | State hypotheses and look for evidence |
| An image that is easy to share | The settings that produced it | Archive the corpus, cleaning, and settings |

Word size is not importance in the broad sense. It encodes a value defined by the chosen processing. An infrequent notion may be decisive for one group or risk; a frequent notion may be vague, required by the question, or repeated by a single source.

## Five checks before presenting the result

### Is the corpus identifiable?

Can you explain where the texts came from, which responses are included, and which are not? If the scope is unclear, the cloud cannot be interpreted responsibly.

### Is the cleaning reversible?

Did you keep the original text and transformation list? A removal decision should be discussable, correctable, and repeatable.

### Have the important terms been read in their sentences?

A term list is not enough. Check negations, pronouns, reported quotations, conditions, and ironic or ambiguous wording.

### Are minority voices still visible?

A global cloud naturally gives more space to frequent words. Inspect relevant subgroups and unusual responses before announcing a collective finding.

### Does the conclusion go beyond the visual?

Write what the cloud suggests, then what it cannot establish. A hypothesis becomes a useful line of work when it is confronted with source sentences, other data, or a follow-up question.

## When to move from a cloud to thematic analysis

A cloud is useful for opening an exploration. It becomes insufficient when you need to group different phrasings, distinguish causes, explain exceptions, or justify a decision.

At that point, move from presence to meaning: define themes from the responses, keep their inclusion criteria, connect them to supporting quotations, and make visible the cases that do not fit neatly. A frequency can help prioritise reading; it does not replace the reasoning that connects an observation to an action.

To go further without repeating a full coding method, read [Read 200 Open Responses Without Lying to Yourself](/en/blog/analyzing-200-open-responses-without-bias-a-method-for-actionable-decisions) and [Open Questions: Get Actionable Answers Without Bias](/en/blog/open-ended-questions-get-actionable-data-without-bias). The [Open Questions](/en/product/open-questions) surface can then support a collection process that gives respondents room to describe their situations.

In that workflow, Harmate can help connect open responses with themes and inspectable verbatims, while leaving the team to examine differences and decide what to retain. This does not turn a cloud into an automatic diagnosis: the visual remains a starting point, and a decision should remain connected to its evidence.

## FAQ

### Is a word cloud qualitative analysis?

It is an exploratory visualisation of terms under a counting or relevance rule. It can support qualitative analysis, but it is not by itself an interpretation of experiences or causes.

### Should very frequent words always be removed?

No. A frequent word may be central to the phenomenon. Remove only what does not serve your question and document the choice; compare a version with and without the removal when useful.

### Can a cloud measure agreement?

Not directly. A word may appear in approval, criticism, a question, or a negation. To discuss agreement, define what you are measuring and read the responses in context.

### Which tool should you choose?

Choose a tool that accepts your corpus format, lets you control ignored terms, and allows you to retrieve the result. Interpretation depends mostly on the corpus, transformations, and checks you preserve.

### Are open-ended questions essential?

They let people describe situations, reasons, or exceptions that predefined choices did not anticipate. They still require careful treatment and do not by themselves guarantee validity, representativeness, or comparability.

## Sources

- [Participatory visualization with Wordle — Viégas, Wattenberg, and Feinberg (IEEE, 2009)](https://research.ibm.com/publications/participatory-visualization-with-wordle)
- [Taking Word Clouds Apart: An Empirical Investigation of the Design Space for Keyword Summaries — Felix, Franconeri, and Bertini (IEEE, 2018)](https://nyuvis.github.io/word-cloud/paper.pdf)
- [Lifting the Fog on Word Clouds: An Evaluation of Interpretability in 234 Individuals — Maurits, Boers, and Knevel (VISIGRAPP, 2022)](https://www.scitepress.org/Papers/2022/107787/107787.pdf)
- [Voyant Tools — Corpus Terms](https://docs.voyant-tools.org/docs/tutorial-corpusterms.html) and [Contexts](https://docs.voyant-tools.org/docs/tutorial-contexts.html)