Qualitative Disaggregation — A Job Aid

In plain language: we sort what people told us by who they are, then read each group’s words.

We are looking for whether the service works differently for different people.

Page 1 of 2 — The Check You Can Run Today

Next: Page 2 of 2 — The Method Behind the Test

Here’s an aid to help you compare how different groups experience your service, in their own words.

Open it when you have a pile of comments and a research question you’d like to answer.

It carries the check we ran in the workshop, the four questions that go with it, and the method around them. In our course, we asked “Who does this service fail?”

Have We Heard Enough? The Hold-Back Test

Hold about one in five of your customer experience data back from your initial read, and never fewer than five items. When you are done, read the hold-back items. Do they add a reason you have not yet encountered?

In the workshop: hold some back, read them last, ask whether they add a reason you didn’t already have.

  1. Set aside a random portion of the comments: about one in five, and never fewer than five. Do not read them.
  2. Analyze the rest. Write down every distinct reason the service worked or failed for someone.
  3. Now read the ones you set aside.
  4. Ask one question: does this contain a reason we did not already have?

Nothing new? We have evidence, not a feeling, that we have heard the range of what this data holds. For this group, on this question.

Something new? We are not done. And now we know it, instead of assuming it.

That is the difference between a claim and a test. A claim we assert. A test we can fail. That is the only reason passing it means anything.

Two honest limits.

Small piles. If we only have a dozen comments, the test is weak no matter how we split them. Run it anyway, and say so rather than dress it up.

New reasons only. This test is good at catching a reason we had not heard. It is blunter about deeper understanding of a reason we already had. That is real learning too, and harder to check.

Researchers call the state this test looks for saturation: the point where more data stops teaching us anything new. The set we hold back is a holdout.

The Four Saturation Questions

People will offer a number of interviews that counts as enough. Someone else’s number is not evidence about our data. The right number depends on how alike our people are and how narrow our question is. So do not ask “how many is enough?” It has no answer.

Ask these four questions instead. We can answer all of them with the data we already have.

Question 1 — Saturated on What?

We never saturate in general. We saturate on one question, about one group.

“Have we heard the ways this process fails people who work for themselves?” Answerable.
“Have we heard how people who work for themselves experience government?” Not answerable, ever.

Narrow question, similar group: we will get there quickly. Broad question, mixed group: we may never get there, and we should say so.

Write down the question before you start. If it will not fit in one sentence, it is not yet a question we can saturate on.

Question 2 — When Did the Last New Thing Appear?

We can only answer this if we read in batches and wrote down what each batch added.

Read six comments. Write down every distinct reason the service worked or failed for someone. Read six more. Write down anything new. Keep going.

If we read the whole pile in one go, we cannot answer this question. Not because we did bad work. Saturation is about the order we read in, not the size of the pile. There is no way to recover it afterward.

This is the cheapest discipline in the whole method, and almost nobody does it.

Question 3 — If We Read More, Would We Learn Something New? Test It.

Run the Hold-Back Test above. Do not guess at the answer. A guess is a claim. The test can fail, which is why passing it counts.

Question 4 — What Could We Never Have Learned from This Data?

We can be perfectly saturated and perfectly blind.

Saturation tells us what our data has run out of. It tells us nothing about who never got into our data. If our survey only reached people with an email address, we can hear the same thing from every one of them and still miss an entire experience. That experience belongs only to the people without one.

Saturation is never evidence of coverage. Say so, every time.

Count Reasons, Not Themes

Themes multiply with how finely we code. Reasons do not.

When we ask the saturation question, the thing we need the full range of is the distinct reasons the service works or fails for someone. Each distinct reason needs a different response.

The test for whether two comments are the same reason: would the same change address both?
Same change, same reason. Different change, new reason.

Count the ones that work, not only the ones that fail. A person who got through easily is telling us something new if they got through for a reason we had not yet heard. That is often the most useful item in the pile. It tells us what the real barrier was.

Where this sits. Affinity mapping produces themes, and themes are the right unit there. The saturation question in disaggregation is a different question: have we heard the reasons this works or fails for this group? Different question, different unit. Both are correct in their place.

The Statement to Write

“We read [the comments] in batches and stopped finding new reasons after [point]. We held back [portion] and read them last. They contained no reason we had not already found. That gives us reasonable confidence we have described the range of reasons this process works and fails for [group].

It gives us no basis for saying how many people are affected, and no information about the people who never reached us.”

That statement is our permission and our limit in one breath. It is also a statement we can defend to a skeptical director, because everything in it is something we did.

Our statement, for our data:

Write-on lines. Print this page, or write your statement in your own notes.

Saturation Is Not Prevalence

These get mixed up because both arrive as numbers. They are not the same kind of claim. Prevalence means how common something is across a whole population.

Saturation and Prevalence, Compared
QuestionSaturationPrevalence
AnswersHave we heard the range of things there are to hear?How many people does this affect?
Is a claim aboutThe set of experiencesThe population
We get there byListening until people stop telling us new thingsHaving a sample we can defend
More time with the same commentsHelpsDoes not help at all

Think of it this way: saturation is a vocabulary list. Prevalence is a census.

If we want to know what words are spoken in a town, we can stop asking once people stop giving us new ones. If we want to know how many people speak each word, no amount of careful listening will get us there.

Next: The Method Behind the Test

Print or save the whole job aid, both pages, as a PDF