Every qualitative researcher hits the same wall around interview twelve or fifteen: nothing new seems to be coming up. The temptation is to call it saturation and stop. Sometimes that's right. Often it's just fatigue talking, or a codebook that's too coarse to notice what's still changing underneath.
The problem with "I think we're done"
Saturation is usually described as the point where additional data stops producing new codes or new understanding. That's a fine definition in principle, and useless in practice unless you're actually counting something. Most researchers judge it by feel — how tired they are of transcribing, how repetitive the last few interviews sounded. Feel is a bad instrument. It's swayed by interview length, by how articulate your last two participants were, by how far behind you are on your timeline.
Saturation should be a measurement, not a mood. The measurement is simple: for each new transcript, count how many codes in your codebook are genuinely new — not a rephrasing of an existing code, but a concept you hadn't captured before. Plot that count against interview number. When it flattens to zero or near-zero for three or four consecutive transcripts, you have a real signal.
True saturation vs. false saturation
Not all flat lines mean the same thing. A few patterns are worth telling apart before you stop recruiting:
Signal | What it looks like | What it usually means |
|---|---|---|
True saturation | New-code rate drops steadily across a diverse run of participants | The concept space is genuinely covered |
Sampling saturation (false) | New-code rate drops, but recent participants are demographically similar to each other | You've saturated one subgroup, not the population |
Fatigue saturation (false) | New-code rate drops and interview length/depth is also dropping | You or your participants are running out of steam, not topics |
Codebook saturation (false) | New-code rate drops but codes are broad and vague | The codebook isn't fine-grained enough to detect real novelty |
The second and third rows are the ones that quietly wreck studies. A team stops at interview 14 because nothing new is showing up, without noticing that interviews 10 through 14 were all the same age bracket, or that fatigue had crept into both the interviewing and the coding. The fourth row is sneakier still — a codebook with categories like "communication issues" will look saturated almost immediately, because it's too blunt to register that interview 12 raised a genuinely different kind of communication issue than interview 3.
A practical routine
You don't need software to do this well, but you do need discipline. After each interview: code it against the existing codebook, log every code that's new versus reused, and keep a running tally next to basic participant metadata (role, tenure, location, whatever your key variables are). Every few interviews, look at the new-code curve split by those variables, not just in aggregate. If the aggregate curve is flat but one subgroup keeps producing new codes, you're not done — you're just under-sampled in that subgroup.
This is exactly the kind of bookkeeping that's tedious to do by hand across a codebook of any real size, and it's one of the reasons we built Paideias to track new-versus-reused codes automatically as you go, transcript by transcript, so the saturation curve is something you can actually look at rather than something you have to reconstruct from memory at the end.
The honest version of the answer
Saturation isn't a single moment you arrive at and announce. It's a claim you make about a specific codebook, a specific sample, and a specific research question — and it's only as credible as the tracking behind it. "We stopped because nothing new was coming up" is a weak sentence in a methods section. "New codes dropped to zero across the last five interviews, spanning three participant subgroups" is a strong one. The interviews are the same either way. The difference is whether you kept count.
Discussion
or sign in to comment with your account