Not always — and demanding it for every study misunderstands what qualitative coding claims to be. If your codes function as counts or categories, agreement between coders is real evidence of rigour. If your codes are interpretive readings of meaning, a reliability statistic can be theatre. The honest answer depends on what your analysis claims, and this post gives you a working rule for deciding.
What does intercoder reliability actually measure?
It measures one narrow thing: whether two people applying the same codebook to the same data make the same decisions. Nothing about truth, insight, or depth — just consistency of application.
Raw percent agreement flatters you, because two coders agree some of the time by chance alone. Cohen's kappa corrects for that, which is why reviewers ask for it. By the widely used Landis and Koch benchmarks, a kappa of .61–.80 counts as substantial agreement and .81 or above as almost perfect; below .60, your codebook definitions are probably doing less work than you think.
That last point is the useful one: a low kappa is rarely a coder problem. It is a codebook problem — vague definitions, overlapping codes, missing decision rules. Measuring agreement early is one of the fastest ways to find out whether your codebook actually says what you think it says.
When is it worth doing?
When your codes will be treated as categories that get counted, compared, or handed between people. Content analysis reporting frequencies. A team of three coding two hundred transcripts. Longitudinal designs where wave two must be coded the same way as wave one. In these designs, coder drift is real measurement error, and O'Connor and Joffe's 2020 guidelines offer a practical recipe: double-code a 10–25% sample, report the statistic, and describe how disagreements were resolved.
There is also a quieter benefit for teams. The disagreement sessions matter more than the number — arguing about why one coder saw "resignation" where another saw "acceptance" is exactly how a codebook gets sharp. Teams that skip the argument and just report the kappa are keeping the ritual and discarding the value.
When does it not make sense?
When coding is the analysis, not preparation for it. In reflexive thematic analysis, Braun and Clarke are explicit that codes are interpretive readings produced by a researcher's engagement with the data — there is no single correct coding against which to score agreement, so demanding a kappa imports an assumption the method rejects.
McDonald and colleagues' 2019 review of CSCW and HCI practice reached the same conclusion from the field: agreement statistics are one legitimate way to establish rigour, suited to some designs and not others. The alternatives are not softer — an audit trail of coding decisions, systematic memoing, and peer debriefing document interpretive judgement the way kappa documents categorical consistency.
The mistake to avoid is the decorative kappa: running a reliability exercise on deeply interpretive codes because a reviewer might want a number. It signals rigour while measuring almost nothing.
How do you report it either way?
Report the process, not just the outcome. If you measured agreement: what percentage of the data was double-coded, which statistic you used and why, what the value was, and — most importantly — how disagreements were resolved and what changed in the codebook as a result. A kappa with no process description tells a reviewer very little.
If you deliberately did not measure agreement, say so and say why, in one honest sentence: coding was interpretive and consistency was established through an audit trail and regular peer debriefing rather than inter-coder statistics. Reviewers reject silence and hand-waving far more often than they reject documented judgement.
The bottom line
Match the evidence to the claim. Codes that get counted need demonstrated consistency: double-code 10–25%, report kappa, describe the resolution process. Codes that carry interpretation need documented judgement: audit trail, memos, debriefing. Choosing the right one — and saying why — is itself the mark of a rigorous study.
Discussion
or sign in to comment with your account