AI & Methodology

How Accurate Is AI at Coding Qualitative Data?

What 75% agreement actually means when the baseline is human interpretation.

· 4 min read· 12 views
How AccurateIs AI at CodingQualitative Data?What 75% agreement actually means

You have heard the pitch: AI can code your transcripts, find your themes, and deliver your findings in minutes. It sounds too good to be true, and in some ways it is. But the real question is not whether AI coding is accurate. It is what accuracy even means when the ground truth is a human interpretation.

What We Mean by Accuracy

When a researcher asks how accurate AI coding is, they usually mean: does the AI assign the same codes I would? That question assumes there is a single correct answer — one right code for every passage. Anyone who has done qualitative analysis knows this is not how coding works. Two skilled researchers coding the same transcript will agree on 70-80% of codes at best. The rest is legitimate disagreement about emphasis, framing, and interpretation.

So when an AI tool claims 90% accuracy, ask: 90% agreement with which human? Because the baseline is not a perfect standard. It is one person's judgement.

What the Research Actually Shows

The studies that exist on AI coding accuracy point to a consistent range. For first-cycle descriptive coding — the kind where you label a passage as 'describes first day at university' — AI tools match human coders at around 75-85% agreement. That is within the range of human inter-coder reliability for the same task.

For second-cycle pattern coding — where you group descriptive codes into themes — the accuracy drops to 60-70%. AI can cluster similar passages together reliably, but the conceptual leap from cluster to theme still favours the human.

For axial coding — connecting categories and identifying causal relationships — most AI tools perform poorly. Twenty to thirty percent agreement is not unusual. The AI simply does not have the lived experience of the research context to make those connections.

The Variables That Matter

Accuracy is not a single number. It depends on three things you control.

Transcript quality. AI performs best on clean, well-structured transcripts where speakers are clearly identified and the audio was clear. Heavy accents, overlapping speech, or transcripts with transcription errors reduce accuracy.

Codebook specificity. You will get better results if you give the AI a well-defined codebook with examples than if you ask it to generate codes from scratch. A codebook with definitions and boundary markers can push agreement from 70% to 85%.

Domain complexity. Technical or specialist domains reduce AI accuracy. A tool trained on general academic text will code a physics interview less reliably than one fine-tuned on science communication. If your research is in a narrow field, budget extra time for manual review of AI-generated codes.

Where AI Is Actually Good Enough

For certain tasks, AI coding accuracy is already good enough to be useful.

Descriptive coding at scale. If you have 50 transcripts and you need to label every passage with a basic topic descriptor, AI does this reliably. A researcher reviewing and adjusting AI-generated descriptive codes saves roughly 60% of the time compared to coding from scratch.

Pattern spotting. AI is excellent at identifying passages that share similar language. It will not always name the pattern correctly, but it will reliably surface candidate groupings for you to review. This is the single most useful application of AI in qualitative analysis right now.

Negative case detection. Because AI reads every transcript with equal attention, it catches the passage that contradicts your emerging theme — the one participant who said something different when everyone else agreed. Humans tend to notice patterns; AI tends to notice anomalies.

Where AI Falls Short

The weaknesses are as important as the strengths.

Context-dependent meaning. A word can mean different things in different parts of the same interview. Human coders track this. AI tools treat the word as the same token everywhere it appears.

Irony, sarcasm, and indirection. Participants say things they do not mean. They use humour to deflect. They imply rather than state. AI has almost no ability to detect these layers.

Emotional nuance. Two passages that describe similar events — one with anger, one with resignation — might receive the same AI code because the content is similar. The emotional difference is invisible to a model trained on word patterns.

Absence as data. The most telling moment in an interview is sometimes the thing a participant refuses to discuss. AI does not code silences.

The Practical Bottom Line

AI coding accuracy is good enough for first-cycle descriptive coding at scale. It is not good enough for interpretive work without human review. The workflow that works is: let the AI do the first pass, then read every coded cluster yourself, adjusting codes and splitting categories as you go. The AI cuts the mechanical time. You still own the analysis.

A platform like Paideias supports this workflow — AI-assisted first pass, human review and refinement, transparent audit trail of whose code is whose. The goal is not AI that codes perfectly. It is a system where you code better because the AI handled the grunt work.

The question is not 'is AI accurate enough?' It is 'are you using it accurately enough for the stage of analysis you are in?'

#ai#coding#accuracy#qualitative-analysis#reliability
Share

Discussion

or sign in to comment with your account

Keep reading

Coding & Analysis

Do I Have Too Many Codes? A Sanity Check for Your Codebook

Most applied projects land between 30 and 70 codes, and past 100 a consolidation pass is usually overdue. The real test, though, is whether you can still tell your codes apart. Here are the warning signs of codebook bloat and how to merge, nest, and prune without losing nuance.

5 min read