You are planning a qualitative research project. The question of sample size comes up — and suddenly you are doing maths. Ten interviews? Twenty? A hundred? You start wondering: if an AI could do the interviewing, would anyone notice? More importantly, would the data suffer?
The question is no longer hypothetical. In 2024, LSE researchers Friedrich Geiecke and Xavier Jaravel launched an open-source platform that uses LLMs to conduct thousands of qualitative interviews in hours. When they asked a panel of sociology PhD students from Harvard and LSE to evaluate the transcripts, the AI-led interviews scored "approximately comparable to an average human expert." Participants, meanwhile, reported enjoying the interaction more than filling out open-text surveys and wrote significantly more words.
That finding turns heads. But "comparable under controlled conditions" is not the same as "appropriate for every study." Here is what the evidence actually says about choosing between AI and human interviewers — and how to make the decision for your specific research context.
Where do AI interviewers win?
Scale. This is the headline advantage. A human researcher can conduct at most four or five quality interviews in a day before fatigue sets in. An AI interviewer can run five hundred in the same period. For research questions that need breadth — employee sentiment across a global organisation, customer attitudes across demographic segments, market-wide attitude studies — AI interviewing does not just make things faster. It makes entirely new research designs possible.
Consistency. Human interviewers vary. Different people emphasise different questions, use different phrasing, and apply different amounts of prompting. These variations introduce noise that is hard to account for in analysis. AI interviewers apply the same guide and probing logic to every participant. When comparability across participants matters, this is a genuine methodological advantage — your data reflects participant differences, not interviewer differences.
Always-on availability. Participants complete AI interviews at 11 pm, on weekends, between meetings. No calendar wrangling, no rescheduling emails. The low friction translates into higher completion rates and more diverse samples — people who would not book a 60-minute Zoom slot will type out their thoughts while commuting.
Reduced social desirability bias. Counterintuitive but well-documented: people sometimes disclose more honestly to a machine than to a person. Research on medical and mental health self-disclosure consistently finds more accurate reporting through computer-mediated interviews, especially on topics where participants might fear judgement. The absence of a human face can be an advantage when what you are studying is stigmatised, sensitive, or personally difficult.
Cost. A human expert interview costs roughly €100–300 in researcher time before you factor in recruitment and transcription. AI interviews cost a fraction of that. Research programmes that were previously cost-prohibitive — longitudinal studies with weekly check-ins, or large-N qualitative designs — become viable.
Where do human interviewers remain essential?
Sensitive and high-stakes topics. Trauma, grief, discrimination, serious illness — these are domains where the participant needs to feel genuinely heard, not merely recorded. A skilled human interviewer creates a relational container for difficult conversations. In these contexts, using an AI interviewer is not just methodologically questionable; it is ethically problematic. The risk of causing distress without being able to manage it in real time is real.
Complex probing and domain expertise. Subject-matter experts present a different challenge. When an interviewee says something technically dense, a domain-literate human can probe intelligently. "Wait — if that mechanism operates that way, what explains the anomaly in your Q3 data?" This kind of conversational intelligence requires understanding what you are hearing. Current AI interviewers vary widely in their probing quality — some ask generic follow-ups regardless of content, while better tools maintain a coherent thread. None match a skilled domain expert.
Observation-dependent research. Ethnographic fieldwork, usability testing where you watch someone use a product, diary studies that combine self-report with ambient data — any research where the interview is accompanied by observation requires a human in the room. The conversation alone is incomplete without what you see.
Relationship-building. In customer research and stakeholder studies, the interview doubles as a relationship moment. A human conversation signals that the organisation values the participant as an individual. An AI interview signals efficiency. Both are valid, but they communicate different things about your relationship with the people you study.
Low-trust contexts. Participants who are sceptical of AI, worried about how their data will be used, or from communities with historical reasons to distrust research institutions may engage far more openly with a human they can read and assess for themselves. Trust is earned in person.
How do you decide for your study?
The real question is not "AI or human?" but "what kind of data does my research question actually need?"
Use AI when: your primary need is breadth. You want to hear from many people across segments, detect patterns in a population, or collect structured qualitative data at a scale that supports cross-case comparison alongside thematic interpretation. Your topic benefits from anonymity — people speak more freely when they know no one is watching. Your budget does not stretch to a large human interview team.
Use a human when: your participants are in vulnerable circumstances. Your research questions require adaptive domain expertise mid-conversation. You need to observe behaviour, not just hear self-report. The interview itself serves a relational purpose beyond data collection.
Combine both when: your question has multiple layers. Start with AI for the wide net — the initial sweep across hundreds of participants that surfaces patterns and divergent cases. Then follow up with human interviews on the questions that emerge. AI for breadth, humans for depth. This hybrid model is where the most sophisticated research programmes are heading.
The bottom line
The LSE team's finding that AI interviews are "approximately comparable to an average human expert" is important. It means the baseline is good. But a good baseline is not the ceiling, and the ceiling depends on what you need.
A shallow AI interviewer with generic follow-ups will produce shallow data. A well-designed one — with prompting that draws on established interview methodology, adapts to what participants say, and maintains a coherent thread — will produce richer material. The tool matters as much as the choice to use it.
The honest answer is this: AI interviewers have earned a place in the qualitative toolkit, especially for breadth, scale, and reduced social desirability effects. But they are not replacements for humans in every context. Match the method to the question, not the other way around. When you do, you get better data either way.
Discussion
or sign in to comment with your account