Peer reviewers are getting stricter about how we write up our methods. If you submit a manuscript that simply says "we used AI to analyze the data", you are going to get desk-rejected. Reviewers want to know exactly how you used it, what it did, and most importantly, what it did not do.
What do peer reviewers expect when you use AI?
Peer reviewers expect a complete audit trail that proves you used AI as an assistant, not a replacement for your own analytical thinking. They want to see the specific tools named, the prompts documented, and concrete evidence that a human researcher verified the outputs.
For a long time, the rules were murky. However, in the last 12 to 18 months, major publishers like SAGE, Elsevier, Wiley, and Springer Nature have published specific regulations regarding AI use in research. They all converge on one point: transparency. You cannot just name the tool. You need to specify its exact role. Did it generate initial codes? Did it summarize transcripts? Did it cluster themes? You need to document configuration settings, metadata variables used for filtering, and core prompts.
How much of the analysis can AI actually do?
AI tools excel at mechanical tasks like summarization and initial clustering, but they cannot perform the deep interpretive labor required for true theoretical understanding. You can rely on AI to process massive datasets quickly, but you must manually validate the themes it generates.
Traditional manual coding is generally manageable up to about 10 to 20 interviews before it becomes a tedious chore. However, sample sizes in leading qualitative journals have grown, with many now expecting 50 to 100 interviews for a robust study. AI-assisted tools excel here because they can systematically process 100 or more interviews without fatigue.
But scale does not equal depth. The AI is doing the heavy lifting of sorting, but you are still responsible for the meaning. If you are doing Reflexive Thematic Analysis, you still need to map your procedure to the 6 phases outlined by Braun and Clarke (2006). The AI might help you with familiarization and initial coding, but the later phases of defining and naming themes require your theoretical lens.
What are the most common mistakes researchers make?
The biggest failure mode is accepting AI-generated themes uncritically, which often leads to fabricated quotes and sycophantic over-interpretation of your data. If you simply paste AI results into your manuscript, reviewers will flag it immediately.
Large language models suffer from sycophancy. They want to please you. If you prompt an AI to find examples of organizational conflict, it will over-interpret neutral statements to satisfy your request.
Let us look at a worked example. Imagine you are studying workplace dynamics and you ask an AI to find conflict. The AI might flag quotes like "We had different perspectives on the timeline" or "We debated the strategy extensively" as dysfunctional conflict. But as a human researcher, you know the difference between healthy debate and actual interpersonal friction. If you just accept the AI's tags, your findings will be skewed. A better approach is to use software to systematically isolate all passages where participants discuss disagreements, and then manually review them to apply theoretical nuance. You need to distinguish between strategic disagreement and interpersonal friction yourself.
Another critical risk is fabricated quotes. LLMs occasionally generate plausible-sounding quotes that are not verbatim from the actual transcript. Using these in a manuscript is a major research integrity issue, and blaming the AI is not a valid defense.
The 20-Question Documentation Framework
To ensure your methodology chapter is bulletproof, you can structure your pre-publication check around a 20-question framework divided into four core categories. This is the exact kind of structured rigor that reviewers are looking for.
Category | Core Focus | Example Question to Address in Your Methods |
|---|---|---|
1. Documentation | Naming the specific tools and their precise functions. | Did we specify the exact version of the AI tool and document the prompts or configuration settings used? |
2. Traceability | Linking findings back to the raw data. | Can every quote and theme be traced directly back to a timestamped original transcript? |
3. Quality & Bias | Checking for AI hallucinations and over-interpretation. | Did we actively search for negative cases that the AI's pattern-matching might have missed? |
4. Human Oversight | Proving the researcher made the final interpretive calls. | Have we documented instances where we merged, split, or rejected AI-generated themes? |
How do you prove human engagement to reviewers?
You prove human engagement by explicitly describing where you disagreed with the AI and how you modified its initial suggestions. Documenting your decisions to merge, split, or rename AI-generated themes shows that you maintained intellectual control.
Reviewers want proof that the human researcher made the final interpretive judgments. You can work these exact validation decisions into your methods write-up. Explain how you renamed or redefined some of the AI-generated themes. Document instances where you merged themes the AI separated, or split themes the AI combined. Show where you added themes the AI completely missed. Explain why you rejected certain themes proposed by the AI as being insufficiently grounded in the data.
This is where your choice of software matters. If you are using generalist chatbots, keeping this audit trail is a nightmare. This is why purpose-built tools are becoming the standard. For example, using a tool like Paideias keeps the researcher in the loop by design, making it much easier to pull the exact documentation you need for your methodology chapter. If you are curious about what this looks like in practice, you can read more about how Researcher-in-the-Loop Isn't a Buzzword. In Paideias, It's the Whole Design.
Frequently Asked Questions
Should I include my exact AI prompts in the appendix?
Yes, including your prompts or configuration settings in an appendix or supplemental materials is highly recommended. It allows other researchers to understand exactly how you directed the AI and provides a level of transparency that peer reviewers increasingly demand. If your prompts are proprietary or too complex, you should at least provide a structural overview of the instructions you gave the tool.
Does using AI for coding violate data privacy rules?
It certainly can if you are not careful. Many common chatbots store data on external servers or use uploaded data to train their models, which violates strict GDPR requirements for non-anonymized qualitative interview data. While anonymizing transcripts before AI processing is the safest approach, be aware that anonymization has physical limits. Certain unique combinations of demographic details can still re-identify participants. Always use tools with enterprise-grade security and explicit data privacy agreements.
Do I need to report negative case analysis?
Absolutely. Researchers must actively search for and report data that challenges or complicates their emerging themes. This is vital when using AI, because pattern-matching tools naturally prioritize data that fits patterns over data that disrupts them. Documenting your search for negative cases proves to reviewers that you did not just blindly accept the AI's sanitized version of your data.
Discussion
or sign in to comment with your account