
AI-generated follow-up questions can produce new information, but their value depends on the initial question and the probing instruction.
In the University of Nottingham study, one AI follow-up increased average response length by about 30 words and length-adjusted lexical diversity by 9%. Some probes made previously rare topics much more salient. Other probes mostly repeated or reinforced what participants had already said. AI probing worked best after broad or unfamiliar questions and added less when the original question already asked for reasons.
Nottingham recruited 296 UK participants for a study about public perceptions of dairy calf welfare. The sample was representative by gender, age and region. Participants answered six researcher-written open questions, each followed by one AI-generated probe.
Researchers compared the initial answer with the combined initial and follow-up answer. They measured word count and corrected type-token ratio, then used keyness analysis to identify words and topic clusters that became more prominent after the probe.
This design evaluates the incremental contribution of a follow-up rather than comparing two separate groups.
They produced more material and more varied language.
The model-adjusted word count increased from 41.16 for the initial answer to 71.11 after adding the AI follow-up response. The difference was 29.95 words, p < .001.
Corrected type-token ratio increased from 3.43 to 3.75 after controlling for word count, a 9% gain. Because the lexical-diversity model adjusted for length, the result is evidence against simple repetition as the only explanation.
Whether the extra information was genuinely new depended on the question.
The clearest expansion followed broad questions or probes that sought information not already requested.
These examples show that the probe can make an under-articulated dimension visible across many participants.
AI probes were less informative when the initial question already requested the same reasoning. For "Why is it important that dairy calves are provided with a good life?", no new common topic appeared after the probe. The static question already asked why, while the probing instruction also focused on reasons and motivations.
A question about cow-calf separation produced nearly identical prevalence for a wellbeing topic, 71 initial responses and 72 follow-up responses. The probe reinforced the topic rather than expanding it. For a question about disbudding, legal or standards-related considerations increased from 14 to 24 respondents, but the difference was not statistically significant.
No. Nottingham found both expansion and reinforcement
A follow-up can:
These outcomes should be distinguished during evaluation. More words are not the same as a new analytical direction.
Broad questions give participants room to choose what is salient, but their first answers may remain thin. A targeted follow-up can ask them to connect the answer to consequences, examples or a neglected dimension. Nottingham also notes that calf welfare was unfamiliar to much of the public. Participants may have needed a prompt to articulate considerations they did not spontaneously structure in their first reply.
The study therefore recommends AI follow-ups when the topic is unfamiliar and the respondent has not yet explained their reasoning.
A second request for reasons can become tautological. Nottingham observed this pattern on questions that already required justification.
The design lesson is to give the static question and follow-up different jobs. If the main question asks why, the probe might seek an example, a condition under which the answer changes or an area not yet covered. Repeating "why" is unlikely to create much incremental value.
Yes.
The wellbeing example rose from 10 to 110 respondents after the probe. This may reveal a consideration that was present but under-articulated. It also demonstrates that the instruction directed attention toward a domain. Researchers should preserve the distinction between spontaneous and prompted content. They should not report a prompted topic as if it emerged unaided.
In Nottingham, yes on average. Corrected type-token ratio rose by 0.318 after controlling for total word count. The paper concludes that respondents used richer and more varied vocabulary rather than simply repeating the initial answer.
This is an aggregate result. Individual probes can still be repetitive, especially when the prompt duplicates the static question.
Human Highway's qualitative analysis found that conversational AI responses more often included explanations, personal episodes, exceptions and decision criteria than traditional questionnaire responses. Nottingham shows that prompts can explicitly seek reasons or examples.
The value comes from selecting the missing form of information. If the participant has already supplied a reason, another request for reasons has low marginal value.
Nottingham tested one follow-up per static question. Mannheim allowed dynamic probing beyond two turns but analyzed only the first two for comparability. Responsive Research used controlled probing parameters.
None of the studies identifies an optimal number. The evidence supports evaluating incremental value after each probe rather than assuming that more turns produce more insight.
No.
A probe is most justified when the answer is incomplete relative to the research objective. Nottingham found that some questions gained new topics and others did not. Applying the same probing rule everywhere can add respondent burden and analytical volume without equivalent value.
The studies suggest comparing each probe with what came before:
A probe has reached diminishing returns when it adds little new content or starts to restate the participant's answer.
Nottingham provides a replicable framework.
A human qualitative review remains important because frequency alone cannot determine usefulness.
The research shows that dynamic probes can be combined with fixed questions and closed items. Nottingham describes the method as an AI-powered survey, although its fieldwork was hosted on Glaut. Specific third-party integrations were not tested in these papers.
Nottingham compared six questions within one unfamiliar and potentially sensitive topic. The results varied by question wording, which is strong evidence that probe performance is context dependent. A broader cross-topic benchmark would need more studies. Current results should not be generalized to every category or respondent population.
No. Some expanded the topic, while others reinforced existing reasoning.
About 30 words, from a model-adjusted 41.16 to 71.11 for the combined response.
Yes. Length-adjusted lexical diversity increased by 9%.
A probing instruction that repeats the task already given in the static question.
Collect, analyze, and report research from any source with more depth, speed, and control.
Schedule a free demo
The AI-native research platform for modern researchers. Deliver insights 5x deeper, 20x faster with AI-moderated voice interviews and agentic analysis, in 50+ languages.

Lorem ipsum dolor sit amet, consectetur adipiscing elit, sed do eiusmod tempor incididunt ut labore et dolore magna aliqua. Ut enim ad minim veniam, quis nostrud exercitation ullamco laboris nisi ut aliquip ex ea commodo consequat. Duis aute irure dolor in reprehenderit in voluptate velit esse cillum dolore eu fugiat nulla pariatur.
Block quote
Ordered list
Unordered list
Bold text
Emphasis
Superscript
Subscript