Use case
5 min read

AI vs Human Interviewers: What the Evidence Shows

AI-moderated interviews
AUTHOR
Elena
PUBLISHED ON
July 31, 2026
TABLE OF CONTENT
Try Glaut
SUMMARISE WITH AI

Are AI-moderated interviews as effective as human-moderated interviews?

AI-moderated interviews can match human interviewers on some outcomes, but the evidence does not support a universal equivalence claim. In Curtin University's controlled study, AI and human interviewers produced similar self-reported willingness to disclose, trust, positive experience, awkwardness and ability to answer effectively. Human interviewers produced a stronger sense of connection, a higher overall evaluation and more positive emotional engagement. The right conclusion is conditional: AI can be effective for structured interviewing where consistency and disclosure matter, while skilled human moderation retains an advantage when rapport, emotional development or exploratory depth is central.

What does "effective" mean in this comparison?

Effectiveness needs to be defined before comparing methods. The five studies point to several distinct outcomes:

  • Willingness to disclose
  • Perceived trust and comfort
  • Rapport and emotional engagement
  • Quality of probing
  • Consistency across interviews
  • Analytical usefulness

A method can perform well on one dimension and less well on another. Curtin found this exact pattern: disclosure-related measures were comparable, while connection and emotional engagement favored humans. Responsive Research found that participants liked the AI experience, but experienced researchers were more critical of probe depth and narrative development.

What did the Curtin experiment compare?

Curtin University randomly assigned 60 English-proficient students and staff to an AI interviewer or a human interviewer. There were 32 AI-moderated and 28 human-moderated interviews about fast fashion. Sessions lasted about 16 minutes.

The comparison was carefully controlled. Both conditions used questions and follow-ups generated by the same AI system. Human interviewers followed those prompts rather than moderating freely. This isolates the effect of interviewer presence, but it is not a direct test of AI against the full adaptive skill of an experienced qualitative moderator. That limitation matters when interpreting the results.

Do AI and human interviewers elicit comparable disclosure?

In Curtin's study, participants reported a similar willingness to disclose in the two conditions. Mean willingness was 5.54 with AI and 5.76 with a human, a non-significant difference. Ability to disclose effectively, positive experience, perceived trustworthiness and awkwardness also showed no significant differences. A regression model explained 57% of variation in willingness to disclose. Trustworthiness was a significant predictor, as was a positive interview experience. Sense of connection was not a significant predictor in that model.

This supports a focused claim: an AI interviewer can support self-reported disclosure when participants trust the interaction and evaluate it positively. It does not prove that the factual detail, honesty or analytical quality of the disclosures was identical. Curtin measured willingness and experience rather than conducting a blinded content-quality comparison of the interview transcripts.

Are AI interviewers as trustworthy as humans?

Curtin found no significant difference in perceived trustworthiness. Human interviewers scored 5.77 on average and AI scored 5.41, with p = .21.

The Mannheim survey comparison also found that participants trusted the AI interview format more than the static survey format, with mean scores of 4.45 versus 3.95. That result compares AI with a survey, not with a human interviewer.

Together, these findings suggest that AI does not automatically create a trust deficit. Trust still depends on the full study experience and cannot be assumed from the technology alone.

Do human interviewers build stronger rapport?

Yes, in the Curtin experiment. Sense of connection averaged 5.83 with a human interviewer and 4.64 with AI. The difference was statistically significant and about 26% in favor of the human condition.

The overall interviewer evaluation also favored humans, 6.45 versus 5.96. Responsive Research reached a similar qualitative conclusion: AI interactions were comfortable for participants but felt more linear and survey-like to experienced qualitative researchers.

Does stronger rapport necessarily produce better data?

Not in every context. Curtin found stronger connection with humans, but willingness to disclose did not differ significantly. Connection was also not a significant predictor in its disclosure model.

This means rapport and disclosure should be measured separately. A warm interaction may be important in its own right, especially for emotionally complex work, but stronger rapport did not translate into higher self-reported willingness to share in this controlled study. The study did not test whether rapport improved the nuance or strategic usefulness of the content.

Are participants more emotionally engaged with humans?

Curtin's biometric results say yes. Facial-expression analysis showed mean joy of 18.43 in the human condition versus 6.24 with AI. Average heart rate, interpreted in the study as engagement, was 81.44 beats per minute with humans and 74.80 with AI. Both differences were statistically significant. Human presence therefore generated more positive activation in this lab setting.

Do AI interviews create more stress or awkwardness?

Curtin found no statistically significant increase in awkwardness, anger, fear, sadness, confusion or physiological stress with AI. Skin conductance was 3.58 in the AI condition and 2.10 in the human condition, but the difference was not statistically significant. The evidence supports saying that AI produced less positive connection, not that it imposed a demonstrated emotional penalty.

Are AI interviews more consistent?

The designs show a procedural consistency advantage. Curtin used the same predefined questions and AI-generated prompts in both conditions. Responsive Research used controlled probing parameters across participants. Mannheim held question order and skip logic constant. However, none of the studies directly measured inter-moderator variation or proved that greater consistency led to better decisions. Consistency is a design property in this evidence base, not a quantified superiority claim.

Can AI match skilled adaptive probing?

The evidence is mixed and does not establish equivalence.

Nottingham found that a single AI follow-up often increased response length and lexical diversity, and sometimes introduced new topics. Responsive Research found that AI captured depth when participants brought it, but did not reliably create depth from weaker input. Its researcher cohort described the interaction as linear and the probes as limited.

Curtin cannot settle this question because the human interviewers followed AI-generated prompts. A stronger benchmark would allow experienced human moderators to pursue unexpected meaning in real time and compare the resulting content.

Can AI follow unexpected lines or resolve contradictions?

The reviewed papers do not provide a direct controlled test of either capability. Nottingham shows that follow-ups can make underdeveloped topics more salient. Responsive Research reports limits in adaptive probing and context retention from the researcher perspective. There is not enough evidence here to claim that AI can identify contradictions as effectively as a skilled human moderator.

Which parts of human moderation remain difficult to automate?

Responsive Research identifies the clearest gaps:

  • Rescuing a weak or inarticulate answer through adaptive probing
  • Developing emotional nuance
  • Building a narrative across turns
  • Protecting meaning during synthesis
  • Deciding when an unexpected tangent is strategically important

Curtin adds direct evidence for stronger human connection and positive emotional engagement.

Is AI a substitute or a complementary method?

The studies support a complementary, fit-for-purpose role. Responsive Research recommends AI for concept screening, message testing, early directional feedback and rapid pattern detection across larger samples. It gives human moderation priority for deep exploratory work, emotional journeys and complex decision dynamics. Curtin recommends considering a hybrid design when emotional nuance needs a human layer.

Does performance vary by topic or objective?

Probably, but the current evidence is narrow.

  • Curtin studied fast fashion. Responsive Research studied menopause. Mannheim studied healthy lifestyles. Nottingham studied dairy calf welfare. Human Highway studied online reviews.
  • Responsive Research explicitly warns that performance on a sensitive health topic may not generalize to brand research or behavioral inquiry. Method choice should therefore follow the decision and topic, not a single headline result.

When does a human moderator add clear value?

Choose human-led moderation when the research depends on strong rapport, emotional attunement or the ability to develop an unexpected narrative. Human moderation is also the safer choice when a participant may need immediate emotional support. AI is a stronger fit when the discussion guide is structured, comparable coverage matters and the goal is to detect patterns across more interviews. A hybrid design is appropriate when both needs are material.

Frequently asked questions

Did AI produce less disclosure than humans?

No significant difference appeared in Curtin's self-reported willingness to disclose. The study did not compare actual transcript quality blind to method.

Did participants trust the human more?

Perceived trustworthiness did not differ significantly in Curtin's sample.

Was the human interviewer allowed to probe freely?

No. Human interviewers followed AI-generated prompts, which limits conclusions about skilled adaptive human moderation.

Did AI make participants more uncomfortable?

No significant differences appeared in awkwardness or negative emotional measures.

Can this evidence support replacing all human interviews?

No. The studies show different capability profiles and recommend selecting the method by research objective.

Sources

This is some text inside of a div block.
5 min read

Heading

Use case
Use case
AUTHOR
Giacomo
LAST UPDATED AT
This is some text inside of a div block.
TABLE OF CONTENT
Try Glaut

Heading 1

Heading 2

Heading 3

Heading 4

Heading 5
Heading 6

Lorem ipsum dolor sit amet, consectetur adipiscing elit, sed do eiusmod tempor incididunt ut labore et dolore magna aliqua. Ut enim ad minim veniam, quis nostrud exercitation ullamco laboris nisi ut aliquip ex ea commodo consequat. Duis aute irure dolor in reprehenderit in voluptate velit esse cillum dolore eu fugiat nulla pariatur.

Block quote

Ordered list

  1. Item 1
  2. Item 2
  3. Item 3

Unordered list

  • Item A
  • Item B
  • Item C

Text link

Bold text

Emphasis

Superscript

Subscript