ObGyn Intelligence · Infographics
Infographics/ Clinical AI/Bot or Not: Readers Preferred the AI Story and Could Not Identify It
Clinical AI

Bot or Not: Readers Preferred the AI Story and Could Not Identify It

Three studies, 2,587 US adults, and a bias that points the opposite way from the blindness it accompanies.

Click the figure to open it full size.

Published August 5, 2026  ·  Category Clinical AI  ·  Author Amos Grünebaum, MD

Summary

Sears and Weisberg ran three studies asking whether ordinary readers can tell AI-generated short fiction from human-written short fiction, and how being told the author's identity changes their judgment. Six roughly 1,000-word stories were used: three published by human authors, three generated by ChatGPT 4.0 from prompts built to match each human story's premise, symbol, and point of view.

In Study 1, 1,682 readers each read one story and rated it. The ChatGPT stories scored higher on both absorption (1.42 vs 1.00) and perceived quality (1.54 vs 0.97) on scales running −3 to +3. Independently, being told a story was human-written raised ratings regardless of who actually wrote it (absorption 1.32 vs 1.10; quality 1.40 vs 1.12). Readers preferred the machine text while penalizing the machine label.

In Studies 2 and 3, 905 readers saw both stories side by side and were told outright that one was human and one was AI. Study 2 participants were correct 39.4% of the time — significantly below chance. Study 3 participants were correct 52.0% of the time — indistinguishable from chance. Self-reported AI expertise predicted accuracy in both; expertise with fiction predicted nothing. Readers who said they judged by the writing's language did worse than those who did not.

Clinical bottom line

  1. Preference and attribution came apart: the same readers rated AI text higher and rated the AI label lower.
  2. In a forced two-way choice with the base rate disclosed — an easier task than anything real life offers — accuracy was at or below chance in both studies.
  3. Confidence did not track accuracy, and neither did time spent reading. Feeling certain about spotting AI text carries no information.
  4. AI literacy was the only measured predictor of getting it right. Literary expertise was not, which argues that detection is a trained skill rather than a taste.
  5. Scope limits worth naming: six short realistic-fiction stories, Prolific participants, no preregistration, and two studies months apart that disagreed on whether accuracy was below chance or merely at it.

Key references

  1. Sears S, Weisberg DS. Bot or not: can people tell the difference between stories written by a human or by an AI system? Judgm Decis Mak. 2026;21:e21.Primary source. All figures in this infographic are taken from this paper.doi:10.1017/jdm.2026.10042 →
  2. Porter B, Machery E. AI-generated poetry is indistinguishable from human-written poetry and is rated more favorably. Sci Rep. 2024;14(1):26133.Prior study using the same 2×2 design in poetry, with concordant results.doi:10.1038/s41598-024-76900-1 →
  3. Köbis N, Mossink LD. Artificial intelligence versus Maya Angelou: experimental evidence that people cannot differentiate AI-generated from human-written poetry. Comput Human Behav. 2021;114:106553.Shows detection collapses once a human selects the best AI output.doi:10.1016/j.chb.2020.106553 →
  4. Dietvorst BJ, Simmons JP, Massey C. Algorithm aversion: people erroneously avoid algorithms after seeing them err. J Exp Psychol Gen. 2015;144(1):114–26.The algorithm-aversion framework behind the labeling effect.doi:10.1037/xge0000033 →
artificial intelligencelarge language modelsChatGPTTuring testalgorithm aversionAI literacydetectioncreative writing

Related infographics

This figure reproduces data as published in the cited sources. It is educational and decision-support content and does not replace individualized clinical judgment. Where the source study's design limits how far its numbers travel, that limit is stated on the figure rather than left to the reader.