Bot or Not: Readers Preferred the AI Story and Could Not Identify It
Three studies, 2,587 US adults, and a bias that points the opposite way from the blindness it accompanies.
Summary
Sears and Weisberg ran three studies asking whether ordinary readers can tell AI-generated short fiction from human-written short fiction, and how being told the author's identity changes their judgment. Six roughly 1,000-word stories were used: three published by human authors, three generated by ChatGPT 4.0 from prompts built to match each human story's premise, symbol, and point of view.
In Study 1, 1,682 readers each read one story and rated it. The ChatGPT stories scored higher on both absorption (1.42 vs 1.00) and perceived quality (1.54 vs 0.97) on scales running −3 to +3. Independently, being told a story was human-written raised ratings regardless of who actually wrote it (absorption 1.32 vs 1.10; quality 1.40 vs 1.12). Readers preferred the machine text while penalizing the machine label.
In Studies 2 and 3, 905 readers saw both stories side by side and were told outright that one was human and one was AI. Study 2 participants were correct 39.4% of the time — significantly below chance. Study 3 participants were correct 52.0% of the time — indistinguishable from chance. Self-reported AI expertise predicted accuracy in both; expertise with fiction predicted nothing. Readers who said they judged by the writing's language did worse than those who did not.
Clinical bottom line
- Preference and attribution came apart: the same readers rated AI text higher and rated the AI label lower.
- In a forced two-way choice with the base rate disclosed — an easier task than anything real life offers — accuracy was at or below chance in both studies.
- Confidence did not track accuracy, and neither did time spent reading. Feeling certain about spotting AI text carries no information.
- AI literacy was the only measured predictor of getting it right. Literary expertise was not, which argues that detection is a trained skill rather than a taste.
- Scope limits worth naming: six short realistic-fiction stories, Prolific participants, no preregistration, and two studies months apart that disagreed on whether accuracy was below chance or merely at it.
Key references
- Sears S, Weisberg DS. Bot or not: can people tell the difference between stories written by a human or by an AI system? Judgm Decis Mak. 2026;21:e21.Primary source. All figures in this infographic are taken from this paper.doi:10.1017/jdm.2026.10042 →
- Porter B, Machery E. AI-generated poetry is indistinguishable from human-written poetry and is rated more favorably. Sci Rep. 2024;14(1):26133.Prior study using the same 2×2 design in poetry, with concordant results.doi:10.1038/s41598-024-76900-1 →
- Köbis N, Mossink LD. Artificial intelligence versus Maya Angelou: experimental evidence that people cannot differentiate AI-generated from human-written poetry. Comput Human Behav. 2021;114:106553.Shows detection collapses once a human selects the best AI output.doi:10.1016/j.chb.2020.106553 →
- Dietvorst BJ, Simmons JP, Massey C. Algorithm aversion: people erroneously avoid algorithms after seeing them err. J Exp Psychol Gen. 2015;144(1):114–26.The algorithm-aversion framework behind the labeling effect.doi:10.1037/xge0000033 →
Related infographics

Clinical AI Competence in Obstetrics and Gynecology
Adoption has outrun training: 81% of physicians report awareness or use of AI, while fewer than 15% report formal expertise in its clinical application.

Placental Volume: 2D EPV vs 3D VOCAL™
A 30-second handheld measurement, checked against the gold standard in 58 paired scans.

Vaginal Progesterone in Twins: No Neurodevelopmental Signal at 6-9 Years
A blinded follow-up of a three-arm randomized trial found no effect of 200 or 400 mg/day vaginal progesterone on behavior or cognition in 206 dichorionic twins aged 6 to 9.