Clinical AI Competence in Obstetrics and Gynecology
Adoption has outrun training: 81% of physicians report awareness or use of AI, while fewer than 15% report formal expertise in its clinical application.
Summary
This Clinical Opinion argues that competent, accountable use of generative artificial intelligence is becoming a component of safe professional practice in obstetrics and gynecology. The authors define clinical AI competence as the ability to use AI tools for appropriate clinical tasks while critically appraising output, recognizing hallucinations and outdated guidance, verifying sources, preserving confidentiality, communicating uncertainty, and maintaining independent clinical judgment. Their central claim is not that clinicians will use AI, but that the risk lies in using it poorly.
The evidentiary tension is a competence gap. Survey data cited in the paper place physician awareness or use of AI at 81%, while fewer than 15% of students and faculty report formal expertise in its clinical application; roughly one in three US adults now report using large language models for health information. On the benefit side, a 70-clinician randomized trial found structured clinician-LLM workflows reached 82% to 85% diagnostic accuracy versus 75% with conventional resources, and documentation, readability, and evidence-retrieval tasks show the clearest efficiency evidence. Autonomous clinical decision-making without physician review is described as inconsistent with professional responsibility and unsupported by current evidence.
The argument is framed historically. Obstetric ultrasound, electronic fetal monitoring, cell-free DNA screening, and cervical-length screening each provoked early resistance before becoming standard, a pattern the authors term primum non mutare (first do no change), which they propose replacing with primum benefacere (first do good), under structured use, source verification, and explicit physician oversight. They are explicit that direct evidence showing improved maternal, fetal, or neonatal outcomes from clinical AI use has not been established and that outcome-based evaluation should be a research priority.
Clinical bottom line
- This is a Clinical Opinion, not a trial: it synthesizes survey data, diagnostic-accuracy benchmarks, and historical parallels rather than reporting a primary outcome study. Its central figures are drawn from cited third-party sources.
- The framing statistic is a gap between use and competence — 81% awareness or use versus fewer than 15% with formal expertise — reported from two independent surveys with different denominators, so the derived percentage-point gap is illustrative rather than a within-sample comparison.
- The strongest efficacy signal cited is a 70-clinician randomized trial (82–85% vs 75% diagnostic accuracy for structured clinician-LLM workflows); benefits are most established for documentation, readability, and evidence retrieval.
- The authors draw a firm line at autonomous AI clinical decision-making without physician review, which they hold to be inconsistent with professional responsibility and unsupported by current evidence.
- Direct evidence that AI integration improves maternal, fetal, or neonatal outcomes has not been established; whether AI narrows the documented 17-year evidence-to-practice gap remains an open research question.
Key references
- Grünebaum A, Dudenhausen J, Chervenak FA. Clinical artificial intelligence competence in obstetrics and gynecology: patient safety, physician accountability, and responsible use. Am J Obstet Gynecol. 2026.Primary source. All figures in this infographic are taken from this paper.doi:10.1016/j.ajog.2026.06.011 →
- Ke Y, Jin L, Ong JCL, et al. AI-induced never-skilling in medical education. Nat Med. 2026;32:1997–2006.Cited source for the finding that fewer than 15% of students and faculty report formal expertise in clinical AI use.
- Everett SS, Bunning BJ, Jain P, et al. From tool to teammate in a randomized controlled trial of clinician-AI collaborative workflows for diagnosis. NPJ Digit Med. 2026;9:409.The 70-clinician randomized trial cited for 82–85% vs 75% diagnostic accuracy.
Related infographics

Bot or Not: Readers Preferred the AI Story and Could Not Identify It
Three studies, 2,587 US adults, and a bias that points the opposite way from the blindness it accompanies.

Placental Volume: 2D EPV vs 3D VOCAL™
A 30-second handheld measurement, checked against the gold standard in 58 paired scans.

Vaginal Progesterone in Twins: No Neurodevelopmental Signal at 6-9 Years
A blinded follow-up of a three-arm randomized trial found no effect of 200 or 400 mg/day vaginal progesterone on behavior or cognition in 206 dichorionic twins aged 6 to 9.