The three new tests that were actually put to a trial
One was compared against the standard and failed to show it was not worse. One has never been randomised at all. One made scans faster without finding more.
Most of what is sold in this space has never been tested against anything. Three technologies have real comparative data. Here is what it showed.
1. Testing an embryo without touching it. Instead of biopsying cells, the laboratory reads the DNA the embryo has shed into the drop of fluid it was grown in. The largest study collected 1,301 blastocysts across eight IVF centres on four continents.
A false positive here can mean a normal embryo is set aside. Roughly one result in eight in that study was a false positive, in the largest and most favourable study of the method.
A smaller study went further and compared each sample against the inner cell mass, the part of the embryo that actually becomes the baby. In 35 donated research embryos, the fluid sample matched the inner cell mass in 14 of 24 comparisons. The standard biopsy did better but not by much: 22 of 32. Those are small numbers and one case moving changes them, which is why they are given as fractions. They make one point clearly: the biopsy everyone treats as the reference is not itself a perfect picture of the baby.
The current systematic review concludes: General concordance rates in comparison with biopsy-based PGT-A are promising, but it is clear that additional research and understanding are needed before adopting noninvasive and minimally invasive PGT-A as a widely used tool with strong clinical utility.
No randomised trial has reported whether using it changes anyone's chance of having a baby. Three are registered and recruiting.
2. Artificial intelligence choosing the embryo. This one has a completed randomised trial: 14 IVF clinics in Australia and Europe, 1,066 women, double blind, comparing a deep learning algorithm against an embryologist looking down a microscope.
| Outcome | AI | Embryologist | Difference (95% confidence interval) |
|---|---|---|---|
| Clinical pregnancy | 46.5% | 48.2% | 1.7 points lower (7.7 lower to 4.3 higher) |
| Live birth | 39.8% | 43.5% | 3.9 points lower (9.9 lower to 2.2 higher) |
The trial's own conclusion: This study was not able to demonstrate noninferiority of deep learning for clinical pregnancy rate when compared to standard morphology and a predefined prioritization scheme.
Be precise about what that means. The trial did not show the software is worse. It failed to show that it is not worse. The interval for live birth still allows a real loss of nearly ten percentage points. Every point estimate favoured the human. The accompanying editorial in the field's own journal: Good blastocyst morphology has again proven itself as a high bar in predicting live birth.
3. Artificial intelligence in the anomaly scan. One randomised trial exists. It is important to say how this guide knows its numbers: the published journal version could not be opened, so every figure below comes from the trial's preprint, which has not been peer reviewed. Treat them as provisional.
The trial randomised 78 pregnancies at a single hospital, deliberately enriched so that 26 of them, about a third, involved congenital heart disease. Scans were much faster with the software, a median of 11.4 minutes against 19.7 minutes, and the sonographers' measured workload was lower. Sensitivity for picking up a malformation was 88.9% with the software and 81.5% without, a difference that was not statistically significant.
The authors' own conclusion claims a time saving and a workload reduction without a reduction in diagnostic performance.
That is not a claim that more was found.
Everything else in this area is retrospective. In the largest review of AI for spotting fetal heart problems, 15 studies covering pictures from 30,121 fetuses, 14 of the 15 teams built and tested their own model, only one model was checked by anyone else, the risk of bias was rated moderate to high, and adherence to AI reporting standards was low. No study was found in which AI improved anomaly detection in a real screening population, and none reporting whether it changed anything for a baby.
"If the answer is agreement rates, agreement with what, and how often is it wrong in each direction?"
"Does faster or easier for the operator mean better for my baby, and how would we know?"