AI Embryo Grading: What It Can and Cannot Tell You

Two embryologists grade the same blastocyst. One calls it 4AA, the other 4AB. Both are experienced, both are looking at the same image, and both are right in the sense that the grading scheme they are applying has a subjective component that training reduces but does not remove. Multiply that across a lab, across a year, and the grading data a clinic uses to make transfer decisions and to measure itself carries a variability nobody quite quantifies.
This is the honest case for algorithmic grading, and it is a better case than the one usually made.
What does AI embryo grading actually do?
It scores embryo images or time-lapse video against patterns learnt from historical embryos whose outcomes are known, and returns a ranking or a probability.
What it produces is a consistent number. The same embryo, assessed twice, gets the same score, which is not true of human assessment. Depending on the system it may work from a single static image at a defined time point, or from continuous time-lapse imaging that captures the timing of divisions — information a human observer cannot see without watching continuously.
What it does not produce is a new kind of knowledge. It is learning correlations between appearance and outcome in the data it was trained on, which means its usefulness is bounded by how much appearance predicts outcome in the first place, and by how much its training population resembles yours.
What it is genuinely good at
Consistency and ranking, which happen to be what the transfer decision needs most.
- Deselection. Identifying embryos unlikely to be viable is easier than identifying which one will implant, and it is clinically useful on its own.
- Ranking within a cohort. When a couple has six blastocysts, the practical question is which to transfer first, and a consistent ordering is worth having.
- Standardising the record. A score that does not drift with who was on shift makes the lab's own data usable for the first time.
- Surfacing timing. Time-lapse systems capture division timings that correlate with development and that nobody was recording before.
That last point deserves emphasis. A clinic that adopts algorithmic scoring often finds the largest immediate benefit is not the score at all. It is that grading finally became structured data, so fertilisation rates, blastocyst conversion and implantation per grade can be computed without a manual audit.
What it cannot tell you
It cannot tell you an embryo is chromosomally normal, and it cannot tell you a transfer will succeed.
The first is the misunderstanding that causes real harm in a consultation room. Morphology and ploidy are related but not equivalent. A high-scoring embryo can be aneuploid and a lower-scoring one euploid. Any system presented as a substitute for genetic testing is being oversold, and the counselling implication is direct: a patient told their embryo scored well may hear a promise that was not made.
The second is a matter of what implantation depends on. Endometrial receptivity, the transfer itself, and factors nobody has characterised all sit between a good embryo and a pregnancy. A model scoring the embryo is scoring one input to a multi-factor outcome.
A grading model ranks embryos. It does not predict babies, and a clinic that lets the distinction blur in its counselling will regret it.
The questions to ask a vendor
Most of the difference between systems is in the answers to five questions, and none of them is about the headline accuracy figure.
- What population was it trained on, in terms of age distribution, clinic geography and time period? A model built on a different population may transfer poorly to yours.
- What is the outcome label? Implantation, clinical pregnancy and live birth are different targets, and a model trained on one should not be described using another.
- Has it been validated prospectively, and independently of the developer?
- What exactly does the score mean? A number between one and ten is not a probability unless the vendor says it is and shows the calibration.
- How does it handle images it has not seen the like of before? A model that returns a confident score for an out-of-distribution image is more dangerous than one that declines.
Ask also, plainly, whether the vendor considers it a medical device and what regulatory position it holds. The answer tells you something about the vendor regardless of what it is.
How to introduce it without disrupting the lab
Run it alongside human grading before it influences anything, and keep the embryologist's assessment first.
The pattern that preserves clinical judgement is the same one that works for AI elsewhere in medicine: the embryologist grades, records their grade, and then sees the score. That order matters. Reversed, the human grade quietly becomes an endorsement of the machine's, and within a few months the lab has lost the independent assessment it would need to detect a problem.
Record both. The disagreements are the most valuable data the deployment produces — they show where the model diverges from your lab's practice, and over enough cycles, with outcomes attached, they show which of the two was closer.
Set a review point before you start, with a defined question: after this many cycles, does the score add information beyond our own grading in our own patients. Then answer it honestly, including the possibility that the answer is no.
What it means for the patient conversation
Consistency is easier to explain than a black box, and patients ask better questions than clinics expect.
Be able to say what the tool does, that it assists the embryologist rather than replacing them, that a score is a ranking and not a guarantee, and that it does not assess chromosomes. Where a score influenced which embryo was transferred, that is worth being able to explain afterwards, particularly if the cycle fails.
Algorithmic grading is a genuine improvement to a genuinely subjective task. It earns its place by making the lab's assessments consistent and its data countable — not by knowing something about an embryo that nobody else does.


