ART Success Rates: Always Ask for the Denominator

A hoarding near the highway advertises a fertility centre with a seventy-eight percent success rate. Down the road, another claims sixty-five. A couple sitting in a consultation room has seen both. They ask the obvious question, which is why anyone would choose the second clinic, and the honest answer is that the two numbers are not measuring the same thing and neither of them has told you what they divided by.
This is not mainly a marketing problem. It is a measurement problem, and it starts inside the clinic, where the same ambiguity makes it impossible to tell whether this year was better than last.
Every rate is a fraction, and the argument is always about the bottom
A success rate is a numerator over a denominator. In ART the numerator is usually clear enough. The denominator is where the disagreement lives, and moving it changes the answer dramatically without anybody stating an untruth.
Take clinical pregnancy. Per embryo transfer, the denominator excludes every cycle that was cancelled, every retrieval that produced nothing to transfer, and every freeze-all that has not come back yet. Per oocyte retrieval, it includes the failed fertilisations. Per cycle started, it includes the cancellations. The same clinic, the same year, the same patients: three quite different percentages, all correct.
- Per transfer flatters clinics that cancel freely and transfer selectively.
- Per retrieval is harder, and closer to what a patient going through a retrieval experiences.
- Per cycle started is hardest, and closest to the question a couple is actually asking.
- Cumulative live birth per retrieval, counting the fresh transfer and every subsequent frozen transfer from that cohort, is the most honest of all and the slowest to compute, because it needs the follow-up to be complete.
A rate without its denominator is not a statistic. It is a claim.
The indicators worth watching inside the clinic
Headline pregnancy rates are for the outside world. The numbers that tell a clinic something it can act on are further up the chain, because they isolate a step.
- Cycle cancellation rate, split by reason: poor response, over-response and risk of hyperstimulation, personal reasons.
- Oocyte yield against the antral follicle count that was expected, which tests both stimulation and the honesty of the counting.
- Maturity rate: mature oocytes as a share of those retrieved.
- Fertilisation rate, calculated on mature oocytes for ICSI and on inseminated oocytes for conventional IVF. Conflating the two makes the number meaningless.
- Blastocyst conversion, as a share of normally fertilised embryos.
- Usable embryo rate per retrieval, which is what determines whether a couple gets more than one attempt from one stimulation.
- Implantation rate, sacs per embryo transferred, which is the closest thing to a measure of embryo quality.
- Miscarriage rate, because a pregnancy rate uncorrected by this can hide a real problem.
- Freeze-thaw survival, which is a direct measure of laboratory technique.
A shift in any one of these localises the issue. Fertilisation down while maturity holds points at the andrology side or the injection technique. Blastocyst conversion down while fertilisation holds points at culture conditions. Implantation down while grading is unchanged points at the transfer, or at grading that has quietly drifted.
Small numbers lie, confidently
A clinic doing forty retrievals a month is looking at small samples the moment you segment by age band, protocol and clinician. Nine cycles in a cell is not a trend, but it will render as a percentage that looks exactly as authoritative as one computed from nine hundred.
The discipline that fixes this is simple and rarely implemented: set a minimum denominator, and have the system decline to report below it. Not display zero, not display a dash — say plainly that there are too few records to report, and show the count. It feels like a loss of functionality. It is the opposite: it stops the monthly meeting spending twenty minutes on a swing that was two patients.
The same applies to comparisons between clinicians. Case mix differs, referral patterns differ, and the doctor who takes the difficult cases will look worse on every raw measure. Comparison without adjustment for age and reserve is not analysis; it is a way of making colleagues defensive.
Data quality is the number nobody puts on the dashboard
Every rate above assumes the underlying data exists. In practice a clinic's dashboard is quietly built on whatever was recorded, and the gaps do not announce themselves.
The cycles that disappear are always the same kind. Cancelled cycles that were never closed in the system and sit as active forever. Couples who went elsewhere after a failure, so no outcome was ever recorded. Freeze-alls whose subsequent transfers were entered as fresh cycles. Beta hCG results that came back to a phone and were never entered.
Each of these biases the result in the same direction, because the missing cases are disproportionately the failures. A clinic with poor follow-up capture will always look better than it is, and will never know by how much.
So the first panel on a KPI dashboard should not be a success rate. It should be completeness: how many cycles in this period have a recorded outcome, how many are still open past their expected close, how many are missing a fertilisation check. A clinic reporting a pregnancy rate on sixty percent outcome capture is reporting on a sample it did not choose.
Building something people actually use
The dashboards that get used share a few properties, and they are mostly about trust rather than visualisation.
Every figure states its numerator and denominator on the face of it, not in a tooltip. Every figure is clickable through to the underlying cycles, because the first reaction of any clinician to a surprising number is to want to see the cases, and a dashboard that cannot show them gets dismissed once and never revisited. Definitions are written down and visible, so that "clinical pregnancy" means the same thing in March as it did in January. Filters cover age band, cycle type, protocol, fresh versus frozen, and clinician, because the aggregate is rarely the interesting view.
And the whole thing refuses to invent a number. If the denominator is below the threshold, it says so.
What to do with the numbers once you have them
The point of measuring is to change something, and the change is usually less dramatic than expected.
Look at the step, not the headline. Compare cohorts, not months, since seasonal variation in a small clinic is mostly noise. Change one thing at a time, because two simultaneous changes to the lab protocol produce an uninterpretable result. Give it enough cycles to mean something before deciding. And when a number improves, check the completeness panel first, because the commonest cause of a sudden improvement is a change in what got recorded rather than in what happened.
The clinics that improve fastest are not the ones with the most sophisticated analytics. They are the ones that record outcomes completely, define their terms once, and are willing to look at a number that is worse than last quarter without immediately explaining it away.


