Key Takeaways
- Raw pregnancy rates, clinical pregnancy rates, and live birth rates measure completely different outcomes and cannot be directly compared
- Per-cycle, per-transfer, and per-patient denominators produce dramatically different success rates from the same data
- Age stratification is mandatory for meaningful benchmarking, not optional
- Case-mix adjustment accounts for patient profile differences between clinics
- Without these corrections, your benchmarking is comparing unlike things and drawing wrong conclusions
If I show you two clinics, one reporting a 45% success rate and another reporting 32%, which one is performing better?
The honest answer is: you have no idea. And neither do most medical directors who are asked this question.
Not because the data is wrong. But because "success rate" is not one number. It's at least nine different numbers, depending on what outcome you're measuring and what denominator you're dividing by. And unless both clinics are calculating it the same way, on comparable patient populations, that comparison is meaningless.
This matters more than it sounds. Because when UAE fertility clinics benchmark themselves against published literature, against competitor marketing, or against their own historical performance, the majority are comparing numbers that were never designed to be compared. And then making clinical or commercial decisions based on a comparison that doesn't hold.
The three outcomes that get called the same thing
Let's start with the outcome. What are you actually measuring when you say "success"?
Positive pregnancy test (biochemical pregnancy). A rising hCG level two weeks after embryo transfer. This is the earliest measurable sign of pregnancy. It's also the least clinically meaningful, because a significant proportion of biochemical pregnancies miscarry before a gestational sac is ever visible on ultrasound. If your clinic reports this number as your headline success rate, you're overstating your clinical pregnancy rate by 10 to 20 percentage points.
Clinical pregnancy. An intrauterine gestational sac visible on ultrasound, usually confirmed around 6 to 7 weeks. This is the standard reporting metric in most peer-reviewed IVF literature and the one used by registries like ESHRE. It's more clinically stable than a biochemical pregnancy. But it still isn't a baby. A clinical pregnancy that miscarries at 8 weeks is counted as a "success" under this definition, even though the patient does not take home a child.
Live birth. A delivery of at least one live-born infant. This is what patients actually care about. It's the hardest outcome to achieve and the one that survival bias affects least. It's also the metric that takes the longest to collect, because you have to wait 9 months after the transfer to know the result. Many clinics don't track this systematically. And even fewer report it as their primary success metric, because it's 8 to 12 percentage points lower than their clinical pregnancy rate.
Three definitions. A clinic reporting biochemical pregnancy per transfer will have a headline number that's 15 percentage points higher than a clinic reporting live birth per started cycle. Both might call it their "IVF success rate" in patient-facing marketing. And a patient comparing them has no way to know they're looking at completely different things.
The denominator problem nobody talks about
Now let's talk about what you're dividing by. Because even if two clinics are measuring the same outcome, the denominator determines whether the numbers are comparable.
Per started cycle. Every patient who begins ovarian stimulation is included in the denominator, even if they cancel before egg retrieval, retrieve no eggs, produce no embryos, or never make it to transfer. This gives you the lowest success rate, because you're counting all the cycles that failed early. It's the most conservative denominator and the one that most accurately reflects a patient's chance of success if they walk into your clinic and start a cycle.
Per egg retrieval. Only cycles that proceeded to egg collection are counted. Cycles that cancelled during stimulation are excluded. This inflates your success rate relative to per-started-cycle, because you've removed the worst-performing patients from the denominator. If your cancellation rate is 12%, your per-retrieval success rate will be about 12% higher than your per-cycle rate, even though nothing about your lab or clinical performance has changed.
Per embryo transfer. Only cycles that reached embryo transfer are counted. This excludes cancellations and retrieval cycles that produced no viable embryos. It gives the highest reported success rate of the three, because the denominator now contains only patients who made it past every early-stage filter. A clinic with high early attrition will see a massive gap between its per-cycle rate and its per-transfer rate. The per-transfer number is not wrong. But it's answering a different question: "If you make it to transfer, what are your chances?" rather than "If you start a cycle, what are your chances?"
Same outcome, three denominators. A clinic reporting clinical pregnancy per transfer might publish 52%. The same clinic reporting live birth per started cycle would be at 34%. Both numbers are correct. Neither is comparable to a competitor clinic unless that clinic is using the same combination of outcome and denominator.
Why age stratification is not optional
Here's the thing that makes all of this worse: even if two clinics are measuring the same outcome with the same denominator, the comparison still doesn't work unless they have comparable patient populations.
A 28-year-old patient with normal ovarian reserve has roughly double the per-cycle live birth rate of a 41-year-old patient with diminished reserve. If Clinic A treats primarily younger patients and Clinic B treats primarily older patients, Clinic A will have a higher aggregate success rate for reasons that have nothing to do with clinical quality.
This is why ESHRE's IVF benchmarking indicators stratify success rates by age group. The standard categories are under 35, 35 to 37, 38 to 39, 40 to 42, and over 42. When you break your data this way, you can benchmark your under-35 rate against another clinic's under-35 rate. That comparison is fair, because the populations are biologically similar.
An aggregate success rate is not. If you're reporting an overall clinic success rate of 42% and comparing it to a published benchmark of 38%, you need to know what the age distribution was in that benchmark population before you conclude that you're outperforming. If your patient population is 5 years younger on average, your 4-percentage-point advantage is not performance. It's selection bias.
Case-mix adjustment: the step most clinics skip
Age stratification solves part of the problem. But patient populations differ in more ways than just age. AMH levels, BMI, diagnosis, number of previous failed cycles, use of donor eggs, fresh versus frozen transfers. all of these factors independently affect success rates.
If you're benchmarking against another clinic or against your own historical data, and the case mix has changed, your raw success rate comparison is confounded. The way to fix this is case-mix adjustment: a statistical model that estimates what your success rate would have been if you'd treated the same patient population as the comparator.
In a well-executed benchmarking analysis, you'd run a logistic regression model with live birth as the outcome and patient-level covariates like age, AMH, BMI, embryo quality, and transfer day as predictors. Then you'd use that model to generate an expected success rate for each patient, sum those expectations, and compare the sum to your observed success rate. The difference is your performance gap, adjusted for case mix.
This is not standard practice in most UAE fertility clinics. It should be. Because without it, you're attributing differences in success rates to clinical performance when a significant portion of the difference is patient selection.
What correct benchmarking actually looks like
If you want to benchmark your IVF success rates in a way that produces actionable, valid comparisons, here's what you need to do:
- Pick one outcome definition and stick to it. Live birth per started cycle is the gold standard for patient-facing reporting.
- Stratify by age using standard categories. under 35, 35–37, 38–39, 40–42, over 42. Report each separately.
- If comparing to published benchmarks, verify that the benchmark is using the same outcome, the same denominator, and a comparable time period. A 2018 ESHRE benchmark is not a valid comparator for your 2026 data unless you account for protocol changes and population shifts.
- Apply case-mix adjustment when comparing across clinics or across time periods with different patient profiles. This requires logistic regression modeling. It's not optional if you want the comparison to be valid.
- Separately track cycles using donor eggs, frozen embryos, and PGT-tested embryos. These subgroups have systematically different success rates and should not be aggregated into a single headline number.
This is more work than pulling a single percentage from your database and putting it in a report. But it's the only way to know whether the number you're looking at actually means what you think it means.
Where most clinics are getting it wrong
The mistakes I see most often in UAE clinic benchmarking are:
Comparing clinical pregnancy rates from your clinic to live birth rates from a competitor. You will always look better. The comparison is not valid.
Using per-transfer success rates in patient-facing marketing without clarifying the denominator. Patients assume "success rate" means "if I start treatment, what are my chances." If you're reporting per-transfer, you're answering a different question. And you're overstating the probability by the size of your cancellation and attrition rate.
Benchmarking an aggregate success rate against age-stratified literature without adjusting for your own age distribution. If your patient population skews younger, your higher aggregate rate is expected. It does not indicate superior clinical performance.
Comparing this year's success rate to last year's without accounting for changes in patient case mix. If your average patient AMH dropped by 0.4 ng/mL between 2025 and 2026, a flat success rate is actually an improvement. A raw year-on-year comparison misses this.
All of these errors are common. All of them lead to wrong conclusions. And all of them are fixable with the right analytic approach.
Want to benchmark your IVF outcomes correctly?
StatZen Analytics helps UAE fertility clinics calculate and compare success rates the right way. See our IVF & Fertility Analytics service or book a free 30-minute consultation to discuss your benchmarking challenges.
About StatZen Analytics
StatZen Analytics is a UAE-based healthcare data consultancy specialising in biostatistics, clinical analytics, and performance monitoring for fertility clinics, hospitals, and polyclinics. Founded by a biostatistician with over 3 years of hands-on UAE fertility hospital experience.
