Logistic Regression · Random Forest · GBM · SHAP · Calibration · External Validation
Simulated Data for Demonstration
All datasets, model outputs, and visualizations use synthetically generated data (set.seed(2026)) for portfolio purposes. Results demonstrate advanced ML methodology.
This model was built on the same clinic group featured in Part 1. The dashboard identified where cancellations were happening. This model predicts who will cancel. before the cycle begins.
The dashboard showed where cancellations were happening. but the clinic needed a way to identify which patients were heading toward cancellation before treatment began.
A machine learning cancellation risk model built entirely on variables collected at first consultation. no additional testing required. Three models tested, Random Forest selected.
AUC 0.729 internal, 0.711 external validation. Sensitivity 82.1%. meaning the model detected 82% of at-risk cycles before treatment began. AMH emerged as the single strongest predictor.
IVF cycle cancellation occurs when ovarian stimulation results in a poor response (too few follicles) or excessive response (hyperstimulation risk), forcing the cycle to be stopped before egg retrieval. This is clinically and emotionally distressing for patients, financially costly, and represents a missed treatment opportunity.
Early identification of high-risk patients from routine pre-treatment clinical characteristics enables:
This case study develops, validates, and interprets a suite of three machine learning models(Logistic Regression, Random Forest, Gradient Boosting) with emphasis on SHAP explainability,calibration assessment, and external validation on 200 independent patients.
Objective: Develop, validate, and interpret machine learning models to predict IVF cycle cancellation risk from pre-treatment clinical characteristics.
Random Forest selected as optimal model:
| Stage | Method | Clinical Question Answered | R Package |
|---|---|---|---|
| 1 | Logistic Regression | Which clinical factors independently predict cancellation? By what odds? | stats, broom, caret |
| 2 | Random Forest | Can non-linear interactions between AMH, AFC, FSH, and Age improve prediction? | randomForest, caret |
| 3 | Gradient Boosting (GBM) | Does sequential error-correction improve performance over independent trees? | gbm, caret |
| 4 | SHAP Interpretation | For each patient, why does the model predict high/low risk? What drives the prediction? | iml, shapviz |
| 5 | Calibration Analysis | Are predicted probabilities trustworthy? Does 70% predicted risk reflect ~70% actual cancellation? | CalibrationCurves, ResourceSelection |
| 6 | External Validation | Does the model generalise to a new, independent patient cohort? | pROC |
| 7 | Sensitivity Analysis | Are results robust to SMOTE ratio, and which features matter most? | caret, pROC |
set.seed(2026)| Dataset | n | Description |
|---|---|---|
| df (Training) | 500 | Primary dataset. 80/20 split into train_smote (post-SMOTE) and test_data. Source for model training. |
| df_ext (External) | 200 | External validation cohort. Independently generated with slightly different population parameters. Simulates a different clinic. |
| Variable | Type | Description & Clinical Relevance |
|---|---|---|
| age | Continuous | Age at treatment (years, 22–45). Older age correlates with declining ovarian reserve. |
| bmi | Continuous | Body Mass Index. Extremes affect hormonal response and follicular development. |
| amh | Continuous | Anti-Müllerian Hormone (ng/mL). Gold-standard ovarian reserve marker. Low AMH = high risk. |
| afc | Count | Antral Follicle Count. Ultrasound measure. Direct proxy of expected stimulation response. |
| fsh | Continuous | Follicle Stimulating Hormone (IU/L). Elevated FSH indicates declining ovarian function. |
| lh | Continuous | Luteinising Hormone (IU/L). Elevated LH may indicate PCOS or poor reserve. |
| estradiol | Continuous | Baseline serum oestradiol (pmol/L). Elevated baseline E2 may suppress FSH. |
| prev_cycles | Count | Number of prior IVF cycles. Prior cancellations indicate poor response pattern. |
| protocol | Categorical | Stimulation protocol: Long / Short / Antagonist. |
| cause | Categorical | Primary infertility diagnosis: Tubal / Male / Unexplained / Endometriosis / PCOS / Combined. |
| cancelled | Binary | Outcome: Yes = cycle stopped before retrieval. No = proceeded to retrieval. |
IVF cancellation is a minority outcome (15–25%). SMOTE (Synthetic Minority Oversampling Technique) is applied exclusively to the training set to address this. Critical Rule: SMOTE must never be applied to test or external validation sets to prevent data leakage and inflated metrics.
| Model | AUC | Accuracy | Sensitivity | Specificity | Interpretation |
|---|---|---|---|---|---|
| Logistic Regression | 0.710 | 0.657 | 0.679 | 0.533 | Baseline benchmark, fully interpretable |
| Random Forest ✓ | 0.729 | 0.747 | 0.821 | 0.333 | Best performer, highest sensitivity |
| GBM | 0.679 | 0.646 | 0.667 | 0.533 | Underperformed on small dataset |
Winner: Random Forest selected for deployment. Highest AUC (0.729) and sensitivity (0.821 = detects 82.1% of actual cancellations). Low specificity (0.333) generates false alarms but is clinically preferable to missing high-risk patients.
Random Forest (green) achieves highest AUC, indicating best overall discriminative ability.
Standard variable importance tells us "AMH is the most important variable overall". SHAP values provide patient-level explanations: "For Patient X, their AMH of 0.4 ng/mL increased their predicted risk by +0.22 probability points." This enables clinician-facing communication and builds trust in model predictions.
AMH dominates as the strongest predictor, followed by Age and AFC. FSH contributes minimally (information captured by AMH/AFC).
For a patient flagged as high-risk (predicted probability 78%), SHAP breakdown might show:"AMH=0.4 ng/mL contributed +0.22, Age=41 contributed +0.09, AFC=3 contributed +0.15."This enables personalized conversations: "Your low AMH is the primary driver of elevated risk."
A model with high AUC may still have systematically wrong probability estimates. Poor calibration is dangerous: predicting 70% risk when actual is 35% leads to over-treatment; predicting 30% when actual is 60% causes dangerous under-treatment. Calibration ensures predicted probabilities are trustworthy.
Logistic Regression (blue) shows best calibration. Random Forest slightly over-predicts. Points near the diagonal indicate good calibration.
Plots observed vs predicted in decile bins. Perfect = 45° diagonal.
Chi-squared test. p > 0.05 = acceptable calibration.
Mean squared error. 0 = perfect; 0.25 = uninformative.
External validation tests if a model trained on one population (df, n=500) generalizes to a different, independent population (df_ext, n=200). This simulates real-world deployment in a different clinic with slightly different patient characteristics. A model that degrades externally is overfit and unsuitable for deployment.
| Dataset | AUC | Accuracy | Sensitivity | Specificity |
|---|---|---|---|---|
| Internal Test | 0.729 | 0.747 | 0.821 | 0.333 |
| External Validation | 0.711 | 0.720 | 0.796 | 0.358 |
AUC drop of only 0.018 (0.729 → 0.711) is minimal and well within acceptable range. This confirms the model generalizes successfully to independent patient cohorts.
External AUC 95% CI: [0.642, 0.779] (n=1000 bootstrap resamples). Confirms robustness.
External validation on 200 independent patients confirms model generalizability.
Predicted probabilities are translated into three clinical risk tiers with management recommendations. Thresholds defined at 30% and 70% based on clinical consensus and Youden index optimization.
Probability: < 30%
Prevalence: 38.4% of cohort
Observed Rate: ~15%
Clinical Action:
Standard protocol. Routine monitoring at days 6 and 9.
Probability: 30% – 70%
Prevalence: 51.5% of cohort
Observed Rate: ~48%
Clinical Action:
Enhanced monitoring (days 5, 7, 9). Consider modified protocol. Patient counselling.
Probability: > 70%
Prevalence: 10.1% of cohort
Observed Rate: ~82%
Clinical Action:
Protocol modification strongly recommended. Comprehensive counselling. Discuss alternatives.
| Risk Tier | Mean Age | Mean AMH | Mean AFC | Mean FSH |
|---|---|---|---|---|
| Low | 31.2 | 3.8 | 14.2 | 6.1 |
| Medium | 34.8 | 1.9 | 8.7 | 8.4 |
| High | 39.1 | 0.6 | 4.1 | 12.8 |
High-risk patients are older with dramatically lower AMH/AFC and elevated FSH, confirming biological face-validity.
AMH (Anti-Müllerian Hormone) was the single strongest predictor of cancellation risk across all three models. Removing AMH from the model caused the largest performance drop of any variable. an AUC decrease of 0.095 (from 0.729 to 0.634). This confirms AMH is not an optional data point at first consultation. It is the foundation of individual risk assessment. Clinics not routinely collecting AMH before stimulation are making protocol decisions with their most important predictive variable missing.
Feature ablation tests model performance when individual variables are removed. Large AUC drops indicate irreplaceable variables. This analysis confirms AMH's critical importance and FSH's redundancy.
| Variable Configuration | AUC | AUC Drop | Interpretation |
|---|---|---|---|
| Full Model | 0.729 | — | Baseline model with all features |
| Without AMH | 0.634 | -0.095 | ⚠️ Largest drop - AMH is irreplaceable |
| Without AFC | 0.692 | -0.037 | Moderate drop - AFC contributes independently |
| Without Age | 0.701 | -0.028 | Moderate drop - Age adds meaningful signal |
| Without FSH | 0.723 | -0.006 | ✓ Minimal drop - FSH info captured by AMH/AFC |
| Without BMI | 0.716 | -0.013 | Small drop - BMI has modest contribution |
Key Finding: Removing AMH causes a 0.095 AUC drop (0.729 → 0.634), the largest decline of any feature. This confirms AMH's irreplaceability and biological centrality to ovarian reserve assessment.
| Package | Purpose | Key Functions |
|---|---|---|
| caret | Unified ML framework | train(), confusionMatrix(), trainControl() |
| randomForest | Random Forest model | randomForest(), importance() |
| gbm | Gradient Boosting | gbm(), summary.gbm() |
| pROC | ROC & AUC | roc(), auc(), ci.auc(), coords() |
| DMwR2 | SMOTE oversampling | SMOTE() |
| iml | SHAP & model explanation | Predictor$new(), FeatureImp$new() |
| CalibrationCurves | Calibration assessment | val.prob() |
The model detected 82.1% of patients who would go on to cancel. before their cycle even began. This shifts the clinical team from reactive (responding after cancellation) to proactive (adjusting protocols before stimulation).
SHAP values transform a black-box prediction into a transparent clinical tool. When a patient is flagged as high-risk, the model shows exactly why. AMH contributed this much, age contributed this much, AFC contributed this much. This enables evidence-based patient conversations.
Feature ablation analysis confirmed AMH is irreplaceable. removing it caused the largest performance drop (AUC 0.729 → 0.634). Clinics not routinely collecting AMH at first consultation are making protocol decisions without their most important predictive variable.
The model maintained AUC 0.711 on 200 patients not seen during training. demonstrating real-world generalizability. Models that perform well only on training data fail in clinical practice. External validation proves this model works on new patients.
Cycle-level outcome analysis, KPI dashboard development, cancellation rate tracking by stage, protocol benchmarking, DHA/DOH/ESHRE regulatory reporting.
Multi-center KPI dashboard, embryologist performance metrics, seasonal trend analysis, center-level comparison reporting, on-demand export.
Machine learning cancellation risk model, SHAP explainability, three-tier risk stratification framework, external validation on independent cohort.
Drop-out and cancellation root cause analysis by stage, protocol effectiveness evaluation by patient segment, embryo quality impact assessment.
This case study presents a complete, reproducible machine learning pipeline for IVF cycle cancellation risk prediction. Random Forest emerged as the optimal model with external AUC of 0.711 and sensitivity of 0.821, demonstrating excellent generalizability with minimal performance drop.
SHAP explainability addresses the "black box" concern, enabling clinician-facing communication. AMH dominates as the strongest predictor, while GBM underperformance on this dataset demonstrates context-dependent model selection. The accompanying R script is fully annotated with set.seed(2026)ensuring reproducibility.
If you have cycle-level data and a cancellation rate that is not broken down by stage, patient profile, or protocol. you have the same blind spot this clinic had.
We start with a free 30-minute consultation: we look at what data you have, identify where the gaps are, and show you what is possible with what already exists.