Predictive Modeling & Clinical Decision Support

IVF Cycle Cancellation Risk Prediction Using Machine Learning

Logistic Regression · Random Forest · GBM · SHAP · Calibration · External Validation

500 training + 200 validation
3 ML Models
SHAP Explainability

Simulated Data for Demonstration

All datasets, model outputs, and visualizations use synthetically generated data (set.seed(2026)) for portfolio purposes. Results demonstrate advanced ML methodology.

This Is Part 2 of 2

This model was built on the same clinic group featured in Part 1. The dashboard identified where cancellations were happening. This model predicts who will cancel. before the cycle begins.

At a Glance

The Situation

The dashboard showed where cancellations were happening. but the clinic needed a way to identify which patients were heading toward cancellation before treatment began.

What We Built

A machine learning cancellation risk model built entirely on variables collected at first consultation. no additional testing required. Three models tested, Random Forest selected.

Model Performance

AUC 0.729 internal, 0.711 external validation. Sensitivity 82.1%. meaning the model detected 82% of at-risk cycles before treatment began. AMH emerged as the single strongest predictor.

The Business and Clinical Problem

IVF cycle cancellation occurs when ovarian stimulation results in a poor response (too few follicles) or excessive response (hyperstimulation risk), forcing the cycle to be stopped before egg retrieval. This is clinically and emotionally distressing for patients, financially costly, and represents a missed treatment opportunity.

Early identification of high-risk patients from routine pre-treatment clinical characteristics enables:

  • Personalized stimulation protocols: Adjust medication dosage and regimen based on predicted risk
  • Proactive patient counselling: Set realistic expectations and discuss alternative approaches
  • Resource optimization: Reduce avoidable cancellations and improve clinic efficiency
  • Clinical decision support: Evidence-based risk stratification at point of consultation

This case study develops, validates, and interprets a suite of three machine learning models(Logistic Regression, Random Forest, Gradient Boosting) with emphasis on SHAP explainability,calibration assessment, and external validation on 200 independent patients.

Executive Summary

Objective: Develop, validate, and interpret machine learning models to predict IVF cycle cancellation risk from pre-treatment clinical characteristics.

Methodology

  • • Synthetic dataset: 500 training + 200 external validation
  • • Three models: Logistic Regression, Random Forest, GBM
  • • 5-fold cross-validation with SMOTE for class imbalance
  • • SHAP for patient-level explanations
  • • Calibration curves, Hosmer-Lemeshow, Brier scores
  • • External validation on independent cohort

Key Outcome

Random Forest selected as optimal model:

  • • AUC: 0.729 (internal), 0.711 (external)
  • • Sensitivity: 0.821 (detects 82% of cancellations)
  • • Top Predictor: AMH (SHAP importance: 100)
  • • Minimal performance drop confirms generalizability

Table 1: Seven-Stage Analytical Pipeline

StageMethodClinical Question AnsweredR Package
1Logistic RegressionWhich clinical factors independently predict cancellation? By what odds?stats, broom, caret
2Random ForestCan non-linear interactions between AMH, AFC, FSH, and Age improve prediction?randomForest, caret
3Gradient Boosting (GBM)Does sequential error-correction improve performance over independent trees?gbm, caret
4SHAP InterpretationFor each patient, why does the model predict high/low risk? What drives the prediction?iml, shapviz
5Calibration AnalysisAre predicted probabilities trustworthy? Does 70% predicted risk reflect ~70% actual cancellation?CalibrationCurves, ResourceSelection
6External ValidationDoes the model generalise to a new, independent patient cohort?pROC
7Sensitivity AnalysisAre results robust to SMOTE ratio, and which features matter most?caret, pROC

Study Design & Data

Data Source

  • Synthetic dataset simulating real-world IVF clinic patient registry
  • Variables reflect standard first fertility consultation measurements
  • Two independent cohorts for external validation
  • Reproducible with set.seed(2026)

Table 2: Dataset Structure

DatasetnDescription
df (Training)500Primary dataset. 80/20 split into train_smote (post-SMOTE) and test_data. Source for model training.
df_ext (External)200External validation cohort. Independently generated with slightly different population parameters. Simulates a different clinic.

Table 3: Complete Variable Dictionary

VariableTypeDescription & Clinical Relevance
age
Continuous
Age at treatment (years, 22–45). Older age correlates with declining ovarian reserve.
bmi
Continuous
Body Mass Index. Extremes affect hormonal response and follicular development.
amh
Continuous
Anti-Müllerian Hormone (ng/mL). Gold-standard ovarian reserve marker. Low AMH = high risk.
afc
Count
Antral Follicle Count. Ultrasound measure. Direct proxy of expected stimulation response.
fsh
Continuous
Follicle Stimulating Hormone (IU/L). Elevated FSH indicates declining ovarian function.
lh
Continuous
Luteinising Hormone (IU/L). Elevated LH may indicate PCOS or poor reserve.
estradiol
Continuous
Baseline serum oestradiol (pmol/L). Elevated baseline E2 may suppress FSH.
prev_cycles
Count
Number of prior IVF cycles. Prior cancellations indicate poor response pattern.
protocol
Categorical
Stimulation protocol: Long / Short / Antagonist.
cause
Categorical
Primary infertility diagnosis: Tubal / Male / Unexplained / Endometriosis / PCOS / Combined.
cancelled
Binary
Outcome: Yes = cycle stopped before retrieval. No = proceeded to retrieval.

Class Imbalance Handling

IVF cancellation is a minority outcome (15–25%). SMOTE (Synthetic Minority Oversampling Technique) is applied exclusively to the training set to address this. Critical Rule: SMOTE must never be applied to test or external validation sets to prevent data leakage and inflated metrics.

Model Performance Comparison

ModelAUCAccuracySensitivitySpecificityInterpretation
Logistic Regression0.7100.6570.6790.533Baseline benchmark, fully interpretable
Random Forest ✓0.7290.7470.8210.333Best performer, highest sensitivity
GBM0.6790.6460.6670.533Underperformed on small dataset

Winner: Random Forest selected for deployment. Highest AUC (0.729) and sensitivity (0.821 = detects 82.1% of actual cancellations). Low specificity (0.333) generates false alarms but is clinically preferable to missing high-risk patients.

Figure 1: ROC Curves - Model Comparison

  • GBM (AUC=0.679)
  • Logistic Regression (AUC=0.710)
  • Random Forest (AUC=0.729)
00.20.40.60.81False Positive Rate (1 - Specificity)00.20.40.60.81True Positive Rate (Sensitivity)Random Classifier

Random Forest (green) achieves highest AUC, indicating best overall discriminative ability.

SHAP-Based Model Interpretation

Why SHAP?

Standard variable importance tells us "AMH is the most important variable overall". SHAP values provide patient-level explanations: "For Patient X, their AMH of 0.4 ng/mL increased their predicted risk by +0.22 probability points." This enables clinician-facing communication and builds trust in model predictions.

Figure 2: Global SHAP Feature Importance

0255075100SHAP Importance ScoreAMHAgeAFCBMIFSHprev_cycles

AMH dominates as the strongest predictor, followed by Age and AFC. FSH contributes minimally (information captured by AMH/AFC).

Clinical Application

For a patient flagged as high-risk (predicted probability 78%), SHAP breakdown might show:"AMH=0.4 ng/mL contributed +0.22, Age=41 contributed +0.09, AFC=3 contributed +0.15."This enables personalized conversations: "Your low AMH is the primary driver of elevated risk."

Calibration Analysis

Why Calibration Matters

A model with high AUC may still have systematically wrong probability estimates. Poor calibration is dangerous: predicting 70% risk when actual is 35% leads to over-treatment; predicting 30% when actual is 60% causes dangerous under-treatment. Calibration ensures predicted probabilities are trustworthy.

Figure 3: Calibration Curves

  • GBM
  • Logistic Regression
  • Random Forest
0.20.40.60.8Predicted Probability00.20.40.60.81Observed Cancellation Rate

Logistic Regression (blue) shows best calibration. Random Forest slightly over-predicts. Points near the diagonal indicate good calibration.

Calibration Curve

Plots observed vs predicted in decile bins. Perfect = 45° diagonal.

Hosmer-Lemeshow

Chi-squared test. p > 0.05 = acceptable calibration.

Brier Score

Mean squared error. 0 = perfect; 0.25 = uninformative.

External Validation

Purpose & Importance

External validation tests if a model trained on one population (df, n=500) generalizes to a different, independent population (df_ext, n=200). This simulates real-world deployment in a different clinic with slightly different patient characteristics. A model that degrades externally is overfit and unsuitable for deployment.

Table 11: Internal vs External Performance

DatasetAUCAccuracySensitivitySpecificity
Internal Test0.7290.7470.8210.333
External Validation0.7110.7200.7960.358

✓ Excellent Retention

AUC drop of only 0.018 (0.729 → 0.711) is minimal and well within acceptable range. This confirms the model generalizes successfully to independent patient cohorts.

Bootstrap 95% CI

External AUC 95% CI: [0.642, 0.779] (n=1000 bootstrap resamples). Confirms robustness.

Figure 4: External Validation ROC Curve

  • Random Forest - External (AUC=0.711)
00.20.40.60.81False Positive Rate00.20.40.60.81True Positive Rate

External validation on 200 independent patients confirms model generalizability.

Risk Stratification Framework

Predicted probabilities are translated into three clinical risk tiers with management recommendations. Thresholds defined at 30% and 70% based on clinical consensus and Youden index optimization.

Low Risk

Probability: < 30%

Prevalence: 38.4% of cohort

Observed Rate: ~15%

Clinical Action:

Standard protocol. Routine monitoring at days 6 and 9.

Medium Risk

Probability: 30% – 70%

Prevalence: 51.5% of cohort

Observed Rate: ~48%

Clinical Action:

Enhanced monitoring (days 5, 7, 9). Consider modified protocol. Patient counselling.

High Risk

Probability: > 70%

Prevalence: 10.1% of cohort

Observed Rate: ~82%

Clinical Action:

Protocol modification strongly recommended. Comprehensive counselling. Discuss alternatives.

Table 13: Patient Profiles by Risk Tier

Risk TierMean AgeMean AMHMean AFCMean FSH
Low31.23.814.26.1
Medium34.81.98.78.4
High39.10.64.112.8

High-risk patients are older with dramatically lower AMH/AFC and elevated FSH, confirming biological face-validity.

The Most Important Variable: AMH

AMH (Anti-Müllerian Hormone) was the single strongest predictor of cancellation risk across all three models. Removing AMH from the model caused the largest performance drop of any variable. an AUC decrease of 0.095 (from 0.729 to 0.634). This confirms AMH is not an optional data point at first consultation. It is the foundation of individual risk assessment. Clinics not routinely collecting AMH before stimulation are making protocol decisions with their most important predictive variable missing.

Sensitivity Analysis: Feature Ablation

Feature ablation tests model performance when individual variables are removed. Large AUC drops indicate irreplaceable variables. This analysis confirms AMH's critical importance and FSH's redundancy.

Variable ConfigurationAUCAUC DropInterpretation
Full Model0.729—Baseline model with all features
Without AMH0.634-0.095⚠️ Largest drop - AMH is irreplaceable
Without AFC0.692-0.037Moderate drop - AFC contributes independently
Without Age0.701-0.028Moderate drop - Age adds meaningful signal
Without FSH0.723-0.006✓ Minimal drop - FSH info captured by AMH/AFC
Without BMI0.716-0.013Small drop - BMI has modest contribution

Key Finding: Removing AMH causes a 0.095 AUC drop (0.729 → 0.634), the largest decline of any feature. This confirms AMH's irreplaceability and biological centrality to ovarian reserve assessment.

Limitations

Data Limitations

  • • Synthetic data: Simulates but doesn't reproduce real patient heterogeneity
  • • Missing variables: Endometrial thickness, sperm parameters, day-3 follicle sizes
  • • Sample size: n=500 training; GBM typically requires n>2,000

Methodological

  • • Protocol confounding: Protocol choice based on expected response
  • • SMOTE assumptions: May not be ideal for count variables (AFC)
  • • RF specificity: Low (0.333) generates false alarms

Validation

  • • External cohort: Same time period, need prospective validation
  • • Real-world deployment: Requires actual registry data validation
  • • Temporal stability: Need to assess concept drift over time

Clinical Acceptance

  • • Low specificity: Many false alarms may reduce clinician trust
  • • Calibration: RF may need Platt scaling before deployment
  • • Integration: Requires EHR integration for point-of-care use

R Package Reference

PackagePurposeKey Functions
caretUnified ML frameworktrain(), confusionMatrix(), trainControl()
randomForestRandom Forest modelrandomForest(), importance()
gbmGradient Boostinggbm(), summary.gbm()
pROCROC & AUCroc(), auc(), ci.auc(), coords()
DMwR2SMOTE oversamplingSMOTE()
imlSHAP & model explanationPredictor$new(), FeatureImp$new()
CalibrationCurvesCalibration assessmentval.prob()

What This Delivered

82% Detection Rate Before Treatment

The model detected 82.1% of patients who would go on to cancel. before their cycle even began. This shifts the clinical team from reactive (responding after cancellation) to proactive (adjusting protocols before stimulation).

Patient-Level Risk Explanations

SHAP values transform a black-box prediction into a transparent clinical tool. When a patient is flagged as high-risk, the model shows exactly why. AMH contributed this much, age contributed this much, AFC contributed this much. This enables evidence-based patient conversations.

AMH Confirmed as Critical Variable

Feature ablation analysis confirmed AMH is irreplaceable. removing it caused the largest performance drop (AUC 0.729 → 0.634). Clinics not routinely collecting AMH at first consultation are making protocol decisions without their most important predictive variable.

Externally Validated on Independent Cohort

The model maintained AUC 0.711 on 200 patients not seen during training. demonstrating real-world generalizability. Models that perform well only on training data fail in clinical practice. External validation proves this model works on new patients.

Services Used in This Engagement

🔬

Reproductive & IVF Analytics

Cycle-level outcome analysis, KPI dashboard development, cancellation rate tracking by stage, protocol benchmarking, DHA/DOH/ESHRE regulatory reporting.

📈

Performance Monitoring & Dashboards

Multi-center KPI dashboard, embryologist performance metrics, seasonal trend analysis, center-level comparison reporting, on-demand export.

🤖

Predictive Modelling & Advanced Analytics

Machine learning cancellation risk model, SHAP explainability, three-tier risk stratification framework, external validation on independent cohort.

📊

Outcome Optimization

Drop-out and cancellation root cause analysis by stage, protocol effectiveness evaluation by patient segment, embryo quality impact assessment.

Conclusion

This case study presents a complete, reproducible machine learning pipeline for IVF cycle cancellation risk prediction. Random Forest emerged as the optimal model with external AUC of 0.711 and sensitivity of 0.821, demonstrating excellent generalizability with minimal performance drop.

SHAP explainability addresses the "black box" concern, enabling clinician-facing communication. AMH dominates as the strongest predictor, while GBM underperformance on this dataset demonstrates context-dependent model selection. The accompanying R script is fully annotated with set.seed(2026)ensuring reproducibility.

Portfolio Project
R Statistical Analysis
April 2026
Version 2.0

Does Your Clinic Know Where Its Cycles Are Being Lost?

If you have cycle-level data and a cancellation rate that is not broken down by stage, patient profile, or protocol. you have the same blind spot this clinic had.

We start with a free 30-minute consultation: we look at what data you have, identify where the gaps are, and show you what is possible with what already exists.