
Abstract
One theme worth emphasizing is that “ECMO prediction” is not a single task. Some models are intended to inform decisions before cannulation, whereas others incorporate variables such as ECMO duration, 24-hour flow, sweep gas, oxygenation, or length of stay [1]. In prediction modeling, information leakage occurs when model development uses data that would not be available at the intended time of prediction, which can produce overly optimistic performance estimates. Thus, those later variables would constitute leakage in a pre-cannulation triage model, but may be entirely appropriate inputs for a dynamic monitoring model that is explicitly updated after ECMO initiation. Separating pre-cannulation selection, early complication surveillance, liberation planning, and post-discharge risk may make this literature easier to interpret and compare.
A second distinction is between prognosis and treatment benefit. Most models in the review estimate outcomes among patients who already received ECMO [1]. That information may still be useful for counseling, benchmarking, and surveillance, but it does not answer the counterfactual question that often matters most at the bedside: who is more likely to do better with ECMO than with continued conventional care? A treatment-benefit model, by contrast, would estimate for an individual patient the difference in outcome under ECMO versus continued conventional care, requiring a causal estimand and appropriate comparator data. Prediction research and causal inference address different clinical questions, and conflating them risks turning prognostic tools into de facto gatekeepers for a scarce therapy [5,6]. In a setting as resource-intensive and ethically sensitive as ECMO, that difference is not semantic.
The review also highlights how often machine learning rediscovers familiar physiology. Age, lactate, acid-base variables, hemodynamic instability, bilirubin, creatinine, and early ECMO flow variables recur across these models, with substantial overlap with traditional tools such as the Respiratory Extracorporeal Membrane Oxygenation Survival Prediction (RESP) score and the Survival After Veno-Arterial ECMO (SAVE) score [1–3]. That pattern also mirrors the broader ECMO prognostic literature, in which apparent performance often attenuates on external validation [4]. From one perspective, this is reassuring: the models are learning clinically recognizable signals. From another, it suggests that the incremental value of machine learning may lie less in identifying wholly novel predictors and more in modeling nonlinearity, interactions, and time-updated recalibration around variables clinicians already use. This may explain why hybrid approaches—anchoring machine learning to interpretable clinical variables and established scores—feel more plausible than wholesale replacement of conventional prognostic frameworks.
Another useful perspective concerns what the model may actually be learning. Predictors such as center volume, insurance status, ECMO duration, and post-cannulation flow variables do not reflect baseline biology alone; they also encode institutional expertise, care pathways, and social context [1]. That may limit transportability across settings, but it also reminds us that ECMO outcomes are co-produced by patient factors and the system delivering care. A model that performs well in one registry may therefore be capturing local practice as much as underlying disease severity. External validation matters for this reason, but so do calibration and threshold-specific performance. The area under the receiver operating characteristic curve (AUROC) can summarize discrimination, yet it says less about whether predicted probabilities are trustworthy or whether acting on them would help clinicians and families [7,8].
Finally, the field may now be ready to move beyond retrospective model comparison. A risk estimate becomes more clinically meaningful when it is linked to a response: intensified neuromonitoring, anticoagulation review, transplant evaluation or escalation, or a structured weaning discussion. Prospective impact studies evaluating whether these tools change clinical decisions or improve patient-centered outcomes remain largely absent from the ECMO artificial intelligence literature and should be a priority. Transparent reporting and early-stage clinical evaluation frameworks for artificial intelligence offer a useful path from promising models toward credible bedside tools [9,10]. The review by Keane et al. makes clear that the next advance in ECMO artificial intelligence is unlikely to come from a more complex algorithm alone. It is more likely to come from tighter alignment between the prediction target, the time of intended use, and the decision that follows.
We use cookies to provide you with the best possible user experience. By continuing to use our site, you agree to their use. Learn more