Viyan

Viyan AI

Medical AI Benchmarks Often Rely on Data Leakage

High-performing medical AI models often ingest post-diagnostic data that acts as a lookup key for outcomes, masking poor predictive capability.

Cardiovascular screening models trained on survey data often achieve high performance by ingesting features that act as proxies for existing diagnoses. An audit of 2022 Behavioral Risk Factor Surveillance System data shows that reported AUROC scores near 0.89 are largely an artifact of target leakage. The two features responsible for this inflation were 'taking medicine for high blood pressure' and 'taking cholesterol medication.' Because these features are clinical consequences of a pre-existing heart condition rather than precursors to a future one, they allow a model to essentially look up a diagnosis that has already occurred. When these two post-diagnostic variables were removed, the performance of ten distinct models dropped by 0.049 to 0.051 AUROC and converged into a narrow, near-identical performance band. This collapse in performance suggests that what was once interpreted as predictive intelligence was merely the detection of secondary clinical markers.

The Failure of Complexity

The study compared ten classifiers across five tiers of leakage risk, evaluating everything from linear models to tabular foundation models. The "headroom" often cited in literature regarding foundation models for tabular data disappeared once researchers controlled for leakage. Removing these features reduced the gap between the strongest and weakest models to a range of only 0.0045 AUROC. The table below illustrates how removing two variables affects performance and inference overhead.

Model Class Reported AUROC AUROC After Removal Inference Speed vs. EBM
Tabular Foundation Models ~0.89 ~0.84 ~104x Slower
Glass-Box (EBM) ~0.89 ~0.84 1x (Baseline)

Architecture Versus Transparency

Model choice is secondary to data lineage. The glass-box Explainable Boosting Machine (EBM) performed on par with all other models, maintaining non-inferiority within a 0.005 margin while remaining significantly faster. EBMs achieve interpretability through shape functions, which map a specific risk score to every possible value of a single variable. Unlike a linear coefficient that assigns a fixed weight, a shape function allows the model to capture non-linear, idiosyncratic risk levels for individual data points. This makes it possible for a developer to visualize and manually adjust an unfair spike in risk.

Conversely, tabular foundation models utilize high-dimensional embeddings and attention layers that combine thousands of features into a single latent space. Because these layers calculate interactions across all inputs simultaneously, the final decision is mathematically inseparable; there is no way to isolate how much a single feature contributed to a specific prediction. In the original data, one threshold detected 75.4% of women's infarctions against 89.0% of men's. Researchers used the transparency of the EBM to manually edit these shape functions, narrowing the infarction detection gap between men and women to 0.010. This level of auditability is not feasible with the opaque architectures of tabular foundation models.

Data Hygiene as a Predictive Requirement

It remains unknown how much of current tabular foundation model performance in other clinical domains is similarly dependent on leakage rather than genuine pattern recognition. If the target labels in a dataset contain even faint signals of post-diagnostic information, these models will prioritize those shortcuts over causal features. This problem is exacerbated by the lack of temporal constraints in many public datasets, where future information is frequently included in training sets. For instance, if a dataset includes a 'medication start date' that occurs after the 'cardiovascular event' date, the model has an unfair advantage that does not exist in clinical practice.

To move forward, the field needs to adopt standardized, leakage-tiered auditing. Any reported performance gain from a new model architecture should be viewed as a potential artifact of the evaluation pipeline until the data is cleaned of temporal overlap. The primary challenge is not the computational capacity of the models, but the rigor of the feature engineering and the enforcement of temporal ordering. Future work must prioritize datasets where predictive features strictly precede the clinical event in time, ensuring that the model is learning to predict a future risk rather than classifying a past state.

Sources