Contrastive Learning for Clinical Outcome Prediction with Partial Data Sources.

Meng Xia, Jonathan Wilson, Benjamin Goldstein, Ricardo Henao

Proceedings of machine learning research 2024 Jul

filter terms:

data sources (5)

research (1)

Sizes of these terms reflect their relevance to your search.

The use of machine learning models to predict clinical outcomes from (longitudinal) electronic health record (EHR) data is becoming increasingly popular due to advances in deep architectures, representation learning, and the growing availability of large EHR datasets. Existing models generally assume access to the same data sources during both training and inference stages. However, this assumption is often challenged by the fact that real-world clinical datasets originate from various data sources (with distinct sets of covariates), which though can be available for training (in a research or retrospective setting), are more realistically only partially available (a subset of such sets) for inference when deployed. So motivated, we introduce Contrastive Learning for clinical Outcome Prediction with Partial data Sources (CLOPPS), that trains encoders to capture information across different data sources and then leverages them to build classifiers restricting access to a single data source. This approach can be used with existing cross-sectional or longitudinal outcome classification models. We present experiments on two real-world datasets demonstrating that CLOPPS consistently outperforms strong baselines in several practical scenarios.

Citation

Meng Xia, Jonathan Wilson, Benjamin Goldstein, Ricardo Henao. Contrastive Learning for Clinical Outcome Prediction with Partial Data Sources. Proceedings of machine learning research. 2024 Jul;235:54156-54177

PMID: 39148511

View Full Text