Semi-supervised Calibration of Risk with Noisy Event Times (SCORNET) Using Electronic Health Record Data

Abstract Leveraging large-scale electronic health record (EHR) data to estimate survival curves for clinical events can enable more powerful risk estimation and comparative effectiveness research. However, use of EHR data is hindered by a lack of direct event times observations. Occurrence times of relevant diagnostic codes or target disease mentions in clinical notes are at best a good approximation of the true disease onset time. On the other hand, extracting precise information on the exact event time requires laborious manual chart review and is sometimes altogether infeasible due to a lack of detailed documentation. Current status labels – binary indicators of phenotype status during follow up – are significantly more efficient and feasible to compile, enabling more precise survival curve estimation given limited resources. Existing survival analysis methods using current status labels focus almost entirely on supervised estimation, and naive incorporation of unlabeled data into these methods may lead to biased results. In this paper we propose Semi-supervised Calibration of Risk with Noisy Event Times (SCORNET), which yields a consistent and efficient survival curve estimator by leveraging a small size of current status labels and a large size of imperfect surrogate features. In addition to providing theoretical justification of SCORNET, we demonstrate in both simulation and real-world EHR settings that SCORNET achieves efficiency akin to the parametric Weibull regression model, while also exhibiting non-parametric flexibility and relatively low empirical bias in a variety of generative settings..

Medienart:

Preprint

Erscheinungsjahr:

2021

Erschienen:

2021

Enthalten in:

bioRxiv.org - (2021) vom: 19. Apr. Zur Gesamtaufnahme - year:2021

Sprache:

Englisch

Beteiligte Personen:

Ahuja, Yuri [VerfasserIn]
Liang, Liang [VerfasserIn]
Huang, Selena [VerfasserIn]
Cai, Tianxi [VerfasserIn]

Links:

Volltext [kostenfrei]

doi:

10.1101/2021.01.08.425976

funding:

Förderinstitution / Projekttitel:

PPN (Katalog-ID):

XBI019700741