Details
Studies based on electronic health records (EHR) often rely on structured data, which may incompletely capture important clinical phenotypes in EHR notes. The purpose of this study was to assess two natural language processing (NLP) tools to extract phenotypes from unstructured EHR notes, and to evaluate the added value of integrating NLP-derived phenotypes with structured EHR data at a health system scale.
This retrospective study is based on inpatient and outpatient EHR data from the Mass General Brigham healthcare system between January 1, 2019 and December 31, 2020. Two established rule-based NLP tools were applied to extract smoking and obesity information from 19 215 303 clinical notes of 503 025 patients. NLP performance was evaluated through a manual review of stratified samples. Phenotype prevalence was estimated using structured EHR data alone and compared with prevalence estimates obtained by supplementing structured data with NLP-derived features.