Published February 2018 | Version v1
Journal article

Using Natural Language Processing of Free-Text Radiology Reports to Identify Type 1 Modic Endplate Changes

  • 1. Radia, Inc. (United States)
  • 2. University of Washington, Department of Biostatistics (United States)
  • 3. University of Washington, Department of Rehabilitation Medicine (United States)
  • 4. Emory University School of Medicine, Department of Radiology and Imaging Sciences (United States)
  • 5. University of Washington, Comparative Effectiveness, Cost and Outcomes Research Center (United States)
  • 6. Division of Research, Kaiser Permanente Northern California (United States)
  • 7. Brigham and Women's Hospital and Spine Unit, Department of Anesthesiology, Perioperative and Pain Medicine, Harvard Vanguard Medical Associates (United States)
  • 8. Neuroscience Institute, Henry Ford Hospital (United States)
  • 9. Mayo Clinic, Department of Radiology (United States)
  • 10. Kaiser Permanente of Washington Research Institute (United States)
  • 11. Henry Ford Hospital, Department of Radiology (United States)
  • 12. Stanford University, Department of Radiology (United States)
  • 13. Dartmouth College, Department of Biomedical Data Science (United States)

Description

Electronic medical record (EMR) systems provide easy access to radiology reports and offer great potential to support quality improvement efforts and clinical research. Harnessing the full potential of the EMR requires scalable approaches such as natural language processing (NLP) to convert text into variables used for evaluation or analysis. Our goal was to determine the feasibility of using NLP to identify patients with Type 1 Modic endplate changes using clinical reports of magnetic resonance (MR) imaging examinations of the spine. Identifying patients with Type 1 Modic change who may be eligible for clinical trials is important as these findings may be important targets for intervention. Four annotators identified all reports that contained Type 1 Modic change, using N = 458 randomly selected lumbar spine MR reports. We then implemented a rule-based NLP algorithm in Java using regular expressions. The prevalence of Type 1 Modic change in the annotated dataset was 10%. Results were recall (sensitivity) 35/50 = 0.70 (95% confidence interval (C.I.) 0.52–0.82), specificity 404/408 = 0.99 (0.97–1.0), precision (positive predictive value) 35/39 = 0.90 (0.75–0.97), negative predictive value 404/419 = 0.96 (0.94–0.98), and F1-score 0.79 (0.43–1.0). Our evaluation shows the efficacy of rule-based NLP approach for identifying patients with Type 1 Modic change if the emphasis is on identifying only relevant cases with low concern regarding false negatives. As expected, our results show that specificity is higher than recall. This is due to the inherent difficulty of eliciting all possible keywords given the enormous variability of lumbar spine reporting, which decreases recall, while availability of good negation algorithms improves specificity.

Additional details

Identifiers

Publishing Information

Journal Title
Journal of Digital Imaging (Online)
Journal Volume
31
Journal Issue
1
Journal Page Range
p. 84-90
ISSN
1618-727X

INIS

Country of Publication
United States
Country of Input or Organization
International Atomic Energy Agency (IAEA)
INIS RN
50039920
Subject category
S62: RADIOLOGY AND NUCLEAR MEDICINE;
Descriptors DEI
ACCURACY; ALGORITHMS; AVAILABILITY; CLINICAL TRIALS; DATASETS; MEDICAL RECORDS; RADIOLOGY; SENSITIVITY; VERTEBRAE
Descriptors DEC
BODY; DOCUMENT TYPES; MATHEMATICAL LOGIC; MEDICINE; NUCLEAR MEDICINE; ORGANS; SKELETON; TESTING

Optional Information

Copyright
Copyright (c) 2018 Society for Imaging Informatics in Medicine