Ad Hoc Classification of Radiology Reports

Open Access

1 September 1999

journal article
Published by Oxford University Press (OUP) in Journal of the American Medical Informatics Association

Vol. 6 (5) , 393-411
https://doi.org/10.1136/jamia.1999.0060393

Abstract

Objective: The task of ad hoc classification is to automatically place a large number of text documents into nonstandard categories that are determined by a user. The authors examine the use of statistical information retrieval techniques for ad hoc classification of dictated mammography reports. Design: The authors' approach is the automated generation of a classification algorithm based on positive and negative evidence that is extracted from relevance-judged documents. Test documents are sorted into three conceptual bins: membership in a user-defined class, exclusion from the user-defined class, and uncertain. Documentation of absent findings through the use of negation and conjunction, a hallmark of interpretive test results, is managed by expansion and tokenization of these phrases. Measurements Classifier performance is evaluated using a single measure, the F measure, which provides a weighted combination of recall and precision of document sorting into true positive and true negative bins. Results: Single terms are the most effective text feature in the classification profile, with some improvement provided by the addition of pairs of unordered terms to the profile. Excessive iterations of automated classifier enhancement degrade performance because of overtraining. Performance is best when the proportions of relevant and irrelevant documents in the training collection are close to equal. Special handling of negation phrases improves performance when the number of terms in the classification profile is limited. Conclusions: The ad hoc classifier system is a promising approach for the classification of large collections of medical documents. NegExpander can distinguish positive evidence from negative evidence when the negative evidence plays an important role in the classification.

Keywords

This publication has 11 references indexed in Scilit:

A Reliability Study for Evaluating Information Extraction from Radiology Reports
Journal of the American Medical Informatics Association, 1999
An Experiment Comparing Lexical and Statistical Methods for Extracting MeSH Terms from Clinical Free Text
Journal of the American Medical Informatics Association, 1998
Automatic prediction of trauma registry procedure codes from emergency room dictations.
1998
Puya: a method of attracting attention to relevant physical findings.
1997
Identification of findings suspicious for breast cancer based on natural language processing of mammogram reports.
1997
Identification of suspected tuberculosis patients based on natural language processing of chest radiograph reports.
1996
Unlocking Clinical Data from Narrative Reports: A Study of Natural Language Processing
Annals of Internal Medicine, 1995
Automated identification of episodes of asthma exacerbation for quality measurement in a computer-based medical record.
1995
Inductive text classification for medical applications
Journal of Experimental & Theoretical Artificial Intelligence, 1995
Splines in Statistics
Journal of the American Statistical Association, 1983