Identification of patient name references within medical documents using semantic selectional restrictions.

1 January 2002

journal article

p. 757-61

Abstract

De-identification of a patient's personal data from medical records is a protective legal requirement imposed before medical documents can be used for research purposes or transferred to other healthcare providers (e.g., teachers, students, tele-consultations). This de-identification process is tedious if performed manually, and is known to be quite faulty in direct search and replace strategies [9]. In this paper, we report on the identification step of this process. The proposed algorithm is based on estimating the fitness of candidate patient name references to a set of semantic selectional restrictions. The semantic restrictions place tight contextual requirements upon candidate words in the report text and are determined automatically from a manually tagged corpus of training reports. Maximum entropy classifiers are used to provide a probabilistic measure of the belief of a given candidate token to a given semantic restriction. We report on the design and preliminary evaluation of the system within the do-main of pediatric urology.

This publication has 3 references indexed in Scilit:

DataServer
Academic Radiology, 2002
Automatic Record Hash Coding and Linkage for Epidemiological Follow-up Data Confidentiality
Methods of Information in Medicine, 1998
Basic principles of ROC analysis
Seminars in Nuclear Medicine, 1978