Identification of patient name references within medical documents using semantic selectional restrictions.
- 1 January 2002
- journal article
- p. 757-61
Abstract
De-identification of a patient's personal data from medical records is a protective legal requirement imposed before medical documents can be used for research purposes or transferred to other healthcare providers (e.g., teachers, students, tele-consultations). This de-identification process is tedious if performed manually, and is known to be quite faulty in direct search and replace strategies [9]. In this paper, we report on the identification step of this process. The proposed algorithm is based on estimating the fitness of candidate patient name references to a set of semantic selectional restrictions. The semantic restrictions place tight contextual requirements upon candidate words in the report text and are determined automatically from a manually tagged corpus of training reports. Maximum entropy classifiers are used to provide a probabilistic measure of the belief of a given candidate token to a given semantic restriction. We report on the design and preliminary evaluation of the system within the do-main of pediatric urology.This publication has 3 references indexed in Scilit:
- DataServerAcademic Radiology, 2002
- Automatic Record Hash Coding and Linkage for Epidemiological Follow-up Data ConfidentialityMethods of Information in Medicine, 1998
- Basic principles of ROC analysisSeminars in Nuclear Medicine, 1978