Empirical data on corpus design and usage in biomedical natural language processing.

1 January 2005

journal article

Vol. 2005, 156-60

Abstract

This paper describes the design of six publicly available biomedical corpora. We then present usage data for the six corpora. We show that corpora that are carefully annotated with respect to structural and linguistic characteristics and that are distributed in standard formats are more widely used than corpora that are not. These findings have implications for the design of the next generation of biomedical corpora.

This publication has 4 references indexed in Scilit:

GENETAG: a tagged corpus for gene/protein named entity recognition
BMC Bioinformatics, 2005
Protein names and how to find them
International Journal of Medical Informatics, 2002
Automatic extraction of biological information from scientific text: protein-protein interactions.
1999
Constructing biological knowledge bases by extracting information from text sources.
1999