Subclass Mapping: Identifying Common Subtypes in Independent Disease Data Sets
Top Cited Papers
Open Access
- 21 November 2007
- journal article
- research article
- Published by Public Library of Science (PLoS) in PLOS ONE
- Vol. 2 (11) , e1195
- https://doi.org/10.1371/journal.pone.0001195
Abstract
Whole genome expression profiles are widely used to discover molecular subtypes of diseases. A remaining challenge is to identify the correspondence or commonality of subtypes found in multiple, independent data sets generated on various platforms. While model-based supervised learning is often used to make these connections, the models can be biased to the training data set and thus miss inherent, relevant substructure in the test data. Here we describe an unsupervised subclass mapping method (SubMap), which reveals common subtypes between independent data sets. The subtypes within a data set can be determined by unsupervised clustering or given by predetermined phenotypes before applying SubMap. We define a measure of correspondence for subtypes and evaluate its significance building on our previous work on gene set enrichment analysis. The strength of the SubMap method is that it does not impose the structure of one data set upon another, but rather uses a bi-directional approach to highlight the common substructures in both. We show how this method can reveal the correspondence between several cancer-related data sets. Notably, it identifies common subtypes of breast cancer associated with estrogen receptor status, and a subgroup of lymphoma patients who share similar survival patterns, thus improving the accuracy of a clinical outcome predictor.Keywords
This publication has 19 references indexed in Scilit:
- Are clusters found in one dataset present in another dataset?Biostatistics, 2006
- Gene set enrichment analysis: A knowledge-based approach for interpreting genome-wide expression profilesProceedings of the National Academy of Sciences, 2005
- Independence and reproducibility across microarray platformsNature Methods, 2005
- The Use of Molecular Profiling to Predict Survival after Chemotherapy for Diffuse Large-B-Cell LymphomaNew England Journal of Medicine, 2002
- Large-scale analysis of the human and mouse transcriptomesProceedings of the National Academy of Sciences, 2002
- Gene expression profiling predicts clinical outcome of breast cancerNature, 2002
- Diffuse large B-cell lymphoma outcome prediction by gene-expression profiling and supervised machine learningNature Medicine, 2002
- Multiclass cancer diagnosis using tumor gene expression signaturesProceedings of the National Academy of Sciences, 2001
- Predicting the clinical status of human breast cancer by using gene expression profilesProceedings of the National Academy of Sciences, 2001
- Significance analysis of microarrays applied to the ionizing radiation responseProceedings of the National Academy of Sciences, 2001