The Protein Information Resource

1 January 2003

journal article
Published by Oxford University Press (OUP) in Nucleic Acids Research

Vol. 31 (1) , 345-347
https://doi.org/10.1093/nar/gkg040

Abstract

The Protein Information Resource (PIR) is an integrated public resource of protein informatics that supports genomic and proteomic research and scientific discovery. PIR maintains the Protein Sequence Database (PSD), an annotated protein database containing over 283 000 sequences covering the entire taxonomic range. Family classification is used for sensitive identification, consistent annotation, and detection of annotation errors. The superfamily curation defines signature domain architecture and categorizes memberships to improve automated classification. To increase the amount of experimental annotation, the PIR has developed a bibliography system for literature searching, mapping, and user submission, and has conducted retrospective attribution of citations for experimental features. PIR also maintains NREF, a non-redundant reference database, and iProClass, an integrated database of protein family, function, and structure information. PIR-NREF provides a timely and comprehensive collection of protein sequences, currently consisting of more than 1 000 000 entries from PIR-PSD, SWISS-PROT, TrEMBL, RefSeq, GenPept, and PDB. The PIR web site (http://pir.georgetown.edu) connects data analysis tools to underlying databases for information retrieval and knowledge discovery, with functionalities for interactive queries, combinations of sequence and text searches, and sorting and visual exploration of search results. The FTP site provides free download for PSD and NREF biweekly releases and auxiliary databases and files.

Keywords

This publication has 9 references indexed in Scilit:

The Protein Data Bank: unifying the archive
Nucleic Acids Research, 2002
iProClass: an integrated, comprehensive and annotated protein classification database
Nucleic Acids Research, 2001
RefSeq and LocusLink: NCBI gene-centered resources
Nucleic Acids Research, 2001
The SWISS-PROT protein sequence database and its supplement TrEMBL in 2000
Nucleic Acids Research, 2000
Gapped BLAST and PSI-BLAST: a new generation of protein database search programs
Nucleic Acids Research, 1997
Superfamily classification in PIR-international protein sequence database
Published by Elsevier ,1996
Maximum Discrimination Hidden Markov Models of Sequence Consensus
Journal of Computational Biology, 1995
CLUSTAL W: improving the sensitivity of progressive multiple sequence alignment through sequence weighting, position-specific gap penalties and weight matrix choice
Nucleic Acids Research, 1994
Improved tools for biological sequence comparison.
Proceedings of the National Academy of Sciences, 1988