Public knowledge document

UniProt: the universal protein knowledgebase.

The UniProt knowledgebase is a large resource of protein sequences and associated detailed annotation. The database contains over 60 million sequences, of which over half a million sequences have been curated by experts who critically review experimental and predicted data for each protein. The remainder are automatically annotated based on rule systems that rely on the expert curated knowledge. Since our last update in 2014, we have more than doubled the number of reference proteomes to 5631, giving a greater coverage of taxonomic diversity. We implemented a pipeline to remove redundant highly similar proteomes that were causing excessive redundancy in UniProt. The initial run of this pipeline reduced the number of sequences in UniProt by 47 million. For our users interested in the accessory proteomes, we have made available sets of pan proteome sequences that cover the diversity of sequences for each species that is found in its strains and sub-strains. To help interpretation of geno

UniProt: the universal protein knowledgebase.

> 商业许可源文 · EUROPE_PMC · [CC-BY](https://creativecommons.org/licenses/by/)

书目信息

  • 引用:The UniProt Consortium. (2017). UniProt: the universal protein knowledgebase. Nucleic acids research. PMID 27899622 · PMC5210571 · DOI 10.1093/nar/gkw1099
  • 证据类型:PRIMARY_RESEARCH
  • 主题:proteomics
  • 被引次数(采集时):3485
  • 原始记录:[Europe PMC](https://europepmc.org/article/MED/27899622)
  • 来源许可:[CC-BY](https://creativecommons.org/licenses/by/)
  • 作者摘要(按来源许可复用)

    The UniProt knowledgebase is a large resource of protein sequences and associated detailed annotation. The database contains over 60 million sequences, of which over half a million sequences have been curated by experts who critically review experimental and predicted data for each protein. The remainder are automatically annotated based on rule systems that rely on the expert curated knowledge. Since our last update in 2014, we have more than doubled the number of reference proteomes to 5631, giving a greater coverage of taxonomic diversity. We implemented a pipeline to remove redundant highly similar proteomes that were causing excessive redundancy in UniProt. The initial run of this pipeline reduced the number of sequences in UniProt by 47 million. For our users interested in the accessory proteomes, we have made available sets of pan proteome sequences that cover the diversity of sequences for each species that is found in its strains and sub-strains. To help interpretation of genomic variants, we provide tracks of detailed protein information for the major genome browsers. We provide a SPARQL endpoint that allows complex queries of the more than 22 billion triples of data in UniProt (http://sparql.uniprot.org/). UniProt resources can be accessed via the website at http://www.uniprot.org/.

    合规说明

    本页保存的是来源文献书目信息及其在 CC-BY 许可下公开的作者摘要。除去除来源 HTML 标签和规范化空白外,摘要未作内容改写。本页不代表 GeniOmics 的医学建议;原文版权、署名和许可仍归原权利人,请通过原始记录核对最新版本、更正或撤稿状态。

    UniProt: the universal protein knowledgebase. · GeniOmics