Public knowledge document

Evaluation of statistical methods for normalization and differential expression in mRNA-Seq experiments.

Background High-throughput sequencing technologies, such as the Illumina Genome Analyzer, are powerful new tools for investigating a wide range of biological and medical questions. Statistical and computational methods are key for drawing meaningful and accurate conclusions from the massive and complex datasets generated by the sequencers. We provide a detailed evaluation of statistical methods for normalization and differential expression (DE) analysis of Illumina transcriptome sequencing (mRNA-Seq) data. Results We compare statistical methods for detecting genes that are significantly DE between two types of biological samples and find that there are substantial differences in how the test statistics handle low-count genes. We evaluate how DE results are affected by features of the sequencing platform, such as, varying gene lengths, base-calling calibration method (with and without phi X control lane), and flow-cell/library preparation effects. We investigate the impact of the read c

Evaluation of statistical methods for normalization and differential expression in mRNA-Seq experiments.

> 商业许可源文 · EUROPE_PMC · [CC-BY](https://creativecommons.org/licenses/by/)

书目信息

  • 引用:Bullard JH, Purdom E, Hansen KD, Dudoit S. (2010). Evaluation of statistical methods for normalization and differential expression in mRNA-Seq experiments. BMC bioinformatics. PMID 20167110 · PMC2838869 · DOI 10.1186/1471-2105-11-94
  • 证据类型:BENCHMARK
  • 主题:rna-seq
  • 被引次数(采集时):1240
  • 原始记录:[Europe PMC](https://europepmc.org/article/MED/20167110)
  • 来源许可:[CC-BY](https://creativecommons.org/licenses/by/)
  • 作者摘要(按来源许可复用)

    Background High-throughput sequencing technologies, such as the Illumina Genome Analyzer, are powerful new tools for investigating a wide range of biological and medical questions. Statistical and computational methods are key for drawing meaningful and accurate conclusions from the massive and complex datasets generated by the sequencers. We provide a detailed evaluation of statistical methods for normalization and differential expression (DE) analysis of Illumina transcriptome sequencing (mRNA-Seq) data. Results We compare statistical methods for detecting genes that are significantly DE between two types of biological samples and find that there are substantial differences in how the test statistics handle low-count genes. We evaluate how DE results are affected by features of the sequencing platform, such as, varying gene lengths, base-calling calibration method (with and without phi X control lane), and flow-cell/library preparation effects. We investigate the impact of the read count normalization method on DE results and show that the standard approach of scaling by total lane counts (e.g., RPKM) can bias estimates of DE. We propose more general quantile-based normalization procedures and demonstrate an improvement in DE detection. Conclusions Our results have significant practical and methodological implications for the design and analysis of mRNA-Seq experiments. They highlight the importance of appropriate statistical methods for normalization and DE inference, to account for features of the sequencing platform that could impact the accuracy of results. They also reveal the need for further research in the development of statistical and computational methods for mRNA-Seq.

    合规说明

    本页保存的是来源文献书目信息及其在 CC-BY 许可下公开的作者摘要。除去除来源 HTML 标签和规范化空白外,摘要未作内容改写。本页不代表 GeniOmics 的医学建议;原文版权、署名和许可仍归原权利人,请通过原始记录核对最新版本、更正或撤稿状态。

    Evaluation of statistical methods for normalization and differential expression in mRNA-Seq experiments. · GeniOmics