Topic books

RNA-seq循证手册

按质量评分、证据类型和发表时间组织;优先收录明确允许商业复用的来源文献。

328 chapters

  1. 01Meta-analysis of tumor- and T cell-intrinsic mechanisms of sensitization to checkpoint inhibition.Checkpoint inhibitors (CPIs) augment adaptive immunity. Systematic pan-tumor analyses may reveal the relative importance of tumor-cell-intrinsic and microenvironmental features underpinning CPI sensitization. Here, we collated whole-exome and transcriptomic data for >1,000 CPI-treated patients across seven tumor types, utilizing standardized bioinformatics workflows and clinical outcome criteria to validate multivariable predictors of CPI sensitization. Clonal tumor mutation burden (TMB) was the strongest predictor of CPI response, followed by total TMB and CXCL9 expression. Subclonal TMB, somatic copy alteration burden, and histocompatibility leukocyte antigen (HLA) evolutionary divergence failed to attain pan-cancer significance. Dinucleotide variants were identified as a source of immunogenic epitopes associated with radical amino acid substitutions and enhanced peptide hydrophobicity/immunogenicity. Copy-number analysis revealed two additional determinants of CPI outcome supported
  2. 02Interpretation of T cell states from single-cell transcriptomics data using reference atlases.Single-cell RNA sequencing (scRNA-seq) has revealed an unprecedented degree of immune cell diversity. However, consistent definition of cell subtypes and cell states across studies and diseases remains a major challenge. Here we generate reference T cell atlases for cancer and viral infection by multi-study integration, and develop ProjecTILs, an algorithm for reference atlas projection. In contrast to other methods, ProjecTILs allows not only accurate embedding of new scRNA-seq data into a reference without altering its structure, but also characterizing previously unknown cell states that "deviate" from the reference. ProjecTILs accurately predicts the effects of cell perturbations and identifies gene programs that are altered in different conditions and tissues. A meta-analysis of tumor-infiltrating T cells from several cohorts reveals a strong conservation of T cell subtypes between human and mouse, providing a consistent basis to describe T cell heterogeneity across studies, disea
  3. 03Single-cell and bulk transcriptome sequencing identifies two epithelial tumor cell states and refines the consensus molecular classification of colorectal cancer.The consensus molecular subtype (CMS) classification of colorectal cancer is based on bulk transcriptomics. The underlying epithelial cell diversity remains unclear. We analyzed 373,058 single-cell transcriptomes from 63 patients, focusing on 49,155 epithelial cells. We identified a pervasive genetic and transcriptomic dichotomy of malignant cells, based on distinct gene expression, DNA copy number and gene regulatory network. We recapitulated these subtypes in bulk transcriptomes from 3,614 patients. The two intrinsic subtypes, iCMS2 and iCMS3, refine CMS. iCMS3 comprises microsatellite unstable (MSI-H) cancers and one-third of microsatellite-stable (MSS) tumors. iCMS3 MSS cancers are transcriptomically more similar to MSI-H cancers than to other MSS cancers. CMS4 cancers had either iCMS2 or iCMS3 epithelium; the latter had the worst prognosis. We defined the intrinsic epithelial axis of colorectal cancer and propose a refined 'IMF' classification with five subtypes, combining intrins
  4. 04A comprehensive benchmarking with practical guidelines for cellular deconvolution of spatial transcriptomics.Spatial transcriptomics technologies are used to profile transcriptomes while preserving spatial information, which enables high-resolution characterization of transcriptional patterns and reconstruction of tissue architecture. Due to the existence of low-resolution spots in recent spatial transcriptomics technologies, uncovering cellular heterogeneity is crucial for disentangling the spatial patterns of cell types, and many related methods have been proposed. Here, we benchmark 18 existing methods resolving a cellular deconvolution task with 50 real-world and simulated datasets by evaluating the accuracy, robustness, and usability of the methods. We compare these methods comprehensively using different metrics, resolutions, spatial transcriptomics technologies, spot numbers, and gene numbers. In terms of performance, CARD, Cell2location, and Tangram are the best methods for conducting the cellular deconvolution task. To refine our comparative results, we provide decision-tree-style gu
  5. 05Profiling the heterogeneity of colorectal cancer consensus molecular subtypes using spatial transcriptomics.The consensus molecular subtypes (CMS) of colorectal cancer (CRC) is the most widely-used gene expression-based classification and has contributed to a better understanding of disease heterogeneity and prognosis. Nevertheless, CMS intratumoral heterogeneity restricts its clinical application, stressing the necessity of further characterizing the composition and architecture of CRC. Here, we used Spatial Transcriptomics (ST) in combination with single-cell RNA sequencing (scRNA-seq) to decipher the spatially resolved cellular and molecular composition of CRC. In addition to mapping the intratumoral heterogeneity of CMS and their microenvironment, we identified cell communication events in the tumor-stroma interface of CMS2 carcinomas. This includes tumor growth-inhibiting as well as -activating signals, such as the potential regulation of the ETV4 transcriptional activity by DCN or the PLAU-PLAUR ligand-receptor interaction. Our study illustrates the potential of ST to resolve CRC molec
  6. 06Intestinal dysbiosis in preterm infants preceding necrotizing enterocolitis: a systematic review and meta-analysis.Background Necrotizing enterocolitis (NEC) is a catastrophic disease of preterm infants, and microbial dysbiosis has been implicated in its pathogenesis. Studies evaluating the microbiome in NEC and preterm infants lack power and have reported inconsistent results. Methods and results Our objectives were to perform a systematic review and meta-analyses of stool microbiome profiles in preterm infants to discern and describe microbial dysbiosis prior to the onset of NEC and to explore heterogeneity among studies. We searched MEDLINE, PubMed, CINAHL, and conference abstracts from the proceedings of Pediatric Academic Societies and reference lists of relevant identified articles in April 2016. Studies comparing the intestinal microbiome in preterm infants who developed NEC to those of controls, using culture-independent molecular techniques and reported α and β-diversity metrics, and microbial profiles were included. In addition, 16S ribosomal ribonucleic acid (rRNA) sequence data with cli
  7. 07Identifying gene targets for brain-related traits using transcriptomic and methylomic data from blood.Understanding the difference in genetic regulation of gene expression between brain and blood is important for discovering genes for brain-related traits and disorders. Here, we estimate the correlation of genetic effects at the top-associated cis-expression or -DNA methylation (DNAm) quantitative trait loci (cis-eQTLs or cis-mQTLs) between brain and blood (r b ). Using publicly available data, we find that genetic effects at the top cis-eQTLs or mQTLs are highly correlated between independent brain and blood samples ([Formula: see text] for cis-eQTLs and [Formula: see text] for cis-mQTLs). Using meta-analyzed brain cis-eQTL/mQTL data (n = 526 to 1194), we identify 61 genes and 167 DNAm sites associated with four brain-related phenotypes, most of which are a subset of the discoveries (97 genes and 295 DNAm sites) using data from blood with larger sample sizes (n = 1980 to 14,115). Our results demonstrate the gain of power in gene discovery for brain-related phenotypes using blood cis-e
  8. 08Guidelines for bioinformatics of single-cell sequencing data analysis in Alzheimer's disease: review, recommendation, implementation and application.Alzheimer's disease (AD) is the most common form of dementia, characterized by progressive cognitive impairment and neurodegeneration. Extensive clinical and genomic studies have revealed biomarkers, risk factors, pathways, and targets of AD in the past decade. However, the exact molecular basis of AD development and progression remains elusive. The emerging single-cell sequencing technology can potentially provide cell-level insights into the disease. Here we systematically review the state-of-the-art bioinformatics approaches to analyze single-cell sequencing data and their applications to AD in 14 major directions, including 1) quality control and normalization, 2) dimension reduction and feature extraction, 3) cell clustering analysis, 4) cell type inference and annotation, 5) differential expression, 6) trajectory inference, 7) copy number variation analysis, 8) integration of single-cell multi-omics, 9) epigenomic analysis, 10) gene network inference, 11) prioritization of cell s
  9. 09Systematic comparison of sequencing-based spatial transcriptomic methods.Recent developments of sequencing-based spatial transcriptomics (sST) have catalyzed important advancements by facilitating transcriptome-scale spatial gene expression measurement. Despite this progress, efforts to comprehensively benchmark different platforms are currently lacking. The extant variability across technologies and datasets poses challenges in formulating standardized evaluation metrics. In this study, we established a collection of reference tissues and regions characterized by well-defined histological architectures, and used them to generate data to compare 11 sST methods. We highlighted molecular diffusion as a variable parameter across different methods and tissues, significantly affecting the effective resolutions. Furthermore, we observed that spatial transcriptomic data demonstrate unique attributes beyond merely adding a spatial axis to single-cell data, including an enhanced ability to capture patterned rare cell states along with specific markers, albeit being
  10. 10STOmicsDB: a comprehensive database for spatial transcriptomics data sharing, analysis and visualization.Recent technological developments in spatial transcriptomics allow researchers to measure gene expression of cells and their spatial locations at the single-cell level, generating detailed biological insight into biological processes. A comprehensive database could facilitate the sharing of spatial transcriptomic data and streamline the data acquisition process for researchers. Here, we present the Spatial TranscriptOmics DataBase (STOmicsDB), a database that serves as a one-stop hub for spatial transcriptomics. STOmicsDB integrates 218 manually curated datasets representing 17 species. We annotated cell types, identified spatial regions and genes, and performed cell-cell interaction analysis for these datasets. STOmicsDB features a user-friendly interface for the rapid visualization of millions of cells. To further facilitate the reusability and interoperability of spatial transcriptomic data, we developed standards for spatial transcriptomic data archiving and constructed a spatial t
  11. 11A human embryonic limb cell atlas resolved in space and time.Human limbs emerge during the fourth post-conception week as mesenchymal buds, which develop into fully formed limbs over the subsequent months 1 . This process is orchestrated by numerous temporally and spatially restricted gene expression programmes, making congenital alterations in phenotype common 2 . Decades of work with model organisms have defined the fundamental mechanisms underlying vertebrate limb development, but an in-depth characterization of this process in humans has yet to be performed. Here we detail human embryonic limb development across space and time using single-cell and spatial transcriptomics. We demonstrate extensive diversification of cells from a few multipotent progenitors to myriad differentiated cell states, including several novel cell populations. We uncover two waves of human muscle development, each characterized by different cell states regulated by separate gene expression programmes, and identify musculin (MSC) as a key transcriptional repressor mai
  12. 12Benchmarking clustering, alignment, and integration methods for spatial transcriptomics.Background Spatial transcriptomics (ST) is advancing our understanding of complex tissues and organisms. However, building a robust clustering algorithm to define spatially coherent regions in a single tissue slice and aligning or integrating multiple tissue slices originating from diverse sources for essential downstream analyses remains challenging. Numerous clustering, alignment, and integration methods have been specifically designed for ST data by leveraging its spatial information. The absence of comprehensive benchmark studies complicates the selection of methods and future method development. Results In this study, we systematically benchmark a variety of state-of-the-art algorithms with a wide range of real and simulated datasets of varying sizes, technologies, species, and complexity. We analyze the strengths and weaknesses of each method using diverse quantitative and qualitative metrics and analyses, including eight metrics for spatial clustering accuracy and contiguity, un
  13. 13A high-resolution transcriptomic and spatial atlas of cell types in the whole mouse brain.The mammalian brain consists of millions to billions of cells that are organized into many cell types with specific spatial distribution patterns and structural and functional properties 1-3 . Here we report a comprehensive and high-resolution transcriptomic and spatial cell-type atlas for the whole adult mouse brain. The cell-type atlas was created by combining a single-cell RNA-sequencing (scRNA-seq) dataset of around 7 million cells profiled (approximately 4.0 million cells passing quality control), and a spatial transcriptomic dataset of approximately 4.3 million cells using multiplexed error-robust fluorescence in situ hybridization (MERFISH). The atlas is hierarchically organized into 4 nested levels of classification: 34 classes, 338 subclasses, 1,201 supertypes and 5,322 clusters. We present an online platform, Allen Brain Cell Atlas, to visualize the mouse whole-brain cell-type atlas along with the single-cell RNA-sequencing and MERFISH datasets. We systematically analysed the
  14. 14CellMarker 2.0: an updated database of manually curated cell markers in human/mouse and web tools based on scRNA-seq data.CellMarker 2.0 (http://bio-bigdata.hrbmu.edu.cn/CellMarker or http://117.50.127.228/CellMarker/) is an updated database that provides a manually curated collection of experimentally supported markers of various cell types in different tissues of human and mouse. In addition, web tools for analyzing single cell sequencing data are described. We have updated CellMarker 2.0 with more data and several new features, including (i) Appending 36 300 tissue-cell type-maker entries, 474 tissues, 1901 cell types and 4566 markers over the previous version. The current release recruits 26 915 cell markers, 2578 cell types and 656 tissues, resulting in a total of 83 361 tissue-cell type-maker entries. (ii) There is new marker information from 48 sequencing technology sources, including 10X Chromium, Smart-Seq2 and Drop-seq, etc. (iii) Adding 29 types of cell markers, including protein-coding gene lncRNA and processed pseudogene, etc. Additionally, six flexible web tools, including cell annotation, c
  15. 15Comparison of methods and resources for cell-cell communication inference from single-cell RNA-Seq data.The growing availability of single-cell data, especially transcriptomics, has sparked an increased interest in the inference of cell-cell communication. Many computational tools were developed for this purpose. Each of them consists of a resource of intercellular interactions prior knowledge and a method to predict potential cell-cell communication events. Yet the impact of the choice of resource and method on the resulting predictions is largely unknown. To shed light on this, we systematically compare 16 cell-cell communication inference resources and 7 methods, plus the consensus between the methods' predictions. Among the resources, we find few unique interactions, a varying degree of overlap, and an uneven coverage of specific pathways and tissue-enriched proteins. We then examine all possible combinations of methods and resources and show that both strongly influence the predicted intercellular interactions. Finally, we assess the agreement of cell-cell communication methods with
  16. 16A multimodal cell census and atlas of the mammalian primary motor cortex.Here we report the generation of a multimodal cell census and atlas of the mammalian primary motor cortex as the initial product of the BRAIN Initiative Cell Census Network (BICCN). This was achieved by coordinated large-scale analyses of single-cell transcriptomes, chromatin accessibility, DNA methylomes, spatially resolved single-cell transcriptomes, morphological and electrophysiological properties and cellular resolution input-output mapping, integrated through cross-modal computational analysis. Our results advance the collective knowledge and understanding of brain cell-type organization 1-5 . First, our study reveals a unified molecular genetic landscape of cortical cell types that integrates their transcriptome, open chromatin and DNA methylation maps. Second, cross-species analysis achieves a consensus taxonomy of transcriptomic types and their hierarchical organization that is conserved from mouse to marmoset and human. Third, in situ single-cell transcriptomics provides a sp
  17. 17Molecularly defined and spatially resolved cell atlas of the whole mouse brain.In mammalian brains, millions to billions of cells form complex interaction networks to enable a wide range of functions. The enormous diversity and intricate organization of cells have impeded our understanding of the molecular and cellular basis of brain function. Recent advances in spatially resolved single-cell transcriptomics have enabled systematic mapping of the spatial organization of molecularly defined cell types in complex tissues 1-3 , including several brain regions (for example, refs. 1-11 ). However, a comprehensive cell atlas of the whole brain is still missing. Here we imaged a panel of more than 1,100 genes in approximately 10 million cells across the entire adult mouse brains using multiplexed error-robust fluorescence in situ hybridization 12 and performed spatially resolved, single-cell expression profiling at the whole-transcriptome scale by integrating multiplexed error-robust fluorescence in situ hybridization and single-cell RNA sequencing data. Using this appr
  18. 18A transcriptomic and epigenomic cell atlas of the mouse primary motor cortex.Single-cell transcriptomics can provide quantitative molecular signatures for large, unbiased samples of the diverse cell types in the brain 1-3 . With the proliferation of multi-omics datasets, a major challenge is to validate and integrate results into a biological understanding of cell-type organization. Here we generated transcriptomes and epigenomes from more than 500,000 individual cells in the mouse primary motor cortex, a structure that has an evolutionarily conserved role in locomotion. We developed computational and statistical methods to integrate multimodal data and quantitatively validate cell-type reproducibility. The resulting reference atlas-containing over 56 neuronal cell types that are highly replicable across analysis methods, sequencing technologies and modalities-is a comprehensive molecular and genomic account of the diverse neuronal and non-neuronal cell types in the mouse primary motor cortex. The atlas includes a population of excitatory neurons that resemble
  19. 19A single cell atlas of human cornea that defines its development, limbal progenitor cells and their interactions with the immune cells.Purpose Single cell (sc) analyses of key embryonic, fetal and adult stages were performed to generate a comprehensive single cell atlas of all the corneal and adjacent conjunctival cell types from development to adulthood. Methods Four human adult and seventeen embryonic and fetal corneas from 10 to 21 post conception week (PCW) specimens were dissociated to single cells and subjected to scRNA- and/or ATAC-Seq using the 10x Genomics platform. These were embedded using Uniform Manifold Approximation and Projection (UMAP) and clustered using Seurat graph-based clustering. Cluster identification was performed based on marker gene expression, bioinformatic data mining and immunofluorescence (IF) analysis. RNA interference, IF, colony forming efficiency and clonal assays were performed on cultured limbal epithelial cells (LECs). Results scRNA-Seq analysis of 21,343 cells from four adult human corneas and adjacent conjunctivas revealed the presence of 21 cell clusters, representing the proge
  20. 20Single-cell sequencing to multi-omics: technologies and applications.Cells, as the fundamental units of life, contain multidimensional spatiotemporal information. Single-cell RNA sequencing (scRNA-seq) is revolutionizing biomedical science by analyzing cellular state and intercellular heterogeneity. Undoubtedly, single-cell transcriptomics has emerged as one of the most vibrant research fields today. With the optimization and innovation of single-cell sequencing technologies, the intricate multidimensional details concealed within cells are gradually unveiled. The combination of scRNA-seq and other multi-omics is at the forefront of the single-cell field. This involves simultaneously measuring various omics data within individual cells, expanding our understanding across a broader spectrum of dimensions. Single-cell multi-omics precisely captures the multidimensional aspects of single-cell transcriptomes, immune repertoire, spatial information, temporal information, epitopes, and other omics in diverse spatiotemporal contexts. In addition to depicting t
  21. 21Estimation of cell lineages in tumors from spatial transcriptomics data.Spatial transcriptomics (ST) technology through in situ capturing has enabled topographical gene expression profiling of tumor tissues. However, each capturing spot may contain diverse immune and malignant cells, with different cell densities across tissue regions. Cell type deconvolution in tumor ST data remains challenging for existing methods designed to decompose general ST or bulk tumor data. We develop the Spatial Cellular Estimator for Tumors (SpaCET) to infer cell identities from tumor ST data. SpaCET first estimates cancer cell abundance by integrating a gene pattern dictionary of copy number alterations and expression changes in common malignancies. A constrained regression model then calibrates local cell densities and determines immune and stromal cell lineage fractions. SpaCET provides higher accuracy than existing methods based on simulation and real ST data with matched double-blind histopathology annotations as ground truth. Further, coupling cell fractions with ligand-
  22. 22ArchR is a scalable software package for integrative single-cell chromatin accessibility analysis.The advent of single-cell chromatin accessibility profiling has accelerated the ability to map gene regulatory landscapes but has outpaced the development of scalable software to rapidly extract biological meaning from these data. Here we present a software suite for single-cell analysis of regulatory chromatin in R (ArchR; https://www.archrproject.com/ ) that enables fast and comprehensive analysis of single-cell chromatin accessibility data. ArchR provides an intuitive, user-focused interface for complex single-cell analyses, including doublet removal, single-cell clustering and cell type identification, unified peak set generation, cellular trajectory identification, DNA element-to-gene linkage, transcription factor footprinting, mRNA expression level prediction from chromatin accessibility and multi-omic integration with single-cell RNA sequencing (scRNA-seq). Enabling the analysis of over 1.2 million single cells within 8 h on a standard Unix laptop, ArchR is a comprehensive softw
  23. 23Single-cell RNA sequencing technologies and applications: A brief overview.Single-cell RNA sequencing (scRNA-seq) technology has become the state-of-the-art approach for unravelling the heterogeneity and complexity of RNA transcripts within individual cells, as well as revealing the composition of different cell types and functions within highly organized tissues/organs/organisms. Since its first discovery in 2009, studies based on scRNA-seq provide massive information across different fields making exciting new discoveries in better understanding the composition and interaction of cells within humans, model animals and plants. In this review, we provide a concise overview about the scRNA-seq technology, experimental and computational procedures for transforming the biological and molecular processes into computational and statistical data. We also provide an explanation of the key technological steps in implementing the technology. We highlight a few examples on how scRNA-seq can provide unique information for better understanding health and diseases. One im
  24. 24Molecular subgroups of medulloblastoma: an international meta-analysis of transcriptome, genetic aberrations, and clinical data of WNT, SHH, Group 3, and Group 4 medulloblastomas.Medulloblastoma is the most common malignant brain tumor in childhood. Molecular studies from several groups around the world demonstrated that medulloblastoma is not one disease but comprises a collection of distinct molecular subgroups. However, all these studies reported on different numbers of subgroups. The current consensus is that there are only four core subgroups, which should be termed WNT, SHH, Group 3 and Group 4. Based on this, we performed a meta-analysis of all molecular and clinical data of 550 medulloblastomas brought together from seven independent studies. All cases were analyzed by gene expression profiling and for most cases SNP or array-CGH data were available. Data are presented for all medulloblastomas together and for each subgroup separately. For validation purposes, we compared the results of this meta-analysis with another large medulloblastoma cohort (n = 402) for which subgroup information was obtained by immunohistochemistry. Results from both cohorts are
  25. 25The transcriptional landscape of age in human peripheral blood.Disease incidences increase with age, but the molecular characteristics of ageing that lead to increased disease susceptibility remain inadequately understood. Here we perform a whole-blood gene expression meta-analysis in 14,983 individuals of European ancestry (including replication) and identify 1,497 genes that are differentially expressed with chronological age. The age-associated genes do not harbor more age-associated CpG-methylation sites than other genes, but are instead enriched for the presence of potentially functional CpG-methylation sites in enhancer and insulator regions that associate with both chronological age and gene expression levels. We further used the gene expression profiles to calculate the 'transcriptomic age' of an individual, and show that differences between transcriptomic age and chronological age are associated with biological features linked to ageing, such as blood pressure, cholesterol levels, fasting glucose, and body mass index. The transcriptomic p
  26. 26An introduction to spatial transcriptomics for biomedical research.Single-cell transcriptomics (scRNA-seq) has become essential for biomedical research over the past decade, particularly in developmental biology, cancer, immunology, and neuroscience. Most commercially available scRNA-seq protocols require cells to be recovered intact and viable from tissue. This has precluded many cell types from study and largely destroys the spatial context that could otherwise inform analyses of cell identity and function. An increasing number of commercially available platforms now facilitate spatially resolved, high-dimensional assessment of gene transcription, known as 'spatial transcriptomics'. Here, we introduce different classes of method, which either record the locations of hybridized mRNA molecules in tissue, image the positions of cells themselves prior to assessment, or employ spatial arrays of mRNA probes of pre-determined location. We review sizes of tissue area that can be assessed, their spatial resolution, and the number and types of genes that can
  27. 27Statistics or biology: the zero-inflation controversy about scRNA-seq data.Researchers view vast zeros in single-cell RNA-seq data differently: some regard zeros as biological signals representing no or low gene expression, while others regard zeros as missing data to be corrected. To help address the controversy, here we discuss the sources of biological and non-biological zeros; introduce five mechanisms of adding non-biological zeros in computational benchmarking; evaluate the impacts of non-biological zeros on data analysis; benchmark three input data types: observed counts, imputed counts, and binarized counts; discuss the open questions regarding non-biological zeros; and advocate the importance of transparent analysis.
  28. 28Standardizing workflows in imaging transcriptomics with the abagen toolbox.Gene expression fundamentally shapes the structural and functional architecture of the human brain. Open-access transcriptomic datasets like the Allen Human Brain Atlas provide an unprecedented ability to examine these mechanisms in vivo; however, a lack of standardization across research groups has given rise to myriad processing pipelines for using these data. Here, we develop the abagen toolbox, an open-access software package for working with transcriptomic data, and use it to examine how methodological variability influences the outcomes of research using the Allen Human Brain Atlas. Applying three prototypical analyses to the outputs of 750,000 unique processing pipelines, we find that choice of pipeline has a large impact on research findings, with parameters commonly varied in the literature influencing correlations between derived gene expression and other imaging phenotypes by as much as ρ ≥ 1.0. Our results further reveal an ordering of parameter importance, with processing
  29. 29From bulk, single-cell to spatial RNA sequencing.RNA sequencing (RNAseq) can reveal gene fusions, splicing variants, mutations/indels in addition to differential gene expression, thus providing a more complete genetic picture than DNA sequencing. This most widely used technology in genomics tool box has evolved from classic bulk RNA sequencing (RNAseq), popular single cell RNA sequencing (scRNAseq) to newly emerged spatial RNA sequencing (spRNAseq). Bulk RNAseq studies average global gene expression, scRNAseq investigates single cell RNA biology up to 20,000 individual cells simultaneously, while spRNAseq has ability to dissect RNA activities spatially, representing next generation of RNA sequencing. This article highlights these technologies, characteristic features and suitable applications in precision oncology.
  30. 30Current challenges and best-practice protocols for microbiome analysis.Analyzing the microbiome of diverse species and environments using next-generation sequencing techniques has significantly enhanced our understanding on metabolic, physiological and ecological roles of environmental microorganisms. However, the analysis of the microbiome is affected by experimental conditions (e.g. sequencing errors and genomic repeats) and computationally intensive and cumbersome downstream analysis (e.g. quality control, assembly, binning and statistical analyses). Moreover, the introduction of new sequencing technologies and protocols led to a flood of new methodologies, which also have an immediate effect on the results of the analyses. The aim of this work is to review the most important workflows for 16S rRNA sequencing and shotgun and long-read metagenomics, as well as to provide best-practice protocols on experimental design, sample processing, sequencing, assembly, binning, annotation and visualization. To simplify and standardize the computational analysis, w
  31. 31Nicotinamide Riboside Augments the Aged Human Skeletal Muscle NAD<sup>+</sup> Metabolome and Induces Transcriptomic and Anti-inflammatory Signatures.Nicotinamide adenine dinucleotide (NAD + ) is modulated by conditions of metabolic stress and has been reported to decline with aging in preclinical models, but human data are sparse. Nicotinamide riboside (NR) supplementation ameliorates metabolic dysfunction in rodents. We aimed to establish whether oral NR supplementation in aged participants can increase the skeletal muscle NAD + metabolome and if it can alter muscle mitochondrial bioenergetics. We supplemented 12 aged men with 1 g NR per day for 21 days in a placebo-controlled, randomized, double-blind, crossover trial. Targeted metabolomics showed that NR elevated the muscle NAD + metabolome, evident by increased nicotinic acid adenine dinucleotide and nicotinamide clearance products. Muscle RNA sequencing revealed NR-mediated downregulation of energy metabolism and mitochondria pathways, without altering mitochondrial bioenergetics. NR also depressed levels of circulating inflammatory cytokines. Our data establish that oral NR i
  32. 32Single-cell RNA sequencing in cancer research.Single-cell RNA sequencing (scRNA-seq), a technology that analyzes transcriptomes of complex tissues at single-cell levels, can identify differential gene expression and epigenetic factors caused by mutations in unicellular genomes, as well as new cell-specific markers and cell types. scRNA-seq plays an important role in various aspects of tumor research. It reveals the heterogeneity of tumor cells and monitors the progress of tumor development, thereby preventing further cellular deterioration. Furthermore, the transcriptome analysis of immune cells in tumor tissue can be used to classify immune cells, their immune escape mechanisms and drug resistance mechanisms, and to develop effective clinical targeted therapies combined with immunotherapy. Moreover, this method enables the study of intercellular communication and the interaction of tumor cells and non-malignant cells to reveal their role in carcinogenesis. scRNA-seq provides new technical means for further development of tumor re
  33. 33Applications of multi-omics analysis in human diseases.Multi-omics usually refers to the crossover application of multiple high-throughput screening technologies represented by genomics, transcriptomics, single-cell transcriptomics, proteomics and metabolomics, spatial transcriptomics, and so on, which play a great role in promoting the study of human diseases. Most of the current reviews focus on describing the development of multi-omics technologies, data integration, and application to a particular disease; however, few of them provide a comprehensive and systematic introduction of multi-omics. This review outlines the existing technical categories of multi-omics, cautions for experimental design, focuses on the integrated analysis methods of multi-omics, especially the approach of machine learning and deep learning in multi-omics data integration and the corresponding tools, and the application of multi-omics in medical researches (e.g., cancer, neurodegenerative diseases, aging, and drug target discovery) as well as the corresponding
  34. 34Advances in spatial transcriptomic data analysis.Spatial transcriptomics is a rapidly growing field that promises to comprehensively characterize tissue organization and architecture at the single-cell or subcellular resolution. Such information provides a solid foundation for mechanistic understanding of many biological processes in both health and disease that cannot be obtained by using traditional technologies. The development of computational methods plays important roles in extracting biological signals from raw data. Various approaches have been developed to overcome technology-specific limitations such as spatial resolution, gene coverage, sensitivity, and technical biases. Downstream analysis tools formulate spatial organization and cell-cell communications as quantifiable properties, and provide algorithms to derive such properties. Integrative pipelines further assemble multiple tools in one package, allowing biologists to conveniently analyze data from beginning to end. In this review, we summarize the state of the art of
  35. 35Clinical and translational values of spatial transcriptomics.The combination of spatial transcriptomics (ST) and single cell RNA sequencing (scRNA-seq) acts as a pivotal component to bridge the pathological phenomes of human tissues with molecular alterations, defining in situ intercellular molecular communications and knowledge on spatiotemporal molecular medicine. The present article overviews the development of ST and aims to evaluate clinical and translational values for understanding molecular pathogenesis and uncovering disease-specific biomarkers. We compare the advantages and disadvantages of sequencing- and imaging-based technologies and highlight opportunities and challenges of ST. We also describe the bioinformatics tools necessary on dissecting spatial patterns of gene expression and cellular interactions and the potential applications of ST in human diseases for clinical practice as one of important issues in clinical and translational medicine, including neurology, embryo development, oncology, and inflammation. Thus, clear clinica
  36. 36jMorp updates in 2020: large enhancement of multi-omics data resources on the general Japanese population.In the Tohoku Medical Megabank project, genome and omics analyses of participants in two cohort studies were performed. A part of the data is available at the Japanese Multi Omics Reference Panel (jMorp; https://jmorp.megabank.tohoku.ac.jp) as a web-based database, as reported in our previous manuscript published in Nucleic Acid Research in 2018. At that time, jMorp mainly consisted of metabolome data; however, now genome, methylome, and transcriptome data have been integrated in addition to the enhancement of the number of samples for the metabolome data. For genomic data, jMorp provides a Japanese reference sequence obtained using de novo assembly of sequences from three Japanese individuals and allele frequencies obtained using whole-genome sequencing of 8,380 Japanese individuals. In addition, the omics data include methylome and transcriptome data from ∼300 samples and distribution of concentrations of more than 755 metabolites obtained using high-throughput nuclear magnetic reson
  37. 37Spatial transcriptomics atlas reveals the crosstalk between cancer-associated fibroblasts and tumor microenvironment components in colorectal cancer.Background The tumor-promoting role of tumor microenvironment (TME) in colorectal cancer has been widely investigated in cancer biology. Cancer-associated fibroblasts (CAFs), as the main stromal component in TME, play an important role in promoting tumor progression and metastasis. Hence, we explored the crosstalk between CAFs and microenvironment in the pathogenesis of colorectal cancer in order to provide basis for precision therapy. Methods We integrated spatial transcriptomics (ST) and bulk-RNA sequencing datasets to explore the functions of CAFs in the microenvironment of CRC. In detail, single sample gene set enrichment analysis (ssGSEA), gene set variation analysis (GSVA), pseudotime analysis and cell proportion analysis were utilized to identify the cell types and functions of each cell cluster. Immunofluorescence and immunohistochemistry were applied to confirm the results based on bioinformatics analysis. Results We profiled the tumor heterogeneity landscape and identified tw
  38. 38Integrating Molecular Perspectives: Strategies for Comprehensive Multi-Omics Integrative Data Analysis and Machine Learning Applications in Transcriptomics, Proteomics, and Metabolomics.With the advent of high-throughput technologies, the field of omics has made significant strides in characterizing biological systems at various levels of complexity. Transcriptomics, proteomics, and metabolomics are the three most widely used omics technologies, each providing unique insights into different layers of a biological system. However, analyzing each omics data set separately may not provide a comprehensive understanding of the subject under study. Therefore, integrating multi-omics data has become increasingly important in bioinformatics research. In this article, we review strategies for integrating transcriptomics, proteomics, and metabolomics data, including co-expression analysis, metabolite-gene networks, constraint-based models, pathway enrichment analysis, and interactome analysis. We discuss combined omics integration approaches, correlation-based strategies, and machine learning techniques that utilize one or more types of omics data. By presenting these methods,
  39. 39Optimizing Xenium In Situ data utility by quality assessment and best-practice analysis workflows.The Xenium In Situ platform is a new spatial transcriptomics product commercialized by 10x Genomics, capable of mapping hundreds of genes in situ at subcellular resolution. Given the multitude of commercially available spatial transcriptomics technologies, recommendations in choice of platform and analysis guidelines are increasingly important. Herein, we explore 25 Xenium datasets generated from multiple tissues and species, comparing scalability, resolution, data quality, capacities and limitations with eight other spatially resolved transcriptomics technologies and commercial platforms. In addition, we benchmark the performance of multiple open-source computational tools, when applied to Xenium datasets, in tasks including preprocessing, cell segmentation, selection of spatially variable features and domain identification. This study serves as an independent analysis of the performance of Xenium, and provides best practices and recommendations for analysis of such datasets.
  40. 40Advances in spatial transcriptomics and its applications in cancer research.Malignant tumors have increasing morbidity and high mortality, and their occurrence and development is a complicate process. The development of sequencing technologies enabled us to gain a better understanding of the underlying genetic and molecular mechanisms in tumors. In recent years, the spatial transcriptomics sequencing technologies have been developed rapidly and allow the quantification and illustration of gene expression in the spatial context of tissues. Compared with the traditional transcriptomics technologies, spatial transcriptomics technologies not only detect gene expression levels in cells, but also inform the spatial location of genes within tissues, cell composition of biological tissues, and interaction between cells. Here we summarize the development of spatial transcriptomics technologies, spatial transcriptomics tools and its application in cancer research. We also discuss the limitations and challenges of current spatial transcriptomics approaches, as well as fu
  41. 41Normalization and variance stabilization of single-cell RNA-seq data using regularized negative binomial regression.Single-cell RNA-seq (scRNA-seq) data exhibits significant cell-to-cell variation due to technical factors, including the number of molecules detected in each cell, which can confound biological heterogeneity with technical effects. To address this, we present a modeling framework for the normalization and variance stabilization of molecular count data from scRNA-seq experiments. We propose that the Pearson residuals from "regularized negative binomial regression," where cellular sequencing depth is utilized as a covariate in a generalized linear model, successfully remove the influence of technical characteristics from downstream analyses while preserving biological heterogeneity. Importantly, we show that an unconstrained negative binomial model may overfit scRNA-seq data, and overcome this by pooling information across genes with similar abundances to obtain stable parameter estimates. Our procedure omits the need for heuristic steps including pseudocount addition or log-transformati
  42. 42Transcriptome assembly from long-read RNA-seq alignments with StringTie2.RNA sequencing using the latest single-molecule sequencing instruments produces reads that are thousands of nucleotides long. The ability to assemble these long reads can greatly improve the sensitivity of long-read analyses. Here we present StringTie2, a reference-guided transcriptome assembler that works with both short and long reads. StringTie2 includes new methods to handle the high error rate of long reads and offers the ability to work with full-length super-reads assembled from short reads, which further improves the quality of short-read assemblies. StringTie2 is more accurate and faster and uses less memory than all comparable short-read and long-read analysis tools.
  43. 43A benchmark of batch-effect correction methods for single-cell RNA sequencing data.Background Large-scale single-cell transcriptomic datasets generated using different technologies contain batch-specific systematic variations that present a challenge to batch-effect removal and data integration. With continued growth expected in scRNA-seq data, achieving effective batch integration with available computational resources is crucial. Here, we perform an in-depth benchmark study on available batch correction methods to determine the most suitable method for batch-effect removal. Results We compare 14 methods in terms of computational runtime, the ability to handle large datasets, and batch-effect correction efficacy while preserving cell type purity. Five scenarios are designed for the study: identical cell types with different technologies, non-identical cell types, multiple batches, big data, and simulated data. Performance is evaluated using four benchmarking metrics including kBET, LISI, ASW, and ARI. We also investigate the use of batch-corrected data to study diff
  44. 44Cross-validation of survival associated biomarkers in gastric cancer using transcriptomic data of 1,065 patients.Introduction Multiple gene expression based prognostic biomarkers have been repeatedly identified in gastric carcinoma. However, without confirmation in an independent validation study, their clinical utility is limited. Our goal was to establish a robust database enabling the swift validation of previous and future gastric cancer survival biomarker candidates. Results The entire database incorporates 1,065 gastric carcinoma samples, gene expression data. Out of 29 established markers, higher expression of BECN1 (HR = 0.68, p = 1.5E-05), CASP3 (HR = 0.5, p = 6E-14), COX2 (HR = 0.72, p = 0.0013), CTGF (HR = 0.72, p = 0.00051), CTNNB1 (HR = 0.47, p = 4.3E-15), MET (HR = 0.63, p = 1.3E-05), and SIRT1 (HR = 0.64, p = 2.2E-07) correlated to longer OS. Higher expression of BIRC5 (HR = 1.45, p = 1E-04), CNTN1 (HR = 1.44, p = 3.5E- 05), EGFR (HR = 1.86, p = 8.5E-11), ERCC1 (HR = 1.36, p = 0.0012), HER2 (HR = 1.41, p = 0.00011), MMP2 (HR = 1.78, p = 2.6E-09), PFKB4 (HR = 1.56, p = 3.2E-07), SPH
  45. 45A single-cell and single-nucleus RNA-Seq toolbox for fresh and frozen human tumors.Single-cell genomics is essential to chart tumor ecosystems. Although single-cell RNA-Seq (scRNA-Seq) profiles RNA from cells dissociated from fresh tumors, single-nucleus RNA-Seq (snRNA-Seq) is needed to profile frozen or hard-to-dissociate tumors. Each requires customization to different tissue and tumor types, posing a barrier to adoption. Here, we have developed a systematic toolbox for profiling fresh and frozen clinical tumor samples using scRNA-Seq and snRNA-Seq, respectively. We analyzed 216,490 cells and nuclei from 40 samples across 23 specimens spanning eight tumor types of varying tissue and sample characteristics. We evaluated protocols by cell and nucleus quality, recovery rate and cellular composition. scRNA-Seq and snRNA-Seq from matched samples recovered the same cell types, but at different proportions. Our work provides guidance for studies in a broad range of tumors, including criteria for testing and selecting methods from the toolbox for other tumors, thus paving
  46. 46Prediction of functional microRNA targets by integrative modeling of microRNA binding and target expression data.We perform a large-scale RNA sequencing study to experimentally identify genes that are downregulated by 25 miRNAs. This RNA-seq dataset is combined with public miRNA target binding data to systematically identify miRNA targeting features that are characteristic of both miRNA binding and target downregulation. By integrating these common features in a machine learning framework, we develop and validate an improved computational model for genome-wide miRNA target prediction. All prediction data can be accessed at miRDB ( http://mirdb.org ).
  47. 47SUPPA2: fast, accurate, and uncertainty-aware differential splicing analysis across multiple conditions.Despite the many approaches to study differential splicing from RNA-seq, many challenges remain unsolved, including computing capacity and sequencing depth requirements. Here we present SUPPA2, a new method that addresses these challenges, and enables streamlined analysis across multiple conditions taking into account biological variability. Using experimental and simulated data, we show that SUPPA2 achieves higher accuracy compared to other methods, especially at low sequencing depth and short read length. We use SUPPA2 to identify novel Transformer2-regulated exons, novel microexons induced during differentiation of bipolar neurons, and novel intron retention events during erythroblast differentiation.
  48. 48Systematic assessment of tissue dissociation and storage biases in single-cell and single-nucleus RNA-seq workflows.Background Single-cell RNA sequencing has been widely adopted to estimate the cellular composition of heterogeneous tissues and obtain transcriptional profiles of individual cells. Multiple approaches for optimal sample dissociation and storage of single cells have been proposed as have single-nuclei profiling methods. What has been lacking is a systematic comparison of their relative biases and benefits. Results Here, we compare gene expression and cellular composition of single-cell suspensions prepared from adult mouse kidney using two tissue dissociation protocols. For each sample, we also compare fresh cells to cryopreserved and methanol-fixed cells. Lastly, we compare this single-cell data to that generated using three single-nucleus RNA sequencing workflows. Our data confirms prior reports that digestion on ice avoids the stress response observed with 37 °C dissociation. It also reveals cell types more abundant either in the cold or warm dissociations that may represent populati
  49. 49An accurate and robust imputation method scImpute for single-cell RNA-seq data.The emerging single-cell RNA sequencing (scRNA-seq) technologies enable the investigation of transcriptomic landscapes at the single-cell resolution. ScRNA-seq data analysis is complicated by excess zero counts, the so-called dropouts due to low amounts of mRNA sequenced within individual cells. We introduce scImpute, a statistical method to accurately and robustly impute the dropouts in scRNA-seq data. scImpute automatically identifies likely dropouts, and only perform imputation on these values without introducing new biases to the rest data. scImpute also detects outlier cells and excludes them from imputation. Evaluation based on both simulated and real human and mouse scRNA-seq data suggests that scImpute is an effective tool to recover transcriptome dynamics masked by dropouts. scImpute is shown to identify likely dropouts, enhance the clustering of cell subpopulations, improve the accuracy of differential expression analysis, and aid the study of gene expression dynamics.
  50. 50Accuracy assessment of fusion transcript detection via read-mapping and de novo fusion transcript assembly-based methods.Background Accurate fusion transcript detection is essential for comprehensive characterization of cancer transcriptomes. Over the last decade, multiple bioinformatic tools have been developed to predict fusions from RNA-seq, based on either read mapping or de novo fusion transcript assembly. Results We benchmark 23 different methods including applications we develop, STAR-Fusion and TrinityFusion, leveraging both simulated and real RNA-seq. Overall, STAR-Fusion, Arriba, and STAR-SEQR are the most accurate and fastest for fusion detection on cancer transcriptomes. Conclusion The lower accuracy of de novo assembly-based methods notwithstanding, they are useful for reconstructing fusion isoforms and tumor viruses, both of which are important in cancer research.
  51. 51scNMT-seq enables joint profiling of chromatin accessibility DNA methylation and transcription in single cells.Parallel single-cell sequencing protocols represent powerful methods for investigating regulatory relationships, including epigenome-transcriptome interactions. Here, we report a single-cell method for parallel chromatin accessibility, DNA methylation and transcriptome profiling. scNMT-seq (single-cell nucleosome, methylation and transcription sequencing) uses a GpC methyltransferase to label open chromatin followed by bisulfite and RNA sequencing. We validate scNMT-seq by applying it to differentiating mouse embryonic stem cells, finding links between all three molecular layers and revealing dynamic coupling between epigenomic layers during differentiation.
  52. 52A comparison of automatic cell identification methods for single-cell RNA sequencing data.Background Single-cell transcriptomics is rapidly advancing our understanding of the cellular composition of complex tissues and organisms. A major limitation in most analysis pipelines is the reliance on manual annotations to determine cell identities, which are time-consuming and irreproducible. The exponential growth in the number of cells and samples has prompted the adaptation and development of supervised classification methods for automatic cell identification. Results Here, we benchmarked 22 classification methods that automatically assign cell identities including single-cell-specific and general-purpose classifiers. The performance of the methods is evaluated using 27 publicly available single-cell RNA sequencing datasets of different sizes, technologies, species, and levels of complexity. We use 2 experimental setups to evaluate the performance of each method for within dataset predictions (intra-dataset) and across datasets (inter-dataset) based on accuracy, percentage of u
  53. 53scPred: accurate supervised method for cell-type classification from single-cell RNA-seq data.Single-cell RNA sequencing has enabled the characterization of highly specific cell types in many tissues, as well as both primary and stem cell-derived cell lines. An important facet of these studies is the ability to identify the transcriptional signatures that define a cell type or state. In theory, this information can be used to classify an individual cell based on its transcriptional profile. Here, we present scPred, a new generalizable method that is able to provide highly accurate classification of single cells, using a combination of unbiased feature selection from a reduced-dimension space, and machine-learning probability-based prediction method. We apply scPred to scRNA-seq data from pancreatic tissue, mononuclear cells, colorectal tumor biopsies, and circulating dendritic cells and show that scPred is able to classify individual cells with high accuracy. The generalized method is available at https://github.com/powellgenomicslab/scPred/.
  54. 54Feature selection and dimension reduction for single-cell RNA-Seq based on a multinomial model.Single-cell RNA-Seq (scRNA-Seq) profiles gene expression of individual cells. Recent scRNA-Seq datasets have incorporated unique molecular identifiers (UMIs). Using negative controls, we show UMI counts follow multinomial sampling with no zero inflation. Current normalization procedures such as log of counts per million and feature selection by highly variable genes produce false variability in dimension reduction. We propose simple multinomial methods, including generalized principal component analysis (GLM-PCA) for non-normal distributions, and feature selection using deviance. These methods outperform the current practice in a downstream clustering assessment using ground truth datasets.
  55. 55Benchmarking of cell type deconvolution pipelines for transcriptomics data.Many computational methods have been developed to infer cell type proportions from bulk transcriptomics data. However, an evaluation of the impact of data transformation, pre-processing, marker selection, cell type composition and choice of methodology on the deconvolution results is still lacking. Using five single-cell RNA-sequencing (scRNA-seq) datasets, we generate pseudo-bulk mixtures to evaluate the combined impact of these factors. Both bulk deconvolution methodologies and those that use scRNA-seq data as reference perform best when applied to data in linear scale and the choice of normalization has a dramatic impact on some, but not all methods. Overall, methods that use scRNA-seq data have comparable performance to the best performing bulk methods whereas semi-supervised approaches show higher error values. Moreover, failure to include cell types in the reference that are present in a mixture leads to substantially worse results, regardless of the previous choices. Altogether,
  56. 56Vireo: Bayesian demultiplexing of pooled single-cell RNA-seq data without genotype reference.Multiplexed single-cell RNA-seq analysis of multiple samples using pooling is a promising experimental design, offering increased throughput while allowing to overcome batch variation. To reconstruct the sample identify of each cell, genetic variants that segregate between the samples in the pool have been proposed as natural barcode for cell demultiplexing. Existing demultiplexing strategies rely on availability of complete genotype data from the pooled samples, which limits the applicability of such methods, in particular when genetic variation is not the primary object of study. To address this, we here present Vireo, a computationally efficient Bayesian model to demultiplex single-cell data from pooled experimental designs. Uniquely, our model can be applied in settings when only partial or no genotype information is available. Using pools based on synthetic mixtures and results on real data, we demonstrate the robustness of Vireo and illustrate the utility of multiplexed experimen
  57. 57MetaCell: analysis of single-cell RNA-seq data using K-nn graph partitions.scRNA-seq profiles each represent a highly partial sample of mRNA molecules from a unique cell that can never be resampled, and robust analysis must separate the sampling effect from biological variance. We describe a methodology for partitioning scRNA-seq datasets into metacells: disjoint and homogenous groups of profiles that could have been resampled from the same cell. Unlike clustering analysis, our algorithm specializes at obtaining granular as opposed to maximal groups. We show how to use metacells as building blocks for complex quantitative transcriptional maps while avoiding data smoothing. Our algorithms are implemented in the MetaCell R/C++ software package.
  58. 58scRNA-seq assessment of the human lung, spleen, and esophagus tissue stability after cold preservation.Background The Human Cell Atlas is a large international collaborative effort to map all cell types of the human body. Single-cell RNA sequencing can generate high-quality data for the delivery of such an atlas. However, delays between fresh sample collection and processing may lead to poor data and difficulties in experimental design. Results This study assesses the effect of cold storage on fresh healthy spleen, esophagus, and lung from ≥ 5 donors over 72 h. We collect 240,000 high-quality single-cell transcriptomes with detailed cell type annotations and whole genome sequences of donors, enabling future eQTL studies. Our data provide a valuable resource for the study of these 3 organs and will allow cross-organ comparison of cell types. We see little effect of cold ischemic time on cell yield, total number of reads per cell, and other quality control metrics in any of the tissues within the first 24 h. However, we observe a decrease in the proportions of lung T cells at 72 h, higher
  59. 59Unsupervised spatially embedded deep representation of spatial transcriptomics.Optimal integration of transcriptomics data and associated spatial information is essential towards fully exploiting spatial transcriptomics to dissect tissue heterogeneity and map out inter-cellular communications. We present SEDR, which uses a deep autoencoder coupled with a masked self-supervised learning mechanism to construct a low-dimensional latent representation of gene expression, which is then simultaneously embedded with the corresponding spatial information through a variational graph autoencoder. SEDR achieved higher clustering performance on manually annotated 10 × Visium datasets and better scalability on high-resolution spatial transcriptomics datasets than existing methods. Additionally, we show SEDR's ability to impute and denoise gene expression (URL: https://github.com/JinmiaoChenLab/SEDR/ ).
  60. 60Integrative analyses of single-cell transcriptome and regulome using MAESTRO.We present Model-based AnalysEs of Transcriptome and RegulOme (MAESTRO), a comprehensive open-source computational workflow ( http://github.com/liulab-dfci/MAESTRO ) for the integrative analyses of single-cell RNA-seq (scRNA-seq) and ATAC-seq (scATAC-seq) data from multiple platforms. MAESTRO provides functions for pre-processing, alignment, quality control, expression and chromatin accessibility quantification, clustering, differential analysis, and annotation. By modeling gene regulatory potential from chromatin accessibilities at the single-cell level, MAESTRO outperforms the existing methods for integrating the cell clusters between scRNA-seq and scATAC-seq. Furthermore, MAESTRO supports automatic cell-type annotation using predefined cell type marker genes and identifies driver regulators from differential scRNA-seq genes and scATAC-seq peaks.
  61. 61Implementation of next generation sequencing into pediatric hematology-oncology practice: moving beyond actionable alterations.Background Molecular characterization has the potential to advance the management of pediatric cancer and high-risk hematologic disease. The clinical integration of genome sequencing into standard clinical practice has been limited and the potential utility of genome sequencing to identify clinically impactful information beyond targetable alterations has been underestimated. Methods The Precision in Pediatric Sequencing (PIPseq) Program at Columbia University Medical Center instituted prospective clinical next generation sequencing (NGS) for pediatric cancer and hematologic disorders at risk for treatment failure. We performed cancer whole exome sequencing (WES) of patient-matched tumor-normal samples and RNA sequencing (RNA-seq) of tumor to identify sequence variants, fusion transcripts, relative gene expression, and copy number variation (CNV). A directed cancer gene panel assay was used when sample adequacy was a concern. Constitutional WES of patients and parents was performed whe
  62. 62Statistical and machine learning methods for spatially resolved transcriptomics data analysis.The recent advancement in spatial transcriptomics technology has enabled multiplexed profiling of cellular transcriptomes and spatial locations. As the capacity and efficiency of the experimental technologies continue to improve, there is an emerging need for the development of analytical approaches. Furthermore, with the continuous evolution of sequencing protocols, the underlying assumptions of current analytical methods need to be re-evaluated and adjusted to harness the increasing data complexity. To motivate and aid future model development, we herein review the recent development of statistical and machine learning methods in spatial transcriptomics, summarize useful resources, and highlight the challenges and opportunities ahead.
  63. 63Advances in spatial transcriptomics and related data analysis strategies.Spatial transcriptomics technologies developed in recent years can provide various information including tissue heterogeneity, which is fundamental in biological and medical research, and have been making significant breakthroughs. Single-cell RNA sequencing (scRNA-seq) cannot provide spatial information, while spatial transcriptomics technologies allow gene expression information to be obtained from intact tissue sections in the original physiological context at a spatial resolution. Various biological insights can be generated into tissue architecture and further the elucidation of the interaction between cells and the microenvironment. Thus, we can gain a general understanding of histogenesis processes and disease pathogenesis, etc. Furthermore, in silico methods involving the widely distributed R and Python packages for data analysis play essential roles in deriving indispensable bioinformation and eliminating technological limitations. In this review, we summarize available techno
  64. 64Evaluation of cell-cell interaction methods by integrating single-cell RNA sequencing data with spatial information.Background Cell-cell interactions are important for information exchange between different cells, which are the fundamental basis of many biological processes. Recent advances in single-cell RNA sequencing (scRNA-seq) enable the characterization of cell-cell interactions using computational methods. However, it is hard to evaluate these methods since no ground truth is provided. Spatial transcriptomics (ST) data profiles the relative position of different cells. We propose that the spatial distance suggests the interaction tendency of different cell types, thus could be used for evaluating cell-cell interaction tools. Results We benchmark 16 cell-cell interaction methods by integrating scRNA-seq with ST data. We characterize cell-cell interactions into short-range and long-range interactions using spatial distance distributions between ligands and receptors. Based on this classification, we define the distance enrichment score and apply an evaluation workflow to 16 cell-cell interactio
  65. 65Psoriatic Arthritis: Pathogenesis and Targeted Therapies.Psoriatic arthritis (PsA), a heterogeneous chronic inflammatory immune-mediated disease characterized by musculoskeletal inflammation (arthritis, enthesitis, spondylitis, and dactylitis), generally occurs in patients with psoriasis. PsA is also associated with uveitis and inflammatory bowel disease (Crohn's disease and ulcerative colitis). To capture these manifestations as well as the associated comorbidities, and to recognize their underlining common pathogenesis, the name of psoriatic disease was coined. The pathogenesis of PsA is complex and multifaceted, with an interplay of genetic predisposition, triggering environmental factors, and activation of the innate and adaptive immune system, although autoinflammation has also been implicated. Research has identified several immune-inflammatory pathways defined by cytokines (IL-23/IL-17, TNF), leading to the development of efficacious therapeutic targets. However, heterogeneous responses to these drugs occur in different patients and i
  66. 66Computational Approaches and Challenges in Spatial Transcriptomics.The development of spatial transcriptomics (ST) technologies has transformed genetic research from a single-cell data level to a two-dimensional spatial coordinate system and facilitated the study of the composition and function of various cell subsets in different environments and organs. The large-scale data generated by these ST technologies, which contain spatial gene expression information, have elicited the need for spatially resolved approaches to meet the requirements of computational and biological data interpretation. These requirements include dealing with the explosive growth of data to determine the cell-level and gene-level expression, correcting the inner batch effect and loss of expression to improve the data quality, conducting efficient interpretation and in-depth knowledge mining both at the single-cell and tissue-wide levels, and conducting multi-omics integration analysis to provide an extensible framework toward the in-depth understanding of biological processes.
  67. 67Spatial omics: Navigating to the golden era of cancer research.The idea that tumour microenvironment (TME) is organised in a spatial manner will not surprise many cancer biologists; however, systematically capturing spatial architecture of TME is still not possible until recent decade. The past five years have witnessed a boom in the research of high-throughput spatial techniques and algorithms to delineate TME at an unprecedented level. Here, we review the technological progress of spatial omics and how advanced computation methods boost multi-modal spatial data analysis. Then, we discussed the potential clinical translations of spatial omics research in precision oncology, and proposed a transfer of spatial ecological principles to cancer biology in spatial data interpretation. So far, spatial omics is placing us in the golden age of spatial cancer research. Further development and application of spatial omics may lead to a comprehensive decoding of the TME ecosystem and bring the current spatiotemporal molecular medical research into an entirel
  68. 68Evaluating spatially variable gene detection methods for spatial transcriptomics data.Background The identification of genes that vary across spatial domains in tissues and cells is an essential step for spatial transcriptomics data analysis. Given the critical role it serves for downstream data interpretations, various methods for detecting spatially variable genes (SVGs) have been proposed. However, the lack of benchmarking complicates the selection of a suitable method. Results Here we systematically evaluate a panel of popular SVG detection methods on a large collection of spatial transcriptomics datasets, covering various tissue types, biotechnologies, and spatial resolutions. We address questions including whether different methods select a similar set of SVGs, how reliable is the reported statistical significance from each method, how accurate and robust is each method in terms of SVG detection, and how well the selected SVGs perform in downstream applications such as clustering of spatial domains. Besides these, practical considerations such as computational tim
  69. 69Collagen-producing lung cell atlas identifies multiple subsets with distinct localization and relevance to fibrosis.Collagen-producing cells maintain the complex architecture of the lung and drive pathologic scarring in pulmonary fibrosis. Here we perform single-cell RNA-sequencing to identify all collagen-producing cells in normal and fibrotic lungs. We characterize multiple collagen-producing subpopulations with distinct anatomical localizations in different compartments of murine lungs. One subpopulation, characterized by expression of Cthrc1 (collagen triple helix repeat containing 1), emerges in fibrotic lungs and expresses the highest levels of collagens. Single-cell RNA-sequencing of human lungs, including those from idiopathic pulmonary fibrosis and scleroderma patients, demonstrate similar heterogeneity and CTHRC1-expressing fibroblasts present uniquely in fibrotic lungs. Immunostaining and in situ hybridization show that these cells are concentrated within fibroblastic foci. We purify collagen-producing subpopulations and find disease-relevant phenotypes of Cthrc1-expressing fibroblasts in
  70. 70A deep proteome and transcriptome abundance atlas of 29 healthy human tissues.Genome-, transcriptome- and proteome-wide measurements provide insights into how biological systems are regulated. However, fundamental aspects relating to which human proteins exist, where they are expressed and in which quantities are not fully understood. Therefore, we generated a quantitative proteome and transcriptome abundance atlas of 29 paired healthy human tissues from the Human Protein Atlas project representing human genes by 18,072 transcripts and 13,640 proteins including 37 without prior protein-level evidence. The analysis revealed that hundreds of proteins, particularly in testis, could not be detected even for highly expressed mRNAs, that few proteins show tissue-specific expression, that strong differences between mRNA and protein quantities within and across tissues exist and that protein expression is often more stable across tissues than that of transcripts. Only 238 of 9,848 amino acid variants found by exome sequencing could be confidently detected at the protein
  71. 71An atlas of the aging lung mapped by single cell transcriptomics and deep tissue proteomics.Aging promotes lung function decline and susceptibility to chronic lung diseases, which are the third leading cause of death worldwide. Here, we use single cell transcriptomics and mass spectrometry-based proteomics to quantify changes in cellular activity states across 30 cell types and chart the lung proteome of young and old mice. We show that aging leads to increased transcriptional noise, indicating deregulated epigenetic control. We observe cell type-specific effects of aging, uncovering increased cholesterol biosynthesis in type-2 pneumocytes and lipofibroblasts and altered relative frequency of airway epithelial cells as hallmarks of lung aging. Proteomic profiling reveals extracellular matrix remodeling in old mice, including increased collagen IV and XVI and decreased Fraser syndrome complex proteins and collagen XIV. Computational integration of the aging proteome with the single cell transcriptomes predicts the cellular source of regulated proteins and creates an unbiased r
  72. 72Expression Atlas update: from tissues to single cells.Expression Atlas is EMBL-EBI's resource for gene and protein expression. It sources and compiles data on the abundance and localisation of RNA and proteins in various biological systems and contexts and provides open access to this data for the research community. With the increased availability of single cell RNA-Seq datasets in the public archives, we have now extended Expression Atlas with a new added-value service to display gene expression in single cells. Single Cell Expression Atlas was launched in 2018 and currently includes 123 single cell RNA-Seq studies from 12 species. The website can be searched by genes within or across species to reveal experiments, tissues and cell types where this gene is expressed or under which conditions it is a marker gene. Within each study, cells can be visualized using a pre-calculated t-SNE plot and can be coloured by different features or by cell clusters based on gene expression. Within each experiment, there are links to downloadable files,
  73. 73Revealing the Critical Regulators of Cell Identity in the Mouse Cell Atlas.Recent progress in single-cell technologies has enabled the identification of all major cell types in mouse. However, for most cell types, the regulatory mechanism underlying their identity remains poorly understood. By computational analysis of the recently published mouse cell atlas data, we have identified 202 regulons whose activities are highly variable across different cell types, and more importantly, predicted a small set of essential regulators for each major cell type in mouse. Systematic validation by automated literature and data mining provides strong additional support for our predictions. Thus, these predictions serve as a valuable resource that would be useful for the broad biological community. Finally, we have built a user-friendly, interactive web portal to enable users to navigate this mouse cell network atlas.
  74. 74FungiDB: An Integrated Bioinformatic Resource for Fungi and Oomycetes.FungiDB (fungidb.org) is a free online resource for data mining and functional genomics analysis for fungal and oomycete species. FungiDB is part of the Eukaryotic Pathogen Genomics Database Resource (EuPathDB, eupathdb.org) platform that integrates genomic, transcriptomic, proteomic, and phenotypic datasets, and other types of data for pathogenic and nonpathogenic, free-living and parasitic organisms. FungiDB is one of the largest EuPathDB databases containing nearly 100 genomes obtained from GenBank, Aspergillus Genome Database (AspGD), The Broad Institute, Joint Genome Institute (JGI), Ensembl, and other sources. FungiDB offers a user-friendly web interface with embedded bioinformatics tools that support custom in silico experiments that leverage FungiDB-integrated data. In addition, a Galaxy-based workspace enables users to generate custom pipelines for large-scale data analysis (e.g., RNA-Seq, variant calling, etc.). This review provides an introduction to the FungiDB resources an
  75. 75Multiomic spatial landscape of innate immune cells at human central nervous system borders.The innate immune compartment of the human central nervous system (CNS) is highly diverse and includes several immune-cell populations such as macrophages that are frequent in the brain parenchyma (microglia) and less numerous at the brain interfaces as CNS-associated macrophages (CAMs). Due to their scantiness and particular location, little is known about the presence of temporally and spatially restricted CAM subclasses during development, health and perturbation. Here we combined single-cell RNA sequencing, time-of-flight mass cytometry and single-cell spatial transcriptomics with fate mapping and advanced immunohistochemistry to comprehensively characterize the immune system at human CNS interfaces with over 356,000 analyzed transcriptomes from 102 individuals. We also provide a comprehensive analysis of resident and engrafted myeloid cells in the brains of 15 individuals with peripheral blood stem cell transplantation, revealing compartment-specific engraftment rates across diffe
  76. 76Spatial multimodal analysis of transcriptomes and metabolomes in tissues.We present a spatial omics approach that combines histology, mass spectrometry imaging and spatial transcriptomics to facilitate precise measurements of mRNA transcripts and low-molecular-weight metabolites across tissue regions. The workflow is compatible with commercially available Visium glass slides. We demonstrate the potential of our method using mouse and human brain samples in the context of dopamine and Parkinson's disease.
  77. 77Single-cell and spatial transcriptomics analysis of non-small cell lung cancer.Lung cancer is the second most frequently diagnosed cancer and the leading cause of cancer-related mortality worldwide. Tumour ecosystems feature diverse immune cell types. Myeloid cells, in particular, are prevalent and have a well-established role in promoting the disease. In our study, we profile approximately 900,000 cells from 25 treatment-naive patients with adenocarcinoma and squamous-cell carcinoma by single-cell and spatial transcriptomics. We note an inverse relationship between anti-inflammatory macrophages and NK cells/T cells, and with reduced NK cell cytotoxicity within the tumour. While we observe a similar cell type composition in both adenocarcinoma and squamous-cell carcinoma, we detect significant differences in the co-expression of various immune checkpoint inhibitors. Moreover, we reveal evidence of a transcriptional "reprogramming" of macrophages in tumours, shifting them towards cholesterol export and adopting a foetal-like transcriptional signature which promote
  78. 78Unveiling inflammatory and prehypertrophic cell populations as key contributors to knee cartilage degeneration in osteoarthritis using multi-omics data integration.Objectives Single-cell and spatial transcriptomics analysis of human knee articular cartilage tissue to present a comprehensive transcriptome landscape and osteoarthritis (OA)-critical cell populations. Methods Single-cell RNA sequencing and spatially resolved transcriptomic technology have been applied to characterise the cellular heterogeneity of human knee articular cartilage which were collected from 8 OA donors, and 3 non-OA control donors, and a total of 19 samples. The novel chondrocyte population and marker genes of interest were validated by immunohistochemistry staining, quantitative real-time PCR, etc. The OA-critical cell populations were validated through integrative analyses of publicly available bulk RNA sequencing data and large-scale genome-wide association studies. Results We identified 33 cell population-specific marker genes that define 11 chondrocyte populations, including 9 known populations and 2 new populations, that is, pre-inflammatory chondrocyte population (
  79. 79Spatial transcriptomics reveal neuron-astrocyte synergy in long-term memory.Memory encodes past experiences, thereby enabling future plans. The basolateral amygdala is a centre of salience networks that underlie emotional experiences and thus has a key role in long-term fear memory formation 1 . Here we used spatial and single-cell transcriptomics to illuminate the cellular and molecular architecture of the role of the basolateral amygdala in long-term memory. We identified transcriptional signatures in subpopulations of neurons and astrocytes that were memory-specific and persisted for weeks. These transcriptional signatures implicate neuropeptide and BDNF signalling, MAPK and CREB activation, ubiquitination pathways, and synaptic connectivity as key components of long-term memory. Notably, upon long-term memory formation, a neuronal subpopulation defined by increased Penk and decreased Tac expression constituted the most prominent component of the memory engram of the basolateral amygdala. These transcriptional changes were observed both with single-cell RNA
  80. 80Deciphering tissue structure and function using spatial transcriptomics.The rapid development of spatial transcriptomics (ST) techniques has allowed the measurement of transcriptional levels across many genes together with the spatial positions of cells. This has led to an explosion of interest in computational methods and techniques for harnessing both spatial and transcriptional information in analysis of ST datasets. The wide diversity of approaches in aim, methodology and technology for ST provides great challenges in dissecting cellular functions in spatial contexts. Here, we synthesize and review the key problems in analysis of ST data and methods that are currently applied, while also expanding on open questions and areas of future development.
  81. 81stPlus: a reference-based method for the accurate enhancement of spatial transcriptomics.Motivation Single-cell RNA sequencing (scRNA-seq) techniques have revolutionized the investigation of transcriptomic landscape in individual cells. Recent advancements in spatial transcriptomic technologies further enable gene expression profiling and spatial organization mapping of cells simultaneously. Among the technologies, imaging-based methods can offer higher spatial resolutions, while they are limited by either the small number of genes imaged or the low gene detection sensitivity. Although several methods have been proposed for enhancing spatially resolved transcriptomics, inadequate accuracy of gene expression prediction and insufficient ability of cell-population identification still impede the applications of these methods. Results We propose stPlus, a reference-based method that leverages information in scRNA-seq data to enhance spatial transcriptomics. Based on an auto-encoder with a carefully tailored loss function, stPlus performs joint embedding and predicts spatial ge
  82. 82Benchmarking and integration of methods for deconvoluting spatial transcriptomic data.Motivation The rapid development of spatial transcriptomics (ST) approaches has provided new insights into understanding tissue architecture and function. However, the gene expressions measured at a spot may contain contributions from multiple cells due to the low-resolution of current ST technologies. Although many computational methods have been developed to disentangle discrete cell types from spatial mixtures, the community lacks a thorough evaluation of the performance of those deconvolution methods. Results Here, we present a comprehensive benchmarking of 14 deconvolution methods on four datasets. Furthermore, we investigate the robustness of different methods to sequencing depth, spot size and the choice of normalization. Moreover, we propose a new ensemble learning-based deconvolution method (EnDecon) by integrating multiple individual methods for more accurate deconvolution. The major new findings include: (i) cell2loction, RCTD and spatialDWLS are more accurate than other ST
  83. 83Spatial Transcriptomics: A Powerful Tool in Disease Understanding and Drug Discovery.Recent advancements in modern science have provided robust tools for drug discovery. The rapid development of transcriptome sequencing technologies has given rise to single-cell transcriptomics and single-nucleus transcriptomics, increasing the accuracy of sequencing and accelerating the drug discovery process. With the evolution of single-cell transcriptomics, spatial transcriptomics (ST) technology has emerged as a derivative approach. Spatial transcriptomics has emerged as a hot topic in the field of omics research in recent years; it not only provides information on gene expression levels but also offers spatial information on gene expression. This technology has shown tremendous potential in research on disease understanding and drug discovery. In this article, we introduce the analytical strategies of spatial transcriptomics and review its applications in novel target discovery and drug mechanism unravelling. Moreover, we discuss the current challenges and issues in this research
  84. 84Advancements in single-cell RNA sequencing and spatial transcriptomics: transforming biomedical research.In recent years, significant advancements in biochemistry, materials science, engineering, and computer-aided testing have driven the development of high-throughput tools for profiling genetic information. Single-cell RNA sequencing (scRNA-seq) technologies have established themselves as key tools for dissecting genetic sequences at the level of single cells. These technologies reveal cellular diversity and allow for the exploration of cell states and transformations with exceptional resolution. Unlike bulk sequencing, which provides population-averaged data, scRNA-seq can detect cell subtypes or gene expression variations that would otherwise be overlooked. However, a key limitation of scRNA-seq is its inability to preserve spatial information about the RNA transcriptome, as the process requires tissue dissociation and cell isolation. Spatial transcriptomics is a pivotal advancement in medical biotechnology, facilitating the identification of molecules such as RNA in their original sp
  85. 85SAW: an efficient and accurate data analysis workflow for Stereo-seq spatial transcriptomics.The basic analysis steps of spatial transcriptomics require obtaining gene expression information from both space and cells. The existing tools for these analyses incur performance issues when dealing with large datasets. These issues involve computationally intensive spatial localization, RNA genome alignment, and excessive memory usage in large chip scenarios. These problems affect the applicability and efficiency of the analysis. Here, a high-performance and accurate spatial transcriptomics data analysis workflow, called Stereo-seq Analysis Workflow (SAW), was developed for the Stereo-seq technology developed at BGI. SAW includes mRNA spatial position reconstruction, genome alignment, gene expression matrix generation, and clustering. The workflow outputs files in a universal format for subsequent personalized analysis. The execution time for the entire analysis is ∼148 min with 1 GB reads 1 × 1 cm chip test data, 1.8 times faster than with an unoptimized workflow.
  86. 86Current state and future prospects of spatial biology in colorectal cancer.Over the past century, colorectal cancer (CRC) has become one of the most devastating cancers impacting the human population. To gain a deeper understanding of the molecular mechanisms driving this solid tumor, researchers have increasingly turned their attention to the tumor microenvironment (TME). Spatial transcriptomics and proteomics have emerged as a particularly powerful technology for deciphering the complexity of CRC tumors, given that the TME and its spatial organization are critical determinants of disease progression and treatment response. Spatial transcriptomics enables high-resolution mapping of the whole transcriptome. While spatial proteomics maps protein expression and function across tissue sections. Together, they provide a detailed view of the molecular landscape and cellular interactions within the TME. In this review, we delve into recent advances in spatial biology technologies applied to CRC research, highlighting both the methodologies and the challenges associ
  87. 87Single Cell Atlas: a single-cell multi-omics human cell encyclopedia.Single-cell sequencing datasets are key in biology and medicine for unraveling insights into heterogeneous cell populations with unprecedented resolution. Here, we construct a single-cell multi-omics map of human tissues through in-depth characterizations of datasets from five single-cell omics, spatial transcriptomics, and two bulk omics across 125 healthy adult and fetal tissues. We construct its complement web-based platform, the Single Cell Atlas (SCA, www.singlecellatlas.org ), to enable vast interactive data exploration of deep multi-omics signatures across human fetal and adult tissues. The atlas resources and database queries aspire to serve as a one-stop, comprehensive, and time-effective resource for various omics studies.
  88. 88Inference and analysis of cell-cell communication using CellChat.Understanding global communications among cells requires accurate representation of cell-cell signaling links and effective systems-level analyses of those links. We construct a database of interactions among ligands, receptors and their cofactors that accurately represent known heteromeric molecular complexes. We then develop CellChat, a tool that is able to quantitatively infer and analyze intercellular communication networks from single-cell RNA-sequencing (scRNA-seq) data. CellChat predicts major signaling inputs and outputs for cells and how those cells and signals coordinate for functions using network analysis and pattern recognition approaches. Through manifold learning and quantitative contrasts, CellChat classifies signaling pathways and delineates conserved and context-specific pathways across different datasets. Applying CellChat to mouse and human skin datasets shows its ability to extract complex signaling patterns. Our versatile and easy-to-use toolkit CellChat and a web
  89. 89The history and advances in cancer immunotherapy: understanding the characteristics of tumor-infiltrating immune cells and their therapeutic implications.Immunotherapy has revolutionized cancer treatment and rejuvenated the field of tumor immunology. Several types of immunotherapy, including adoptive cell transfer (ACT) and immune checkpoint inhibitors (ICIs), have obtained durable clinical responses, but their efficacies vary, and only subsets of cancer patients can benefit from them. Immune infiltrates in the tumor microenvironment (TME) have been shown to play a key role in tumor development and will affect the clinical outcomes of cancer patients. Comprehensive profiling of tumor-infiltrating immune cells would shed light on the mechanisms of cancer-immune evasion, thus providing opportunities for the development of novel therapeutic strategies. However, the highly heterogeneous and dynamic nature of the TME impedes the precise dissection of intratumoral immune cells. With recent advances in single-cell technologies such as single-cell RNA sequencing (scRNA-seq) and mass cytometry, systematic interrogation of the TME is feasible and
  90. 90Current best practices in single-cell RNA-seq analysis: a tutorial.Single-cell RNA-seq has enabled gene expression to be studied at an unprecedented resolution. The promise of this technology is attracting a growing user base for single-cell analysis methods. As more analysis tools are becoming available, it is becoming increasingly difficult to navigate this landscape and produce an up-to-date workflow to analyse one's data. Here, we detail the steps of a typical single-cell RNA-seq analysis, including pre-processing (quality control, normalization, data correction, feature selection, and dimensionality reduction) and cell- and gene-level downstream analysis. We formulate current best-practice recommendations for these steps based on independent comparison studies. We have integrated these best-practice recommendations into a workflow, which we apply to a public dataset to further illustrate how these steps work in practice. Our documented case study can be found at https://www.github.com/theislab/single-cell-tutorial This review will serve as a work
  91. 91A survey of best practices for RNA-seq data analysis.RNA-sequencing (RNA-seq) has a wide variety of applications, but no single analysis pipeline can be used in all cases. We review all of the major steps in RNA-seq data analysis, including experimental design, quality control, read alignment, quantification of gene and transcript levels, visualization, differential gene expression, alternative splicing, functional analysis, gene fusion detection and eQTL mapping. We highlight the challenges associated with each step. We discuss the analysis of small RNAs and the integration of RNA-seq with other functional genomics techniques. Finally, we discuss the outlook for novel technologies that are changing the state of the art in transcriptomics.
  92. 92Dynamic transcriptomic m 6 A decoration: writers, erasers, readers and functions in RNA metabolism.N 6 -methyladenosine (m 6 A) is a chemical modification present in multiple RNA species, being most abundant in mRNAs. Studies on enzymes or factors that catalyze, recognize, and remove m 6 A have revealed its comprehensive roles in almost every aspect of mRNA metabolism, as well as in a variety of physiological processes. This review describes the current understanding of the m 6 A modification, particularly the functions of its writers, erasers, readers in RNA metabolism, with an emphasis on its role in regulating the isoform dosage of mRNAs.
  93. 93A step-by-step workflow for low-level analysis of single-cell RNA-seq data with Bioconductor.Single-cell RNA sequencing (scRNA-seq) is widely used to profile the transcriptome of individual cells. This provides biological resolution that cannot be matched by bulk RNA sequencing, at the cost of increased technical noise and data complexity. The differences between scRNA-seq and bulk RNA-seq data mean that the analysis of the former cannot be performed by recycling bioinformatics pipelines for the latter. Rather, dedicated single-cell methods are required at various steps to exploit the cellular resolution while accounting for technical noise. This article describes a computational workflow for low-level analyses of scRNA-seq data, based primarily on software packages from the open-source Bioconductor project. It covers basic steps including quality control, data exploration and normalization, as well as more complex procedures such as cell cycle phase assignment, identification of highly variable and correlated genes, clustering into subpopulations and marker gene detection. An
  94. 94IOBR: Multi-Omics Immuno-Oncology Biological Research to Decode Tumor Microenvironment and Signatures.Recent advances in next-generation sequencing (NGS) technologies have triggered the rapid accumulation of publicly available multi-omics datasets. The application of integrated omics to explore robust signatures for clinical translation is increasingly emphasized, and this is attributed to the clinical success of immune checkpoint blockades in diverse malignancies. However, effective tools for comprehensively interpreting multi-omics data are still warranted to provide increased granularity into the intrinsic mechanism of oncogenesis and immunotherapeutic sensitivity. Therefore, we developed a computational tool for effective Immuno-Oncology Biological Research (IOBR), providing a comprehensive investigation of the estimation of reported or user-built signatures, TME deconvolution, and signature construction based on multi-omics data. Notably, IOBR offers batch analyses of these signatures and their correlations with clinical phenotypes, long non-coding RNA (lncRNA) profiling, genomic
  95. 95Validation of miRNA prognostic power in hepatocellular carcinoma using expression data of independent datasets.Multiple studies suggested using different miRNAs as biomarkers for prognosis of hepatocellular carcinoma (HCC). We aimed to assemble a miRNA expression database from independent datasets to enable an independent validation of previously published prognostic biomarkers of HCC. A miRNA expression database was established by searching the TCGA (RNA-seq) and GEO (microarray) repositories to identify miRNA datasets with available expression and clinical data. A PubMed search was performed to identify prognostic miRNAs for HCC. We performed a uni- and multivariate Cox regression analysis to validate the prognostic significance of these miRNAs. The Limma R package was applied to compare the expression of miRNAs between tumor and normal tissues. We uncovered 214 publications containing 223 miRNAs identified as potential prognostic biomarkers for HCC. In the survival analysis, the expression levels of 55 and 84 miRNAs were significantly correlated with overall survival in RNA-seq and gene chip
  96. 96SARTools: A DESeq2- and EdgeR-Based R Pipeline for Comprehensive Differential Analysis of RNA-Seq Data.Background Several R packages exist for the detection of differentially expressed genes from RNA-Seq data. The analysis process includes three main steps, namely normalization, dispersion estimation and test for differential expression. Quality control steps along this process are recommended but not mandatory, and failing to check the characteristics of the dataset may lead to spurious results. In addition, normalization methods and statistical models are not exchangeable across the packages without adequate transformations the users are often not aware of. Thus, dedicated analysis pipelines are needed to include systematic quality control steps and prevent errors from misusing the proposed methods. Results SARTools is an R pipeline for differential analysis of RNA-Seq count data. It can handle designs involving two or more conditions of a single biological factor with or without a blocking factor (such as a batch effect or a sample pairing). It is based on DESeq2 and edgeR and is com
  97. 97From reads to genes to pathways: differential expression analysis of RNA-Seq experiments using Rsubread and the edgeR quasi-likelihood pipeline.In recent years, RNA sequencing (RNA-seq) has become a very widely used technology for profiling gene expression. One of the most common aims of RNA-seq profiling is to identify genes or molecular pathways that are differentially expressed (DE) between two or more biological conditions. This article demonstrates a computational workflow for the detection of DE genes and pathways from RNA-seq data by providing a complete analysis of an RNA-seq experiment profiling epithelial cell subsets in the mouse mammary gland. The workflow uses R software packages from the open-source Bioconductor project and covers all steps of the analysis pipeline, including alignment of read sequences, data exploration, differential expression analysis, visualization and pathway analysis. Read alignment and count quantification is conducted using the Rsubread package and the statistical analyses are performed using the edgeR package. The differential expression analysis uses the quasi-likelihood functionality o
  98. 98TNMplot.com: A Web Tool for the Comparison of Gene Expression in Normal, Tumor and Metastatic Tissues.Genes showing higher expression in either tumor or metastatic tissues can help in better understanding tumor formation and can serve as biomarkers of progression or as potential therapy targets. Our goal was to establish an integrated database using available transcriptome-level datasets and to create a web platform which enables the mining of this database by comparing normal, tumor and metastatic data across all genes in real time. We utilized data generated by either gene arrays from the Gene Expression Omnibus of the National Center for Biotechnology Information (NCBI-GEO) or RNA-seq from The Cancer Genome Atlas (TCGA), Therapeutically Applicable Research to Generate Effective Treatments (TARGET), and The Genotype-Tissue Expression (GTEx) repositories. The altered expression within different platforms was analyzed separately. Statistical significance was computed using Mann-Whitney or Kruskal-Wallis tests. False Discovery Rate (FDR) was computed using the Benjamini-Hochberg method.
  99. 99A practical guide to single-cell RNA-sequencing for biomedical research and clinical applications.RNA sequencing (RNA-seq) is a genomic approach for the detection and quantitative analysis of messenger RNA molecules in a biological sample and is useful for studying cellular responses. RNA-seq has fueled much discovery and innovation in medicine over recent years. For practical reasons, the technique is usually conducted on samples comprising thousands to millions of cells. However, this has hindered direct assessment of the fundamental unit of biology-the cell. Since the first single-cell RNA-sequencing (scRNA-seq) study was published in 2009, many more have been conducted, mostly by specialist laboratories with unique skills in wet-lab single-cell genomics, bioinformatics, and computation. However, with the increasing commercial availability of scRNA-seq platforms, and the rapid ongoing maturation of bioinformatics approaches, a point has been reached where any biomedical researcher or clinician can use scRNA-seq to make exciting discoveries. In this review, we present a practical
  100. 100Deep learning and alignment of spatially resolved single-cell transcriptomes with Tangram.Charting an organs' biological atlas requires us to spatially resolve the entire single-cell transcriptome, and to relate such cellular features to the anatomical scale. Single-cell and single-nucleus RNA-seq (sc/snRNA-seq) can profile cells comprehensively, but lose spatial information. Spatial transcriptomics allows for spatial measurements, but at lower resolution and with limited sensitivity. Targeted in situ technologies solve both issues, but are limited in gene throughput. To overcome these limitations we present Tangram, a method that aligns sc/snRNA-seq data to various forms of spatial data collected from the same region, including MERFISH, STARmap, smFISH, Spatial Transcriptomics (Visium) and histological images. Tangram can map any type of sc/snRNA-seq data, including multimodal data such as those from SHARE-seq, which we used to reveal spatial patterns of chromatin accessibility. We demonstrate Tangram on healthy mouse brain tissue, by reconstructing a genome-wide anatomica
  101. 101Survival analysis across the entire transcriptome identifies biomarkers with the highest prognostic power in breast cancer.Introduction Extensive research is directed to uncover new biomarkers capable to stratify breast cancer patients into clinically relevant cohorts. However, the overall performance ranking of such marker candidates compared to other genes is virtually absent. Here, we present the ranking of all survival related genes in chemotherapy treated basal and estrogen positive/HER2 negative breast cancer. Methods We searched the GEO repository to uncover transcriptomic datasets with available follow-up and clinical data. After quality control and normalization, samples entered an integrated database. Molecular subtypes were designated using gene expression data. Relapse-free survival analysis was performed using Cox proportional hazards regression. False discovery rate was computed to combat multiple hypothesis testing. Kaplan-Meier plots were drawn to visualize the best performing genes. Results The entire database includes 7,830 unique samples from 55 independent datasets. Of those with availa
  102. 102Comparative cellular analysis of motor cortex in human, marmoset and mouse.The primary motor cortex (M1) is essential for voluntary fine-motor control and is functionally conserved across mammals 1 . Here, using high-throughput transcriptomic and epigenomic profiling of more than 450,000 single nuclei in humans, marmoset monkeys and mice, we demonstrate a broadly conserved cellular makeup of this region, with similarities that mirror evolutionary distance and are consistent between the transcriptome and epigenome. The core conserved molecular identities of neuronal and non-neuronal cell types allow us to generate a cross-species consensus classification of cell types, and to infer conserved properties of cell types across species. Despite the overall conservation, however, many species-dependent specializations are apparent, including differences in cell-type proportions, gene expression, DNA methylation and chromatin state. Few cell-type marker genes are conserved across species, revealing a short list of candidate genes and regulatory mechanisms that are re
  103. 103Single-Cell RNA-Seq Technologies and Related Computational Data Analysis.Single-cell RNA sequencing (scRNA-seq) technologies allow the dissection of gene expression at single-cell resolution, which greatly revolutionizes transcriptomic studies. A number of scRNA-seq protocols have been developed, and these methods possess their unique features with distinct advantages and disadvantages. Due to technical limitations and biological factors, scRNA-seq data are noisier and more complex than bulk RNA-seq data. The high variability of scRNA-seq data raises computational challenges in data analysis. Although an increasing number of bioinformatics methods are proposed for analyzing and interpreting scRNA-seq data, novel algorithms are required to ensure the accuracy and reproducibility of results. In this review, we provide an overview of currently available single-cell isolation protocols and scRNA-seq technologies, and discuss the methods for diverse scRNA-seq data analyses including quality control, read mapping, gene expression quantification, batch effect corr
  104. 104Mapping information-rich genotype-phenotype landscapes with genome-scale Perturb-seq.A central goal of genetics is to define the relationships between genotypes and phenotypes. High-content phenotypic screens such as Perturb-seq (CRISPR-based screens with single-cell RNA-sequencing readouts) enable massively parallel functional genomic mapping but, to date, have been used at limited scales. Here, we perform genome-scale Perturb-seq targeting all expressed genes with CRISPR interference (CRISPRi) across >2.5 million human cells. We use transcriptional phenotypes to predict the function of poorly characterized genes, uncovering new regulators of ribosome biogenesis (including CCDC86, ZNF236, and SPATA5L1), transcription (C7orf26), and mitochondrial respiration (TMEM242). In addition to assigning gene function, single-cell transcriptional phenotypes allow for in-depth dissection of complex cellular phenomena-from RNA processing to differentiation. We leverage this ability to systematically identify genetic drivers and consequences of aneuploidy and to discover an unantici
  105. 105Fully-automated and ultra-fast cell-type identification using specific marker combinations from single-cell transcriptomic data.Identification of cell populations often relies on manual annotation of cell clusters using established marker genes. However, the selection of marker genes is a time-consuming process that may lead to sub-optimal annotations as the markers must be informative of both the individual cell clusters and various cell types present in the sample. Here, we developed a computational platform, ScType, which enables a fully-automated and ultra-fast cell-type identification based solely on a given scRNA-seq data, along with a comprehensive cell marker database as background information. Using six scRNA-seq datasets from various human and mouse tissues, we show how ScType provides unbiased and accurate cell type annotations by guaranteeing the specificity of positive and negative marker genes across cell clusters and cell types. We also demonstrate how ScType distinguishes between healthy and malignant cell populations, based on single-cell calling of single-nucleotide variants, making it a versa
  106. 106Effect of the intratumoral microbiota on spatial and cellular heterogeneity in cancer.The tumour-associated microbiota is an intrinsic component of the tumour microenvironment across human cancer types 1,2 . Intratumoral host-microbiota studies have so far largely relied on bulk tissue analysis 1-3 , which obscures the spatial distribution and localized effect of the microbiota within tumours. Here, by applying in situ spatial-profiling technologies 4 and single-cell RNA sequencing 5 to oral squamous cell carcinoma and colorectal cancer, we reveal spatial, cellular and molecular host-microbe interactions. We adapted 10x Visium spatial transcriptomics to determine the identity and in situ location of intratumoral microbial communities within patient tissues. Using GeoMx digital spatial profiling 6 , we show that bacterial communities populate microniches that are less vascularized, highly immuno‑suppressive and associated with malignant cells with lower levels of Ki-67 as compared to bacteria-negative tumour regions. We developed a single-cell RNA-sequencing method that
  107. 107Probabilistic harmonization and annotation of single-cell transcriptomics data with deep generative models.As the number of single-cell transcriptomics datasets grows, the natural next step is to integrate the accumulating data to achieve a common ontology of cell types and states. However, it is not straightforward to compare gene expression levels across datasets and to automatically assign cell type labels in a new dataset based on existing annotations. In this manuscript, we demonstrate that our previously developed method, scVI, provides an effective and fully probabilistic approach for joint representation and analysis of scRNA-seq data, while accounting for uncertainty caused by biological and measurement noise. We also introduce single-cell ANnotation using Variational Inference (scANVI), a semi-supervised variant of scVI designed to leverage existing cell state annotations. We demonstrate that scVI and scANVI compare favorably to state-of-the-art methods for data integration and cell state annotation in terms of accuracy, scalability, and adaptability to challenging settings. In co
  108. 108Comparison and evaluation of statistical error models for scRNA-seq.Background Heterogeneity in single-cell RNA-seq (scRNA-seq) data is driven by multiple sources, including biological variation in cellular state as well as technical variation introduced during experimental processing. Deconvolving these effects is a key challenge for preprocessing workflows. Recent work has demonstrated the importance and utility of count models for scRNA-seq analysis, but there is a lack of consensus on which statistical distributions and parameter settings are appropriate. Results Here, we analyze 59 scRNA-seq datasets that span a wide range of technologies, systems, and sequencing depths in order to evaluate the performance of different error models. We find that while a Poisson error model appears appropriate for sparse datasets, we observe clear evidence of overdispersion for genes with sufficient sequencing depth in all biological systems, necessitating the use of a negative binomial model. Moreover, we find that the degree of overdispersion varies widely across
  109. 109hdWGCNA identifies co-expression networks in high-dimensional transcriptomics data.Biological systems are immensely complex, organized into a multi-scale hierarchy of functional units based on tightly regulated interactions between distinct molecules, cells, organs, and organisms. While experimental methods enable transcriptome-wide measurements across millions of cells, popular bioinformatic tools do not support systems-level analysis. Here we present hdWGCNA, a comprehensive framework for analyzing co-expression networks in high-dimensional transcriptomics data such as single-cell and spatial RNA sequencing (RNA-seq). hdWGCNA provides functions for network inference, gene module identification, gene enrichment analysis, statistical tests, and data visualization. Beyond conventional single-cell RNA-seq, hdWGCNA is capable of performing isoform-level network analysis using long-read single-cell data. We showcase hdWGCNA using data from autism spectrum disorder and Alzheimer's disease brain samples, identifying disease-relevant co-expression network modules. hdWGCNA i
  110. 110Redefining Tumor-Associated Macrophage Subpopulations and Functions in the Tumor Microenvironment.The immunosuppressive status of the tumor microenvironment (TME) remains poorly defined due to a lack of understanding regarding the function of tumor-associated macrophages (TAMs), which are abundant in the TME. TAMs are crucial drivers of tumor progression, metastasis, and resistance to therapy. Intra- and inter-tumoral spatial heterogeneities are potential keys to understanding the relationships between subpopulations of TAMs and their functions. Antitumor M1-like and pro-tumor M2-like TAMs coexist within tumors, and the opposing effects of these M1/M2 subpopulations on tumors directly impact current strategies to improve antitumor immune responses. Recent studies have found significant differences among monocytes or macrophages from distinct tumors, and other investigations have explored the existence of diverse TAM subsets at the molecular level. In this review, we discuss emerging evidence highlighting the redefinition of TAM subpopulations and functions in the TME and the possib
  111. 111Deciphering spatial domains from spatially resolved transcriptomics with an adaptive graph attention auto-encoder.Recent advances in spatially resolved transcriptomics have enabled comprehensive measurements of gene expression patterns while retaining the spatial context of the tissue microenvironment. Deciphering the spatial context of spots in a tissue needs to use their spatial information carefully. To this end, we develop a graph attention auto-encoder framework STAGATE to accurately identify spatial domains by learning low-dimensional latent embeddings via integrating spatial information and gene expression profiles. To better characterize the spatial similarity at the boundary of spatial domains, STAGATE adopts an attention mechanism to adaptively learn the similarity of neighboring spots, and an optional cell type-aware module through integrating the pre-clustering of gene expressions. We validate STAGATE on diverse spatial transcriptomics datasets generated by different platforms with different spatial resolutions. STAGATE could substantially improve the identification accuracy of spatial
  112. 112Cell type and gene expression deconvolution with BayesPrism enables Bayesian integrative analysis across bulk and single-cell RNA sequencing in oncology.Inferring single-cell compositions and their contributions to global gene expression changes from bulk RNA sequencing (RNA-seq) datasets is a major challenge in oncology. Here we develop Bayesian cell proportion reconstruction inferred using statistical marginalization (BayesPrism), a Bayesian method to predict cellular composition and gene expression in individual cell types from bulk RNA-seq, using patient-derived, scRNA-seq as prior information. We conduct integrative analyses in primary glioblastoma, head and neck squamous cell carcinoma and skin cutaneous melanoma to correlate cell type composition with clinical outcomes across tumor types, and explore spatial heterogeneity in malignant and nonmalignant cell states. We refine current cancer subtypes using gene expression annotation after exclusion of confounding nonmalignant cells. Finally, we identify genes whose expression in malignant cells correlates with macrophage infiltration, T cells, fibroblasts and endothelial cells acro
  113. 113A general and flexible method for signal extraction from single-cell RNA-seq data.Single-cell RNA-sequencing (scRNA-seq) is a powerful high-throughput technique that enables researchers to measure genome-wide transcription levels at the resolution of single cells. Because of the low amount of RNA present in a single cell, some genes may fail to be detected even though they are expressed; these genes are usually referred to as dropouts. Here, we present a general and flexible zero-inflated negative binomial model (ZINB-WaVE), which leads to low-dimensional representations of the data that account for zero inflation (dropouts), over-dispersion, and the count nature of the data. We demonstrate, with simulated and real data, that the model and its associated estimation procedure are able to give a more stable and accurate low-dimensional representation of the data than principal component analysis (PCA) and zero-inflated factor analysis (ZIFA), without the need for a preliminary normalization step.
  114. 114Single-cell sequencing of human midbrain reveals glial activation and a Parkinson-specific neuronal state.Idiopathic Parkinson's disease is characterized by a progressive loss of dopaminergic neurons, but the exact disease aetiology remains largely unknown. To date, Parkinson's disease research has mainly focused on nigral dopaminergic neurons, although recent studies suggest disease-related changes also in non-neuronal cells and in midbrain regions beyond the substantia nigra. While there is some evidence for glial involvement in Parkinson's disease, the molecular mechanisms remain poorly understood. The aim of this study was to characterize the contribution of all cell types of the midbrain to Parkinson's disease pathology by single-nuclei RNA sequencing and to assess the cell type-specific risk for Parkinson's disease using the latest genome-wide association study. We profiled >41 000 single-nuclei transcriptomes of post-mortem midbrain from six idiopathic Parkinson's disease patients and five age-/sex-matched controls. To validate our findings in a spatial context, we utilized immunola
  115. 115Fecal microbiota transplantation protects rotenone-induced Parkinson's disease mice via suppressing inflammation mediated by the lipopolysaccharide-TLR4 signaling pathway through the microbiota-gut-brain axis.Background Parkinson's disease (PD) is a prevalent neurodegenerative disorder, displaying not only well-known motor deficits but also gastrointestinal dysfunctions. Consistently, it has been increasingly evident that gut microbiota affects the communication between the gut and the brain in PD pathogenesis, known as the microbiota-gut-brain axis. As an approach to re-establishing a normal microbiota community, fecal microbiota transplantation (FMT) has exerted beneficial effects on PD in recent studies. Here, in this study, we established a chronic rotenone-induced PD mouse model to evaluate the protective effects of FMT treatment on PD and to explore the underlying mechanisms, which also proves the involvement of gut microbiota dysbiosis in PD pathogenesis via the microbiota-gut-brain axis. Results We demonstrated that gut microbiota dysbiosis induced by rotenone administration caused gastrointestinal function impairment and poor behavioral performances in the PD mice. Moreover, 16S RN
  116. 116Genome, transcriptome and proteome: the rise of omics data and their integration in biomedical sciences.Advances in the technologies and informatics used to generate and process large biological data sets (omics data) are promoting a critical shift in the study of biomedical sciences. While genomics, transcriptomics and proteinomics, coupled with bioinformatics and biostatistics, are gaining momentum, they are still, for the most part, assessed individually with distinct approaches generating monothematic rather than integrated knowledge. As other areas of biomedical sciences, including metabolomics, epigenomics and pharmacogenomics, are moving towards the omics scale, we are witnessing the rise of inter-disciplinary data integration strategies to support a better understanding of biological systems and eventually the development of successful precision medicine. This review cuts across the boundaries between genomics, transcriptomics and proteomics, summarizing how omics data are generated, analysed and shared, and provides an overview of the current strengths and weaknesses of this glo
  117. 117Spatiotemporal analysis of human intestinal development at single-cell resolution.Development of the human intestine is not well understood. Here, we link single-cell RNA sequencing and spatial transcriptomics to characterize intestinal morphogenesis through time. We identify 101 cell states including epithelial and mesenchymal progenitor populations and programs linked to key morphogenetic milestones. We describe principles of crypt-villus axis formation; neural, vascular, mesenchymal morphogenesis, and immune population of the developing gut. We identify the differentiation hierarchies of developing fibroblast and myofibroblast subtypes and describe diverse functions for these including as vascular niche cells. We pinpoint the origins of Peyer's patches and gut-associated lymphoid tissue (GALT) and describe location-specific immune programs. We use our resource to present an unbiased analysis of morphogen gradients that direct sequential waves of cellular differentiation and define cells and locations linked to rare developmental intestinal disorders. We compile a
  118. 118Single cell transcriptional and chromatin accessibility profiling redefine cellular heterogeneity in the adult human kidney.The integration of single cell transcriptome and chromatin accessibility datasets enables a deeper understanding of cell heterogeneity. We performed single nucleus ATAC (snATAC-seq) and RNA (snRNA-seq) sequencing to generate paired, cell-type-specific chromatin accessibility and transcriptional profiles of the adult human kidney. We demonstrate that snATAC-seq is comparable to snRNA-seq in the assignment of cell identity and can further refine our understanding of functional heterogeneity in the nephron. The majority of differentially accessible chromatin regions are localized to promoters and a significant proportion are closely associated with differentially expressed genes. Cell-type-specific enrichment of transcription factor binding motifs implicates the activation of NF-κB that promotes VCAM1 expression and drives transition between a subpopulation of proximal tubule epithelial cells. Our multi-omics approach improves the ability to detect unique cell states within the kidney and
  119. 119Single-cell proteomic and transcriptomic analysis of macrophage heterogeneity using SCoPE2.Background Macrophages are innate immune cells with diverse functional and molecular phenotypes. This diversity is largely unexplored at the level of single-cell proteomes because of the limitations of quantitative single-cell protein analysis. Results To overcome this limitation, we develop SCoPE2, which substantially increases quantitative accuracy and throughput while lowering cost and hands-on time by introducing automated and miniaturized sample preparation. These advances enable us to analyze the emergence of cellular heterogeneity as homogeneous monocytes differentiate into macrophage-like cells in the absence of polarizing cytokines. SCoPE2 quantifies over 3042 proteins in 1490 single monocytes and macrophages in 10 days of instrument time, and the quantified proteins allow us to discern single cells by cell type. Furthermore, the data uncover a continuous gradient of proteome states for the macrophages, suggesting that macrophage heterogeneity may emerge in the absence of pola
  120. 120Accurate and efficient detection of gene fusions from RNA sequencing data.The identification of gene fusions from RNA sequencing data is a routine task in cancer research and precision oncology. However, despite the availability of many computational tools, fusion detection remains challenging. Existing methods suffer from poor prediction accuracy and are computationally demanding. We developed Arriba, a novel fusion detection algorithm with high sensitivity and short runtime. When applied to a large collection of published pancreatic cancer samples ( n = 803), Arriba identified a variety of driver fusions, many of which affected druggable proteins, including ALK, BRAF, FGFR2, NRG1, NTRK1, NTRK3, RET, and ROS1. The fusions were significantly associated with KRAS wild-type tumors and involved proteins stimulating the MAPK signaling pathway, suggesting that they substitute for activating mutations in KRAS In addition, we confirmed the transforming potential of two novel fusions, RRBP1 - RAF1 and RASGRP1 - ATP1A1 , in cellular assays. These results show Arriba'
  121. 121Spatially informed clustering, integration, and deconvolution of spatial transcriptomics with GraphST.Spatial transcriptomics technologies generate gene expression profiles with spatial context, requiring spatially informed analysis tools for three key tasks, spatial clustering, multisample integration, and cell-type deconvolution. We present GraphST, a graph self-supervised contrastive learning method that fully exploits spatial transcriptomics data to outperform existing methods. It combines graph neural networks with self-supervised contrastive learning to learn informative and discriminative spot representations by minimizing the embedding distance between spatially adjacent spots and vice versa. We demonstrated GraphST on multiple tissue types and technology platforms. GraphST achieved 10% higher clustering accuracy and better delineated fine-grained tissue structures in brain and embryo tissues. GraphST is also the only method that can jointly analyze multiple tissue slices in vertical or horizontal integration while correcting batch effects. Lastly, GraphST demonstrated superior
  122. 122Screening cell-cell communication in spatial transcriptomics via collective optimal transport.Spatial transcriptomic technologies and spatially annotated single-cell RNA sequencing datasets provide unprecedented opportunities to dissect cell-cell communication (CCC). However, incorporation of the spatial information and complex biochemical processes required in the reconstruction of CCC remains a major challenge. Here, we present COMMOT (COMMunication analysis by Optimal Transport) to infer CCC in spatial transcriptomics, which accounts for the competition between different ligand and receptor species as well as spatial distances between cells. A collective optimal transport method is developed to handle complex molecular interactions and spatial constraints. Furthermore, we introduce downstream analysis tools to infer spatial signaling directionality and genes regulated by signaling using machine learning models. We apply COMMOT to simulation data and eight spatial datasets acquired with five different technologies to show its effectiveness and robustness in identifying spatia
  123. 123Single-cell RNA-seq reveals fibroblast heterogeneity and increased mesenchymal fibroblasts in human fibrotic skin diseases.Fibrotic skin disease represents a major global healthcare burden, characterized by fibroblast hyperproliferation and excessive accumulation of extracellular matrix. Fibroblasts are found to be heterogeneous in multiple fibrotic diseases, but fibroblast heterogeneity in fibrotic skin diseases is not well characterized. In this study, we explore fibroblast heterogeneity in keloid, a paradigm of fibrotic skin diseases, by using single-cell RNA-seq. Our results indicate that keloid fibroblasts can be divided into 4 subpopulations: secretory-papillary, secretory-reticular, mesenchymal and pro-inflammatory. Interestingly, the percentage of mesenchymal fibroblast subpopulation is significantly increased in keloid compared to normal scar. Functional studies indicate that mesenchymal fibroblasts are crucial for collagen overexpression in keloid. Increased mesenchymal fibroblast subpopulation is also found in another fibrotic skin disease, scleroderma, suggesting this is a broad mechanism for s
  124. 124Mapping transcriptomic vector fields of single cells.Single-cell (sc)RNA-seq, together with RNA velocity and metabolic labeling, reveals cellular states and transitions at unprecedented resolution. Fully exploiting these data, however, requires kinetic models capable of unveiling governing regulatory functions. Here, we introduce an analytical framework dynamo (https://github.com/aristoteleo/dynamo-release), which infers absolute RNA velocity, reconstructs continuous vector fields that predict cell fates, employs differential geometry to extract underlying regulations, and ultimately predicts optimal reprogramming paths and perturbation outcomes. We highlight dynamo's power to overcome fundamental limitations of conventional splicing-based RNA velocity analyses to enable accurate velocity estimations on a metabolically labeled human hematopoiesis scRNA-seq dataset. Furthermore, differential geometry analyses reveal mechanisms driving early megakaryocyte appearance and elucidate asymmetrical regulation within the PU.1-GATA1 circuit. Lever
  125. 125Molecular and spatial signatures of mouse brain aging at single-cell resolution.The diversity and complex organization of cells in the brain have hindered systematic characterization of age-related changes in its cellular and molecular architecture, limiting our ability to understand the mechanisms underlying its functional decline during aging. Here, we generated a high-resolution cell atlas of brain aging within the frontal cortex and striatum using spatially resolved single-cell transcriptomics and quantified changes in gene expression and spatial organization of major cell types in these regions over the mouse lifespan. We observed substantially more pronounced changes in cell state, gene expression, and spatial organization of non-neuronal cells over neurons. Our data revealed molecular and spatial signatures of glial and immune cell activation during aging, particularly enriched in the subcortical white matter, and identified both similarities and notable differences in cell-activation patterns induced by aging and systemic inflammatory challenge. These resu
  126. 126Single-cell transcriptomics reveals cell-type-specific diversification in human heart failure.Heart failure represents a major cause of morbidity and mortality worldwide. Single-cell transcriptomics have revolutionized our understanding of cell composition and associated gene expression. Through integrated analysis of single-cell and single-nucleus RNA-sequencing data generated from 27 healthy donors and 18 individuals with dilated cardiomyopathy, here we define the cell composition of the healthy and failing human heart. We identify cell-specific transcriptional signatures associated with age and heart failure and reveal the emergence of disease-associated cell states. Notably, cardiomyocytes converge toward common disease-associated cell states, whereas fibroblasts and myeloid cells undergo dramatic diversification. Endothelial cells and pericytes display global transcriptional shifts without changes in cell complexity. Collectively, our findings provide a comprehensive analysis of the cellular and transcriptomic landscape of human heart failure, identify cell type-specific t
  127. 127Single-cell RNA sequencing reveals functional heterogeneity of glioma-associated brain macrophages.Microglia are resident myeloid cells in the central nervous system (CNS) that control homeostasis and protect CNS from damage and infections. Microglia and peripheral myeloid cells accumulate and adapt tumor supporting roles in human glioblastomas that show prevalence in men. Cell heterogeneity and functional phenotypes of myeloid subpopulations in gliomas remain elusive. Here we show single-cell RNA sequencing (scRNA-seq) of CD11b + myeloid cells in naïve and GL261 glioma-bearing mice that reveal distinct profiles of microglia, infiltrating monocytes/macrophages and CNS border-associated macrophages. We demonstrate an unforeseen molecular heterogeneity among myeloid cells in naïve and glioma-bearing brains, validate selected marker proteins and show distinct spatial distribution of identified subsets in experimental gliomas. We find higher expression of MHCII encoding genes in glioma-activated male microglia, which was corroborated in bulk and scRNA-seq data from human diffuse gliomas
  128. 128Morphological diversity of single neurons in molecularly defined cell types.Dendritic and axonal morphology reflects the input and output of neurons and is a defining feature of neuronal types 1,2 , yet our knowledge of its diversity remains limited. Here, to systematically examine complete single-neuron morphologies on a brain-wide scale, we established a pipeline encompassing sparse labelling, whole-brain imaging, reconstruction, registration and analysis. We fully reconstructed 1,741 neurons from cortex, claustrum, thalamus, striatum and other brain regions in mice. We identified 11 major projection neuron types with distinct morphological features and corresponding transcriptomic identities. Extensive projectional diversity was found within each of these major types, on the basis of which some types were clustered into more refined subtypes. This diversity follows a set of generalizable principles that govern long-range axonal projections at different levels, including molecular correspondence, divergent or convergent projection, axon termination pattern,
  129. 129Single cell transcriptomic landscape of diabetic foot ulcers.Diabetic foot ulceration (DFU) is a devastating complication of diabetes whose pathogenesis remains incompletely understood. Here, we profile 174,962 single cells from the foot, forearm, and peripheral blood mononuclear cells using single-cell RNA sequencing. Our analysis shows enrichment of a unique population of fibroblasts overexpressing MMP1, MMP3, MMP11, HIF1A, CHI3L1, and TNFAIP6 and increased M1 macrophage polarization in the DFU patients with healing wounds. Further, analysis of spatially separated samples from the same patient and spatial transcriptomics reveal preferential localization of these healing associated fibroblasts toward the wound bed as compared to the wound edge or unwounded skin. Spatial transcriptomics also validates our findings of higher abundance of M1 macrophages in healers and M2 macrophages in non-healers. Our analysis provides deep insights into the wound healing microenvironment, identifying cell types that could be critical in promoting DFU healing, an
  130. 130Spatial epigenome-transcriptome co-profiling of mammalian tissues.Emerging spatial technologies, including spatial transcriptomics and spatial epigenomics, are becoming powerful tools for profiling of cellular states in the tissue context 1-5 . However, current methods capture only one layer of omics information at a time, precluding the possibility of examining the mechanistic relationship across the central dogma of molecular biology. Here, we present two technologies for spatially resolved, genome-wide, joint profiling of the epigenome and transcriptome by cosequencing chromatin accessibility and gene expression, or histone modifications (H3K27me3, H3K27ac or H3K4me3) and gene expression on the same tissue section at near-single-cell resolution. These were applied to embryonic and juvenile mouse brain, as well as adult human brain, to map how epigenetic mechanisms control transcriptional phenotype and cell dynamics in tissue. Although highly concordant tissue features were identified by either spatial epigenome or spatial transcriptome we also obs
  131. 131Design and computational analysis of single-cell RNA-sequencing experiments.Single-cell RNA-sequencing (scRNA-seq) has emerged as a revolutionary tool that allows us to address scientific questions that eluded examination just a few years ago. With the advantages of scRNA-seq come computational challenges that are just beginning to be addressed. In this article, we highlight the computational methods available for the design and analysis of scRNA-seq experiments, their advantages and disadvantages in various settings, the open questions for which novel methods are needed, and expected future developments in this exciting area.
  132. 132zUMIs - A fast and flexible pipeline to process RNA sequencing data with UMIs.Background Single-cell RNA-sequencing (scRNA-seq) experiments typically analyze hundreds or thousands of cells after amplification of the cDNA. The high throughput is made possible by the early introduction of sample-specific bar codes (BCs), and the amplification bias is alleviated by unique molecular identifiers (UMIs). Thus, the ideal analysis pipeline for scRNA-seq data needs to efficiently tabulate reads according to both BC and UMI. Findings zUMIs is a pipeline that can handle both known and random BCs and also efficiently collapse UMIs, either just for exon mapping reads or for both exon and intron mapping reads. If BC annotation is missing, zUMIs can accurately detect intact cells from the distribution of sequencing reads. Another unique feature of zUMIs is the adaptive downsampling function that facilitates dealing with hugely varying library sizes but also allows the user to evaluate whether the library has been sequenced to saturation. To illustrate the utility of zUMIs, we
  133. 133Impaired local intrinsic immunity to SARS-CoV-2 infection in severe COVID-19.SARS-CoV-2 infection can cause severe respiratory COVID-19. However, many individuals present with isolated upper respiratory symptoms, suggesting potential to constrain viral pathology to the nasopharynx. Which cells SARS-CoV-2 primarily targets and how infection influences the respiratory epithelium remains incompletely understood. We performed scRNA-seq on nasopharyngeal swabs from 58 healthy and COVID-19 participants. During COVID-19, we observe expansion of secretory, loss of ciliated, and epithelial cell repopulation via deuterosomal cell expansion. In mild and moderate COVID-19, epithelial cells express anti-viral/interferon-responsive genes, while cells in severe COVID-19 have muted anti-viral responses despite equivalent viral loads. SARS-CoV-2 RNA + host-target cells are highly heterogenous, including developing ciliated, interferon-responsive ciliated, AZGP1 high goblet, and KRT13 + "hillock"-like cells, and we identify genes associated with susceptibility, resistance, or in
  134. 134Cancer-associated fibroblast classification in single-cell and spatial proteomics data.Cancer-associated fibroblasts (CAFs) are a diverse cell population within the tumour microenvironment, where they have critical effects on tumour evolution and patient prognosis. To define CAF phenotypes, we analyse a single-cell RNA sequencing (scRNA-seq) dataset of over 16,000 stromal cells from tumours of 14 breast cancer patients, based on which we define and functionally annotate nine CAF phenotypes and one class of pericytes. We validate this classification system in four additional cancer types and use highly multiplexed imaging mass cytometry on matched breast cancer samples to confirm our defined CAF phenotypes at the protein level and to analyse their spatial distribution within tumours. This general CAF classification scheme will allow comparison of CAF phenotypes across studies, facilitate analysis of their functional roles, and potentially guide development of new treatment strategies in the future.
  135. 135Integration of spatial and single-cell transcriptomic data elucidates mouse organogenesis.Molecular profiling of single cells has advanced our knowledge of the molecular basis of development. However, current approaches mostly rely on dissociating cells from tissues, thereby losing the crucial spatial context of regulatory processes. Here, we apply an image-based single-cell transcriptomics method, sequential fluorescence in situ hybridization (seqFISH), to detect mRNAs for 387 target genes in tissue sections of mouse embryos at the 8-12 somite stage. By integrating spatial context and multiplexed transcriptional measurements with two single-cell transcriptome atlases, we characterize cell types across the embryo and demonstrate that spatially resolved expression of genes not profiled by seqFISH can be imputed. We use this high-resolution spatial map to characterize fundamental steps in the patterning of the midbrain-hindbrain boundary (MHB) and the developing gut tube. We uncover axes of cell differentiation that are not apparent from single-cell RNA-sequencing (scRNA-seq)
  136. 136SCODE: an efficient regulatory network inference algorithm from single-cell RNA-Seq during differentiation.Motivation The analysis of RNA-Seq data from individual differentiating cells enables us to reconstruct the differentiation process and the degree of differentiation (in pseudo-time) of each cell. Such analyses can reveal detailed expression dynamics and functional relationships for differentiation. To further elucidate differentiation processes, more insight into gene regulatory networks is required. The pseudo-time can be regarded as time information and, therefore, single-cell RNA-Seq data are time-course data with high time resolution. Although time-course data are useful for inferring networks, conventional inference algorithms for such data suffer from high time complexity when the number of samples and genes is large. Therefore, a novel algorithm is necessary to infer networks from single-cell RNA-Seq during differentiation. Results In this study, we developed the novel and efficient algorithm SCODE to infer regulatory networks, based on ordinary differential equations. We appli
  137. 137An integrated multi-omics analysis identifies prognostic molecular subtypes of non-muscle-invasive bladder cancer.The molecular landscape in non-muscle-invasive bladder cancer (NMIBC) is characterized by large biological heterogeneity with variable clinical outcomes. Here, we perform an integrative multi-omics analysis of patients diagnosed with NMIBC (n = 834). Transcriptomic analysis identifies four classes (1, 2a, 2b and 3) reflecting tumor biology and disease aggressiveness. Both transcriptome-based subtyping and the level of chromosomal instability provide independent prognostic value beyond established prognostic clinicopathological parameters. High chromosomal instability, p53-pathway disruption and APOBEC-related mutations are significantly associated with transcriptomic class 2a and poor outcome. RNA-derived immune cell infiltration is associated with chromosomally unstable tumors and enriched in class 2b. Spatial proteomics analysis confirms the higher infiltration of class 2b tumors and demonstrates an association between higher immune cell infiltration and lower recurrence rates. Final
  138. 138Comprehensive metabolomics expands precision medicine for triple-negative breast cancer.Metabolic reprogramming is a hallmark of cancer. However, systematic characterizations of metabolites in triple-negative breast cancer (TNBC) are still lacking. Our study profiled the polar metabolome and lipidome in 330 TNBC samples and 149 paired normal breast tissues to construct a large metabolomic atlas of TNBC. Combining with previously established transcriptomic and genomic data of the same cohort, we conducted a comprehensive analysis linking TNBC metabolome to genomics. Our study classified TNBCs into three distinct metabolomic subgroups: C1, characterized by the enrichment of ceramides and fatty acids; C2, featured with the upregulation of metabolites related to oxidation reaction and glycosyl transfer; and C3, having the lowest level of metabolic dysregulation. Based on this newly developed metabolomic dataset, we refined previous TNBC transcriptomic subtypes and identified some crucial subtype-specific metabolites as potential therapeutic targets. The transcriptomic luminal
  139. 139Spatially restricted drivers and transitional cell populations cooperate with the microenvironment in untreated and chemo-resistant pancreatic cancer.Pancreatic ductal adenocarcinoma is a lethal disease with limited treatment options and poor survival. We studied 83 spatial samples from 31 patients (11 treatment-naïve and 20 treated) using single-cell/nucleus RNA sequencing, bulk-proteogenomics, spatial transcriptomics and cellular imaging. Subpopulations of tumor cells exhibited signatures of proliferation, KRAS signaling, cell stress and epithelial-to-mesenchymal transition. Mapping mutations and copy number events distinguished tumor populations from normal and transitional cells, including acinar-to-ductal metaplasia and pancreatic intraepithelial neoplasia. Pathology-assisted deconvolution of spatial transcriptomic data identified tumor and transitional subpopulations with distinct histological features. We showed coordinated expression of TIGIT in exhausted and regulatory T cells and Nectin in tumor cells. Chemo-resistant samples contain a threefold enrichment of inflammatory cancer-associated fibroblasts that upregulate metal
  140. 140Single-cell genomics and spatial transcriptomics: Discovery of novel cell states and cellular interactions in liver physiology and disease biology.Transcriptome analysis enables the study of gene expression in human tissues and is a valuable tool to characterise liver function and gene expression dynamics during liver disease, as well as to identify prognostic markers or signatures, and to facilitate discovery of new therapeutic targets. In contrast to whole tissue RNA sequencing analysis, single-cell RNA-sequencing (scRNA-seq) and spatial transcriptomics enables the study of transcriptional activity at the single cell or spatial level. ScRNA-seq has paved the way for the discovery of previously unknown cell types and subtypes in normal and diseased liver, facilitating the study of rare cells (such as liver progenitor cells) and the functional roles of non-parenchymal cells in chronic liver disease and cancer. By adding spatial information to scRNA-seq data, spatial transcriptomics has transformed our understanding of tissue functional organisation and cell-to-cell interactions in situ. These approaches have recently been applied
  141. 141Seamless integration of image and molecular analysis for spatial transcriptomics workflows.Background Recent advancements in in situ gene expression technologies constitute a new and rapidly evolving field of transcriptomics. With the recent launch of the 10x Genomics Visium platform, such methods have started to become widely adopted. The experimental protocol is conducted on individual tissue sections collected from a larger tissue sample. The two-dimensional nature of this data requires multiple consecutive sections to be collected from the sample in order to construct a comprehensive three-dimensional map of the tissue. However, there is currently no software available that lets the user process the images, align stacked experiments, and finally visualize them together in 3D to create a holistic view of the tissue. Results We have developed an R package named STUtility that takes 10x Genomics Visium data as input and provides features to perform standardized data transformations, alignment of multiple tissue sections, regional annotation, and visualizations of the combin
  142. 142Simultaneous trimodal single-cell measurement of transcripts, epitopes, and chromatin accessibility using TEA-seq.Single-cell measurements of cellular characteristics have been instrumental in understanding the heterogeneous pathways that drive differentiation, cellular responses to signals, and human disease. Recent advances have allowed paired capture of protein abundance and transcriptomic state, but a lack of epigenetic information in these assays has left a missing link to gene regulation. Using the heterogeneous mixture of cells in human peripheral blood as a test case, we developed a novel scATAC-seq workflow that increases signal-to-noise and allows paired measurement of cell surface markers and chromatin accessibility: integrated cellular indexing of chromatin landscape and epitopes, called ICICLE-seq. We extended this approach using a droplet-based multiomics platform to develop a trimodal assay that simultaneously measures transcriptomics (scRNA-seq), epitopes, and chromatin accessibility (scATAC-seq) from thousands of single cells, which we term TEA-seq. Together, these multimodal sing
  143. 143Spatial transcriptomics reveals distinct and conserved tumor core and edge architectures that predict survival and targeted therapy response.The spatial organization of the tumor microenvironment has a profound impact on biology and therapy response. Here, we perform an integrative single-cell and spatial transcriptomic analysis on HPV-negative oral squamous cell carcinoma (OSCC) to comprehensively characterize malignant cells in tumor core (TC) and leading edge (LE) transcriptional architectures. We show that the TC and LE are characterized by unique transcriptional profiles, neighboring cellular compositions, and ligand-receptor interactions. We demonstrate that the gene expression profile associated with the LE is conserved across different cancers while the TC is tissue specific, highlighting common mechanisms underlying tumor progression and invasion. Additionally, we find our LE gene signature is associated with worse clinical outcomes while TC gene signature is associated with improved prognosis across multiple cancer types. Finally, using an in silico modeling approach, we describe spatially-regulated patterns of ce
  144. 144Single-cell sequencing techniques from individual to multiomics analyses.Here, we review single-cell sequencing techniques for individual and multiomics profiling in single cells. We mainly describe single-cell genomic, epigenomic, and transcriptomic methods, and examples of their applications. For the integration of multilayered data sets, such as the transcriptome data derived from single-cell RNA sequencing and chromatin accessibility data derived from single-cell ATAC-seq, there are several computational integration methods. We also describe single-cell experimental methods for the simultaneous measurement of two or more omics layers. We can achieve a detailed understanding of the basic molecular profiles and those associated with disease in each cell by utilizing a large number of single-cell sequencing techniques and the accumulated data sets.
  145. 145DeepST: identifying spatial domains in spatial transcriptomics by deep learning.Recent advances in spatial transcriptomics (ST) have brought unprecedented opportunities to understand tissue organization and function in spatial context. However, it is still challenging to precisely dissect spatial domains with similar gene expression and histology in situ. Here, we present DeepST, an accurate and universal deep learning framework to identify spatial domains, which performs better than the existing state-of-the-art methods on benchmarking datasets of the human dorsolateral prefrontal cortex. Further testing on a breast cancer ST dataset, we showed that DeepST can dissect spatial domains in cancer tissue at a finer scale. Moreover, DeepST can achieve not only effective batch integration of ST data generated from multiple batches or different technologies, but also expandable capabilities for processing other spatial omics data. Together, our results demonstrate that DeepST has the exceptional capacity for identifying spatial domains, making it a desirable tool to gai
  146. 146Spatially aware dimension reduction for spatial transcriptomics.Spatial transcriptomics are a collection of genomic technologies that have enabled transcriptomic profiling on tissues with spatial localization information. Analyzing spatial transcriptomic data is computationally challenging, as the data collected from various spatial transcriptomic technologies are often noisy and display substantial spatial correlation across tissue locations. Here, we develop a spatially-aware dimension reduction method, SpatialPCA, that can extract a low dimensional representation of the spatial transcriptomics data with biological signal and preserved spatial correlation structure, thus unlocking many existing computational tools previously developed in single-cell RNAseq studies for tailored analysis of spatial transcriptomics. We illustrate the benefits of SpatialPCA for spatial domain detection and explores its utility for trajectory inference on the tissue and for high-resolution spatial map construction. In the real data applications, SpatialPCA identifies
  147. 147Reference-free cell type deconvolution of multi-cellular pixel-resolution spatially resolved transcriptomics data.Recent technological advancements have enabled spatially resolved transcriptomic profiling but at multi-cellular pixel resolution, thereby hindering the identification of cell-type-specific spatial patterns and gene expression variation. To address this challenge, we develop STdeconvolve as a reference-free approach to deconvolve underlying cell types comprising such multi-cellular pixel resolution spatial transcriptomics (ST) datasets. Using simulated as well as real ST datasets from diverse spatial transcriptomics technologies comprising a variety of spatial resolutions such as Spatial Transcriptomics, 10X Visium, DBiT-seq, and Slide-seq, we show that STdeconvolve can effectively recover cell-type transcriptional profiles and their proportional representation within pixels without reliance on external single-cell transcriptomics references. STdeconvolve provides comparable performance to existing reference-based methods when suitable single-cell references are available, as well as p
  148. 148The role of ABCA7 in Alzheimer's disease: evidence from genomics, transcriptomics and methylomics.Genome-wide association studies (GWAS) originally identified ATP-binding cassette, sub-family A, member 7 (ABCA7), as a novel risk gene of Alzheimer's disease (AD). Since then, accumulating evidence from in vitro, in vivo, and human-based studies has corroborated and extended this association, promoting ABCA7 as one of the most important risk genes of both early-onset and late-onset AD, harboring both common and rare risk variants with relatively large effect on AD risk. Within this review, we provide a comprehensive assessment of the literature on ABCA7, with a focus on AD-related human -omics studies (e.g. genomics, transcriptomics, and methylomics). In European and African American populations, indirect ABCA7 GWAS associations are explained by expansion of an ABCA7 variable number tandem repeat (VNTR), and a common premature termination codon (PTC) variant, respectively. Rare ABCA7 PTC variants are strongly enriched in AD patients, and some of these have displayed inheritance patter
  149. 149Spatial Transcriptomics to define transcriptional patterns of zonation and structural components in the mouse liver.Reconstruction of heterogeneity through single cell transcriptional profiling has greatly advanced our understanding of the spatial liver transcriptome in recent years. However, global transcriptional differences across lobular units remain elusive in physical space. Here, we apply Spatial Transcriptomics to perform transcriptomic analysis across sectioned liver tissue. We confirm that the heterogeneity in this complex tissue is predominantly determined by lobular zonation. By introducing novel computational approaches, we enable transcriptional gradient measurements between tissue structures, including several lobules in a variety of orientations. Further, our data suggests the presence of previously transcriptionally uncharacterized structures within liver tissue, contributing to the overall spatial heterogeneity of the organ. This study demonstrates how comprehensive spatial transcriptomic technologies can be used to delineate extensive spatial gene expression patterns in the liver,
  150. 150Shotgun transcriptome, spatial omics, and isothermal profiling of SARS-CoV-2 infection reveals unique host responses, viral diversification, and drug interactions.In less than nine months, the Severe Acute Respiratory Syndrome Coronavirus 2 (SARS-CoV-2) killed over a million people, including >25,000 in New York City (NYC) alone. The COVID-19 pandemic caused by SARS-CoV-2 highlights clinical needs to detect infection, track strain evolution, and identify biomarkers of disease course. To address these challenges, we designed a fast (30-minute) colorimetric test (LAMP) for SARS-CoV-2 infection from naso/oropharyngeal swabs and a large-scale shotgun metatranscriptomics platform (total-RNA-seq) for host, viral, and microbial profiling. We applied these methods to clinical specimens gathered from 669 patients in New York City during the first two months of the outbreak, yielding a broad molecular portrait of the emerging COVID-19 disease. We find significant enrichment of a NYC-distinctive clade of the virus (20C), as well as host responses in interferon, ACE, hematological, and olfaction pathways. In addition, we use 50,821 patient records to find t
  151. 151Clinical implementation of RNA sequencing for Mendelian disease diagnostics.Background Lack of functional evidence hampers variant interpretation, leaving a large proportion of individuals with a suspected Mendelian disorder without genetic diagnosis after whole genome or whole exome sequencing (WES). Research studies advocate to further sequence transcriptomes to directly and systematically probe gene expression defects. However, collection of additional biopsies and establishment of lab workflows, analytical pipelines, and defined concepts in clinical interpretation of aberrant gene expression are still needed for adopting RNA sequencing (RNA-seq) in routine diagnostics. Methods We implemented an automated RNA-seq protocol and a computational workflow with which we analyzed skin fibroblasts of 303 individuals with a suspected mitochondrial disease that previously underwent WES. We also assessed through simulations how aberrant expression and mono-allelic expression tests depend on RNA-seq coverage. Results We detected on average 12,500 genes per sample inclu
  152. 152High-resolution alignment of single-cell and spatial transcriptomes with CytoSPACE.Recent studies have emphasized the importance of single-cell spatial biology, yet available assays for spatial transcriptomics have limited gene recovery or low spatial resolution. Here we introduce CytoSPACE, an optimization method for mapping individual cells from a single-cell RNA sequencing atlas to spatial expression profiles. Across diverse platforms and tissue types, we show that CytoSPACE outperforms previous methods with respect to noise tolerance and accuracy, enabling tissue cartography at single-cell resolution.
  153. 153SM-Omics is an automated platform for high-throughput spatial multi-omics.The spatial organization of cells and molecules plays a key role in tissue function in homeostasis and disease. Spatial transcriptomics has recently emerged as a key technique to capture and positionally barcode RNAs directly in tissues. Here, we advance the application of spatial transcriptomics at scale, by presenting Spatial Multi-Omics (SM-Omics) as a fully automated, high-throughput all-sequencing based platform for combined and spatially resolved transcriptomics and antibody-based protein measurements. SM-Omics uses DNA-barcoded antibodies, immunofluorescence or a combination thereof, to scale and combine spatial transcriptomics and spatial antibody-based multiplex protein detection. SM-Omics allows processing of up to 64 in situ spatial reactions or up to 96 sequencing-ready libraries, of high complexity, in a ~2 days process. We demonstrate SM-Omics in the mouse brain, spleen and colorectal cancer model, showing its broad utility as a high-throughput platform for spatial multi-
  154. 154Delineating the dynamic evolution from preneoplasia to invasive lung adenocarcinoma by integrating single-cell RNA sequencing and spatial transcriptomics.The cell ecology and spatial niche implicated in the dynamic and sequential process of lung adenocarcinoma (LUAD) from adenocarcinoma in situ (AIS) to minimally invasive adenocarcinoma (MIA) and subsequent invasive adenocarcinoma (IAC) have not yet been elucidated. Here, we performed an integrative analysis of single-cell RNA sequencing (scRNA-seq) and spatial transcriptomics (ST) to characterize the cell atlas of the invasion trajectory of LUAD. We found that the UBE2C + cancer cell subpopulation constantly increased during the invasive process of LUAD with remarkable elevation in IAC, and its spatial distribution was in the peripheral cancer region of the IAC, representing a more malignant phenotype. Furthermore, analysis of the TME cell type subpopulation showed a constant decrease in mast cells, monocytes, and lymphatic endothelial cells, which were implicated in the whole process of invasive LUAD, accompanied by an increase in NK cells and MALT B cells from AIS to MIA and an incre
  155. 155Single-cell RNA sequencing and spatial transcriptomics reveal cancer-associated fibroblasts in glioblastoma with protumoral effects.Cancer-associated fibroblasts (CAFs) were presumed absent in glioblastoma given the lack of brain fibroblasts. Serial trypsinization of glioblastoma specimens yielded cells with CAF morphology and single-cell transcriptomic profiles based on their lack of copy number variations (CNVs) and elevated individual cell CAF probability scores derived from the expression of 9 CAF markers and absence of 5 markers from non-CAF stromal cells sharing features with CAFs. Cells without CNVs and with high CAF probability scores were identified in single-cell RNA-Seq of 12 patient glioblastomas. Pseudotime reconstruction revealed that immature CAFs evolved into subtypes, with mature CAFs expressing actin alpha 2, smooth muscle (ACTA2). Spatial transcriptomics from 16 patient glioblastomas confirmed CAF proximity to mesenchymal glioblastoma stem cells (GSCs), endothelial cells, and M2 macrophages. CAFs were chemotactically attracted to GSCs, and CAFs enriched GSCs. We created a resource of inferred cro
  156. 156Spatial transcriptomics reveals discrete tumour microenvironments and autocrine loops within ovarian cancer subclones.High-grade serous ovarian carcinoma (HGSOC) is genetically unstable and characterised by the presence of subclones with distinct genotypes. Intratumoural heterogeneity is linked to recurrence, chemotherapy resistance, and poor prognosis. Here, we use spatial transcriptomics to identify HGSOC subclones and study their association with infiltrating cell populations. Visium spatial transcriptomics reveals multiple tumour subclones with different copy number alterations present within individual tumour sections. These subclones differentially express various ligands and receptors and are predicted to differentially associate with different stromal and immune cell populations. In one sample, CosMx single molecule imaging reveals subclones differentially associating with immune cell populations, fibroblasts, and endothelial cells. Cell-to-cell communication analysis identifies subclone-specific signalling to stromal and immune cells and multiple subclone-specific autocrine loops. Our study h
  157. 157Spatial transcriptomics identifies molecular niche dysregulation associated with distal lung remodeling in pulmonary fibrosis.Large-scale changes in the structure and cellular makeup of the distal lung are a hallmark of pulmonary fibrosis (PF), but the spatial contexts that contribute to disease pathogenesis have remained uncertain. Using image-based spatial transcriptomics, we analyzed the gene expression of 1.6 million cells from 35 unique lungs. Through complementary cell-based and innovative cell-agnostic analyses, we characterized the localization of PF-emergent cell types, established the cellular and molecular basis of classical PF histopathologic features and identified a diversity of distinct molecularly defined spatial niches in control and PF lungs. Using machine learning and trajectory analysis to segment and rank airspaces on a gradient of remodeling severity, we identified compositional and molecular changes associated with progressive distal lung pathology, beginning with alveolar epithelial dysregulation and culminating with changes in macrophage polarization. Together, these results provide a
  158. 158Microglial mechanisms drive amyloid-β clearance in immunized patients with Alzheimer's disease.Alzheimer's disease (AD) therapies utilizing amyloid-β (Aβ) immunization have shown potential in clinical trials. Yet, the mechanisms driving Aβ clearance in the immunized AD brain remain unclear. Here, we use spatial transcriptomics to explore the effects of both active and passive Aβ immunization in the AD brain. We compare actively immunized patients with AD with nonimmunized patients with AD and neurologically healthy controls, identifying distinct microglial states associated with Aβ clearance. Using high-resolution spatial transcriptomics alongside single-cell RNA sequencing, we delve deeper into the transcriptional pathways involved in Aβ removal after lecanemab treatment. We uncover spatially distinct microglial responses that vary by brain region. Our analysis reveals upregulation of the triggering receptor expressed on myeloid cells 2 (TREM2) and apolipoprotein E (APOE) in microglia across immunization approaches, which correlate positively with antibody responses and Aβ remo
  159. 159Leveraging deep single-soma RNA sequencing to explore the neural basis of human somatosensation.The versatility of somatosensation arises from heterogeneous dorsal root ganglion (DRG) neurons. However, soma transcriptomes of individual human (h)DRG neurons-critical information to decipher their functions-are lacking due to technical difficulties. In this study, we isolated somata from individual hDRG neurons and conducted deep RNA sequencing (RNA-seq) to detect, on average, over 9,000 unique genes per neuron, and we identified 16 neuronal types. These results were corroborated and validated by spatial transcriptomics and RNAscope in situ hybridization. Cross-species analyses revealed divergence among potential pain-sensing neurons and the likely existence of human-specific neuronal types. Molecular-profile-informed microneurography recordings revealed temperature-sensing properties across human sensory afferent types. In summary, by employing single-soma deep RNA-seq and spatial transcriptomics, we generated an hDRG neuron atlas, which provides insights into human somatosensory p
  160. 160Single-Cell RNA Sequencing with Spatial Transcriptomics of Cancer Tissues.Single-cell RNA sequencing (RNA-seq) techniques can perform analysis of transcriptome at the single-cell level and possess an unprecedented potential for exploring signatures involved in tumor development and progression. These techniques can perform sequence analysis of transcripts with a better resolution that could increase understanding of the cellular diversity found in the tumor microenvironment and how the cells interact with each other in complex heterogeneous cancerous tissues. Identifying the changes occurring in the genome and transcriptome in the spatial context is considered to increase knowledge of molecular factors fueling cancers. It may help develop better monitoring strategies and innovative approaches for cancer treatment. Recently, there has been a growing trend in the integration of RNA-seq techniques with contemporary omics technologies to study the tumor microenvironment. There has been a realization that this area of research has a huge scope of application in t
  161. 161THItoGene: a deep learning method for predicting spatial transcriptomics from histological images.Spatial transcriptomics unveils the complex dynamics of cell regulation and transcriptomes, but it is typically cost-prohibitive. Predicting spatial gene expression from histological images via artificial intelligence offers a more affordable option, yet existing methods fall short in extracting deep-level information from pathological images. In this paper, we present THItoGene, a hybrid neural network that utilizes dynamic convolutional and capsule networks to adaptively sense potential molecular signals in histological images for exploring the relationship between high-resolution pathology image phenotypes and regulation of gene expression. A comprehensive benchmark evaluation using datasets from human breast cancer and cutaneous squamous cell carcinoma has demonstrated the superior performance of THItoGene in spatial gene expression prediction. Moreover, THItoGene has demonstrated its capacity to decipher both the spatial context and enrichment signals within specific tissue region
  162. 162Exploring inflammatory signatures in arthritic joint biopsies with Spatial Transcriptomics.Lately it has become possible to analyze transcriptomic profiles in tissue sections with retained cellular context. We aimed to explore synovial biopsies from rheumatoid arthritis (RA) and spondyloarthritis (SpA) patients, using Spatial Transcriptomics (ST) as a proof of principle approach for unbiased mRNA studies at the site of inflammation in these chronic inflammatory diseases. Synovial tissue biopsies from affected joints were studied with ST. The transcriptome data was subjected to differential gene expression analysis (DEA), pathway analysis, immune cell type identification using Xcell analysis and validation with immunohistochemistry (IHC). The ST technology allows selective analyses on areas of interest, thus we analyzed morphologically distinct areas of mononuclear cell infiltrates. The top differentially expressed genes revealed an adaptive immune response profile and T-B cell interactions in RA, while in SpA, the profiles implicate functions associated with tissue repair. W
  163. 163Spatial transcriptomics technology in cancer research.In recent years, spatial transcriptomics (ST) technologies have developed rapidly and have been widely used in constructing spatial tissue atlases and characterizing spatiotemporal heterogeneity of cancers. Currently, ST has been used to profile spatial heterogeneity in multiple cancer types. Besides, ST is a benefit for identifying and comprehensively understanding special spatial areas such as tumor interface and tertiary lymphoid structures (TLSs), which exhibit unique tumor microenvironments (TMEs). Therefore, ST has also shown great potential to improve pathological diagnosis and identify novel prognostic factors in cancer. This review presents recent advances and prospects of applications on cancer research based on ST technologies as well as the challenges.
  164. 164Spatial Transcriptomic Technologies.Spatial transcriptomic technologies enable measurement of expression levels of genes systematically throughout tissue space, deepening our understanding of cellular organizations and interactions within tissues as well as illuminating biological insights in neuroscience, developmental biology and a range of diseases, including cancer. A variety of spatial technologies have been developed and/or commercialized, differing in spatial resolution, sensitivity, multiplexing capability, throughput and coverage. In this paper, we review key enabling spatial transcriptomic technologies and their applications as well as the perspective of the techniques and new emerging technologies that are developed to address current limitations of spatial methodologies. In addition, we describe how spatial transcriptomics data can be integrated with other omics modalities, complementing other methods in deciphering cellar interactions and phenotypes within tissues as well as providing novel insight into tiss
  165. 165The promising application of cell-cell interaction analysis in cancer from single-cell and spatial transcriptomics.Cell-cell interactions instruct cell fate and function. These interactions are hijacked to promote cancer development. Single-cell transcriptomics and spatial transcriptomics have become powerful new tools for researchers to profile the transcriptional landscape of cancer at unparalleled genetic depth. In this review, we discuss the rapidly growing array of computational tools to infer cell-cell interactions from non-spatial single-cell RNA-sequencing and the limited but growing number of methods for spatial transcriptomics data. Downstream analyses of these computational tools and applications to cancer studies are highlighted. We finish by suggesting several directions for further extensions that anticipate the increasing availability of multi-omics cancer data.
  166. 166Spatial Transcriptomics: Emerging Technologies in Tissue Gene Expression Profiling.In this Perspective, we discuss the current status and advances in spatial transcriptomics technologies, which allow high-resolution mapping of gene expression in intact cell and tissue samples. Spatial transcriptomics enables the creation of high-resolution maps of gene expression patterns within their native spatial context, adding an extra layer of information to the bulk sequencing data. Spatial transcriptomics has expanded significantly in recent years and is making a notable impact on a range of fields, including tissue architecture, developmental biology, cancer, and neurodegenerative and infectious diseases. The latest advancements in spatial transcriptomics have resulted in the development of highly multiplexed methods, transcriptomic-wide analysis, and single-cell resolution utilizing diverse technological approaches. In this Perspective, we provide a detailed analysis of the molecular foundations behind the main spatial transcriptomics technologies, including methods based o
  167. 167Advances and Challenges in Spatial Transcriptomics for Developmental Biology.Development from single cells to multicellular tissues and organs involves more than just the exact replication of cells, which is known as differentiation. The primary focus of research into the mechanism of differentiation has been differences in gene expression profiles between individual cells. However, it has predominantly been conducted at low throughput and bulk levels, challenging the efforts to understand molecular mechanisms of differentiation during the developmental process in animals and humans. During the last decades, rapid methodological advancements in genomics facilitated the ability to study developmental processes at a genome-wide level and finer resolution. Particularly, sequencing transcriptomes at single-cell resolution, enabled by single-cell RNA-sequencing (scRNA-seq), was a breath-taking innovation, allowing scientists to gain a better understanding of differentiation and cell lineage during the developmental process. However, single-cell isolation during scRN
  168. 168Spatial Transcriptomics: Technical Aspects of Recent Developments and Their Applications in Neuroscience and Cancer Research.Spatial transcriptomics is a newly emerging field that enables high-throughput investigation of the spatial localization of transcripts and related analyses in various applications for biological systems. By transitioning from conventional biological studies to "in situ" biology, spatial transcriptomics can provide transcriptome-scale spatial information. Currently, the ability to simultaneously characterize gene expression profiles of cells and relevant cellular environment is a paradigm shift for biological studies. In this review, recent progress in spatial transcriptomics and its applications in neuroscience and cancer studies are highlighted. Technical aspects of existing technologies and future directions of new developments (as of March 2023), computational analysis of spatial transcriptome data, application notes in neuroscience and cancer studies, and discussions regarding future directions of spatial multi-omics and their expanding roles in biomedical applications are emphasi
  169. 169Spatial transcriptomics in cancer research and potential clinical impact: a narrative review.Spatial transcriptomics (ST) provides novel insights into the tumor microenvironment (TME). ST allows the quantification and illustration of gene expression profiles in the spatial context of tissues, including both the cancer cells and the microenvironment in which they are found. In cancer research, ST has already provided novel insights into cancer metastasis, prognosis, and immunotherapy responsiveness. The clinical precision oncology application of next-generation sequencing (NGS) and RNA profiling of tumors relies on bulk methods that lack spatial context. The ability to preserve spatial information is now possible, as it allows us to capture tumor heterogeneity and multifocality. In this narrative review, we summarize precision oncology, discuss tumor sequencing in the clinic, and review the available ST research methods, including seqFISH, MERFISH (Vizgen), CosMx SMI (NanoString), Xenium (10x), Visium (10x), Stereo-seq (STOmics), and GeoMx DSP (NanoString). We then review the c
  170. 170A practical guide for choosing an optimal spatial transcriptomics technology from seven major commercially available options.Spatial transcriptomics technology enables the mapping of gene expression within tissues, allowing researchers to visualize the spatial distribution of RNA molecules and gain insights into cellular organization, interactions, and functions in their native environments. A variety of spatial technologies are now commercially available, each offering distinct technical parameters such as cellular resolution, detection sensitivity, gene coverage, and throughput. This wide range of options can make it challenges or create confusion for researchers to select the most appropriate platform for their specific research objectives. In this paper, we will analyze and compare seven major commercially available spatial platforms to guide researchers in choosing the most suitable option for their needs.
  171. 171Evaluation of statistical methods for normalization and differential expression in mRNA-Seq experiments.Background High-throughput sequencing technologies, such as the Illumina Genome Analyzer, are powerful new tools for investigating a wide range of biological and medical questions. Statistical and computational methods are key for drawing meaningful and accurate conclusions from the massive and complex datasets generated by the sequencers. We provide a detailed evaluation of statistical methods for normalization and differential expression (DE) analysis of Illumina transcriptome sequencing (mRNA-Seq) data. Results We compare statistical methods for detecting genes that are significantly DE between two types of biological samples and find that there are substantial differences in how the test statistics handle low-count genes. We evaluate how DE results are affected by features of the sequencing platform, such as, varying gene lengths, base-calling calibration method (with and without phi X control lane), and flow-cell/library preparation effects. We investigate the impact of the read c
  172. 172Computational analysis of bacterial RNA-Seq data.Recent advances in high-throughput RNA sequencing (RNA-seq) have enabled tremendous leaps forward in our understanding of bacterial transcriptomes. However, computational methods for analysis of bacterial transcriptome data have not kept pace with the large and growing data sets generated by RNA-seq technology. Here, we present new algorithms, specific to bacterial gene structures and transcriptomes, for analysis of RNA-seq data. The algorithms are implemented in an open source software system called Rockhopper that supports various stages of bacterial RNA-seq data analysis, including aligning sequencing reads to a genome, constructing transcriptome maps, quantifying transcript abundance, testing for differential gene expression, determining operon structures and visualizing results. We demonstrate the performance of Rockhopper using 2.1 billion sequenced reads from 75 RNA-seq experiments conducted with Escherichia coli, Neisseria gonorrhoeae, Salmonella enterica, Streptococcus pyogene
  173. 173APOE modulates microglial immunometabolism in response to age, amyloid pathology, and inflammatory challenge.The E4 allele of Apolipoprotein E (APOE) is associated with both metabolic dysfunction and a heightened pro-inflammatory response: two findings that may be intrinsically linked through the concept of immunometabolism. Here, we combined bulk, single-cell, and spatial transcriptomics with cell-specific and spatially resolved metabolic analyses in mice expressing human APOE to systematically address the role of APOE across age, neuroinflammation, and AD pathology. RNA sequencing (RNA-seq) highlighted immunometabolic changes across the APOE4 glial transcriptome, specifically in subsets of metabolically distinct microglia enriched in the E4 brain during aging or following an inflammatory challenge. E4 microglia display increased Hif1α expression and a disrupted tricarboxylic acid (TCA) cycle and are inherently pro-glycolytic, while spatial transcriptomics and mass spectrometry imaging highlight an E4-specific response to amyloid that is characterized by widespread alterations in lipid metab
  174. 174Spatial transcriptomics reveals niche-specific enrichment and vulnerabilities of radial glial stem-like cells in malignant gliomas.Diffuse midline glioma-H3K27M mutant (DMG) and glioblastoma (GBM) are the most lethal brain tumors that primarily occur in pediatric and adult patients, respectively. Both tumors exhibit significant heterogeneity, shaped by distinct genetic/epigenetic drivers, transcriptional programs including RNA splicing, and microenvironmental cues in glioma niches. However, the spatial organization of cellular states and niche-specific regulatory programs remain to be investigated. Here, we perform a spatial profiling of DMG and GBM combining short- and long-read spatial transcriptomics, and single-cell transcriptomic datasets. We identify clinically relevant transcriptional programs, RNA isoform diversity, and multi-cellular ecosystems across different glioma niches. We find that while the tumor core enriches for oligodendrocyte precursor-like cells, radial glial stem-like (RG-like) cells are enriched in the neuron-rich invasive niche in both DMG and GBM. Further, we identify niche-specific regul
  175. 175Integration of spatial and single-cell transcriptomics localizes epithelial cell-immune cross-talk in kidney injury.Single-cell sequencing studies have characterized the transcriptomic signature of cell types within the kidney. However, the spatial distribution of acute kidney injury (AKI) is regional and affects cells heterogeneously. We first optimized coordination of spatial transcriptomics and single-nuclear sequencing data sets, mapping 30 dominant cell types to a human nephrectomy. The predicted cell-type spots corresponded with the underlying histopathology. To study the implications of AKI on transcript expression, we then characterized the spatial transcriptomic signature of 2 murine AKI models: ischemia/reperfusion injury (IRI) and cecal ligation puncture (CLP). Localized regions of reduced overall expression were associated with injury pathways. Using single-cell sequencing, we deconvoluted the signature of each spatial transcriptomic spot, identifying patterns of colocalization between immune and epithelial cells. Neutrophils infiltrated the renal medulla in the ischemia model. Atf3 was
  176. 176Spatial cellular architecture predicts prognosis in glioblastoma.Intra-tumoral heterogeneity and cell-state plasticity are key drivers for the therapeutic resistance of glioblastoma. Here, we investigate the association between spatial cellular organization and glioblastoma prognosis. Leveraging single-cell RNA-seq and spatial transcriptomics data, we develop a deep learning model to predict transcriptional subtypes of glioblastoma cells from histology images. Employing this model, we phenotypically analyze 40 million tissue spots from 410 patients and identify consistent associations between tumor architecture and prognosis across two independent cohorts. Patients with poor prognosis exhibit higher proportions of tumor cells expressing a hypoxia-induced transcriptional program. Furthermore, a clustering pattern of astrocyte-like tumor cells is associated with worse prognosis, while dispersion and connection of the astrocytes with other transcriptional subtypes correlate with decreased risk. To validate these results, we develop a separate deep lear
  177. 177BASS: multi-scale and multi-sample analysis enables accurate cell type clustering and spatial domain detection in spatial transcriptomic studies.Spatial transcriptomic studies are reaching single-cell spatial resolution, with data often collected from multiple tissue sections. Here, we present a computational method, BASS, that enables multi-scale and multi-sample analysis for single-cell resolution spatial transcriptomics. BASS performs cell type clustering at the single-cell scale and spatial domain detection at the tissue regional scale, with the two tasks carried out simultaneously within a Bayesian hierarchical modeling framework. We illustrate the benefits of BASS through comprehensive simulations and applications to three datasets. The substantial power gain brought by BASS allows us to reveal accurate transcriptomic and cellular landscape in both cortex and hypothalamus.
  178. 178The spatial transcriptomic landscape of the healing mouse intestine following damage.The intestinal barrier is composed of a complex cell network defining highly compartmentalized and specialized structures. Here, we use spatial transcriptomics to define how the transcriptomic landscape is spatially organized in the steady state and healing murine colon. At steady state conditions, we demonstrate a previously unappreciated molecular regionalization of the colon, which dramatically changes during mucosal healing. Here, we identified spatially-organized transcriptional programs defining compartmentalized mucosal healing, and regions with dominant wired pathways. Furthermore, we showed that decreased p53 activation defined areas with increased presence of proliferating epithelial stem cells. Finally, we mapped transcriptomics modules associated with human diseases demonstrating the translational potential of our dataset. Overall, we provide a publicly available resource defining principles of transcriptomic regionalization of the colon during mucosal healing and a framewo
  179. 179Single-cell and spatial transcriptomics identify a macrophage population associated with skeletal muscle fibrosis.Macrophages are essential for skeletal muscle homeostasis, but how their dysregulation contributes to the development of fibrosis in muscle disease remains unclear. Here, we used single-cell transcriptomics to determine the molecular attributes of dystrophic and healthy muscle macrophages. We identified six clusters and unexpectedly found that none corresponded to traditional definitions of M1 or M2 macrophages. Rather, the predominant macrophage signature in dystrophic muscle was characterized by high expression of fibrotic factors, galectin-3 (gal-3) and osteopontin ( Spp1 ). Spatial transcriptomics, computational inferences of intercellular communication, and in vitro assays indicated that macrophage-derived Spp1 regulates stromal progenitor differentiation. Gal-3 + macrophages were chronically activated in dystrophic muscle, and adoptive transfer assays showed that the gal-3 + phenotype was the dominant molecular program induced within the dystrophic milieu. Gal-3 + macrophages wer
  180. 180STalign: Alignment of spatial transcriptomics data using diffeomorphic metric mapping.Spatial transcriptomics (ST) technologies enable high throughput gene expression characterization within thin tissue sections. However, comparing spatial observations across sections, samples, and technologies remains challenging. To address this challenge, we develop STalign to align ST datasets in a manner that accounts for partially matched tissue sections and other local non-linear distortions using diffeomorphic metric mapping. We apply STalign to align ST datasets within and across technologies as well as to align ST datasets to a 3D common coordinate framework. We show that STalign achieves high gene expression and cell-type correspondence across matched spatial locations that is significantly improved over landmark-based affine alignments. Applying STalign to align ST datasets of the mouse brain to the 3D common coordinate framework from the Allen Brain Atlas, we highlight how STalign can be used to lift over brain region annotations and enable the interrogation of compositiona
  181. 181Single-Nucleus RNA Sequencing and Spatial Transcriptomics Reveal the Immunological Microenvironment of Cervical Squamous Cell Carcinoma.The effective treatment of advanced cervical cancer remains challenging. Herein, single-nucleus RNA sequencing (snRNA-seq) and SpaTial enhanced resolution omics-sequencing (Stereo-seq) are used to investigate the immunological microenvironment of cervical squamous cell carcinoma (CSCC). The expression levels of most immune suppressive genes in the tumor and inflammation areas of CSCC are not significantly higher than those in the non-cancer samples, except for LGALS9 and IDO1. Stronger signals of CD56 + NK cells and immature dendritic cells are found in the hypermetabolic tumor areas, whereas more eosinophils, immature B cells, and Treg cells are found in the hypometabolic tumor areas. Moreover, a cluster of pro-tumorigenic cancer-associated myofibroblasts (myCAFs) are identified. The myCAFs may support the growth and metastasis of tumors by inhibiting lymphocyte infiltration and remodeling of the tumor extracellular matrix. Furthermore, these myCAFs are associated with poorer survival
  182. 182Combined Single-Cell and Spatial Transcriptomics Reveal the Metabolic Evolvement of Breast Cancer during Early Dissemination.Breast cancer is now the most frequently diagnosed malignancy, and metastasis remains the leading cause of death in breast cancer. However, little is known about the dynamic changes during the evolvement of dissemination. In this study, 65 968 cells from four patients with breast cancer and paired metastatic axillary lymph nodes are profiled using single-cell RNA sequencing (scRNA-seq) and spatial transcriptomics. A disseminated cancer cell cluster with high levels of oxidative phosphorylation (OXPHOS), including the upregulation of cytochrome C oxidase subunit 6C and dehydrogenase/reductase 2, is identified. The transition between glycolysis and OXPHOS when dissemination initiates is noticed. Furthermore, this distinct cell cluster is distributed along the tumor's leading edge. The findings here are verified in three different cohorts of breast cancer patients and an external scRNA-seq dataset, which includes eight patients with breast cancer and paired metastatic axillary lymph nodes
  183. 183Spatial transcriptomics reveals molecular dysfunction associated with cortical Lewy pathology.A key hallmark of Parkinson's disease (PD) is Lewy pathology. Composed of α-synuclein, Lewy pathology is found both in dopaminergic neurons that modulate motor function, and cortical regions that control cognitive function. Recent work has established the molecular identity of dopaminergic neurons susceptible to death, but little is known about cortical neurons susceptible to Lewy pathology or molecular changes induced by aggregates. In the current study, we use spatial transcriptomics to capture whole transcriptome signatures from cortical neurons with α-synuclein pathology compared to neurons without pathology. We find, both in PD and related PD dementia, dementia with Lewy bodies and in the pre-formed fibril α-synucleinopathy mouse model, that specific classes of excitatory neurons are vulnerable to developing Lewy pathology. Further, we identify conserved gene expression changes in aggregate-bearing neurons that we designate the Lewy-associated molecular dysfunction from aggregates
  184. 184Simultaneous CRISPR screening and spatial transcriptomics reveal intracellular, intercellular, and functional transcriptional circuits.Pooled optical screens have enabled the study of cellular interactions, morphology, or dynamics at massive scale, but they have not yet leveraged the power of highly plexed single-cell resolved transcriptomic readouts to inform molecular pathways. Here, we present a combination of imaging spatial transcriptomics with parallel optical detection of in situ amplified guide RNAs (Perturb-FISH). Perturb-FISH recovers intracellular effects that are consistent with single-cell RNA-sequencing-based readouts of perturbation effects (Perturb-seq) in a screen of lipopolysaccharide response in cultured monocytes, and it uncovers intercellular and density-dependent regulation of the innate immune response. Similarly, in three-dimensional xenograft models, Perturb-FISH identifies tumor-immune interactions altered by genetic knockout. When paired with a functional readout in a separate screen of autism spectrum disorder risk genes in human-induced pluripotent stem cell (hIPSC) astrocytes, Perturb-FIS
  185. 185Integration Analysis of Single-Cell Multi-Omics Reveals Prostate Cancer Heterogeneity.Prostate cancer (PCa) is an extensive heterogeneous disease with a complex cellular ecosystem in the tumor microenvironment (TME). However, the manner in which heterogeneity is shaped by tumors and stromal cells, or vice versa, remains poorly understood. In this study, single-cell RNA sequencing, spatial transcriptomics, and bulk ATAC-sequence are integrated from a series of patients with PCa and healthy controls. A stemness subset of club cells marked with SOX9 high AR low expression is identified, which is markedly enriched after neoadjuvant androgen-deprivation therapy (ADT). Furthermore, a subset of CD8 + CXCR6 + T cells that function as effector T cells is markedly reduced in patients with malignant PCa. For spatial transcriptome analysis, machine learning and computational intelligence are comprehensively utilized to identify the cellular diversity of prostate cancer cells and cell-cell communication in situ. Macrophage and neutrophil state transitions along the trajectory of can
  186. 186BIDCell: Biologically-informed self-supervised learning for segmentation of subcellular spatial transcriptomics data.Recent advances in subcellular imaging transcriptomics platforms have enabled high-resolution spatial mapping of gene expression, while also introducing significant analytical challenges in accurately identifying cells and assigning transcripts. Existing methods grapple with cell segmentation, frequently leading to fragmented cells or oversized cells that capture contaminated expression. To this end, we present BIDCell, a self-supervised deep learning-based framework with biologically-informed loss functions that learn relationships between spatially resolved gene expression and cell morphology. BIDCell incorporates cell-type data, including single-cell transcriptomics data from public repositories, with cell morphology information. Using a comprehensive evaluation framework consisting of metrics in five complementary categories for cell segmentation performance, we demonstrate that BIDCell outperforms other state-of-the-art methods according to many metrics across a variety of tissue
  187. 187Single cell and spatial transcriptomics highlight the interaction of club-like cells with immunosuppressive myeloid cells in prostate cancer.Prostate cancer treatment resistance is a significant challenge facing the field. Genomic and transcriptomic profiling have partially elucidated the mechanisms through which cancer cells escape treatment, but their relation toward the tumor microenvironment (TME) remains elusive. Here we present a comprehensive transcriptomic landscape of the prostate TME at multiple points in the standard treatment timeline employing single-cell RNA-sequencing and spatial transcriptomics data from 120 patients. We identify club-like cells as a key epithelial cell subtype that acts as an interface between the prostate and the immune system. Tissue areas enriched with club-like cells have depleted androgen signaling and upregulated expression of luminal progenitor cell markers. Club-like cells display a senescence-associated secretory phenotype and their presence is linked to increased polymorphonuclear myeloid-derived suppressor cell (PMN-MDSC) activity. Our results indicate that club-like cells are as
  188. 188Multiscale topology classifies cells in subcellular spatial transcriptomics.Spatial transcriptomics measures in situ gene expression at millions of locations within a tissue 1 , hitherto with some trade-off between transcriptome depth, spatial resolution and sample size 2 . Although integration of image-based segmentation has enabled impactful work in this context, it is limited by imaging quality and tissue heterogeneity. By contrast, recent array-based technologies offer the ability to measure the entire transcriptome at subcellular resolution across large samples 3-6 . Presently, there exist no approaches for cell type identification that directly leverage this information to annotate individual cells. Here we propose a multiscale approach to automatically classify cell types at this subcellular level, using both transcriptomic information and spatial context. We showcase this on both targeted and whole-transcriptome spatial platforms, improving cell classification and morphology for human kidney tissue and pinpointing individual sparsely distributed renal
  189. 189Spatial Transcriptomic and Metabolomic Landscapes of Oral Submucous Fibrosis-Derived Oral Squamous Cell Carcinoma and its Tumor Microenvironment.In South and Southeast Asia, the habit of chewing betel nuts is prevalent, which leads to oral submucous fibrosis (OSF). OSF is a well-established precancerous lesion, and a portion of OSF cases eventually progress to oral squamous cell carcinoma (OSCC). However, the specific molecular mechanisms underlying the malignant transformation of OSCC from OSF are poorly understood. In this study, the leading-edge techniques of Spatial Transcriptomics (ST) and Spatial Metabolomics (SM) are integrated to obtain spatial location information of cancer cells, fibroblasts, and immune cells, as well as the transcriptomic and metabolomic landscapes in OSF-derived OSCC tissues. This work reveals for the first time that some OSF-derived OSCC cells undergo partial epithelial-mesenchymal transition (pEMT) within the in situ carcinoma (ISC) region, eventually acquiring fibroblast-like phenotypes and participating in collagen deposition. Complex interactions among epithelial cells, fibroblasts, and immune
  190. 190Deciphering cell-cell communication at single-cell resolution for spatial transcriptomics with subgraph-based graph attention network.The inference of cell-cell communication (CCC) is crucial for a better understanding of complex cellular dynamics and regulatory mechanisms in biological systems. However, accurately inferring spatial CCCs at single-cell resolution remains a significant challenge. To address this issue, we present a versatile method, called DeepTalk, to infer spatial CCC at single-cell resolution by integrating single-cell RNA sequencing (scRNA-seq) data and spatial transcriptomics (ST) data. DeepTalk utilizes graph attention network (GAT) to integrate scRNA-seq and ST data, which enables accurate cell-type identification for single-cell ST data and deconvolution for spot-based ST data. Then, DeepTalk can capture the connections among cells at multiple levels using subgraph-based GAT, and further achieve spatially resolved CCC inference at single-cell resolution. DeepTalk achieves excellent performance in discovering meaningful spatial CCCs on multiple cross-platform datasets, which demonstrates its su
  191. 191Spatial transcriptomics reveals human cortical layer and area specification.The human cerebral cortex is composed of six layers and dozens of areas that are molecularly and structurally distinct 1-4 . Although single-cell transcriptomic studies have advanced the molecular characterization of human cortical development, a substantial gap exists owing to the loss of spatial context during cell dissociation 5-8 . Here we used multiplexed error-robust fluorescence in situ hybridization (MERFISH) 9 , augmented with deep-learning-based nucleus segmentation, to examine the molecular, cellular and cytoarchitectural development of the human fetal cortex with spatially resolved single-cell resolution. Our extensive spatial atlas, encompassing more than 18 million single cells, spans eight cortical areas across seven developmental time points. We uncovered the early establishment of the six-layer structure, identifiable by the laminar distribution of excitatory neuron subtypes, 3 months before the emergence of cytoarchitectural layers. Notably, we discovered two distinct
  192. 192Analysis and Visualization of Spatial Transcriptomic Data.Human and animal tissues consist of heterogeneous cell types that organize and interact in highly structured manners. Bulk and single-cell sequencing technologies remove cells from their original microenvironments, resulting in a loss of spatial information. Spatial transcriptomics is a recent technological innovation that measures transcriptomic information while preserving spatial information. Spatial transcriptomic data can be generated in several ways. RNA molecules are measured by in situ sequencing, in situ hybridization, or spatial barcoding to recover original spatial coordinates. The inclusion of spatial information expands the range of possibilities for analysis and visualization, and spurred the development of numerous novel methods. In this review, we summarize the core concepts of spatial genomics technology and provide a comprehensive review of current analysis and visualization methods for spatial transcriptomics.
  193. 193Isoform Age - Splice Isoform Profiling Using Long-Read Technologies.Alternative splicing (AS) of RNA is a key mechanism that results in the expression of multiple transcript isoforms from single genes and leads to an increase in the complexity of both the transcriptome and proteome. Regulation of AS is critical for the correct functioning of many biological pathways, while disruption of AS can be directly pathogenic in diseases such as cancer or cause risk for complex disorders. Current short-read sequencing technologies achieve high read depth but are limited in their ability to resolve complex isoforms. In this review we examine how long-read sequencing (LRS) technologies can address this challenge by covering the entire RNA sequence in a single read and thereby distinguish isoform changes that could impact RNA regulation or protein function. Coupling LRS with technologies such as single cell sequencing, targeted sequencing and spatial transcriptomics is producing a rapidly expanding suite of technological approaches to profile alternative splicing a
  194. 194Spatial transcriptomics in development and disease.The proper functioning of diverse biological systems depends on the spatial organization of their cells, a critical factor for biological processes like shaping intricate tissue functions and precisely determining cell fate. Nonetheless, conventional bulk or single-cell RNA sequencing methods were incapable of simultaneously capturing both gene expression profiles and the spatial locations of cells. Hence, a multitude of spatially resolved technologies have emerged, offering a novel dimension for investigating regional gene expression, spatial domains, and interactions between cells. Spatial transcriptomics (ST) is a method that maps gene expression in tissue while preserving spatial information. It can reveal cellular heterogeneity, spatial organization and functional interactions in complex biological systems. ST can also complement and integrate with other omics methods to provide a more comprehensive and holistic view of biological systems at multiple levels of resolution. Since th
  195. 195Spatial transcriptomics in neuroscience.The brain is one of the most complex living tissue types and is composed of an exceptional diversity of cell types displaying unique functional connectivity. Single-cell RNA sequencing (scRNA-seq) can be used to efficiently map the molecular identities of the various cell types in the brain by providing the transcriptomic profiles of individual cells isolated from the tissue. However, the lack of spatial context in scRNA-seq prevents a comprehensive understanding of how different configurations of cell types give rise to specific functions in individual brain regions and how each distinct cell is connected to form a functional unit. To understand how the various cell types contribute to specific brain functions, it is crucial to correlate the identities of individual cells obtained through scRNA-seq with their spatial information in intact tissue. Spatial transcriptomics (ST) can resolve the complex spatial organization of cell types in the brain and their connectivity. Various ST tool
  196. 196Subcellular Transcriptomics and Proteomics: A Comparative Methods Review.The internal environment of cells is molecularly crowded, which requires spatial organization via subcellular compartmentalization. These compartments harbor specific conditions for molecules to perform their biological functions, such as coordination of the cell cycle, cell survival, and growth. This compartmentalization is also not static, with molecules trafficking between these subcellular neighborhoods to carry out their functions. For example, some biomolecules are multifunctional, requiring an environment with differing conditions or interacting partners, and others traffic to export such molecules. Aberrant localization of proteins or RNA species has been linked to many pathological conditions, such as neurological, cancer, and pulmonary diseases. Differential expression studies in transcriptomics and proteomics are relatively common, but the majority have overlooked the importance of subcellular information. In addition, subcellular transcriptomics and proteomics data do not a
  197. 197Single-cell omics traces the heterogeneity of prostate cancer cells and the tumor microenvironment.Prostate cancer is one of the more heterogeneous tumour types. In recent years, with the rapid development of single-cell sequencing and spatial transcriptome technologies, researchers have gained a more intuitive and comprehensive understanding of the heterogeneity of prostate cancer. Tumour-associated epithelial cells; cancer-associated fibroblasts; the complexity of the immune microenvironment, and the heterogeneity of the spatial distribution of tumour cells and other cancer-promoting molecules play a crucial role in the growth, invasion, and metastasis of prostate cancer. Single-cell multi-omics biotechnology, especially single-cell transcriptome sequencing, reveals the expression level of single cells with higher resolution and finely dissects the molecular characteristics of different tumour cells. We reviewed the recent literature on prostate cancer cells, focusing on single-cell RNA sequencing. And we analysed the heterogeneity and spatial distribution differences of different
  198. 198A review of spatial profiling technologies for characterizing the tumor microenvironment in immuno-oncology.Interpreting the mechanisms and principles that govern gene activity and how these genes work according to -their cellular distribution in organisms has profound implications for cancer research. The latest technological advancements, such as imaging-based approaches and next-generation single-cell sequencing technologies, have established a platform for spatial transcriptomics to systematically quantify the expression of all or most genes in the entire tumor microenvironment and explore an array of disease milieus, particularly in tumors. Spatial profiling technologies permit the study of transcriptional activity at the spatial or single-cell level. This multidimensional classification of the transcriptomic and proteomic signatures of tumors, especially the associated immune and stromal cells, facilitates evaluation of tumor heterogeneity, details of the evolutionary trajectory of each tumor, and multifaceted interactions between each tumor cell and its microenvironment. Therefore, sp
  199. 199Emerging artificial intelligence applications in Spatial Transcriptomics analysis.Spatial transcriptomics (ST) has advanced significantly in the last few years. Such advancement comes with the urgent need for novel computational methods to handle the unique challenges of ST data analysis. Many artificial intelligence (AI) methods have been developed to utilize various machine learning and deep learning techniques for computational ST analysis. This review provides a comprehensive and up-to-date survey of current AI methods for ST analysis.
  200. 200A spatial sequencing atlas of age-induced changes in the lung during influenza infection.Influenza virus infection causes increased morbidity and mortality in the elderly. Aging impairs the immune response to influenza, both intrinsically and because of altered interactions with endothelial and pulmonary epithelial cells. To characterize these changes, we performed single-cell RNA sequencing (scRNA-seq), spatial transcriptomics, and bulk RNA sequencing (bulk RNA-seq) on lung tissue from young and aged female mice at days 0, 3, and 9 post-influenza infection. Our analyses identified dozens of key genes differentially expressed in kinetic, age-dependent, and cell type-specific manners. Aged immune cells exhibited altered inflammatory, memory, and chemotactic profiles. Aged endothelial cells demonstrated characteristics of reduced vascular wound healing and a prothrombotic state. Spatial transcriptomics identified novel profibrotic and antifibrotic markers expressed by epithelial and non-epithelial cells, highlighting the complex networks that promote fibrosis in aged lungs.
  201. 201PATRIC, the bacterial bioinformatics database and analysis resource.The Pathosystems Resource Integration Center (PATRIC) is the all-bacterial Bioinformatics Resource Center (BRC) (http://www.patricbrc.org). A joint effort by two of the original National Institute of Allergy and Infectious Diseases-funded BRCs, PATRIC provides researchers with an online resource that stores and integrates a variety of data types [e.g. genomics, transcriptomics, protein-protein interactions (PPIs), three-dimensional protein structures and sequence typing data] and associated metadata. Datatypes are summarized for individual genomes and across taxonomic levels. All genomes in PATRIC, currently more than 10,000, are consistently annotated using RAST, the Rapid Annotations using Subsystems Technology. Summaries of different data types are also provided for individual genes, where comparisons of different annotations are available, and also include available transcriptomic data. PATRIC provides a variety of ways for researchers to find data of interest and a private workspa
  202. 202A high-resolution anatomical atlas of the transcriptome in the mouse embryo.Ascertaining when and where genes are expressed is of crucial importance to understanding or predicting the physiological role of genes and proteins and how they interact to form the complex networks that underlie organ development and function. It is, therefore, crucial to determine on a genome-wide level, the spatio-temporal gene expression profiles at cellular resolution. This information is provided by colorimetric RNA in situ hybridization that can elucidate expression of genes in their native context and does so at cellular resolution. We generated what is to our knowledge the first genome-wide transcriptome atlas by RNA in situ hybridization of an entire mammalian organism, the developing mouse at embryonic day 14.5. This digital transcriptome atlas, the Eurexpress atlas (http://www.eurexpress.org), consists of a searchable database of annotated images that can be interactively viewed. We generated anatomy-based expression profiles for over 18,000 coding genes and over 400 micro
  203. 203RNA-Seq Atlas of Glycine max: a guide to the soybean transcriptome.Background Next generation sequencing is transforming our understanding of transcriptomes. It can determine the expression level of transcripts with a dynamic range of over six orders of magnitude from multiple tissues, developmental stages or conditions. Patterns of gene expression provide insight into functions of genes with unknown annotation. Results The RNA Seq-Atlas presented here provides a record of high-resolution gene expression in a set of fourteen diverse tissues. Hierarchical clustering of transcriptional profiles for these tissues suggests three clades with similar profiles: aerial, underground and seed tissues. We also investigate the relationship between gene structure and gene expression and find a correlation between gene length and expression. Additionally, we find dramatic tissue-specific gene expression of both the most highly-expressed genes and the genes specific to legumes in seed development and nodule tissues. Analysis of the gene expression profiles of over 2
  204. 204The mouse blood-brain barrier transcriptome: a new resource for understanding the development and function of brain endothelial cells.The blood-brain barrier (BBB) maintains brain homeostasis and limits the entry of toxins and pathogens into the brain. Despite its importance, little is known about the molecular mechanisms regulating the development and function of this crucial barrier. In this study we have developed methods to highly purify and gene profile endothelial cells from different tissues, and by comparing the transcriptional profile of brain endothelial cells with those purified from the liver and lung, we have generated a comprehensive resource of transcripts that are enriched in the BBB forming endothelial cells of the brain. Through this comparison we have identified novel tight junction proteins, transporters, metabolic enzymes, signaling components, and unknown transcripts whose expression is enriched in central nervous system (CNS) endothelial cells. This analysis has identified that RXRalpha signaling cascade is specifically enriched at the BBB, implicating this pathway in regulating this vital barr
  205. 205Effects of diet on resource utilization by a model human gut microbiota containing Bacteroides cellulosilyticus WH2, a symbiont with an extensive glycobiome.The human gut microbiota is an important metabolic organ, yet little is known about how its individual species interact, establish dominant positions, and respond to changes in environmental factors such as diet. In this study, gnotobiotic mice were colonized with an artificial microbiota comprising 12 sequenced human gut bacterial species and fed oscillating diets of disparate composition. Rapid, reproducible, and reversible changes in the structure of this assemblage were observed. Time-series microbial RNA-Seq analyses revealed staggered functional responses to diet shifts throughout the assemblage that were heavily focused on carbohydrate and amino acid metabolism. High-resolution shotgun metaproteomics confirmed many of these responses at a protein level. One member, Bacteroides cellulosilyticus WH2, proved exceptionally fit regardless of diet. Its genome encoded more carbohydrate active enzymes than any previously sequenced member of the Bacteroidetes. Transcriptional profiling i
  206. 206Spatial transcriptomics analysis of neoadjuvant cabozantinib and nivolumab in advanced hepatocellular carcinoma identifies independent mechanisms of resistance and recurrence.Background Novel immunotherapy combination therapies have improved outcomes for patients with hepatocellular carcinoma (HCC), but responses are limited to a subset of patients. Little is known about the inter- and intra-tumor heterogeneity in cellular signaling networks within the HCC tumor microenvironment (TME) that underlie responses to modern systemic therapy. Methods We applied spatial transcriptomics (ST) profiling to characterize the tumor microenvironment in HCC resection specimens from a prospective clinical trial of neoadjuvant cabozantinib, a multi-tyrosine kinase inhibitor that primarily blocks VEGF, and nivolumab, a PD-1 inhibitor in which 5 out of 15 patients were found to have a pathologic response at the time of resection. Results ST profiling demonstrated that the TME of responding tumors was enriched for immune cells and cancer-associated fibroblasts (CAF) with pro-inflammatory signaling relative to the non-responders. The enriched cancer-immune interactions in respon
  207. 207Comprehensive visualization of cell-cell interactions in single-cell and spatial transcriptomics with NICHES.Motivation Recent years have seen the release of several toolsets that reveal cell-cell interactions from single-cell data. However, all existing approaches leverage mean celltype gene expression values, and do not preserve the single-cell fidelity of the original data. Here, we present NICHES (Niche Interactions and Communication Heterogeneity in Extracellular Signaling), a tool to explore extracellular signaling at the truly single-cell level. Results NICHES allows embedding of ligand-receptor signal proxies to visualize heterogeneous signaling archetypes within cell clusters, between cell clusters and across experimental conditions. When applied to spatial transcriptomic data, NICHES can be used to reflect local cellular microenvironment. NICHES can operate with any list of ligand-receptor signaling mechanisms, is compatible with existing single-cell packages, and allows rapid, flexible analysis of cell-cell signaling at single-cell resolution. Availability and implementation NICHES
  208. 208Probabilistic embedding, clustering, and alignment for integrating spatial transcriptomics data with PRECAST.Spatially resolved transcriptomics involves a set of emerging technologies that enable the transcriptomic profiling of tissues with the physical location of expressions. Although a variety of methods have been developed for data integration, most of them are for single-cell RNA-seq datasets without consideration of spatial information. Thus, methods that can integrate spatial transcriptomics data from multiple tissue slides, possibly from multiple individuals, are needed. Here, we present PRECAST, a data integration method for multiple spatial transcriptomics datasets with complex batch effects and/or biological effects between slides. PRECAST unifies spatial factor analysis simultaneously with spatial clustering and embedding alignment, while requiring only partially shared cell/domain clusters across datasets. Using both simulated and four real datasets, we show improved cell/domain detection with outstanding visualization, and the estimated aligned embeddings and cell/domain labels
  209. 209Clonal relations in the mouse brain revealed by single-cell and spatial transcriptomics.The mammalian brain contains many specialized cells that develop from a thin sheet of neuroepithelial progenitor cells. Single-cell transcriptomics revealed hundreds of molecularly diverse cell types in the nervous system, but the lineage relationships between mature cell types and progenitor cells are not well understood. Here we show in vivo barcoding of early progenitors to simultaneously profile cell phenotypes and clonal relations in the mouse brain using single-cell and spatial transcriptomics. By reconstructing thousands of clones, we discovered fate-restricted progenitor cells in the mouse hippocampal neuroepithelium and show that microglia are derived from few primitive myeloid precursors that massively expand to generate widely dispersed progeny. We combined spatial transcriptomics with clonal barcoding and disentangled migration patterns of clonally related cells in densely labeled tissue sections. Our approach enables high-throughput dense reconstruction of cell phenotypes
  210. 210Integration of transcriptomics, proteomics, and metabolomics data to reveal HER2-associated metabolic heterogeneity in gastric cancer with response to immunotherapy and neoadjuvant chemotherapy.Background Currently available prognostic tools and focused therapeutic methods result in unsatisfactory treatment of gastric cancer (GC). A deeper understanding of human epidermal growth factor receptor 2 (HER2)-coexpressed metabolic pathways may offer novel insights into tumour-intrinsic precision medicine. Methods The integrated multi-omics strategies (including transcriptomics, proteomics and metabolomics) were applied to develop a novel metabolic classifier for gastric cancer. We integrated TCGA-STAD cohort (375 GC samples and 56753 genes) and TCPA-STAD cohort (392 GC samples and 218 proteins), and rated them as transcriptomics and proteomics data, resepectively. 224 matched blood samples of GC patients and healthy individuals were collected to carry out untargeted metabolomics analysis. Results In this study, pan-cancer analysis highlighted the crucial role of ERBB2 in the immune microenvironment and metabolic remodelling. In addition, the metabolic landscape of GC indicated that
  211. 211Joint cell segmentation and cell type annotation for spatial transcriptomics.RNA hybridization-based spatial transcriptomics provides unparalleled detection sensitivity. However, inaccuracies in segmentation of image volumes into cells cause misassignment of mRNAs which is a major source of errors. Here, we develop JSTA, a computational framework for joint cell segmentation and cell type annotation that utilizes prior knowledge of cell type-specific gene expression. Simulation results show that leveraging existing cell type taxonomy increases RNA assignment accuracy by more than 45%. Using JSTA, we were able to classify cells in the mouse hippocampus into 133 (sub)types revealing the spatial organization of CA1, CA3, and Sst neuron subtypes. Analysis of within cell subtype spatial differential gene expression of 80 candidate genes identified 63 with statistically significant spatial differential gene expression across 61 (sub)types. Overall, our work demonstrates that known cell type expression patterns can be leveraged to improve the accuracy of RNA hybridizat
  212. 212STRIDE: accurately decomposing and integrating spatial transcriptomics using single-cell RNA sequencing.The recent advances in spatial transcriptomics have brought unprecedented opportunities to understand the cellular heterogeneity in the spatial context. However, the current limitations of spatial technologies hamper the exploration of cellular localizations and interactions at single-cell level. Here, we present spatial transcriptomics deconvolution by topic modeling (STRIDE), a computational method to decompose cell types from spatial mixtures by leveraging topic profiles trained from single-cell transcriptomics. STRIDE accurately estimated the cell-type proportions and showed balanced specificity and sensitivity compared to existing methods. We demonstrated STRIDE's utility by applying it to different spatial platforms and biological systems. Deconvolution by STRIDE not only mapped rare cell types to spatial locations but also improved the identification of spatially localized genes and domains. Moreover, topics discovered by STRIDE were associated with cell-type-specific functions
  213. 213Single-cell multiomics analysis reveals regulatory programs in clear cell renal cell carcinoma.The clear cell renal cell carcinoma (ccRCC) microenvironment consists of many different cell types and structural components that play critical roles in cancer progression and drug resistance, but the cellular architecture and underlying gene regulatory features of ccRCC have not been fully characterized. Here, we applied single-cell RNA sequencing (scRNA-seq) and single-cell assay for transposase-accessible chromatin sequencing (scATAC-seq) to generate transcriptional and epigenomic landscapes of ccRCC. We identified tumor cell-specific regulatory programs mediated by four key transcription factors (TFs) (HOXC5, VENTX, ISL1, and OTP), and these TFs have prognostic significance in The Cancer Genome Atlas (TCGA) database. Targeting these TFs via short hairpin RNAs (shRNAs) or small molecule inhibitors decreased tumor cell proliferation. We next performed an integrative analysis of chromatin accessibility and gene expression for CD8 + T cells and macrophages to reveal the different regul
  214. 214Spatiotemporally deciphering the mysterious mechanism of persistent HPV-induced malignant transition and immune remodelling from HPV-infected normal cervix, precancer to cervical cancer: Integrating single-cell RNA-sequencing and spatial transcriptome.Background The mechanism underlying cervical carcinogenesis that is mediated by persistent human papillomavirus (HPV) infection remains elusive. Aims Here, for the first time, we deciphered both the temporal transition and spatial distribution of cellular subsets during disease progression from normal cervix tissues to precursor lesions to cervical cancer. Materials & methods We generated scRNA-seq profiles and spatial transcriptomics data from nine patient samples, including two HPV-negative normal, two HPV-positive normal, two HPV-positive HSIL and three HPV-positive cancer samples. Results We not only identified three 'HPV-related epithelial clusters' that are unique to normal, high-grade squamous intraepithelial lesions (HSIL) and cervical cancer tissues but also discovered node genes that potentially regulate disease progression. Moreover, we observed the gradual transition of multiple immune cells that exhibited positive immune responses, followed by dysregulation and exhaustion,
  215. 215Single-cell spatial transcriptomics reveals a dynamic control of metabolic zonation and liver regeneration by endothelial cell Wnt2 and Wnt9b.The conclusive identity of Wnts regulating liver zonation (LZ) and regeneration (LR) remains unclear despite an undisputed role of β-catenin. Using single-cell analysis, we identified a conserved Wnt2 and Wnt9b expression in endothelial cells (ECs) in zone 3. EC-elimination of Wnt2 and Wnt9b led to both loss of β-catenin targets in zone 3, and re-appearance of zone 1 genes in zone 3, unraveling dynamicity in the LZ process. Impaired LR observed in the knockouts phenocopied models of defective hepatic Wnt signaling. Administration of a tetravalent antibody to activate Wnt signaling rescued LZ and LR in the knockouts and induced zone 3 gene expression and LR in controls. Administration of the agonist also promoted LR in acetaminophen overdose acute liver failure (ALF) fulfilling an unmet clinical need. Overall, we report an unequivocal role of EC-Wnt2 and Wnt9b in LZ and LR and show the role of Wnt activators as regenerative therapy for ALF.
  216. 216Tissue morphology influences the temporal program of human brain organoid development.Progression through fate decisions determines cellular composition and tissue architecture, but how that same architecture may impact cell fate is less clear. We took advantage of organoids as a tractable model to interrogate this interaction of form and fate. Screening methodological variations revealed that common protocol adjustments impacted various aspects of morphology, from macrostructure to tissue architecture. We examined the impact of morphological perturbations on cell fate through integrated single nuclear RNA sequencing (snRNA-seq) and spatial transcriptomics. Regardless of the specific protocol, organoids with more complex morphology better mimicked in vivo human fetal brain development. Organoids with perturbed tissue architecture displayed aberrant temporal progression, with cells being intermingled in both space and time. Finally, encapsulation to impart a simplified morphology led to disrupted tissue cytoarchitecture and a similar abnormal maturational timing. These d
  217. 217Efficient prediction of a spatial transcriptomics profile better characterizes breast cancer tissue sections without costly experimentation.Spatial transcriptomics is an emerging technology requiring costly reagents and considerable skills, limiting the identification of transcriptional markers related to histology. Here, we show that predicted spatial gene-expression in unmeasured regions and tissues can enhance biologists' histological interpretations. We developed the Deep learning model for Spatial gene Clusters and Expression, DeepSpaCE, and confirmed its performance using the spatial-transcriptome profiles and immunohistochemistry images of consecutive human breast cancer tissue sections. For example, the predicted expression patterns of SPARC, an invasion marker, highlighted a small tumor-invasion region difficult to identify using raw spatial transcriptome data alone because of a lack of measurements. We further developed semi-supervised DeepSpaCE using unlabeled histology images and increased the imputation accuracy of consecutive sections, enhancing applicability for a small sample size. Our method enables users
  218. 218Multi-omics profiles of the intestinal microbiome in irritable bowel syndrome and its bowel habit subtypes.Background Irritable bowel syndrome (IBS) is a common gastrointestinal disorder that is thought to involve alterations in the gut microbiome, but robust microbial signatures have been challenging to identify. As prior studies have primarily focused on composition, we hypothesized that multi-omics assessment of microbial function incorporating both metatranscriptomics and metabolomics would further delineate microbial profiles of IBS and its subtypes. Methods Fecal samples were collected from a racially/ethnically diverse cohort of 495 subjects, including 318 IBS patients and 177 healthy controls, for analysis by 16S rRNA gene sequencing (n = 486), metatranscriptomics (n = 327), and untargeted metabolomics (n = 368). Differentially abundant microbes, predicted genes, transcripts, and metabolites in IBS were identified by multivariate models incorporating age, sex, race/ethnicity, BMI, diet, and HAD-Anxiety. Inter-omic functional relationships were assessed by transcript/gene ratios and
  219. 219Spatial mapping of the total transcriptome by in situ polyadenylation.Spatial transcriptomics reveals the spatial context of gene expression, but current methods are limited to assaying polyadenylated (A-tailed) RNA transcripts. Here we demonstrate that enzymatic in situ polyadenylation of RNA enables detection of the full spectrum of RNAs, expanding the scope of sequencing-based spatial transcriptomics to the total transcriptome. We demonstrate that our spatial total RNA-sequencing (STRS) approach captures coding RNAs, noncoding RNAs and viral RNAs. We apply STRS to study skeletal muscle regeneration and viral-induced myocarditis. Our analyses reveal the spatial patterns of noncoding RNA expression with near-cellular resolution, identify spatially defined expression of noncoding transcripts in skeletal muscle regeneration and highlight host transcriptional responses associated with local viral RNA abundance. STRS requires adding only one step to the widely used Visium spatial total RNA-sequencing protocol from 10x Genomics, and thus could be easily adop
  220. 220Single cell multi-omics reveal intra-cell-line heterogeneity across human cancer cell lines.Human cancer cell lines have long served as tools for cancer research and drug discovery, but the presence and the source of intra-cell-line heterogeneity remain elusive. Here, we perform single-cell RNA-sequencing and ATAC-sequencing on 42 and 39 human cell lines, respectively, to illustrate both transcriptomic and epigenetic heterogeneity within individual cell lines. Our data reveal that transcriptomic heterogeneity is frequently observed in cancer cell lines of different tissue origins, often driven by multiple common transcriptional programs. Copy number variation, as well as epigenetic variation and extrachromosomal DNA distribution all contribute to the detected intra-cell-line heterogeneity. Using hypoxia treatment as an example, we demonstrate that transcriptomic heterogeneity could be reshaped by environmental stress. Overall, our study performs single-cell multi-omics of commonly used human cancer cell lines and offers mechanistic insights into the intra-cell-line heterogene
  221. 221Changing Technologies of RNA Sequencing and Their Applications in Clinical Oncology.RNA sequencing (RNAseq) is one of the most commonly used techniques in life sciences, and has been widely used in cancer research, drug development, and cancer diagnosis and prognosis. Driven by various biological and technical questions, the techniques of RNAseq have progressed rapidly from bulk RNAseq, laser-captured micro-dissected RNAseq, and single-cell RNAseq to digital spatial RNA profiling, spatial transcriptomics, and direct in situ sequencing. These different technologies have their unique strengths, weaknesses, and suitable applications in the field of clinical oncology. To guide cancer researchers to select the most appropriate RNAseq technique for their biological questions, we will discuss each of these technologies, technical features, and clinical applications in cancer. We will help cancer researchers to understand the key differences of these RNAseq technologies and their optimal applications.
  222. 222Integrated proteogenomic characterization of urothelial carcinoma of the bladder.Background Urothelial carcinoma (UC) is the most common pathological type of bladder cancer, a malignant tumor. However, an integrated multi-omics analysis of the Chinese UC patient cohort is lacking. Methods We performed an integrated multi-omics analysis, including whole-exome sequencing, RNA-seq, proteomic, and phosphoproteomic analysis of 116 Chinese UC patients, comprising 45 non-muscle-invasive bladder cancer patients (NMIBCs) and 71 muscle-invasive bladder cancer patients (MIBCs). Result Proteogenomic integration analysis indicated that SND1 and CDK5 amplifications on chromosome 7q were associated with the activation of STAT3, which was relevant to tumor proliferation. Chromosome 5p gain in NMIBC patients was a high-risk factor, through modulating actin cytoskeleton implicating in tumor cells invasion. Phosphoproteomic analysis of tumors and morphologically normal human urothelium produced UC-associated activated kinases, including CDK1 and PRKDC. Proteomic analysis identified t
  223. 223Concordance of MERFISH spatial transcriptomics with bulk and single-cell RNA sequencing.Spatial transcriptomics extends single-cell RNA sequencing (scRNA-seq) by providing spatial context for cell type identification and analysis. Imaging-based spatial technologies such as multiplexed error-robust fluorescence in situ hybridization (MERFISH) can achieve single-cell resolution, directly mapping single-cell identities to spatial positions. MERFISH produces a different data type than scRNA-seq, and a technical comparison between the two modalities is necessary to ascertain how to best integrate them. We performed MERFISH on the mouse liver and kidney and compared the resulting bulk and single-cell RNA statistics with those from the Tabula Muris Senis cell atlas and from two Visium datasets. MERFISH quantitatively reproduced the bulk RNA-seq and scRNA-seq results with improvements in overall dropout rates and sensitivity. Finally, we found that MERFISH independently resolved distinct cell types and spatial structure in both the liver and kidney. Computational integration with
  224. 224Whole-brain comparison of rodent and human brains using spatial transcriptomics.The ever-increasing use of mouse models in preclinical neuroscience research calls for an improvement in the methods used to translate findings between mouse and human brains. Previously, we showed that the brains of primates can be compared in a direct quantitative manner using a common reference space built from white matter tractography data (Mars et al., 2018b). Here, we extend the common space approach to evaluate the similarity of mouse and human brain regions using openly accessible brain-wide transcriptomic data sets. We show that mouse-human homologous genes capture broad patterns of neuroanatomical organization, but the resolution of cross-species correspondences can be improved using a novel supervised machine learning approach. Using this method, we demonstrate that sensorimotor subdivisions of the neocortex exhibit greater similarity between species, compared with supramodal subdivisions, and mouse isocortical regions separate into sensorimotor and supramodal clusters base
  225. 225Spatial transcriptomic characterization of pathologic niches in IPF.Despite advancements in antifibrotic therapy, idiopathic pulmonary fibrosis (IPF) remains a medical condition with unmet needs. Single-cell RNA sequencing (scRNA-seq) has enhanced our understanding of IPF but lacks the cellular tissue context and gene expression localization that spatial transcriptomics provides. To bridge this gap, we profiled IPF and control patient lung tissue using spatial transcriptomics, integrating the data with an IPF scRNA-seq atlas. We identified three disease-associated niches with unique cellular compositions and localizations. These include a fibrotic niche, consisting of myofibroblasts and aberrant basaloid cells, located around airways and adjacent to an airway macrophage niche in the lumen, containing SPP1 + macrophages. In addition, we identified an immune niche, characterized by distinct lymphoid cell foci in fibrotic tissue, surrounded by remodeled endothelial vessels. This spatial characterization of IPF niches will facilitate the identification of
  226. 226Library size confounds biology in spatial transcriptomics data.Spatial molecular data has transformed the study of disease microenvironments, though, larger datasets pose an analytics challenge prompting the direct adoption of single-cell RNA-sequencing tools including normalization methods. Here, we demonstrate that library size is associated with tissue structure and that normalizing these effects out using commonly applied scRNA-seq normalization methods will negatively affect spatial domain identification. Spatial data should not be specifically corrected for library size prior to analysis, and algorithms designed for scRNA-seq data should be adopted with caution.
  227. 227Bento: a toolkit for subcellular analysis of spatial transcriptomics data.The spatial organization of molecules in a cell is essential for their functions. While current methods focus on discerning tissue architecture, cell-cell interactions, and spatial expression patterns, they are limited to the multicellular scale. We present Bento, a Python toolkit that takes advantage of single-molecule information to enable spatial analysis at the subcellular scale. Bento ingests molecular coordinates and segmentation boundaries to perform three analyses: defining subcellular domains, annotating localization patterns, and quantifying gene-gene colocalization. We demonstrate MERFISH, seqFISH + , Molecular Cartography, and Xenium datasets. Bento is part of the open-source Scverse ecosystem, enabling integration with other single-cell analysis tools.
  228. 228Spatial transcriptomics reveals that metabolic characteristics define the tumor immunosuppression microenvironment via iCAF transformation in oral squamous cell carcinoma.Tumor progression is closely related to tumor tissue metabolism and reshaping of the microenvironment. Oral squamous cell carcinoma (OSCC), a representative hypoxic tumor, has a heterogeneous internal metabolic environment. To clarify the relationship between different metabolic regions and the tumor immune microenvironment (TME) in OSCC, Single cell (SC) and spatial transcriptomics (ST) sequencing of OSCC tissues were performed. The proportion of TME in the ST data was obtained through SPOTlight deconvolution using SC and GSE103322 data. The metabolic activity of each spot was calculated using scMetabolism, and k-means clustering was used to classify all spots into hyper-, normal-, or hypometabolic regions. CD4T cell infiltration and TGF-β expression is higher in the hypermetabolic regions than in the others. Through CellPhoneDB and NicheNet cell-cell communication analysis, it was found that in the hypermetabolic region, fibroblasts can utilize the lactate produced by glycolysis of e
  229. 229Niche-DE: niche-differential gene expression analysis in spatial transcriptomics data identifies context-dependent cell-cell interactions.Existing methods for analysis of spatial transcriptomic data focus on delineating the global gene expression variations of cell types across the tissue, rather than local gene expression changes driven by cell-cell interactions. We propose a new statistical procedure called niche-differential expression (niche-DE) analysis that identifies cell-type-specific niche-associated genes, which are differentially expressed within a specific cell type in the context of specific spatial niches. We further develop niche-LR, a method to reveal ligand-receptor signaling mechanisms that underlie niche-differential gene expression patterns. Niche-DE and niche-LR are applicable to low-resolution spot-based spatial transcriptomics data and data that is single-cell or subcellular in resolution.
  230. 230Single-cell and Spatial Transcriptomics Reveals Ferroptosis as The Most Enriched Programmed Cell Death Process in Hemorrhage Stroke-induced Oligodendrocyte-mediated White Matter Injury.Intracerebral hemorrhage (ICH) is a severe stroke subtype with limited therapeutic options. Programmed cell death (PCD) is crucial for immunological balance, and includes necroptosis, pyroptosis, apoptosis, ferroptosis, and necrosis. However, the distinctions between these programmed cell death modalities after ICH remain to be further investigated. We used single-cell transcriptome (single-cell RNA sequencing) and spatial transcriptome (spatial RNA sequencing) techniques to investigate PCD-related gene expression trends in the rat brain following hemorrhagic stroke. Ferroptosis was the main PCD process after ICH, and primarily affected mature oligodendrocytes. Its onset occurred as early as 1 hour post-ICH, peaking at 24 hours post-ICH. Additionally, ferroptosis-related genes were distributed in the hippocampus and choroid plexus. We also elucidated a specific interaction between lipocalin-2 (LCN2)-positive microglia and oligodendrocytes that was mediated by the colony stimulating fac
  231. 231Single-cell tumor heterogeneity landscape of hepatocellular carcinoma: unraveling the pro-metastatic subtype and its interaction loop with fibroblasts.Background Tumor heterogeneity presents a formidable challenge in understanding the mechanisms driving tumor progression and metastasis. The heterogeneity of hepatocellular carcinoma (HCC) in cellular level is not clear. Methods Integration analysis of single-cell RNA sequencing data and spatial transcriptomics data was performed. Multiple methods were applied to investigate the subtype of HCC tumor cells. The functional characteristics, translation factors, clinical implications and microenvironment associations of different subtypes of tumor cells were analyzed. The interaction of subtype and fibroblasts were analyzed. Results We established a heterogeneity landscape of HCC malignant cells by integrated 52 single-cell RNA sequencing data and 5 spatial transcriptomics data. We identified three subtypes in tumor cells, including ARG1 + metabolism subtype (Metab-subtype), TOP2A + proliferation phenotype (Prol-phenotype), and S100A6 + pro-metastatic subtype (EMT-subtype). Enrichment anal
  232. 232Systematic dissection of tumor-normal single-cell ecosystems across a thousand tumors of 30 cancer types.The complexity of the tumor microenvironment poses significant challenges in cancer therapy. Here, to comprehensively investigate the tumor-normal ecosystems, we perform an integrative analysis of 4.9 million single-cell transcriptomes from 1070 tumor and 493 normal samples in combination with pan-cancer 137 spatial transcriptomics, 8887 TCGA, and 1261 checkpoint inhibitor-treated bulk tumors. We define a myriad of cell states constituting the tumor-normal ecosystems and also identify hallmark gene signatures across different cell types and organs. Our atlas characterizes distinctions between inflammatory fibroblasts marked by AKR1C1 or WNT5A in terms of cellular interactions and spatial co-localization patterns. Co-occurrence analysis reveals interferon-enriched community states including tertiary lymphoid structure (TLS) components, which exhibit differential rewiring between tumor, adjacent normal, and healthy normal tissues. The favorable response of interferon-enriched community s
  233. 233A visual-omics foundation model to bridge histopathology with spatial transcriptomics.Artificial intelligence has revolutionized computational biology. Recent developments in omics technologies, including single-cell RNA sequencing and spatial transcriptomics, provide detailed genomic data alongside tissue histology. However, current computational models focus on either omics or image analysis, lacking their integration. To address this, we developed OmiCLIP, a visual-omics foundation model linking hematoxylin and eosin images and transcriptomics using tissue patches from Visium data. We transformed transcriptomic data into 'sentences' by concatenating top-expressed gene symbols from each patch. We curated a dataset of 2.2 million paired tissue images and transcriptomic data across 32 organs to train OmiCLIP integrating histology and transcriptomics. Building on OmiCLIP, our Loki platform offers five key functions: tissue alignment, annotation via bulk RNA sequencing or marker genes, cell-type decomposition, image-transcriptomics retrieval and spatial transcriptomics ge
  234. 234Single-cell and spatial transcriptomics reveal metastasis mechanism and microenvironment remodeling of lymph node in osteosarcoma.Background Osteosarcoma (OS) is the most common primary malignant bone tumor and is highly prone to metastasis. OS can metastasize to the lymph node (LN) through the lymphatics, and the metastasis of tumor cells reestablishes the immune landscape of the LN, which is conducive to the growth of tumor cells. However, the mechanism of LN metastasis of osteosarcoma and remodeling of the metastatic lymph node (MLN) microenvironment is not clear. Methods Single-cell RNA sequencing of 18 samples from paracancerous, primary tumor, and lymph nodes was performed. Then, new signaling axes closely related to metastasis were identified using bioinformatics, in vitro experiments, and immunohistochemistry. The mechanism of remodeling of the LN microenvironment in tumor cells was investigated by integrating single-cell and spatial transcriptomics. Results From 18 single-cell sequencing samples, we obtained 117,964 cells. The pseudotime analysis revealed that osteoblast(OB) cells may follow a differenti
  235. 235Domain generalization enables general cancer cell annotation in single-cell and spatial transcriptomics.Single-cell and spatial transcriptome sequencing, two recently optimized transcriptome sequencing methods, are increasingly used to study cancer and related diseases. Cell annotation, particularly for malignant cell annotation, is essential and crucial for in-depth analyses in these studies. However, current algorithms lack accuracy and generalization, making it difficult to consistently and rapidly infer malignant cells from pan-cancer data. To address this issue, we present Cancer-Finder, a domain generalization-based deep-learning algorithm that can rapidly identify malignant cells in single-cell data with an average accuracy of 95.16%. More importantly, by replacing the single-cell training data with spatial transcriptomic datasets, Cancer-Finder can accurately identify malignant spots on spatial slides. Applying Cancer-Finder to 5 clear cell renal cell carcinoma spatial transcriptomic samples, Cancer-Finder demonstrates a good ability to identify malignant spots and identifies a g
  236. 236Heterogeneous Skeletal Muscle Cell and Nucleus Populations Identified by Single-Cell and Single-Nucleus Resolution Transcriptome Assays.Single-cell RNA-seq (scRNA-seq) has revolutionized modern genomics, but the large size of myotubes and myofibers has restricted use of scRNA-seq in skeletal muscle. For the study of muscle, single-nucleus RNA-seq (snRNA-seq) has emerged not only as an alternative to scRNA-seq, but as a novel method providing valuable insights into multinucleated cells such as myofibers. Nuclei within myofibers specialize at junctions with other cell types such as motor neurons. Nuclear heterogeneity plays important roles in certain diseases such as muscular dystrophies. We survey current methods of high-throughput single cell and subcellular resolution transcriptomics, including single-cell and single-nucleus RNA-seq and spatial transcriptomics, applied to satellite cells, myoblasts, myotubes and myofibers. We summarize the major myonuclei subtypes identified in homeostatic and regenerating tissue including those specific to fiber type or at junctions with other cell types. Disease-specific nucleus pop
  237. 237GSVA: gene set variation analysis for microarray and RNA-seq data.Background Gene set enrichment (GSE) analysis is a popular framework for condensing information from gene expression profiles into a pathway or signature summary. The strengths of this approach over single gene analysis include noise and dimension reduction, as well as greater biological interpretability. As molecular profiling experiments move beyond simple case-control studies, robust and flexible GSE methodologies are needed that can model pathway activity within highly heterogeneous data sets. Results To address this challenge, we introduce Gene Set Variation Analysis (GSVA), a GSE method that estimates variation of pathway activity over a sample population in an unsupervised manner. We demonstrate the robustness of GSVA in a comparison with current state of the art sample-wise enrichment methods. Further, we provide examples of its utility in differential pathway activity and survival analysis. Lastly, we show how GSVA works analogously with data from both microarray and RNA-seq e
  238. 238A scaling normalization method for differential expression analysis of RNA-seq data.The fine detail provided by sequencing-based transcriptome surveys suggests that RNA-seq is likely to become the platform of choice for interrogating steady state RNA. In order to discover biologically important changes in expression, we show that normalization continues to be an essential step in the analysis. We outline a simple and effective method for performing normalization and show dramatically improved results for inferring differential expression in simulated and publicly available data sets.
  239. 239Massively parallel digital transcriptional profiling of single cells.Characterizing the transcriptome of individual cells is fundamental to understanding complex biological systems. We describe a droplet-based system that enables 3' mRNA counting of tens of thousands of single cells per sample. Cell encapsulation, of up to 8 samples at a time, takes place in ∼6 min, with ∼50% cell capture efficiency. To demonstrate the system's technical performance, we collected transcriptome data from ∼250k single cells across 29 samples. We validated the sensitivity of the system and its ability to detect rare populations using cell lines and synthetic RNAs. We profiled 68k peripheral blood mononuclear cells to demonstrate the system's ability to characterize large immune populations. Finally, we used sequence variation in the transcriptome data to determine host and donor chimerism at single-cell resolution from bone marrow mononuclear cells isolated from transplant patients.
  240. 240Differential analyses for RNA-seq: transcript-level estimates improve gene-level inferences.High-throughput sequencing of cDNA (RNA-seq) is used extensively to characterize the transcriptome of cells. Many transcriptomic studies aim at comparing either abundance levels or the transcriptome composition between given conditions, and as a first step, the sequencing reads must be used as the basis for abundance quantification of transcriptomic features of interest, such as genes or transcripts. Various quantification approaches have been proposed, ranging from simple counting of reads that overlap given genomic regions to more complex estimation of underlying transcript abundances. In this paper, we show that gene-level abundance estimates and statistical inference offer advantages over transcript-level analyses, in terms of performance and interpretability. We also illustrate that the presence of differential isoform usage can lead to inflated false discovery rates in differential gene expression analyses on simple count matrices but that this can be addressed by incorporating o
  241. 241Molecular Architecture of the Mouse Nervous System.The mammalian nervous system executes complex behaviors controlled by specialized, precisely positioned, and interacting cell types. Here, we used RNA sequencing of half a million single cells to create a detailed census of cell types in the mouse nervous system. We mapped cell types spatially and derived a hierarchical, data-driven taxonomy. Neurons were the most diverse and were grouped by developmental anatomical units and by the expression of neurotransmitters and neuropeptides. Neuronal diversity was driven by genes encoding cell identity, synaptic connectivity, neurotransmission, and membrane conductance. We discovered seven distinct, regionally restricted astrocyte types that obeyed developmental boundaries and correlated with the spatial distribution of key glutamate and glycine neurotransmitters. In contrast, oligodendrocytes showed a loss of regional identity followed by a secondary diversification. The resource presented here lays a solid foundation for understanding the mol
  242. 242Slingshot: cell lineage and pseudotime inference for single-cell transcriptomics.Background Single-cell transcriptomics allows researchers to investigate complex communities of heterogeneous cells. It can be applied to stem cells and their descendants in order to chart the progression from multipotent progenitors to fully differentiated cells. While a variety of statistical and computational methods have been proposed for inferring cell lineages, the problem of accurately characterizing multiple branching lineages remains difficult to solve. Results We introduce Slingshot, a novel method for inferring cell lineages and pseudotimes from single-cell gene expression data. In previously published datasets, Slingshot correctly identifies the biological signal for one to three branching trajectories. Additionally, our simulation study shows that Slingshot infers more accurate pseudotimes than other leading methods. Conclusions Slingshot is a uniquely robust and flexible tool which combines the highly stable techniques necessary for noisy single-cell data with the ability
  243. 243UMI-tools: modeling sequencing errors in Unique Molecular Identifiers to improve quantification accuracy.Unique Molecular Identifiers (UMIs) are random oligonucleotide barcodes that are increasingly used in high-throughput sequencing experiments. Through a UMI, identical copies arising from distinct molecules can be distinguished from those arising through PCR amplification of the same molecule. However, bioinformatic methods to leverage the information from UMIs have yet to be formalized. In particular, sequencing errors in the UMI sequence are often ignored or else resolved in an ad hoc manner. We show that errors in the UMI sequence are common and introduce network-based methods to account for these errors when identifying PCR duplicates. Using these methods, we demonstrate improved quantification accuracy both under simulated conditions and real iCLIP and single-cell RNA-seq data sets. Reproducibility between iCLIP replicates and single-cell RNA-seq clustering are both improved using our proposed network-based method, demonstrating the value of properly accounting for errors in UMIs.
  244. 244SARS-CoV-2 Receptor ACE2 Is an Interferon-Stimulated Gene in Human Airway Epithelial Cells and Is Detected in Specific Cell Subsets across Tissues.There is pressing urgency to understand the pathogenesis of the severe acute respiratory syndrome coronavirus clade 2 (SARS-CoV-2), which causes the disease COVID-19. SARS-CoV-2 spike (S) protein binds angiotensin-converting enzyme 2 (ACE2), and in concert with host proteases, principally transmembrane serine protease 2 (TMPRSS2), promotes cellular entry. The cell subsets targeted by SARS-CoV-2 in host tissues and the factors that regulate ACE2 expression remain unknown. Here, we leverage human, non-human primate, and mouse single-cell RNA-sequencing (scRNA-seq) datasets across health and disease to uncover putative targets of SARS-CoV-2 among tissue-resident cell subsets. We identify ACE2 and TMPRSS2 co-expressing cells within lung type II pneumocytes, ileal absorptive enterocytes, and nasal goblet secretory cells. Strikingly, we discovered that ACE2 is a human interferon-stimulated gene (ISG) in vitro using airway epithelial cells and extend our findings to in vivo viral infections.
  245. 245High expression of ACE2 receptor of 2019-nCoV on the epithelial cells of oral mucosa.It has been reported that ACE2 is the main host cell receptor of 2019-nCoV and plays a crucial role in the entry of virus into the cell to cause the final infection. To investigate the potential route of 2019-nCov infection on the mucosa of oral cavity, bulk RNA-seq profiles from two public databases including The Cancer Genome Atlas (TCGA) and Functional Annotation of The Mammalian Genome Cap Analysis of Gene Expression (FANTOM5 CAGE) dataset were collected. RNA-seq profiling data of 13 organ types with para-carcinoma normal tissues from TCGA and 14 organ types with normal tissues from FANTOM5 CAGE were analyzed in order to explore and validate the expression of ACE2 on the mucosa of oral cavity. Further, single-cell transcriptomes from an independent data generated in-house were used to identify and confirm the ACE2-expressing cell composition and proportion in oral cavity. The results demonstrated that the ACE2 expressed on the mucosa of oral cavity. Interestingly, this receptor was
  246. 246Single-cell RNA-seq data analysis on the receptor ACE2 expression reveals the potential risk of different human organs vulnerable to 2019-nCoV infection.It has been known that, the novel coronavirus, 2019-nCoV, which is considered similar to SARS-CoV, invades human cells via the receptor angiotensin converting enzyme II (ACE2). Moreover, lung cells that have ACE2 expression may be the main target cells during 2019-nCoV infection. However, some patients also exhibit non-respiratory symptoms, such as kidney failure, implying that 2019-nCoV could also invade other organs. To construct a risk map of different human organs, we analyzed the single-cell RNA sequencing (scRNA-seq) datasets derived from major human physiological systems, including the respiratory, cardiovascular, digestive, and urinary systems. Through scRNA-seq data analyses, we identified the organs at risk, such as lung, heart, esophagus, kidney, bladder, and ileum, and located specific cell types (i.e., type II alveolar cells (AT2), myocardial cells, proximal tubule cells of the kidney, ileum and esophagus epithelial cells, and bladder urothelial cells), which are vulnerabl
  247. 247Scater: pre-processing, quality control, normalization and visualization of single-cell RNA-seq data in R.Motivation Single-cell RNA sequencing (scRNA-seq) is increasingly used to study gene expression at the level of individual cells. However, preparing raw sequence data for further analysis is not a straightforward process. Biases, artifacts and other sources of unwanted variation are present in the data, requiring substantial time and effort to be spent on pre-processing, quality control (QC) and normalization. Results We have developed the R/Bioconductor package scater to facilitate rigorous pre-processing, quality control, normalization and visualization of scRNA-seq data. The package provides a convenient, flexible workflow to process raw sequencing reads into a high-quality expression dataset ready for downstream analysis. scater provides a rich suite of plotting tools for single-cell data and a flexible data structure that is compatible with existing tools and can be used as infrastructure for future software development. Availability and implementation The open-source code, along
  248. 248SoupX removes ambient RNA contamination from droplet-based single-cell RNA sequencing data.Background Droplet-based single-cell RNA sequence analyses assume that all acquired RNAs are endogenous to cells. However, any cell-free RNAs contained within the input solution are also captured by these assays. This sequencing of cell-free RNA constitutes a background contamination that confounds the biological interpretation of single-cell transcriptomic data. Results We demonstrate that contamination from this "soup" of cell-free RNAs is ubiquitous, with experiment-specific variations in composition and magnitude. We present a method, SoupX, for quantifying the extent of the contamination and estimating "background-corrected" cell expression profiles that seamlessly integrate with existing downstream analysis tools. Applying this method to several datasets using multiple droplet sequencing technologies, we demonstrate that its application improves biological interpretation of otherwise misleading data, as well as improving quality control metrics. Conclusions We present SoupX, a to
  249. 249Online survival analysis software to assess the prognostic value of biomarkers using transcriptomic data in non-small-cell lung cancer.In the last decade, optimized treatment for non-small cell lung cancer had lead to improved prognosis, but the overall survival is still very short. To further understand the molecular basis of the disease we have to identify biomarkers related to survival. Here we present the development of an online tool suitable for the real-time meta-analysis of published lung cancer microarray datasets to identify biomarkers related to survival. We searched the caBIG, GEO and TCGA repositories to identify samples with published gene expression data and survival information. Univariate and multivariate Cox regression analysis, Kaplan-Meier survival plot with hazard ratio and logrank P value are calculated and plotted in R. The complete analysis tool can be accessed online at: www.kmplot.com/lung. All together 1,715 samples of ten independent datasets were integrated into the system. As a demonstration, we used the tool to validate 21 previously published survival associated biomarkers. Of these, su
  250. 250iDEP: an integrated web application for differential expression and pathway analysis of RNA-Seq data.Background RNA-seq is widely used for transcriptomic profiling, but the bioinformatics analysis of resultant data can be time-consuming and challenging, especially for biologists. We aim to streamline the bioinformatic analyses of gene-level data by developing a user-friendly, interactive web application for exploratory data analysis, differential expression, and pathway analysis. Results iDEP (integrated Differential Expression and Pathway analysis) seamlessly connects 63 R/Bioconductor packages, 2 web services, and comprehensive annotation and pathway databases for 220 plant and animal species. The workflow can be reproduced by downloading customized R code and related pathway files. As an example, we analyzed an RNA-Seq dataset of lung fibroblasts with Hoxa1 knockdown and revealed the possible roles of SP1 and E2F1 and their target genes, including microRNAs, in blocking G1/S transition. In another example, our analysis shows that in mouse B cells without functional p53, ionizing ra
  251. 251The somatic mutation profiles of 2,433 breast cancers refines their genomic and transcriptomic landscapes.The genomic landscape of breast cancer is complex, and inter- and intra-tumour heterogeneity are important challenges in treating the disease. In this study, we sequence 173 genes in 2,433 primary breast tumours that have copy number aberration (CNA), gene expression and long-term clinical follow-up data. We identify 40 mutation-driver (Mut-driver) genes, and determine associations between mutations, driver CNA profiles, clinical-pathological parameters and survival. We assess the clonal states of Mut-driver mutations, and estimate levels of intra-tumour heterogeneity using mutant-allele fractions. Associations between PIK3CA mutations and reduced survival are identified in three subgroups of ER-positive cancer (defined by amplification of 17q23, 11q13-14 or 8q24). High levels of intra-tumour heterogeneity are in general associated with a worse outcome, but highly aggressive tumours with 11q13-14 amplification have low levels of intra-tumour heterogeneity. These results emphasize the i
  252. 252PanglaoDB: a web server for exploration of mouse and human single-cell RNA sequencing data.Single-cell RNA sequencing is an increasingly used method to measure gene expression at the single cell level and build cell-type atlases of tissues. Hundreds of single-cell sequencing datasets have already been published. However, studies are frequently deposited as raw data, a format difficult to access for biological researchers due to the need for data processing using complex computational pipelines. We have implemented an online database, PanglaoDB, accessible through a user-friendly interface that can be used to explore published mouse and human single cell RNA sequencing studies. PanglaoDB contains pre-processed and pre-computed analyses from more than 1054 single-cell experiments covering most major single cell platforms and protocols, based on more than 4 million cells from a wide range of tissues and organs. The online interface allows users to query and explore cell types, genetic pathways and regulatory networks. In addition, we have established a community-curated cell-ty
  253. 253ComBat-seq : batch effect adjustment for RNA-seq count data.The benefit of integrating batches of genomic data to increase statistical power is often hindered by batch effects, or unwanted variation in data caused by differences in technical factors across batches. It is therefore critical to effectively address batch effects in genomic data to overcome these challenges. Many existing methods for batch effects adjustment assume the data follow a continuous, bell-shaped Gaussian distribution. However in RNA-seq studies the data are typically skewed, over-dispersed counts, so this assumption is not appropriate and may lead to erroneous results. Negative binomial regression models have been used previously to better capture the properties of counts. We developed a batch correction method, ComBat-seq, using a negative binomial regression model that retains the integer nature of count data in RNA-seq studies, making the batch adjusted data compatible with common differential expression software packages that require integer counts. We show in realis
  254. 254One thousand plant transcriptomes and the phylogenomics of green plants.Green plants (Viridiplantae) include around 450,000-500,000 species 1,2 of great diversity and have important roles in terrestrial and aquatic ecosystems. Here, as part of the One Thousand Plant Transcriptomes Initiative, we sequenced the vegetative transcriptomes of 1,124 species that span the diversity of plants in a broad sense (Archaeplastida), including green plants (Viridiplantae), glaucophytes (Glaucophyta) and red algae (Rhodophyta). Our analysis provides a robust phylogenomic framework for examining the evolution of green plants. Most inferred species relationships are well supported across multiple species tree and supermatrix analyses, but discordance among plastid and nuclear gene trees at a few important nodes highlights the complexity of plant genome evolution, including polyploidy, periods of rapid speciation, and extinction. Incomplete sorting of ancestral variation, polyploidization and massive expansions of gene families punctuate the evolutionary history of green pla
  255. 255Single cell RNA sequencing of human liver reveals distinct intrahepatic macrophage populations.The liver is the largest solid organ in the body and is critical for metabolic and immune functions. However, little is known about the cells that make up the human liver and its immune microenvironment. Here we report a map of the cellular landscape of the human liver using single-cell RNA sequencing. We provide the transcriptional profiles of 8444 parenchymal and non-parenchymal cells obtained from the fractionation of fresh hepatic tissue from five human livers. Using gene expression patterns, flow cytometry, and immunohistochemical examinations, we identify 20 discrete cell populations of hepatocytes, endothelial cells, cholangiocytes, hepatic stellate cells, B cells, conventional and non-conventional T cells, NK-like cells, and distinct intrahepatic monocyte/macrophage populations. Together, our study presents a comprehensive view of the human liver at single-cell resolution that outlines the characteristics of resident cells in the liver, and in particular provides a map of the h
  256. 256Molecular and pharmacological modulators of the tumor immune contexture revealed by deconvolution of RNA-seq data.We introduce quanTIseq, a method to quantify the fractions of ten immune cell types from bulk RNA-sequencing data. quanTIseq was extensively validated in blood and tumor samples using simulated, flow cytometry, and immunohistochemistry data.quanTIseq analysis of 8000 tumor samples revealed that cytotoxic T cell infiltration is more strongly associated with the activation of the CXCR3/CXCL9 axis than with mutational load and that deconvolution-based cell scores have prognostic value in several solid cancers. Finally, we used quanTIseq to show how kinase inhibitors modulate the immune contexture and to reveal immune-cell types that underlie differential patients' responses to checkpoint blockers.Availability: quanTIseq is available at http://icbi.at/quantiseq .
  257. 257Single-cell RNA sequencing demonstrates the molecular and cellular reprogramming of metastatic lung adenocarcinoma.Advanced metastatic cancer poses utmost clinical challenges and may present molecular and cellular features distinct from an early-stage cancer. Herein, we present single-cell transcriptome profiling of metastatic lung adenocarcinoma, the most prevalent histological lung cancer type diagnosed at stage IV in over 40% of all cases. From 208,506 cells populating the normal tissues or early to metastatic stage cancer in 44 patients, we identify a cancer cell subtype deviating from the normal differentiation trajectory and dominating the metastatic stage. In all stages, the stromal and immune cell dynamics reveal ontological and functional changes that create a pro-tumoral and immunosuppressive microenvironment. Normal resident myeloid cell populations are gradually replaced with monocyte-derived macrophages and dendritic cells, along with T-cell exhaustion. This extensive single-cell analysis enhances our understanding of molecular and cellular dynamics in metastatic lung cancer and reveal
  258. 258CEL-Seq2: sensitive highly-multiplexed single-cell RNA-Seq.Single-cell transcriptomics requires a method that is sensitive, accurate, and reproducible. Here, we present CEL-Seq2, a modified version of our CEL-Seq method, with threefold higher sensitivity, lower costs, and less hands-on time. We implemented CEL-Seq2 on Fluidigm's C1 system, providing its first single-cell, on-chip barcoding method, and we detected gene expression changes accompanying the progression through the cell cycle in mouse fibroblast cells. We also compare with Smart-Seq to demonstrate CEL-Seq2's increased sensitivity relative to other available methods. Collectively, the improvements make CEL-Seq2 uniquely suited to single-cell RNA-Seq analysis in terms of economics, resolution, and ease of use.
  259. 259EmptyDrops: distinguishing cells from empty droplets in droplet-based single-cell RNA sequencing data.Droplet-based single-cell RNA sequencing protocols have dramatically increased the throughput of single-cell transcriptomics studies. A key computational challenge when processing these data is to distinguish libraries for real cells from empty droplets. Here, we describe a new statistical method for calling cells from droplet-based data, based on detecting significant deviations from the expression profile of the ambient solution. Using simulations, we demonstrate that EmptyDrops has greater power than existing approaches while controlling the false discovery rate among detected cells. Our method also retains distinct cell types that would have been discarded by existing methods in several real data sets.
  260. 260Pooling across cells to normalize single-cell RNA sequencing data with many zero counts.Normalization of single-cell RNA sequencing data is necessary to eliminate cell-specific biases prior to downstream analyses. However, this is not straightforward for noisy single-cell data where many counts are zero. We present a novel approach where expression values are summed across pools of cells, and the summed values are used for normalization. Pool-based size factors are then deconvolved to yield cell-based factors. Our deconvolution approach outperforms existing methods for accurate normalization of cell-specific biases in simulated data. Similar behavior is observed in real data, where deconvolution improves the relevance of results of downstream analyses.
  261. 261Transcriptomic characteristics of bronchoalveolar lavage fluid and peripheral blood mononuclear cells in COVID-19 patients.Circulating in China and 158 other countries and areas, the ongoing COVID-19 outbreak has caused devastating mortality and posed a great threat to public health. However, efforts to identify effectively supportive therapeutic drugs and treatments has been hampered by our limited understanding of host immune response for this fatal disease. To characterize the transcriptional signatures of host inflammatory response to SARS-CoV-2 (HCoV-19) infection, we carried out transcriptome sequencing of the RNAs isolated from the bronchoalveolar lavage fluid (BALF) and peripheral blood mononuclear cells (PBMC) specimens of COVID-19 patients. Our results reveal distinct host inflammatory cytokine profiles to SARS-CoV-2 infection in patients, and highlight the association between COVID-19 pathogenesis and excessive cytokine release such as CCL2/MCP-1, CXCL10/IP-10, CCL3/MIP-1A, and CCL4/MIP1B. Furthermore, SARS-CoV-2 induced activation of apoptosis and P53 signalling pathway in lymphocytes may be th
  262. 262CPGAVAS2, an integrated plastome sequence annotator and analyzer.We previously developed a web server CPGAVAS for annotation, visualization and GenBank submission of plastome sequences. Here, we upgrade the server into CPGAVAS2 to address the following challenges: (i) inaccurate annotation in the reference sequence likely causing the propagation of errors; (ii) difficulty in the annotation of small exons of genes petB, petD and rps16 and trans-splicing gene rps12; (iii) lack of annotation for other genome features and their visualization, such as repeat elements; and (iv) lack of modules for diversity analysis of plastomes. In particular, CPGAVAS2 provides two reference datasets for plastome annotation. The first dataset contains 43 plastomes whose annotation have been validated or corrected by RNA-seq data. The second one contains 2544 plastomes curated with sequence alignment. Two new algorithms are also implemented to correctly annotate small exons and trans-splicing genes. Tandem and dispersed repeats are identified, whose results are displayed
  263. 263Clustering trees: a visualization for evaluating clusterings at multiple resolutions.Clustering techniques are widely used in the analysis of large datasets to group together samples with similar properties. For example, clustering is often used in the field of single-cell RNA-sequencing in order to identify different cell types present in a tissue sample. There are many algorithms for performing clustering, and the results can vary substantially. In particular, the number of groups present in a dataset is often unknown, and the number of clusters identified by an algorithm can change based on the parameters used. To explore and examine the impact of varying clustering resolution, we present clustering trees. This visualization shows the relationships between clusters at multiple resolutions, allowing researchers to see how samples move as the number of clusters increases. In addition, meta-information can be overlaid on the tree to inform the choice of resolution and guide in identification of clusters. We illustrate the features of clustering trees using a series of
  264. 264CIRI: an efficient and unbiased algorithm for de novo circular RNA identification.Recent studies reveal that circular RNAs (circRNAs) are a novel class of abundant, stable and ubiquitous noncoding RNA molecules in animals. Comprehensive detection of circRNAs from high-throughput transcriptome data is an initial and crucial step to study their biogenesis and function. Here, we present a novel chiastic clipping signal-based algorithm, CIRI, to unbiasedly and accurately detect circRNAs from transcriptome data by employing multiple filtration strategies. By applying CIRI to ENCODE RNA-seq data, we for the first time identify and experimentally validate the prevalence of intronic/intergenic circRNAs as well as fragments specific to them in the human transcriptome.
  265. 265RNA-Seq Signatures Normalized by mRNA Abundance Allow Absolute Deconvolution of Human Immune Cell Types.The molecular characterization of immune subsets is important for designing effective strategies to understand and treat diseases. We characterized 29 immune cell types within the peripheral blood mononuclear cell (PBMC) fraction of healthy donors using RNA-seq (RNA sequencing) and flow cytometry. Our dataset was used, first, to identify sets of genes that are specific, are co-expressed, and have housekeeping roles across the 29 cell types. Then, we examined differences in mRNA heterogeneity and mRNA abundance revealing cell type specificity. Last, we performed absolute deconvolution on a suitable set of immune cell types using transcriptomics signatures normalized by mRNA abundance. Absolute deconvolution is ready to use for PBMC transcriptomic data using our Shiny app (https://github.com/giannimonaco/ABIS). We benchmarked different deconvolution and normalization methods and validated the resources in independent cohorts. Our work has research, clinical, and diagnostic value by makin
  266. 266Integrative molecular and clinical modeling of clinical outcomes to PD1 blockade in patients with metastatic melanoma.Immune-checkpoint blockade (ICB) has demonstrated efficacy in many tumor types, but predictors of responsiveness to anti-PD1 ICB are incompletely characterized. In this study, we analyzed a clinically annotated cohort of patients with melanoma (n = 144) treated with anti-PD1 ICB, with whole-exome and whole-transcriptome sequencing of pre-treatment tumors. We found that tumor mutational burden as a predictor of response was confounded by melanoma subtype, whereas multiple novel genomic and transcriptomic features predicted selective response, including features associated with MHC-I and MHC-II antigen presentation. Furthermore, previous anti-CTLA4 ICB exposure was associated with different predictors of response compared to tumors that were naive to ICB, suggesting selective immune effects of previous exposure to anti-CTLA4 ICB. Finally, we developed parsimonious models integrating clinical, genomic and transcriptomic features to predict intrinsic resistance to anti-PD1 ICB in individua
  267. 267Multimodal Analysis of Composition and Spatial Architecture in Human Squamous Cell Carcinoma.To define the cellular composition and architecture of cutaneous squamous cell carcinoma (cSCC), we combined single-cell RNA sequencing with spatial transcriptomics and multiplexed ion beam imaging from a series of human cSCCs and matched normal skin. cSCC exhibited four tumor subpopulations, three recapitulating normal epidermal states, and a tumor-specific keratinocyte (TSK) population unique to cancer, which localized to a fibrovascular niche. Integration of single-cell and spatial data mapped ligand-receptor networks to specific cell types, revealing TSK cells as a hub for intercellular communication. Multiple features of potential immunosuppression were observed, including T regulatory cell (Treg) co-localization with CD8 T cells in compartmentalized tumor stroma. Finally, single-cell characterization of human tumor xenografts and in vivo CRISPR screens identified essential roles for specific tumor subpopulation-enriched gene networks in tumorigenesis. These data define cSCC tumor
  268. 268Bulk tissue cell type deconvolution with multi-subject single-cell expression reference.Knowledge of cell type composition in disease relevant tissues is an important step towards the identification of cellular targets of disease. We present MuSiC, a method that utilizes cell-type specific gene expression from single-cell RNA sequencing (RNA-seq) data to characterize cell type compositions from bulk RNA-seq data in complex tissues. By appropriate weighting of genes showing cross-subject and cross-cell consistency, MuSiC enables the transfer of cell type-specific gene expression information from one dataset to another. When applied to pancreatic islet and whole kidney expression data in human, mouse, and rats, MuSiC outperformed existing methods, especially for tissues with closely related cell types. MuSiC enables the characterization of cellular heterogeneity of complex tissues for understanding of disease mechanisms. As bulk tissue data are more easily accessible than single-cell RNA-seq, MuSiC allows the utilization of the vast amounts of disease relevant bulk tissue R
  269. 269Single-cell RNA-seq enables comprehensive tumour and immune cell profiling in primary breast cancer.Single-cell transcriptome profiling of tumour tissue isolates allows the characterization of heterogeneous tumour cells along with neighbouring stromal and immune cells. Here we adopt this powerful approach to breast cancer and analyse 515 cells from 11 patients. Inferred copy number variations from the single-cell RNA-seq data separate carcinoma cells from non-cancer cells. At a single-cell resolution, carcinoma cells display common signatures within the tumour as well as intratumoral heterogeneity regarding breast cancer subtype and crucial cancer-related pathways. Most of the non-cancer cells are immune cells, with three distinct clusters of T lymphocytes, B lymphocytes and macrophages. T lymphocytes and macrophages both display immunosuppressive characteristics: T cells with a regulatory or an exhausted phenotype and macrophages with an M2 phenotype. These results illustrate that the breast cancer transcriptome has a wide range of intratumoral heterogeneity, which is shaped by the
  270. 270Single-cell RNA-seq denoising using a deep count autoencoder.Single-cell RNA sequencing (scRNA-seq) has enabled researchers to study gene expression at a cellular resolution. However, noise due to amplification and dropout may obstruct analyses, so scalable denoising methods for increasingly large but sparse scRNA-seq data are needed. We propose a deep count autoencoder network (DCA) to denoise scRNA-seq datasets. DCA takes the count distribution, overdispersion and sparsity of the data into account using a negative binomial noise model with or without zero-inflation, and nonlinear gene-gene dependencies are captured. Our method scales linearly with the number of cells and can, therefore, be applied to datasets of millions of cells. We demonstrate that DCA denoising improves a diverse set of typical scRNA-seq data analyses using simulated and real datasets. DCA outperforms existing methods for data imputation in quality and speed, enhancing biological discovery.
  271. 271Pan-cancer genome and transcriptome analyses of 1,699 paediatric leukaemias and solid tumours.Analysis of molecular aberrations across multiple cancer types, known as pan-cancer analysis, identifies commonalities and differences in key biological processes that are dysregulated in cancer cells from diverse lineages. Pan-cancer analyses have been performed for adult but not paediatric cancers, which commonly occur in developing mesodermic rather than adult epithelial tissues. Here we present a pan-cancer study of somatic alterations, including single nucleotide variants, small insertions or deletions, structural variations, copy number alterations, gene fusions and internal tandem duplications in 1,699 paediatric leukaemias and solid tumours across six histotypes, with whole-genome, whole-exome and transcriptome sequencing data processed under a uniform analytical framework. We report 142 driver genes in paediatric cancers, of which only 45% match those found in adult pan-cancer studies; copy number alterations and structural variants constituted the majority (62%) of events. El
  272. 272Single-Cell RNA-Seq Reveals Lineage and X Chromosome Dynamics in Human Preimplantation Embryos.Mouse studies have been instrumental in forming our current understanding of early cell-lineage decisions; however, similar insights into the early human development are severely limited. Here, we present a comprehensive transcriptional map of human embryo development, including the sequenced transcriptomes of 1,529 individual cells from 88 human preimplantation embryos. These data show that cells undergo an intermediate state of co-expression of lineage-specific genes, followed by a concurrent establishment of the trophectoderm, epiblast, and primitive endoderm lineages, which coincide with blastocyst formation. Female cells of all three lineages achieve dosage compensation of X chromosome RNA levels prior to implantation. However, in contrast to the mouse, XIST is transcribed from both alleles throughout the progression of this expression dampening, and X chromosome genes maintain biallelic expression while dosage compensation proceeds. We envision broad utility of this transcription
  273. 273BBKNN: fast batch alignment of single cell transcriptomes.Motivation Increasing numbers of large scale single cell RNA-Seq projects are leading to a data explosion, which can only be fully exploited through data integration. A number of methods have been developed to combine diverse datasets by removing technical batch effects, but most are computationally intensive. To overcome the challenge of enormous datasets, we have developed BBKNN, an extremely fast graph-based data integration algorithm. We illustrate the power of BBKNN on large scale mouse atlasing data, and favourably benchmark its run time against a number of competing methods. Availability and implementation BBKNN is available at https://github.com/Teichlab/bbknn, along with documentation and multiple example notebooks, and can be installed from pip. Supplementary information Supplementary data are available at Bioinformatics online.
  274. 274Molecular Diversity of Midbrain Development in Mouse, Human, and Stem Cells.Understanding human embryonic ventral midbrain is of major interest for Parkinson's disease. However, the cell types, their gene expression dynamics, and their relationship to commonly used rodent models remain to be defined. We performed single-cell RNA sequencing to examine ventral midbrain development in human and mouse. We found 25 molecularly defined human cell types, including five subtypes of radial glia-like cells and four progenitors. In the mouse, two mature fetal dopaminergic neuron subtypes diversified into five adult classes during postnatal development. Cell types and gene expression were generally conserved across species, but with clear differences in cell proliferation, developmental timing, and dopaminergic neuron development. Additionally, we developed a method to quantitatively assess the fidelity of dopaminergic neurons derived from human pluripotent stem cells, at a single-cell level. Thus, our study provides insight into the molecular programs controlling human m
  275. 275Splatter: simulation of single-cell RNA sequencing data.As single-cell RNA sequencing (scRNA-seq) technologies have rapidly developed, so have analysis methods. Many methods have been tested, developed, and validated using simulated datasets. Unfortunately, current simulations are often poorly documented, their similarity to real data is not demonstrated, or reproducible code is not available. Here, we present the Splatter Bioconductor package for simple, reproducible, and well-documented simulation of scRNA-seq data. Splatter provides an interface to multiple simulation methods including Splat, our own simulation, based on a gamma-Poisson distribution. Splat can simulate single populations of cells, populations with multiple cell types, or differentiation paths.
  276. 276Single-cell transcriptomics of human T cells reveals tissue and activation signatures in health and disease.Human T cells coordinate adaptive immunity in diverse anatomic compartments through production of cytokines and effector molecules, but it is unclear how tissue site influences T cell persistence and function. Here, we use single cell RNA-sequencing (scRNA-seq) to define the heterogeneity of human T cells isolated from lungs, lymph nodes, bone marrow and blood, and their functional responses following stimulation. Through analysis of >50,000 resting and activated T cells, we reveal tissue T cell signatures in mucosal and lymphoid sites, and lineage-specific activation states across all sites including distinct effector states for CD8 + T cells and an interferon-response state for CD4 + T cells. Comparing scRNA-seq profiles of tumor-associated T cells to our dataset reveals predominant activated CD8 + compared to CD4 + T cell states within multiple tumor types. Our results therefore establish a high dimensional reference map of human T cell activation in health for analyzing T cells in
  277. 277Spatially and functionally distinct subclasses of breast cancer-associated fibroblasts revealed by single cell RNA sequencing.Cancer-associated fibroblasts (CAFs) are a major constituent of the tumor microenvironment, although their origin and roles in shaping disease initiation, progression and treatment response remain unclear due to significant heterogeneity. Here, following a negative selection strategy combined with single-cell RNA sequencing of 768 transcriptomes of mesenchymal cells from a genetically engineered mouse model of breast cancer, we define three distinct subpopulations of CAFs. Validation at the transcriptional and protein level in several experimental models of cancer and human tumors reveal spatial separation of the CAF subclasses attributable to different origins, including the peri-vascular niche, the mammary fat pad and the transformed epithelium. Gene profiles for each CAF subtype correlate to distinctive functional programs and hold independent prognostic capability in clinical cohorts by association to metastatic disease. In conclusion, the improved resolution of the widely defined
  278. 278METTL3 facilitates tumor progression via an m 6 A-IGF2BP2-dependent mechanism in colorectal carcinoma.Background Colorectal carcinoma (CRC) is one of the most common malignant tumors, and its main cause of death is tumor metastasis. RNA N 6 -methyladenosine (m 6 A) is an emerging regulatory mechanism for gene expression and methyltransferase-like 3 (METTL3) participates in tumor progression in several cancer types. However, its role in CRC remains unexplored. Methods Western blot, quantitative real-time PCR (RT-qPCR) and immunohistochemical (IHC) were used to detect METTL3 expression in cell lines and patient tissues. Methylated RNA immunoprecipitation sequencing (MeRIP-seq) and transcriptomic RNA sequencing (RNA-seq) were used to screen the target genes of METTL3. The biological functions of METTL3 were investigated in vitro and in vivo. RNA pull-down and RNA immunoprecipitation assays were conducted to explore the specific binding of target genes. RNA stability assay was used to detect the half-lives of the downstream genes of METTL3. Results Using TCGA database, higher METTL3 expres
  279. 279The art of using t-SNE for single-cell transcriptomics.Single-cell transcriptomics yields ever growing data sets containing RNA expression levels for thousands of genes from up to millions of cells. Common data analysis pipelines include a dimensionality reduction step for visualising the data in two dimensions, most frequently performed using t-distributed stochastic neighbour embedding (t-SNE). It excels at revealing local structure in high-dimensional data, but naive applications often suffer from severe shortcomings, e.g. the global structure of the data is not represented accurately. Here we describe how to circumvent such pitfalls, and develop a protocol for creating more faithful t-SNE visualisations. It includes PCA initialisation, a high learning rate, and multi-scale similarity kernels; for very large data sets, we additionally use exaggeration and downsampling-based initialisation. We use published single-cell RNA-seq data sets to demonstrate that this protocol yields superior results compared to the naive application of t-SNE.
  280. 280Single-cell RNA-seq: advances and future challenges.Phenotypically identical cells can dramatically vary with respect to behavior during their lifespan and this variation is reflected in their molecular composition such as the transcriptomic landscape. Single-cell transcriptomics using next-generation transcript sequencing (RNA-seq) is now emerging as a powerful tool to profile cell-to-cell variability on a genomic scale. Its application has already greatly impacted our conceptual understanding of diverse biological processes with broad implications for both basic and clinical research. Different single-cell RNA-seq protocols have been introduced and are reviewed here-each one with its own strengths and current limitations. We further provide an overview of the biological questions single-cell RNA-seq has been used to address, the major findings obtained from such studies, and current challenges and expected future developments in this booming field.
  281. 281Single cell RNA sequencing of human microglia uncovers a subset associated with Alzheimer's disease.The extent of microglial heterogeneity in humans remains a central yet poorly explored question in light of the development of therapies targeting this cell type. Here, we investigate the population structure of live microglia purified from human cerebral cortex samples obtained at autopsy and during neurosurgical procedures. Using single cell RNA sequencing, we find that some subsets are enriched for disease-related genes and RNA signatures. We confirm the presence of four of these microglial subpopulations histologically and illustrate the utility of our data by characterizing further microglial cluster 7, enriched for genes depleted in the cortex of individuals with Alzheimer's disease (AD). Histologically, these cluster 7 microglia are reduced in frequency in AD tissue, and we validate this observation in an independent set of single nucleus data. Thus, our live human microglia identify a range of subtypes, and we prioritize one of these as being altered in AD.
  282. 282variancePartition: interpreting drivers of variation in complex gene expression studies.Background As large-scale studies of gene expression with multiple sources of biological and technical variation become widely adopted, characterizing these drivers of variation becomes essential to understanding disease biology and regulatory genetics. Results We describe a statistical and visualization framework, variancePartition, to prioritize drivers of variation based on a genome-wide summary, and identify genes that deviate from the genome-wide trend. Using a linear mixed model, variancePartition quantifies variation in each expression trait attributable to differences in disease status, sex, cell or tissue type, ancestry, genetic background, experimental stimulus, or technical variables. Analysis of four large-scale transcriptome profiling datasets illustrates that variancePartition recovers striking patterns of biological and technical variation that are reproducible across multiple datasets. Conclusions Our open source software, variancePartition, enables rapid interpretation
  283. 283Structural Remodeling of the Human Colonic Mesenchyme in Inflammatory Bowel Disease.Intestinal mesenchymal cells play essential roles in epithelial homeostasis, matrix remodeling, immunity, and inflammation. But the extent of heterogeneity within the colonic mesenchyme in these processes remains unknown. Using unbiased single-cell profiling of over 16,500 colonic mesenchymal cells, we reveal four subsets of fibroblasts expressing divergent transcriptional regulators and functional pathways, in addition to pericytes and myofibroblasts. We identified a niche population located in proximity to epithelial crypts expressing SOX6, F3 (CD142), and WNT genes essential for colonic epithelial stem cell function. In colitis, we observed dysregulation of this niche and emergence of an activated mesenchymal population. This subset expressed TNF superfamily member 14 (TNFSF14), fibroblastic reticular cell-associated genes, IL-33, and Lysyl oxidases. Further, it induced factors that impaired epithelial proliferation and maturation and contributed to oxidative stress and disease seve
  284. 284TopHat-Fusion: an algorithm for discovery of novel fusion transcripts.TopHat-Fusion is an algorithm designed to discover transcripts representing fusion gene products, which result from the breakage and re-joining of two different chromosomes, or from rearrangements within a chromosome. TopHat-Fusion is an enhanced version of TopHat, an efficient program that aligns RNA-seq reads without relying on existing annotation. Because it is independent of gene annotation, TopHat-Fusion can discover fusion products deriving from known genes, unknown genes and unannotated splice variants of known genes. Using RNA-seq data from breast and prostate cancer cell lines, we detected both previously reported and novel fusions with solid supporting evidence. TopHat-Fusion is available at http://tophat-fusion.sourceforge.net/.
  285. 285rnaSPAdes: a de novo transcriptome assembler and its application to RNA-Seq data.Background The possibility of generating large RNA-sequencing datasets has led to development of various reference-based and de novo transcriptome assemblers with their own strengths and limitations. While reference-based tools are widely used in various transcriptomic studies, their application is limited to the organisms with finished and well-annotated genomes. De novo transcriptome reconstruction from short reads remains an open challenging problem, which is complicated by the varying expression levels across different genes, alternative splicing, and paralogous genes. Results Herein we describe the novel transcriptome assembler rnaSPAdes, which has been developed on top of the SPAdes genome assembler and explores computational parallels between assembly of transcriptomes and single-cell genomes. We also present quality assessment reports for rnaSPAdes assemblies, compare it with modern transcriptome assembly tools using several evaluation approaches on various RNA-sequencing datas
  286. 286How many biological replicates are needed in an RNA-seq experiment and which differential expression tool should you use?RNA-seq is now the technology of choice for genome-wide differential gene expression experiments, but it is not clear how many biological replicates are needed to ensure valid biological interpretation of the results or which statistical tools are best for analyzing the data. An RNA-seq experiment with 48 biological replicates in each of two conditions was performed to answer these questions and provide guidelines for experimental design. With three biological replicates, nine of the 11 tools evaluated found only 20%-40% of the significantly differentially expressed (SDE) genes identified with the full set of 42 clean replicates. This rises to >85% for the subset of SDE genes changing in expression by more than fourfold. To achieve >85% for all SDE genes regardless of fold change requires more than 20 biological replicates. The same nine tools successfully control their false discovery rate at ≲5% for all numbers of replicates, while the remaining two tools fail to control their FDR ad
  287. 287Therapy-Induced Evolution of Human Lung Cancer Revealed by Single-Cell RNA Sequencing.Lung cancer, the leading cause of cancer mortality, exhibits heterogeneity that enables adaptability, limits therapeutic success, and remains incompletely understood. Single-cell RNA sequencing (scRNA-seq) of metastatic lung cancer was performed using 49 clinical biopsies obtained from 30 patients before and during targeted therapy. Over 20,000 cancer and tumor microenvironment (TME) single-cell profiles exposed a rich and dynamic tumor ecosystem. scRNA-seq of cancer cells illuminated targetable oncogenes beyond those detected clinically. Cancer cells surviving therapy as residual disease (RD) expressed an alveolar-regenerative cell signature suggesting a therapy-induced primitive cell-state transition, whereas those present at on-therapy progressive disease (PD) upregulated kynurenine, plasminogen, and gap-junction pathways. Active T-lymphocytes and decreased macrophages were present at RD and immunosuppressive cell states characterized PD. Biological features revealed by scRNA-seq we
  288. 288Transcriptomics technologies.Transcriptomics technologies are the techniques used to study an organism's transcriptome, the sum of all of its RNA transcripts. The information content of an organism is recorded in the DNA of its genome and expressed through transcription. Here, mRNA serves as a transient intermediary molecule in the information network, whilst noncoding RNAs perform additional diverse functions. A transcriptome captures a snapshot in time of the total transcripts present in a cell. The first attempts to study the whole transcriptome began in the early 1990s, and technological advances since the late 1990s have made transcriptomics a widespread discipline. Transcriptomics has been defined by repeated technological innovations that transform the field. There are two key contemporary techniques in the field: microarrays, which quantify a set of predetermined sequences, and RNA sequencing (RNA-Seq), which uses high-throughput sequencing to capture all sequences. Measuring the expression of an organism'
  289. 289TransRate: reference-free quality assessment of de novo transcriptome assemblies.TransRate is a tool for reference-free quality assessment of de novo transcriptome assemblies. Using only the sequenced reads and the assembly as input, we show that multiple common artifacts of de novo transcriptome assembly can be readily detected. These include chimeras, structural errors, incomplete assembly, and base errors. TransRate evaluates these errors to produce a diagnostic quality score for each contig, and these contig scores are integrated to evaluate whole assemblies. Thus, TransRate can be used for de novo assembly filtering and optimization as well as comparison of assemblies generated using different methods from the same input reads. Applying the method to a data set of 155 published de novo transcriptome assemblies, we deconstruct the contribution that assembly method, read length, read quantity, and read quality make to the accuracy of de novo transcriptome assemblies and reveal that variance in the quality of the input data explains 43% of the variance in the qua
  290. 290A comparison of methods for differential expression analysis of RNA-seq data.Background Finding genes that are differentially expressed between conditions is an integral part of understanding the molecular basis of phenotypic variation. In the past decades, DNA microarrays have been used extensively to quantify the abundance of mRNA corresponding to different genes, and more recently high-throughput sequencing of cDNA (RNA-seq) has emerged as a powerful competitor. As the cost of sequencing decreases, it is conceivable that the use of RNA-seq for differential expression analysis will increase rapidly. To exploit the possibilities and address the challenges posed by this relatively new type of data, a number of software packages have been developed especially for differential expression analysis of RNA-seq data. Results We conducted an extensive comparison of eleven methods for differential expression analysis of RNA-seq data. All methods are freely available within the R framework and take as input a matrix of counts, i.e. the number of reads mapping to each ge
  291. 291Decontamination of ambient RNA in single-cell RNA-seq with DecontX.Droplet-based microfluidic devices have become widely used to perform single-cell RNA sequencing (scRNA-seq). However, ambient RNA present in the cell suspension can be aberrantly counted along with a cell's native mRNA and result in cross-contamination of transcripts between different cell populations. DecontX is a novel Bayesian method to estimate and remove contamination in individual cells. DecontX accurately predicts contamination levels in a mouse-human mixture dataset and removes aberrant expression of marker genes in PBMC datasets. We also compare the contamination levels between four different scRNA-seq protocols. Overall, DecontX can be incorporated into scRNA-seq workflows to improve downstream analyses.
  292. 292Classification of low quality cells from single-cell RNA-seq data.Single-cell RNA sequencing (scRNA-seq) has broad applications across biomedical research. One of the key challenges is to ensure that only single, live cells are included in downstream analysis, as the inclusion of compromised cells inevitably affects data interpretation. Here, we present a generic approach for processing scRNA-seq data and detecting low quality cells, using a curated set of over 20 biological and technical features. Our approach improves classification accuracy by over 30 % compared to traditional methods when tested on over 5,000 cells, including CD4+ T cells, bone marrow dendritic cells, and mouse embryonic stem cells.
  293. 293Massive mining of publicly available RNA-seq data from human and mouse.RNA sequencing (RNA-seq) is the leading technology for genome-wide transcript quantification. However, publicly available RNA-seq data is currently provided mostly in raw form, a significant barrier for global and integrative retrospective analyses. ARCHS4 is a web resource that makes the majority of published RNA-seq data from human and mouse available at the gene and transcript levels. For developing ARCHS4, available FASTQ files from RNA-seq experiments from the Gene Expression Omnibus (GEO) were aligned using a cloud-based infrastructure. In total 187,946 samples are accessible through ARCHS4 with 103,083 mouse and 84,863 human. Additionally, the ARCHS4 web interface provides intuitive exploration of the processed data through querying tools, interactive visualization, and gene pages that provide average expression across cell lines and tissues, top co-expressed genes for each gene, and predicted biological functions and protein-protein interactions for each gene based on prior kno
  294. 294Identifying gene expression programs of cell-type identity and cellular activity with single-cell RNA-Seq.Identifying gene expression programs underlying both cell-type identity and cellular activities (e.g. life-cycle processes, responses to environmental cues) is crucial for understanding the organization of cells and tissues. Although single-cell RNA-Seq (scRNA-Seq) can quantify transcripts in individual cells, each cell's expression profile may be a mixture of both types of programs, making them difficult to disentangle. Here, we benchmark and enhance the use of matrix factorization to solve this problem. We show with simulations that a method we call consensus non-negative matrix factorization (cNMF) accurately infers identity and activity programs, including their relative contributions in each cell. To illustrate the insights this approach enables, we apply it to published brain organoid and visual cortex scRNA-Seq datasets; cNMF refines cell types and identifies both expected (e.g. cell cycle and hypoxia) and novel activity programs, including programs that may underlie a neurosecr
  295. 295RNA-seq analysis is easy as 1-2-3 with limma, Glimma and edgeR.The ability to easily and efficiently analyse RNA-sequencing data is a key strength of the Bioconductor project. Starting with counts summarised at the gene-level, a typical analysis involves pre-processing, exploratory data analysis, differential expression testing and pathway analysis with the results obtained informing future experiments and validation studies. In this workflow article, we analyse RNA-sequencing data from the mouse mammary gland, demonstrating use of the popular edgeR package to import, organise, filter and normalise the data, followed by the limma package with its voom method, linear modelling and empirical Bayes moderation to assess differential expression and perform gene set testing. This pipeline is further enhanced by the Glimma package which enables interactive exploration of the results so that individual samples and genes can be examined by the user. The complete analysis offered by these three packages highlights the ease with which researchers can turn th
  296. 296Single-cell RNA sequencing highlights the role of inflammatory cancer-associated fibroblasts in bladder urothelial carcinoma.Although substantial progress has been made in cancer biology and treatment, clinical outcomes of bladder carcinoma (BC) patients are still not satisfactory. The tumor microenvironment (TME) is a potential target. Here, by single-cell RNA sequencing on 8 BC tumor samples and 3 para tumor samples, we identify 19 different cell types in the BC microenvironment, indicating high intra-tumoral heterogeneity. We find that tumor cells down regulated MHC-II molecules, suggesting that the downregulated immunogenicity of cancer cells may contribute to the formation of an immunosuppressive microenvironment. We also find that monocytes undergo M2 polarization in the tumor region and differentiate. Furthermore, the LAMP3 + DC subgroup may be able to recruit regulatory T cells, potentially taking part in the formation of an immunosuppressive TME. Through correlation analysis using public datasets containing over 3000 BC samples, we identify a role for inflammatory cancer-associated fibroblasts (iCAF
  297. 297Single-nucleus and single-cell transcriptomes compared in matched cortical cell types.Transcriptomic profiling of complex tissues by single-nucleus RNA-sequencing (snRNA-seq) affords some advantages over single-cell RNA-sequencing (scRNA-seq). snRNA-seq provides less biased cellular coverage, does not appear to suffer cell isolation-based transcriptional artifacts, and can be applied to archived frozen specimens. We used well-matched snRNA-seq and scRNA-seq datasets from mouse visual cortex to compare cell type detection. Although more transcripts are detected in individual whole cells (~11,000 genes) than nuclei (~7,000 genes), we demonstrate that closely related neuronal cell types can be similarly discriminated with both methods if intronic sequences are included in snRNA-seq analysis. We estimate that the nuclear proportion of total cellular mRNA varies from 20% to over 50% for large and small pyramidal neurons, respectively. Together, these results illustrate the high information content of nuclear RNA for characterization of cellular diversity in brain tissues.
  298. 298Comprehensive evaluation of differential gene expression analysis methods for RNA-seq data.A large number of computational methods have been developed for analyzing differential gene expression in RNA-seq data. We describe a comprehensive evaluation of common methods using the SEQC benchmark dataset and ENCODE data. We consider a number of key features, including normalization, accuracy of differential expression detection and differential expression analysis when one condition has no detectable expression. We find significant differences among the methods, but note that array-based methods adapted to RNA-seq data perform comparably to methods designed for RNA-seq. Our results demonstrate that increasing the number of replicate samples significantly improves detection power over increased sequencing depth.
  299. 299TSCAN: Pseudo-time reconstruction and evaluation in single-cell RNA-seq analysis.When analyzing single-cell RNA-seq data, constructing a pseudo-temporal path to order cells based on the gradual transition of their transcriptomes is a useful way to study gene expression dynamics in a heterogeneous cell population. Currently, a limited number of computational tools are available for this task, and quantitative methods for comparing different tools are lacking. Tools for Single Cell Analysis (TSCAN) is a software tool developed to better support in silico pseudo-Time reconstruction in Single-Cell RNA-seq ANalysis. TSCAN uses a cluster-based minimum spanning tree (MST) approach to order cells. Cells are first grouped into clusters and an MST is then constructed to connect cluster centers. Pseudo-time is obtained by projecting each cell onto the tree, and the ordered sequence of cells can be used to study dynamic changes of gene expression along the pseudo-time. Clustering cells before MST construction reduces the complexity of the tree space. This often leads to improv
  300. 300ArrayExpress update - from bulk to single-cell expression data.ArrayExpress (https://www.ebi.ac.uk/arrayexpress) is an archive of functional genomics data from a variety of technologies assaying functional modalities of a genome, such as gene expression or promoter occupancy. The number of experiments based on sequencing technologies, in particular RNA-seq experiments, has been increasing over the last few years and submissions of sequencing data have overtaken microarray experiments in the last 12 months. Additionally, there is a significant increase in experiments investigating single cells, rather than bulk samples, known as single-cell RNA-seq. To accommodate these trends, we have substantially changed our submission tool Annotare which, along with raw and processed data, collects all metadata necessary to interpret these experiments. Selected datasets are re-processed and loaded into our sister resource, the value-added Expression Atlas (and its component Single Cell Expression Atlas), which not only enables users to interpret the data easily
  301. 301Microanatomy of the Human Atherosclerotic Plaque by Single-Cell Transcriptomics.Rationale Atherosclerotic lesions are known for their cellular heterogeneity, yet the molecular complexity within the cells of human plaques has not been fully assessed. Objective Using single-cell transcriptomics and chromatin accessibility, we gained a better understanding of the pathophysiology underlying human atherosclerosis. Methods and results We performed single-cell RNA and single-cell ATAC sequencing on human carotid atherosclerotic plaques to define the cells at play and determine their transcriptomic and epigenomic characteristics. We identified 14 distinct cell populations including endothelial cells, smooth muscle cells, mast cells, B cells, myeloid cells, and T cells and identified multiple cellular activation states and suggested cellular interconversions. Within the endothelial cell population, we defined subsets with angiogenic capacity plus clear signs of endothelial to mesenchymal transition. CD4 + and CD8 + T cells showed activation-based subclasses, each with a gr
  302. 302Integration of mapped RNA-Seq reads into automatic training of eukaryotic gene finding algorithm.We present a new approach to automatic training of a eukaryotic ab initio gene finding algorithm. With the advent of Next-Generation Sequencing, automatic training has become paramount, allowing genome annotation pipelines to keep pace with the speed of genome sequencing. Earlier we developed GeneMark-ES, currently the only gene finding algorithm for eukaryotic genomes that performs automatic training in unsupervised ab initio mode. The new algorithm, GeneMark-ET augments GeneMark-ES with a novel method that integrates RNA-Seq read alignments into the self-training procedure. Use of 'assembled' RNA-Seq transcripts is far from trivial; significant error rate of assembly was revealed in recent assessments. We demonstrated in computational experiments that the proposed method of incorporation of 'unassembled' RNA-Seq reads improves the accuracy of gene prediction; particularly, for the 1.3 GB genome of Aedes aegypti the mean value of prediction Sensitivity and Specificity at the gene leve
  303. 303Single-cell triple omics sequencing reveals genetic, epigenetic, and transcriptomic heterogeneity in hepatocellular carcinomas.Single-cell genome, DNA methylome, and transcriptome sequencing methods have been separately developed. However, to accurately analyze the mechanism by which transcriptome, genome and DNA methylome regulate each other, these omic methods need to be performed in the same single cell. Here we demonstrate a single-cell triple omics sequencing technique, scTrio-seq, that can be used to simultaneously analyze the genomic copy-number variations (CNVs), DNA methylome, and transcriptome of an individual mammalian cell. We show that large-scale CNVs cause proportional changes in RNA expression of genes within the gained or lost genomic regions, whereas these CNVs generally do not affect DNA methylation in these regions. Furthermore, we applied scTrio-seq to 25 single cancer cells derived from a human hepatocellular carcinoma tissue sample. We identified two subpopulations within these cells based on CNVs, DNA methylome, or transcriptome of individual cells. Our work offers a new avenue of disse
  304. 304Heavy Metal Tolerance in Plants: Role of Transcriptomics, Proteomics, Metabolomics, and Ionomics.Heavy metal contamination of soil and water causing toxicity/stress has become one important constraint to crop productivity and quality. This situation has further worsened by the increasing population growth and inherent food demand. It has been reported in several studies that counterbalancing toxicity due to heavy metal requires complex mechanisms at molecular, biochemical, physiological, cellular, tissue, and whole plant level, which might manifest in terms of improved crop productivity. Recent advances in various disciplines of biological sciences such as metabolomics, transcriptomics, proteomics, etc., have assisted in the characterization of metabolites, transcription factors, and stress-inducible proteins involved in heavy metal tolerance, which in turn can be utilized for generating heavy metal-tolerant crops. This review summarizes various tolerance strategies of plants under heavy metal toxicity covering the role of metabolites (metabolomics), trace elements (ionomics), tra
  305. 305Functionally distinct disease-associated fibroblast subsets in rheumatoid arthritis.Fibroblasts regulate tissue homeostasis, coordinate inflammatory responses, and mediate tissue damage. In rheumatoid arthritis (RA), synovial fibroblasts maintain chronic inflammation which leads to joint destruction. Little is known about fibroblast heterogeneity or if aberrations in fibroblast subsets relate to pathology. Here, we show functional and transcriptional differences between fibroblast subsets from human synovial tissues using bulk transcriptomics of targeted subpopulations and single-cell transcriptomics. We identify seven fibroblast subsets with distinct surface protein phenotypes, and collapse them into three subsets by integrating transcriptomic data. One fibroblast subset, characterized by the expression of proteins podoplanin, THY1 membrane glycoprotein and cadherin-11, but lacking CD34, is threefold expanded in patients with RA relative to patients with osteoarthritis. These fibroblasts localize to the perivascular zone in inflamed synovium, secrete proinflammatory
  306. 306Single-cell RNA-seq reveals that glioblastoma recapitulates a normal neurodevelopmental hierarchy.Cancer stem cells are critical for cancer initiation, development, and treatment resistance. Our understanding of these processes, and how they relate to glioblastoma heterogeneity, is limited. To overcome these limitations, we performed single-cell RNA sequencing on 53586 adult glioblastoma cells and 22637 normal human fetal brain cells, and compared the lineage hierarchy of the developing human brain to the transcriptome of cancer cells. We find a conserved neural tri-lineage cancer hierarchy centered around glial progenitor-like cells. We also find that this progenitor population contains the majority of the cancer's cycling cells, and, using RNA velocity, is often the originator of the other cell types. Finally, we show that this hierarchal map can be used to identify therapeutic targets specific to progenitor cancer stem cells. Our analyses show that normal brain development reconciles glioblastoma development, suggests a possible origin for glioblastoma hierarchy, and helps to id
  307. 307Single-cell expression profiling reveals dynamic flux of cardiac stromal, vascular and immune cells in health and injury.Besides cardiomyocytes (CM), the heart contains numerous interstitial cell types which play key roles in heart repair, regeneration and disease, including fibroblast, vascular and immune cells. However, a comprehensive understanding of this interactive cell community is lacking. We performed single-cell RNA-sequencing of the total non-CM fraction and enriched ( Pdgfra -GFP + ) fibroblast lineage cells from murine hearts at days 3 and 7 post-sham or myocardial infarction (MI) surgery. Clustering of >30,000 single cells identified >30 populations representing nine cell lineages, including a previously undescribed fibroblast lineage trajectory present in both sham and MI hearts leading to a uniquely activated cell state defined in part by a strong anti-WNT transcriptome signature. We also uncovered novel myofibroblast subtypes expressing either pro-fibrotic or anti-fibrotic signatures. Our data highlight non-linear dynamics in myeloid and fibroblast lineages after cardiac injury, and prov
  308. 308Single-cell transcriptomes identify human islet cell signatures and reveal cell-type-specific expression changes in type 2 diabetes.Blood glucose levels are tightly controlled by the coordinated action of at least four cell types constituting pancreatic islets. Changes in the proportion and/or function of these cells are associated with genetic and molecular pathophysiology of monogenic, type 1, and type 2 (T2D) diabetes. Cellular heterogeneity impedes precise understanding of the molecular components of each islet cell type that govern islet (dys)function, particularly the less abundant delta and gamma/pancreatic polypeptide (PP) cells. Here, we report single-cell transcriptomes for 638 cells from nondiabetic (ND) and T2D human islet samples. Analyses of ND single-cell transcriptomes identified distinct alpha, beta, delta, and PP/gamma cell-type signatures. Genes linked to rare and common forms of islet dysfunction and diabetes were expressed in the delta and PP/gamma cell types. Moreover, this study revealed that delta cells specifically express receptors that receive and coordinate systemic cues from the leptin,
  309. 309Single-cell analysis reveals fibroblast heterogeneity and myeloid-derived adipocyte progenitors in murine skin wounds.During wound healing in adult mouse skin, hair follicles and then adipocytes regenerate. Adipocytes regenerate from myofibroblasts, a specialized contractile wound fibroblast. Here we study wound fibroblast diversity using single-cell RNA-sequencing. On analysis, wound fibroblasts group into twelve clusters. Pseudotime and RNA velocity analyses reveal that some clusters likely represent consecutive differentiation states toward a contractile phenotype, while others appear to represent distinct fibroblast lineages. One subset of fibroblasts expresses hematopoietic markers, suggesting their myeloid origin. We validate this finding using single-cell western blot and single-cell RNA-sequencing on genetically labeled myofibroblasts. Using bone marrow transplantation and Cre recombinase-based lineage tracing experiments, we rule out cell fusion events and confirm that hematopoietic lineage cells give rise to a subset of myofibroblasts and rare regenerated adipocytes. In conclusion, our study
  310. 310Single-cell transcriptomes of the human skin reveal age-related loss of fibroblast priming.Fibroblasts are an essential cell population for human skin architecture and function. While fibroblast heterogeneity is well established, this phenomenon has not been analyzed systematically yet. We have used single-cell RNA sequencing to analyze the transcriptomes of more than 5,000 fibroblasts from a sun-protected area in healthy human donors. Our results define four main subpopulations that can be spatially localized and show differential secretory, mesenchymal and pro-inflammatory functional annotations. Importantly, we found that this fibroblast 'priming' becomes reduced with age. We also show that aging causes a substantial reduction in the predicted interactions between dermal fibroblasts and other skin cells, including undifferentiated keratinocytes at the dermal-epidermal junction. Our work thus provides evidence for a functional specialization of human dermal fibroblasts and identifies the partial loss of cellular identity as an important age-related change in the human derm
  311. 311Comprehensive analysis of normal adjacent to tumor transcriptomes.Histologically normal tissue adjacent to the tumor (NAT) is commonly used as a control in cancer studies. However, little is known about the transcriptomic profile of NAT, how it is influenced by the tumor, and how the profile compares with non-tumor-bearing tissues. Here, we integrate data from the Genotype-Tissue Expression project and The Cancer Genome Atlas to comprehensively analyze the transcriptomes of healthy, NAT, and tumor tissues in 6506 samples across eight tissues and corresponding tumor types. Our analysis shows that NAT presents a unique intermediate state between healthy and tumor. Differential gene expression and protein-protein interaction analyses reveal altered pathways shared among NATs across tissue types. We characterize a set of 18 genes that are specifically activated in NATs. By applying pathway and tissue composition analyses, we suggest a pan-cancer mechanism of pro-inflammatory signals from the tumor stimulates an inflammatory response in the adjacent endot
  312. 312Plasma exosome microRNAs are indicative of breast cancer.Background microRNAs are promising candidate breast cancer biomarkers due to their cancer-specific expression profiles. However, efforts to develop circulating breast cancer biomarkers are challenged by the heterogeneity of microRNAs in the blood. To overcome this challenge, we aimed to develop a molecular profile of microRNAs specifically secreted from breast cancer cells. Our first step towards this direction relates to capturing and analyzing the contents of exosomes, which are small secretory vesicles that selectively encapsulate microRNAs indicative of their cell of origin. To our knowledge, circulating exosome microRNAs have not been well-evaluated as biomarkers for breast cancer diagnosis or monitoring. Methods Exosomes were collected from the conditioned media of human breast cancer cell lines, mouse plasma of patient-derived orthotopic xenograft models (PDX), and human plasma samples. Exosomes were verified by electron microscopy, nanoparticle tracking analysis, and western bl
  313. 313Accurate detection of m 6 A RNA modifications in native RNA sequences.The epitranscriptomics field has undergone an enormous expansion in the last few years; however, a major limitation is the lack of generic methods to map RNA modifications transcriptome-wide. Here, we show that using direct RNA sequencing, N 6 -methyladenosine (m 6 A) RNA modifications can be detected with high accuracy, in the form of systematic errors and decreased base-calling qualities. Specifically, we find that our algorithm, trained with m 6 A-modified and unmodified synthetic sequences, can predict m 6 A RNA modifications with ~90% accuracy. We then extend our findings to yeast data sets, finding that our method can identify m 6 A RNA modifications in vivo with an accuracy of 87%. Moreover, we further validate our method by showing that these 'errors' are typically not observed in yeast ime4-knockout strains, which lack m 6 A modifications. Our results open avenues to investigate the biological roles of RNA modifications in their native RNA context.
  314. 314METTL14 suppresses proliferation and metastasis of colorectal cancer by down-regulating oncogenic long non-coding RNA XIST.Background N6-methyladenosine (m6A) is the most prevalent RNA epigenetic regulation in eukaryotic cells. However, understanding of m6A in colorectal cancer (CRC) is very limited. We designed this study to investigate the role of m6A in CRC. Methods Expression level of METTL14 was extracted from public database and tissue array to investigate the clinical relevance of METTL14 in CRC. Next, gain/loss of function experiment was used to define the role of METTL14 in the progression of CRC. Moreover, transcriptomic sequencing (RNA-seq) was applied to screen the potential targets of METTL14. The specific binding between METTL14 and presumed target was verified by RNA pull-down and RNA immunoprecipitation (RIP) assay. Furthermore, rescue experiment and methylated RNA immunoprecipitation (Me-RIP) were performed to uncover the mechanism. Results Clinically, loss of METTL14 correlated with unfavorable prognosis of CRC patients. Functionally, knockdown of METTL14 drastically enhanced proliferativ
  315. 315Single-nucleus transcriptome analysis reveals dysregulation of angiogenic endothelial cells and neuroprotective glia in Alzheimer's disease.Alzheimer's disease (AD) is the most common form of dementia but has no effective treatment. A comprehensive investigation of cell type-specific responses and cellular heterogeneity in AD is required to provide precise molecular and cellular targets for therapeutic development. Accordingly, we perform single-nucleus transcriptome analysis of 169,496 nuclei from the prefrontal cortical samples of AD patients and normal control (NC) subjects. Differential analysis shows that the cell type-specific transcriptomic changes in AD are associated with the disruption of biological processes including angiogenesis, immune activation, synaptic signaling, and myelination. Subcluster analysis reveals that compared to NC brains, AD brains contain fewer neuroprotective astrocytes and oligodendrocytes. Importantly, our findings show that a subpopulation of angiogenic endothelial cells is induced in the brain in patients with AD. These angiogenic endothelial cells exhibit increased expression of angiog
  316. 316Spatial maps of prostate cancer transcriptomes reveal an unexplored landscape of heterogeneity.Intra-tumor heterogeneity is one of the biggest challenges in cancer treatment today. Here we investigate tissue-wide gene expression heterogeneity throughout a multifocal prostate cancer using the spatial transcriptomics (ST) technology. Utilizing a novel approach for deconvolution, we analyze the transcriptomes of nearly 6750 tissue regions and extract distinct expression profiles for the different tissue components, such as stroma, normal and PIN glands, immune cells and cancer. We distinguish healthy and diseased areas and thereby provide insight into gene expression changes during the progression of prostate cancer. Compared to pathologist annotations, we delineate the extent of cancer foci more accurately, interestingly without link to histological changes. We identify gene expression gradients in stroma adjacent to tumor regions that allow for re-stratification of the tumor microenvironment. The establishment of these profiles is the first step towards an unbiased view of prosta
  317. 317muscat detects subpopulation-specific state transitions from multi-sample multi-condition single-cell transcriptomics data.Single-cell RNA sequencing (scRNA-seq) has become an empowering technology to profile the transcriptomes of individual cells on a large scale. Early analyses of differential expression have aimed at identifying differences between subpopulations to identify subpopulation markers. More generally, such methods compare expression levels across sets of cells, thus leading to cross-condition analyses. Given the emergence of replicated multi-condition scRNA-seq datasets, an area of increasing focus is making sample-level inferences, termed here as differential state analysis; however, it is not clear which statistical framework best handles this situation. Here, we surveyed methods to perform cross-condition differential state analyses, including cell-level mixed models and methods based on aggregated pseudobulk data. To evaluate method performance, we developed a flexible simulation that mimics multi-sample scRNA-seq data. We analyzed scRNA-seq data from mouse cortex cells to uncover subpop
  318. 318deFuse: an algorithm for gene fusion discovery in tumor RNA-Seq data.Gene fusions created by somatic genomic rearrangements are known to play an important role in the onset and development of some cancers, such as lymphomas and sarcomas. RNA-Seq (whole transcriptome shotgun sequencing) is proving to be a useful tool for the discovery of novel gene fusions in cancer transcriptomes. However, algorithmic methods for the discovery of gene fusions using RNA-Seq data remain underdeveloped. We have developed deFuse, a novel computational method for fusion discovery in tumor RNA-Seq data. Unlike existing methods that use only unique best-hit alignments and consider only fusion boundaries at the ends of known exons, deFuse considers all alignments and all possible locations for fusion boundaries. As a result, deFuse is able to identify fusion sequences with demonstrably better sensitivity than previous approaches. To increase the specificity of our approach, we curated a list of 60 true positive and 61 true negative fusion sequences (as confirmed by RT-PCR), and
  319. 319A diet-induced animal model of non-alcoholic fatty liver disease and hepatocellular cancer.Background & aims The lack of a preclinical model of progressive non-alcoholic steatohepatitis (NASH) that recapitulates human disease is a barrier to therapeutic development. Methods A stable isogenic cross between C57BL/6J (B6) and 129S1/SvImJ (S129) mice were fed a high fat diet with ad libitum consumption of glucose and fructose in physiologically relevant concentrations and compared to mice fed a chow diet and also to both parent strains. Results Following initiation of the obesogenic diet, B6/129 mice developed obesity, insulin resistance, hypertriglyceridemia and increased LDL-cholesterol. They sequentially also developed steatosis (4-8weeks), steatohepatitis (16-24weeks), progressive fibrosis (16weeks onwards) and spontaneous hepatocellular cancer (HCC). There was a strong concordance between the pattern of pathway activation at a transcriptomic level between humans and mice with similar histological phenotypes (FDR 0.02 for early and 0.08 for late time points). Lipogenic, infl
  320. 320A new view of transcriptome complexity and regulation through the lens of local splicing variations.Alternative splicing (AS) can critically affect gene function and disease, yet mapping splicing variations remains a challenge. Here, we propose a new approach to define and quantify mRNA splicing in units of local splicing variations (LSVs). LSVs capture previously defined types of alternative splicing as well as more complex transcript variations. Building the first genome wide map of LSVs from twelve mouse tissues, we find complex LSVs constitute over 30% of tissue dependent transcript variations and affect specific protein families. We show the prevalence of complex LSVs is conserved in humans and identify hundreds of LSVs that are specific to brain subregions or altered in Alzheimer's patients. Amongst those are novel isoforms in the Camk2 family and a novel poison exon in Ptbp1, a key splice factor in neurogenesis. We anticipate the approach presented here will advance the ability to relate tissue-specific splice variation to genetic variation, phenotype, and disease.
  321. 321Gene Regulatory Network Inference from Single-Cell Data Using Multivariate Information Measures.While single-cell gene expression experiments present new challenges for data processing, the cell-to-cell variability observed also reveals statistical relationships that can be used by information theory. Here, we use multivariate information theory to explore the statistical dependencies between triplets of genes in single-cell gene expression datasets. We develop PIDC, a fast, efficient algorithm that uses partial information decomposition (PID) to identify regulatory relationships between genes. We thoroughly evaluate the performance of our algorithm and demonstrate that the higher-order information captured by PIDC allows it to outperform pairwise mutual information-based algorithms when recovering true relationships present in simulated data. We also infer gene regulatory networks from three experimental single-cell datasets and illustrate how network context, choices made during analysis, and sources of variability affect network inference. PIDC tutorials and open-source softwa
  322. 322The Mount Sinai cohort of large-scale genomic, transcriptomic and proteomic data in Alzheimer's disease.Alzheimer's disease (AD) affects half the US population over the age of 85 and is universally fatal following an average course of 10 years of progressive cognitive disability. Genetic and genome-wide association studies (GWAS) have identified about 33 risk factor genes for common, late-onset AD (LOAD), but these risk loci fail to account for the majority of affected cases and can neither provide clinically meaningful prediction of development of AD nor offer actionable mechanisms. This cohort study generated large-scale matched multi-Omics data in AD and control brains for exploring novel molecular underpinnings of AD. Specifically, we generated whole genome sequencing, whole exome sequencing, transcriptome sequencing and proteome profiling data from multiple regions of 364 postmortem control, mild cognitive impaired (MCI) and AD brains with rich clinical and pathophysiological data. All the data went through rigorous quality control. Both the raw and processed data are publicly avail
  323. 323Genetic diagnosis of Mendelian disorders via RNA sequencing.Across a variety of Mendelian disorders, ∼50-75% of patients do not receive a genetic diagnosis by exome sequencing indicating disease-causing variants in non-coding regions. Although genome sequencing in principle reveals all genetic variants, their sizeable number and poorer annotation make prioritization challenging. Here, we demonstrate the power of transcriptome sequencing to molecularly diagnose 10% (5 of 48) of mitochondriopathy patients and identify candidate genes for the remainder. We find a median of one aberrantly expressed gene, five aberrant splicing events and six mono-allelically expressed rare variants in patient-derived fibroblasts and establish disease-causing roles for each kind. Private exons often arise from cryptic splice sites providing an important clue for variant prioritization. One such event is found in the complex I assembly factor TIMMDC1 establishing a novel disease-associated gene. In conclusion, our study expands the diagnostic tools for detecting non-
  324. 324Human whole genome genotype and transcriptome data for Alzheimer's and other neurodegenerative diseases.Previous genome-wide association studies (GWAS), conducted by our group and others, have identified loci that harbor risk variants for neurodegenerative diseases, including Alzheimer's disease (AD). Human disease variants are enriched for polymorphisms that affect gene expression, including some that are known to associate with expression changes in the brain. Postulating that many variants confer risk to neurodegenerative disease via transcriptional regulatory mechanisms, we have analyzed gene expression levels in the brain tissue of subjects with AD and related diseases. Herein, we describe our collective datasets comprised of GWAS data from 2,099 subjects; microarray gene expression data from 773 brain samples, 186 of which also have RNAseq; and an independent cohort of 556 brain samples with RNAseq. We expect that these datasets, which are available to all qualified researchers, will enable investigators to explore and identify transcriptional mechanisms contributing to neurodegene
  325. 325Nanopore direct RNA sequencing maps the complexity of Arabidopsis mRNA processing and m<sup>6</sup>A modification.Understanding genome organization and gene regulation requires insight into RNA transcription, processing and modification. We adapted nanopore direct RNA sequencing to examine RNA from a wild-type accession of the model plant Arabidopsis thaliana and a mutant defective in mRNA methylation (m 6 A). Here we show that m 6 A can be mapped in full-length mRNAs transcriptome-wide and reveal the combinatorial diversity of cap-associated transcription start sites, splicing events, poly(A) site choice and poly(A) tail length. Loss of m 6 A from 3' untranslated regions is associated with decreased relative transcript abundance and defective RNA 3' end formation. A functional consequence of disrupted m 6 A is a lengthening of the circadian period. We conclude that nanopore direct RNA sequencing can reveal the complexity of mRNA processing and modification in full-length single molecule reads. These findings can refine Arabidopsis genome annotation. Further, applying this approach to less well-st
  326. 326scRNA-seq Profiling of Human Testes Reveals the Presence of the ACE2 Receptor, A Target for SARS-CoV-2 Infection in Spermatogonia, Leydig and Sertoli Cells.In December 2019, a novel coronavirus (SARS-CoV-2) was identified in COVID-19 patients in Wuhan, Hubei Province, China. SARS-CoV-2 shares both high sequence similarity and the use of the same cell entry receptor, angiotensin-converting enzyme 2 (ACE2), with severe acute respiratory syndrome coronavirus (SARS-CoV). Several studies have provided bioinformatic evidence of potential routes of SARS-CoV-2 infection in respiratory, cardiovascular, digestive and urinary systems. However, whether the reproductive system is a potential target of SARS-CoV-2 infection has not yet been determined. Here, we investigate the expression pattern of ACE2 in adult human testes at the level of single-cell transcriptomes. The results indicate that ACE2 is predominantly enriched in spermatogonia and Leydig and Sertoli cells. Gene Set Enrichment Analysis (GSEA) indicates that Gene Ontology (GO) categories associated with viral reproduction and transmission are highly enriched in ACE2-positive spermatogonia, w
  327. 327Unravelling subclonal heterogeneity and aggressive disease states in TNBC through single-cell RNA-seq.Triple-negative breast cancer (TNBC) is an aggressive subtype characterized by extensive intratumoral heterogeneity. To investigate the underlying biology, we conducted single-cell RNA-sequencing (scRNA-seq) of >1500 cells from six primary TNBC. Here, we show that intercellular heterogeneity of gene expression programs within each tumor is variable and largely correlates with clonality of inferred genomic copy number changes, suggesting that genotype drives the gene expression phenotype of individual subpopulations. Clustering of gene expression profiles identified distinct subgroups of malignant cells shared by multiple tumors, including a single subpopulation associated with multiple signatures of treatment resistance and metastasis, and characterized functionally by activation of glycosphingolipid metabolism and associated innate immunity pathways. A novel signature defining this subpopulation predicts long-term outcomes for TNBC patients in a large cohort. Collectively, this analys
  328. 328Characterisation of the transcriptome and proteome of SARS-CoV-2 reveals a cell passage induced in-frame deletion of the furin-like cleavage site from the spike glycoprotein.Background SARS-CoV-2 is a recently emerged respiratory pathogen that has significantly impacted global human health. We wanted to rapidly characterise the transcriptomic, proteomic and phosphoproteomic landscape of this novel coronavirus to provide a fundamental description of the virus's genomic and proteomic potential. Methods We used direct RNA sequencing to determine the transcriptome of SARS-CoV-2 grown in Vero E6 cells which is widely used to propagate the novel coronavirus. The viral transcriptome was analysed using a recently developed ORF-centric pipeline. Allied to this, we used tandem mass spectrometry to investigate the proteome and phosphoproteome of the same virally infected cells. Results Our integrated analysis revealed that the viral transcripts (i.e. subgenomic mRNAs) generally fitted the expected transcription model for coronaviruses. Importantly, a 24 nt in-frame deletion was detected in over half of the subgenomic mRNAs encoding the spike (S) glycoprotein and was
RNA-seq循证手册 · GeniOmics