RNA-Seq analysis workshop

Size: px
Start display at page:

Download "RNA-Seq analysis workshop"

Transcription

1 RNA-Seq analysis workshop Zhangjun Fei Boyce Thompson Institute for Plant Research USDA Robert W. Holley Center for Agriculture and Health Cornell University

2 Outline Background of RNA-Seq Application of RNA-Seq (what RNA-Seq can do?) Available sequencing platforms and strategies and which one to choose RNA-Seq data analysis Read processing and quality assessment De novo assembly Alignment to reference genome/transcriptome Differentially expressed gene identification

3 Milestones of Transcriptome analysis Year Milestone 1965 Sequence of the first RNA molecule determined 1977 Development of the Northern blot technique and the Sanger sequencing method 1989 Reports of RT-PCR experiments for transcriptome analysis 1991 First high-throughput EST sequencing study 1992 Introduction of Differential Display for the discovery of differentially expressed genes 1995 Reports of the microarray and Serial Analysis of Gene Expression (SAGE) methods 1996 Suppression subtractive hybridization reported 2005 First next-generation sequencing technology (Roche/454) introduced to the market 2006 First transcriptome sequencing studies using a next-generation technology (Roche/454)

4 New sequencing technologies Next generation sequencing Illumina (HiSeq 2000/2500) Roche/454 Ion Torrent (Ion Proton) ABI/SOLiD Helicos Third generation sequencing Pacific Biosciences Oxford Nanopore Complete Genomics Desktop sequencer Ion Torrent PGM Illumina MiSeq 454 GS Junior

5 RNA-Seq applications

6 RNA-Seq application Accelerating gene discovery and gene family expansion Improving genome annotation identifying novel genes and gene models Identifying tissue/condition specific alternative splicing events

7 RNA-Seq applications Alternative splicing Short reads can t provide the complete structure of an isoform

8 PacBio long reads RNA-Seq applications

9 RNA-Seq applications PacBio long reads error correction

10 Each sample needs four libraries with different insert sizes: 1-2K, 2-3K, 3-5K, >5K RNA-Seq applications

11 RNA-Seq applications

12 RNA-Seq applications Cell 1 Cell 2 No. reads 86,126 80,543 Total base 527,933, ,348,201 Average length 6,129 5,914

13 RNA-Seq applications SNP and SSR marker identification facilitating breeding SNP discovery in RNA-Seq is more challenging than in DNA: Varying levels of coverage depth False discovery around splicing junctions due to incorrect mapping

14 RNA-Seq applications Phylogenetic relationship, population structure, selective sweep

15 RNA-Seq applications Expression QTL Distribution of SNPs (blue) and differentially expressed (DE) genes in IL10-1

16 RNA-Seq applications Mutant gene cloning (BSA RNA-Seq) white fruit x yellow fruit 132 of 189 SNPs in this region F1 F2 kb F3 white pool yellow pool RNA-Seq SNPs and DE genes Feder et al. (2015) A Kelch domain-containing F-box coding gene negatively regulates flavonoid accumulation in Cucumis melo L. Plant Physiol 169:

17 RNA-Seq applications GWAS Distribution of mapped markers associating with the erucic acid trait

18 RNA-Seq applications Genomic imprinting and allele specific expression

19 RNA-Seq applications non-coding RNAs (lncrna, lincrnas )

20 Gene fusion RNA-Seq applications

21 Gene expression profiling RNA-Seq applications

22 RNA-Seq vs microarray Problem of microarray Cross-hybridization Stable probe secondary structures high background (e.g., nonspecific hybridization) limited dynamic range (e.g., nonlinear and saturable hybridization kinetics) RNA-Seq (digital expression analysis) allow direct enumeration of transcript molecules digital expression data are absolute so data can be directly compared across different experiments and laboratories without the need for extensive internal controls or other experimental manipulation provide open systems that allow detection of previously uncharacterized transcripts, as well as rare transcripts

23 RNA-Seq vs microarray high background (e.g., nonspecific hybridization) limited dynamic range (e.g., nonlinear and saturable hybridization kinetics)

24 RNA-Seq applications Summary Accelerating gene discovery and gene family expansion Improving genome annotation identifying novel genes and gene models Identifying tissue/condition specific alternative splicing events SNP and SSR marker identification Phylogenetic relationship, population structure, selective sweep Expression QTL analysis Mutant gene cloning (BSA RNA-Seq) Genome (Transcriptome)-wide associate study Genomic imprinting and allele specific expression analysis Identifying non-coding RNAs (lncrna, lincrnas ) Identifying gene fusion events Gene expression profiling analysis

25 Sequencing platforms and strategies

26 Sequencing platforms Next generation sequencing Illumina (HiSeq 2000/2500) Ion Torrent (Ion Proton) ABI/SOLiD Roche/454 Helicos Third generation sequencing Pacific Biosciences Oxford Nanopore Complete Genomics Desktop sequencer Ion Torrent PGM Illumina MiSeq Illumina NextSeq 454 GS Junior

27 Sequencing platforms Illumina HiSeq 2000/2500 High-output mode ( M reads/ read pairs per lane) Single-end, 50, 100 bp Paired-end, 2 x 125bp Run time: 2-11 days Rapid run mode ( M reads/ read pairs per lane) Single-end, 50, 100, 150 bp Paired-end, 2 x 100 bp Paired-end, 2 x 150 bp Paired-end, 2 x 200 bp Paired-end, 2 x 250 bp Runtime: 7-40 hours Illumina MiSeq 50 bp sequencing kit 300 bp sequencing kit (e.g. 2 x 150 bp) 500 bp sequencing kit (e.g. 2 x 250 bp) 150 bp sequencing kit (e.g. 2 x 75 bp) 600 bp sequencing kit (e.g. 2 x 300 bp) Run time: 5-65 hours

28 Sequencing platforms Single-end or paired-end For gene expression analysis with a reference genome, singleend is enough For de novo assembly, genome annotation, alternative splicing identification, it s better to use paired-end Strand-specific or non strand-specific Always choose strand-specific RNA-Seq if possible

29 Strand-specific RNA sequencing More accurately determine the expression level Significantly reduce false positives in identifying alternatively spliced transcripts Identify antisense transcripts another level of gene regulation in important biological processes Determine the transcribed strand of non-coding RNAs (e.g. lincrnas)

30 Strand-specific RNA-Seq library construction

31 High throughput ssrna-seq Up to 96 libraries in two days Paired-end compatible multiplexing

32 Strand specific RNA sequencing Strand-specific sequencing can produce more accurate digital gene expression data when compared to the conventional Illumina RNA-Seq.

33 Strand specific RNA sequencing

34 Strand specific RNA sequencing Antisense transcript cis-natural antisense transcripts (cis-nat) 1340 cis-nat pairs in Arabidopsis (Wang et al., 2005) 687 cis-nat pairs in rice (Osato et al., 2003) trans-natural antisense transcripts (trans-nat) 1,320 trans-nat pairs in Arabidopsis (Wang et al., 2006) function alternative splicing RNA editing DNA methylation genomic imprinting X-chromosome inactivation

35 Strand specific RNA sequencing Antisense transcript LEFL2040O reads 259 reads LEFL2002DC reads 1189 reads

36 lincrna (determine the sense strand) Strand specific RNA sequencing

37 RNA-Seq strategies Sequencing depth and no. of biological replicates Most frequently asked question How many samples should I multiplex in one lane? or How many reads should I generate for each of my samples? Depend on $$$ Depends on the quality of the library and the reads rrna, trna, organelle, adaptor contamination No. of biological replicates for expression call At least three Effects of read numbers on expression call Mature green fruit library (22M reads) Randomly select , 1-22M reads from the library and calculate gene expression for each dataset (20 different randomizations)

38 RNA-Seq (multiplexing) 0.1M 1M 2M r= r= r= M 5M 10M r= r= r= Mature green fruit, 22M

39 RNA-Seq (multiplexing)

40 RNA-Seq (multiplexing)

41 RNA-Seq data analysis

42 Read quality control (fastqc) Read processing

43 Read quality control (fastqc) Read processing

44 Read quality control (fastqc) Read processing

45 Read processing Remove adaptors and all possible contaminations: rrna, trna, organelle (chloroplast and mitochondrion) RNAs, virus, low quality sequences Arabidopsis 25S ribosomal RNA vs GenBank nr protein database

46 Read processing Remove contaminated sequences Align reads to rrna and organelle sequence database (bowtie or BWA) Affect RPKM values if not removed Trim adaptor and low quality sequences FASTX-Toolkit AdapterRemoval Trimmomatic Cutadapt Condetri ERNE-filter Prinseq SolexaQA-bwa Sickle

47 Read processing

48 RNA-Seq data analysis De novo transcriptome assembly Long reads (454/Sanger) overlap-layout-consensus strategy Short reads (Illumina) de Bruijn graph approach Martin & Wang, 2011

49 De novo transcriptome assembly Long reads (454/Sanger) CAP3 ( TGICL/CAP3 ( MIRA ( Newbler (-cdna) Phrap ( Two major problems in existing EST assembly programs and unigene databases: 1) Large portion of different transcripts (mainly alternative spliced transcripts and paralogs) are incorrectly assembled into same transcripts type I error (false positives) 2) Large portion of nearly identical sequences are not assembled into one transcript type II error (false negatives)

50 Example of type I assembly error (paralog) In DFCI Tomato Gene Index, AW is a member of TC Sequence identity between AW and TC232370: 91.5% AW is aligned to tomato chromosome 4 TC is aligned to tomato chromosome 11

51 Example of type I assembly error (alternative splicing) In DFCI Tomato Gene Index, U95008 is a member of TC226520

52 Example of type II assembly error In DFCI Tomato Gene Index, two unigenes, TC and TC221582, are identical

53 iassembler iterative assemblies (assembly of assemblies) using MIRA and CAP3 (four cycles of MIRA followed by one cycle of CAP3) reduce errors that nearly identical sequences are not assembled Further assembly error identification 1) comparing unigene sequences against themselves to identify nearly identical sequences (type II errors) 2) aligning EST sequences to their corresponding unigene sequences to identify mis-assembled ESTs (type I errors) Both type I and II assembly errors are corrected automatically by the program Unigene base errors are then corrected based on the resulting SAM files

54 iassembler performance A curated Arabidopsis EST dataset, which only contain ESTs that can be perfectly aligned to the TAIR10 cdnas perfectly aligned means that the sequences were aligned to Arabidopsis cdnas in their entire lengths

55 De novo transcriptome assembly Short reads (Illumina) Trinity Trans-ABySS Oases/velvet SOAPdenovo-Trans

56 De novo transcriptome assembly Reference-guided de novo assembly Cufflink IsoLasso Scripture Traph StringTie

57 De novo transcriptome assembly Trinity

58 De novo transcriptome assembly Post processing of de novo assemblies Remove contaminations (bacteria, virus, fungus ) Remove assembly errors (mainly redundancy) Remove errors caused by library preparation (incomplete digestion of dutp containing 2 nd strand during strandspecific RNA-Seq library construction)

59 De novo transcriptome assembly blastx Remove contamination blastn

60 De novo transcriptome assembly Remove contamination DeconSeq SeqClean

61 De novo transcriptome assembly Remove type II assembly error (redundancy) iassembler

62 De novo transcriptome assembly Remove transcripts derived from incomplete 2 nd digestion Gene ID length antisense sense UN comp38294_c0_seq removed

63 De novo transcriptome assembly High number of assembled transcripts Alternative splicing Non-coding RNAs Incomplete coverage of full length transcripts DFCI gene index

64 RNA-Seq data analysis Alignment Align reads to reference genome TopHat HISAT Alignment reads to reference transcriptome bowtie BWA If you have a reference genome, it s not a good idea to align the reads to the predicted CDS or cdna, due to the incomplete prediction of UTRs and alternative splicing

65 RNA-Seq data analysis Visualization tools Integrative Genomics Viewer (IGV)

66 RNA-Seq data analysis Read counting and normalization Read counting htseq-count samtools (samtools view c) Normalization RPKM: reads per kilobase of exon model per million mapped reads FPKM: fragments per kilobase of exon model per million mapped reads

67 RNA-Seq data analysis Quality control biological replicates Sample correlation matrix

68 RNA-Seq data analysis Differentially expressed gene detection Pair-wise comparison DESeq edger Time course data first data transformation using getvariancestabilizeddata function in DESeq (to get normal distribution). Then DE gene identification using F tests in LIMMA Multiple test correction False Discovery Rate (FDR) q value

69 RNA-Seq data analysis Differentially expressed gene detection

RNA-Seq analysis workshop. Zhangjun Fei

RNA-Seq analysis workshop. Zhangjun Fei RNA-Seq analysis workshop Zhangjun Fei Outline Background of RNA-Seq Application of RNA-Seq (what RNA-Seq can do?) Available sequencing platforms and strategies and which one to choose RNA-Seq data analysis

More information

Introduction to RNA-Seq

Introduction to RNA-Seq Introduction to RNA-Seq Monica Britton, Ph.D. Sr. Bioinformatics Analyst March 2015 Workshop Overview of RNA-Seq Activities RNA-Seq Concepts, Terminology, and Work Flows Using Single-End Reads and a Reference

More information

measuring gene expression December 5, 2017

measuring gene expression December 5, 2017 measuring gene expression December 5, 2017 transcription a usually short-lived RNA copy of the DNA is created through transcription RNA is exported to the cytoplasm to encode proteins some types of RNA

More information

Introduction to RNA sequencing

Introduction to RNA sequencing Introduction to RNA sequencing Bioinformatics perspective Olga Dethlefsen NBIS, National Bioinformatics Infrastructure Sweden November 2017 Olga (NBIS) RNA-seq November 2017 1 / 49 Outline Why sequence

More information

Introduction to transcriptome analysis using High Throughput Sequencing technologies. D. Puthier 2012

Introduction to transcriptome analysis using High Throughput Sequencing technologies. D. Puthier 2012 Introduction to transcriptome analysis using High Throughput Sequencing technologies D. Puthier 2012 A typical RNA-Seq experiment Library construction Protocol variations Fragmentation methods RNA: nebulization,

More information

RNA-Seq Workshop AChemS Sunil K Sukumaran Monell Chemical Senses Center Philadelphia

RNA-Seq Workshop AChemS Sunil K Sukumaran Monell Chemical Senses Center Philadelphia RNA-Seq Workshop AChemS 2017 Sunil K Sukumaran Monell Chemical Senses Center Philadelphia Benefits & downsides of RNA-Seq Benefits: High resolution, sensitivity and large dynamic range Independent of prior

More information

Next-Generation Sequencing. Technologies

Next-Generation Sequencing. Technologies Next-Generation Next-Generation Sequencing Technologies Sequencing Technologies Nicholas E. Navin, Ph.D. MD Anderson Cancer Center Dept. Genetics Dept. Bioinformatics Introduction to Bioinformatics GS011062

More information

RNA-Seq Software, Tools, and Workflows

RNA-Seq Software, Tools, and Workflows RNA-Seq Software, Tools, and Workflows Monica Britton, Ph.D. Sr. Bioinformatics Analyst September 1, 2016 Some mrna-seq Applications Differential gene expression analysis Transcriptional profiling Assumption:

More information

RNA-Sequencing analysis

RNA-Sequencing analysis RNA-Sequencing analysis Markus Kreuz 25. 04. 2012 Institut für Medizinische Informatik, Statistik und Epidemiologie Content: Biological background Overview transcriptomics RNA-Seq RNA-Seq technology Challenges

More information

RNA-Seq with the Tuxedo Suite

RNA-Seq with the Tuxedo Suite RNA-Seq with the Tuxedo Suite Monica Britton, Ph.D. Sr. Bioinformatics Analyst September 2015 Workshop The Basic Tuxedo Suite References Trapnell C, et al. 2009 TopHat: discovering splice junctions with

More information

Third Generation Sequencing

Third Generation Sequencing Third Generation Sequencing By Mohammad Hasan Samiee Aref Medical Genetics Laboratory of Dr. Zeinali History of DNA sequencing 1953 : Discovery of DNA structure by Watson and Crick 1973 : First sequence

More information

Gene Expression Technology

Gene Expression Technology Gene Expression Technology Bing Zhang Department of Biomedical Informatics Vanderbilt University bing.zhang@vanderbilt.edu Gene expression Gene expression is the process by which information from a gene

More information

Research school methods seminar Genomics and Transcriptomics

Research school methods seminar Genomics and Transcriptomics Research school methods seminar Genomics and Transcriptomics Stephan Klee 19.11.2014 2 3 4 5 Genetics, Genomics what are we talking about? Genetics and Genomics Study of genes Role of genes in inheritence

More information

Long and short/small RNA-seq data analysis

Long and short/small RNA-seq data analysis Long and short/small RNA-seq data analysis GEF5, 4.9.2015 Sami Heikkinen, PhD, Dos. Topics 1. RNA-seq in a nutshell 2. Long vs short/small RNA-seq 3. Bioinformatic analysis work flows GEF5 / Heikkinen

More information

RNA Seq: Methods and Applica6ons. Prat Thiru

RNA Seq: Methods and Applica6ons. Prat Thiru RNA Seq: Methods and Applica6ons Prat Thiru 1 Outline Intro to RNA Seq Biological Ques6ons Comparison with Other Methods RNA Seq Protocol RNA Seq Applica6ons Annota6on Quan6fica6on Other Applica6ons Expression

More information

Next Generation Sequencing. Jeroen Van Houdt - Leuven 13/10/2017

Next Generation Sequencing. Jeroen Van Houdt - Leuven 13/10/2017 Next Generation Sequencing Jeroen Van Houdt - Leuven 13/10/2017 Landmarks in DNA sequencing 1953 Discovery of DNA double helix structure 1977 A Maxam and W Gilbert "DNA seq by chemical degradation" F Sanger"DNA

More information

Mapping strategies for sequence reads

Mapping strategies for sequence reads Mapping strategies for sequence reads Ernest Turro University of Cambridge 21 Oct 2013 Quantification A basic aim in genomics is working out the contents of a biological sample. 1. What distinct elements

More information

Sanger vs Next-Gen Sequencing

Sanger vs Next-Gen Sequencing Tools and Algorithms in Bioinformatics GCBA815/MCGB815/BMI815, Fall 2017 Week-8: Next-Gen Sequencing RNA-seq Data Analysis Babu Guda, Ph.D. Professor, Genetics, Cell Biology & Anatomy Director, Bioinformatics

More information

Incorporating Molecular ID Technology. Accel-NGS 2S MID Indexing Kits

Incorporating Molecular ID Technology. Accel-NGS 2S MID Indexing Kits Incorporating Molecular ID Technology Accel-NGS 2S MID Indexing Kits Molecular Identifiers (MIDs) MIDs are indices used to label unique library molecules MIDs can assess duplicate molecules in sequencing

More information

Bioinformatics Advice on Experimental Design

Bioinformatics Advice on Experimental Design Bioinformatics Advice on Experimental Design Where do I start? Please refer to the following guide to better plan your experiments for good statistical analysis, best suited for your research needs. Statistics

More information

SMARTer Ultra Low RNA Kit for Illumina Sequencing Two powerful technologies combine to enable sequencing with ultra-low levels of RNA

SMARTer Ultra Low RNA Kit for Illumina Sequencing Two powerful technologies combine to enable sequencing with ultra-low levels of RNA SMARTer Ultra Low RNA Kit for Illumina Sequencing Two powerful technologies combine to enable sequencing with ultra-low levels of RNA The most sensitive cdna synthesis technology, combined with next-generation

More information

Next Gen Sequencing. Expansion of sequencing technology. Contents

Next Gen Sequencing. Expansion of sequencing technology. Contents Next Gen Sequencing Contents 1 Expansion of sequencing technology 2 The Next Generation of Sequencing: High-Throughput Technologies 3 High Throughput Sequencing Applied to Genome Sequencing (TEDed CC BY-NC-ND

More information

RNASEQ WITHOUT A REFERENCE

RNASEQ WITHOUT A REFERENCE RNASEQ WITHOUT A REFERENCE Experimental Design Assembly in Non-Model Organisms And other (hopefully useful) Stuff Meg Staton mstaton1@utk.edu University of Tennessee Knoxville, TN I. Project Design Things

More information

Assessing De-Novo Transcriptome Assemblies

Assessing De-Novo Transcriptome Assemblies Assessing De-Novo Transcriptome Assemblies Shawn T. O Neil Center for Genome Research and Biocomputing Oregon State University Scott J. Emrich University of Notre Dame 100K Contigs, Perfect 1M Contigs,

More information

Analysis of Differential Gene Expression in Cattle Using mrna-seq

Analysis of Differential Gene Expression in Cattle Using mrna-seq Analysis of Differential Gene Expression in Cattle Using mrna-seq mrna-seq A rough guide for green horns Animal and Grassland Research and Innovation Centre Animal and Bioscience Research Department Teagasc,

More information

RNA-seq Data Analysis

RNA-seq Data Analysis Lecture 3. Clustering; Function/Pathway Enrichment analysis RNA-seq Data Analysis Qi Sun Bioinformatics Facility Biotechnology Resource Center Cornell University Lecture 1. Map RNA-seq read to genome Lecture

More information

Intermediate RNA-Seq Tips, Tricks and Non-Human Organisms

Intermediate RNA-Seq Tips, Tricks and Non-Human Organisms Intermediate RNA-Seq Tips, Tricks and Non-Human Organisms Kevin Silverstein PhD, John Garbe PhD and Ying Zhang PhD, Research Informatics Support System (RISS) MSI September 25, 2014 Slides available at

More information

Welcome to the NGS webinar series

Welcome to the NGS webinar series Welcome to the NGS webinar series Webinar 1 NGS: Introduction to technology, and applications NGS Technology Webinar 2 Targeted NGS for Cancer Research NGS in cancer Webinar 3 NGS: Data analysis for genetic

More information

Next Generation Sequencing: An Overview

Next Generation Sequencing: An Overview Next Generation Sequencing: An Overview Cavan Reilly November 13, 2017 Table of contents Next generation sequencing NGS and microarrays Study design Quality assessment Burrows Wheeler transform Next generation

More information

Whole Transcriptome Analysis of Illumina RNA- Seq Data. Ryan Peters Field Application Specialist

Whole Transcriptome Analysis of Illumina RNA- Seq Data. Ryan Peters Field Application Specialist Whole Transcriptome Analysis of Illumina RNA- Seq Data Ryan Peters Field Application Specialist Partek GS in your NGS Pipeline Your Start-to-Finish Solution for Analysis of Next Generation Sequencing Data

More information

Genome Annotation Genome annotation What is the function of each part of the genome? Where are the genes? What is the mrna sequence (transcription, splicing) What is the protein sequence? What does

More information

NOW GENERATION SEQUENCING. Monday, December 5, 11

NOW GENERATION SEQUENCING. Monday, December 5, 11 NOW GENERATION SEQUENCING 1 SEQUENCING TIMELINE 1953: Structure of DNA 1975: Sanger method for sequencing 1985: Human Genome Sequencing Project begins 1990s: Clinical sequencing begins 1998: NHGRI $1000

More information

Shuji Shigenobu. April 3, 2013 Illumina Webinar Series

Shuji Shigenobu. April 3, 2013 Illumina Webinar Series Shuji Shigenobu April 3, 2013 Illumina Webinar Series RNA-seq RNA-seq is a revolutionary tool for transcriptomics using deepsequencing technologies. genome HiSeq2000@NIBB (Wang 2009 with modifications)

More information

less sensitive than RNA-seq but more robust analysis pipelines expensive but quantitiatve standard but typically not high throughput

less sensitive than RNA-seq but more robust analysis pipelines expensive but quantitiatve standard but typically not high throughput Chapter 11: Gene Expression The availability of an annotated genome sequence enables massively parallel analysis of gene expression. The expression of all genes in an organism can be measured in one experiment.

More information

DNA-Sequencing. Technologies & Devices. Matthias Platzer. Genome Analysis Leibniz Institute on Aging - Fritz Lipmann Institute (FLI)

DNA-Sequencing. Technologies & Devices. Matthias Platzer. Genome Analysis Leibniz Institute on Aging - Fritz Lipmann Institute (FLI) DNA-Sequencing Technologies & Devices Matthias Platzer Genome Analysis Leibniz Institute on Aging - Fritz Lipmann Institute (FLI) Genome analysis DNA sequencing platforms ABI 3730xl 4/2004 & 6/2006 1 Mb/day,

More information

RNA-Seq Tutorial 1. Kevin Silverstein, Ying Zhang Research Informatics Solutions, MSI October 18, 2016

RNA-Seq Tutorial 1. Kevin Silverstein, Ying Zhang Research Informatics Solutions, MSI October 18, 2016 RNA-Seq Tutorial 1 Kevin Silverstein, Ying Zhang Research Informatics Solutions, MSI October 18, 2016 Slides available at www.msi.umn.edu/tutorial-materials RNA-Seq Tutorials Lectures RNA-Seq experiment

More information

Sequencing technologies. Jose Blanca COMAV institute bioinf.comav.upv.es

Sequencing technologies. Jose Blanca COMAV institute bioinf.comav.upv.es Sequencing technologies Jose Blanca COMAV institute bioinf.comav.upv.es Outline Sequencing technologies: Sanger 2nd generation sequencing: 3er generation sequencing: 454 Illumina SOLiD Ion Torrent PacBio

More information

DNA-Sequencing. Technologies & Devices. Matthias Platzer. Genome Analysis Leibniz Institute on Aging - Fritz Lipmann Institute (FLI)

DNA-Sequencing. Technologies & Devices. Matthias Platzer. Genome Analysis Leibniz Institute on Aging - Fritz Lipmann Institute (FLI) DNA-Sequencing Technologies & Devices Matthias Platzer Genome Analysis Leibniz Institute on Aging - Fritz Lipmann Institute (FLI) Genome analysis DNA sequencing platforms ABI 3730xl 4/2004 & 6/2006 1 Mb/day,

More information

SCALABLE, REPRODUCIBLE RNA-Seq

SCALABLE, REPRODUCIBLE RNA-Seq SCALABLE, REPRODUCIBLE RNA-Seq SCALABLE, REPRODUCIBLE RNA-Seq Advances in the RNA sequencing workflow, from sample preparation through data analysis, are enabling deeper and more accurate exploration

More information

RNA-seq data analysis with Chipster. Eija Korpelainen CSC IT Center for Science, Finland

RNA-seq data analysis with Chipster. Eija Korpelainen CSC IT Center for Science, Finland RNA-seq data analysis with Chipster Eija Korpelainen CSC IT Center for Science, Finland chipster@csc.fi What will I learn? 1. What you can do with Chipster and how to operate it 2. What RNA-seq can be

More information

De novo metatranscriptome assembly and coral gene expression profile of Montipora capitata with growth anomaly

De novo metatranscriptome assembly and coral gene expression profile of Montipora capitata with growth anomaly Additional File 1 De novo metatranscriptome assembly and coral gene expression profile of Montipora capitata with growth anomaly Monika Frazier, Martin Helmkampf, M. Renee Bellinger, Scott Geib, Misaki

More information

Genomics and Transcriptomics of Spirodela polyrhiza

Genomics and Transcriptomics of Spirodela polyrhiza Genomics and Transcriptomics of Spirodela polyrhiza Doug Bryant Bioinformatics Core Facility & Todd Mockler Group, Donald Danforth Plant Science Center Desired Outcomes High-quality genomic reference sequence

More information

De Novo Assembly of High-throughput Short Read Sequences

De Novo Assembly of High-throughput Short Read Sequences De Novo Assembly of High-throughput Short Read Sequences Chuming Chen Center for Bioinformatics and Computational Biology (CBCB) University of Delaware NECC Third Skate Genome Annotation Workshop May 23,

More information

Sequencing technologies. Jose Blanca COMAV institute bioinf.comav.upv.es

Sequencing technologies. Jose Blanca COMAV institute bioinf.comav.upv.es Sequencing technologies Jose Blanca COMAV institute bioinf.comav.upv.es Outline Sequencing technologies: Sanger 2nd generation sequencing: 3er generation sequencing: 454 Illumina SOLiD Ion Torrent PacBio

More information

Introduction to Bioinformatics and Gene Expression Technologies

Introduction to Bioinformatics and Gene Expression Technologies Introduction to Bioinformatics and Gene Expression Technologies Utah State University Fall 2017 Statistical Bioinformatics (Biomedical Big Data) Notes 1 1 Vocabulary Gene: hereditary DNA sequence at a

More information

Mate-pair library data improves genome assembly

Mate-pair library data improves genome assembly De Novo Sequencing on the Ion Torrent PGM APPLICATION NOTE Mate-pair library data improves genome assembly Highly accurate PGM data allows for de Novo Sequencing and Assembly For a draft assembly, generate

More information

Course Presentation. Ignacio Medina Presentation

Course Presentation. Ignacio Medina Presentation Course Index Introduction Agenda Analysis pipeline Some considerations Introduction Who we are Teachers: Marta Bleda: Computational Biologist and Data Analyst at Department of Medicine, Addenbrooke's Hospital

More information

FGCZ NEWSLETTER FALL Next Generation Sequencing at the Functional Genomics Center Zurich

FGCZ NEWSLETTER FALL Next Generation Sequencing at the Functional Genomics Center Zurich FGCZ NEWSLETTER FALL 2011 newsletter Technologies, Applications, and Access to Support Next Generation Sequencing at the Functional Genomics Center Zurich OVERVIEW 1 NGS AT THE FGCZ Technologies and organization

More information

High Throughput Sequencing Technologies. J Fass UCD Genome Center Bioinformatics Core Monday June 16, 2014

High Throughput Sequencing Technologies. J Fass UCD Genome Center Bioinformatics Core Monday June 16, 2014 High Throughput Sequencing Technologies J Fass UCD Genome Center Bioinformatics Core Monday June 16, 2014 Sequencing Explosion www.genome.gov/sequencingcosts http://t.co/ka5cvghdqo Sequencing Explosion

More information

Post-assembly Data Analysis

Post-assembly Data Analysis Assembled transcriptome Post-assembly Data Analysis Quantification: the expression level of each gene in each sample DE genes: genes differentially expressed between samples Clustering/network analysis

More information

Outline. General principles of clonal sequencing Analysis principles Applications CNV analysis Genome architecture

Outline. General principles of clonal sequencing Analysis principles Applications CNV analysis Genome architecture The use of new sequencing technologies for genome analysis Chris Mattocks National Genetics Reference Laboratory (Wessex) NGRL (Wessex) 2008 Outline General principles of clonal sequencing Analysis principles

More information

Haploid Assembly of Diploid Genomes

Haploid Assembly of Diploid Genomes Haploid Assembly of Diploid Genomes Challenges, Trials, Tribulations 13 October 2011 İnanç Birol Assembly By Short Sequencing IEEE InfoVis 2009 2 3 in Literature ~40 citations on tool comparisons ~20 citations

More information

De novo genome assembly with next generation sequencing data!! "

De novo genome assembly with next generation sequencing data!! De novo genome assembly with next generation sequencing data!! " Jianbin Wang" HMGP 7620 (CPBS 7620, and BMGN 7620)" Genomics lectures" 2/7/12" Outline" The need for de novo genome assembly! The nature

More information

Chapter 15 Gene Technologies and Human Applications

Chapter 15 Gene Technologies and Human Applications Chapter Outline Chapter 15 Gene Technologies and Human Applications Section 1: The Human Genome KEY IDEAS > Why is the Human Genome Project so important? > How do genomics and gene technologies affect

More information

Illumina (Solexa) Throughput: 4 Tbp in one run (5 days) Cheapest sequencing technology. Mismatch errors dominate. Cost: ~$1000 per human genme

Illumina (Solexa) Throughput: 4 Tbp in one run (5 days) Cheapest sequencing technology. Mismatch errors dominate. Cost: ~$1000 per human genme Illumina (Solexa) Current market leader Based on sequencing by synthesis Current read length 100-150bp Paired-end easy, longer matepairs harder Error ~0.1% Mismatch errors dominate Throughput: 4 Tbp in

More information

CM581A2: NEXT GENERATION SEQUENCING PLATFORMS AND LIBRARY GENERATION

CM581A2: NEXT GENERATION SEQUENCING PLATFORMS AND LIBRARY GENERATION CM581A2: NEXT GENERATION SEQUENCING PLATFORMS AND LIBRARY GENERATION Fall 2015 Instructors: Coordinator: Carol Wilusz, Associate Professor MIP, CMB Instructor: Dan Sloan, Assistant Professor, Biology,

More information

Automated size selection of NEBNext Small RNA libraries with the Sage Pippin Prep

Automated size selection of NEBNext Small RNA libraries with the Sage Pippin Prep Automated size selection of NEBNext Small RNA libraries with the Sage Pippin Prep DNA CLONING DNA AMPLIFICATION & PCR EPIGENETICS RNA ANALYSIS LIBRARY PREP FOR NEXT GEN SEQUENCING PROTEIN EXPRESSION &

More information

Genome annotation. Erwin Datema (2011) Sandra Smit (2012, 2013)

Genome annotation. Erwin Datema (2011) Sandra Smit (2012, 2013) Genome annotation Erwin Datema (2011) Sandra Smit (2012, 2013) Genome annotation AGACAAAGATCCGCTAAATTAAATCTGGACTTCACATATTGAAGTGATATCACACGTTTCTCTAAT AATCTCCTCACAATATTATGTTTGGGATGAACTTGTCGTGATTTGCCATTGTAGCAATCACTTGAA

More information

Data Analysis with CASAVA v1.8 and the MiSeq Reporter

Data Analysis with CASAVA v1.8 and the MiSeq Reporter Data Analysis with CASAVA v1.8 and the MiSeq Reporter Eric Smith, PhD Bioinformatics Scientist September 15 th, 2011 2010 Illumina, Inc. All rights reserved. Illumina, illuminadx, Solexa, Making Sense

More information

Post-assembly Data Analysis

Post-assembly Data Analysis Assembled transcriptome Post-assembly Data Analysis Quantification: get expression for each gene in each sample Genes differentially expressed between samples Clustering/network analysis Identifying over-represented

More information

Next Generation Sequencing Lecture Saarbrücken, 19. March Sequencing Platforms

Next Generation Sequencing Lecture Saarbrücken, 19. March Sequencing Platforms Next Generation Sequencing Lecture Saarbrücken, 19. March 2012 Sequencing Platforms Contents Introduction Sequencing Workflow Platforms Roche 454 ABI SOLiD Illumina Genome Anlayzer / HiSeq Problems Quality

More information

RNAseq Differential Gene Expression Analysis Report

RNAseq Differential Gene Expression Analysis Report RNAseq Differential Gene Expression Analysis Report Customer Name: Institute/Company: Project: NGS Data: Bioinformatics Service: IlluminaHiSeq2500 2x126bp PE Differential gene expression analysis Sample

More information

Multiple choice questions (numbers in brackets indicate the number of correct answers)

Multiple choice questions (numbers in brackets indicate the number of correct answers) 1 Multiple choice questions (numbers in brackets indicate the number of correct answers) February 1, 2013 1. Ribose is found in Nucleic acids Proteins Lipids RNA DNA (2) 2. Most RNA in cells is transfer

More information

Machine Learning Methods for RNA-seq-based Transcriptome Reconstruction

Machine Learning Methods for RNA-seq-based Transcriptome Reconstruction Machine Learning Methods for RNA-seq-based Transcriptome Reconstruction Gunnar Rätsch Friedrich Miescher Laboratory Max Planck Society, Tübingen, Germany NGS Bioinformatics Meeting, Paris (March 24, 2010)

More information

Jenny Gu, PhD Strategic Business Development Manager, PacBio

Jenny Gu, PhD Strategic Business Development Manager, PacBio IDT and PacBio joint presentation Characterizing Alzheimer s Disease candidate genes and transcripts with targeted, long-read, single-molecule sequencing Jenny Gu, PhD Strategic Business Development Manager,

More information

IMGM Laboratories GmbH. Sales Manager

IMGM Laboratories GmbH. Sales Manager IMGM Laboratories GmbH Dr. Jennifer K. Kuhn Sales Manager About IMGM Laboratories IMGM Laboratories was founded in 2001 IMGM operates as professional provider of advanced genomic services from research

More information

A Roadmap to the De-novo Assembly of the Banana Slug Genome

A Roadmap to the De-novo Assembly of the Banana Slug Genome A Roadmap to the De-novo Assembly of the Banana Slug Genome Stefan Prost 1 1 Department of Integrative Biology, University of California, Berkeley, United States of America April 6th-10th, 2015 Outline

More information

Microarrays: since we use probes we obviously must know the sequences we are looking at!

Microarrays: since we use probes we obviously must know the sequences we are looking at! These background are needed: 1. - Basic Molecular Biology & Genetics DNA replication Transcription Post-transcriptional RNA processing Translation Post-translational protein modification Gene expression

More information

Molecular Cell Biology - Problem Drill 11: Recombinant DNA

Molecular Cell Biology - Problem Drill 11: Recombinant DNA Molecular Cell Biology - Problem Drill 11: Recombinant DNA Question No. 1 of 10 1. Which of the following statements about the sources of DNA used for molecular cloning is correct? Question #1 (A) cdna

More information

Ecole de Bioinforma(que AVIESAN Roscoff 2014 GALAXY INITIATION. A. Lermine U900 Ins(tut Curie, INSERM, Mines ParisTech

Ecole de Bioinforma(que AVIESAN Roscoff 2014 GALAXY INITIATION. A. Lermine U900 Ins(tut Curie, INSERM, Mines ParisTech GALAXY INITIATION A. Lermine U900 Ins(tut Curie, INSERM, Mines ParisTech How does Next- Gen sequencing work? DNA fragmentation Size selection and clonal amplification Massive parallel sequencing ACCGTTTGCCG

More information

DNA-Sequencing. Technologies & Devices

DNA-Sequencing. Technologies & Devices DNA-Sequencing Technologies & Devices Genome analysis DNA sequencing platforms ABI 3730xl 4/2004 & 6/2006 1 Mb/day, 850 nt reads 2 Mb/day, 550 nt reads Roche/454 GS FLX 12/2006 800 Mb/23h, 800 nt reads

More information

Outline. Annotation of Drosophila Primer. Gene structure nomenclature. Muller element nomenclature. GEP Drosophila annotation projects 01/04/2018

Outline. Annotation of Drosophila Primer. Gene structure nomenclature. Muller element nomenclature. GEP Drosophila annotation projects 01/04/2018 Outline Overview of the GEP annotation projects Annotation of Drosophila Primer January 2018 GEP annotation workflow Practice applying the GEP annotation strategy Wilson Leung and Chris Shaffer AAACAACAATCATAAATAGAGGAAGTTTTCGGAATATACGATAAGTGAAATATCGTTCT

More information

Functional Genomics Overview RORY STARK PRINCIPAL BIOINFORMATICS ANALYST CRUK CAMBRIDGE INSTITUTE 18 SEPTEMBER 2017

Functional Genomics Overview RORY STARK PRINCIPAL BIOINFORMATICS ANALYST CRUK CAMBRIDGE INSTITUTE 18 SEPTEMBER 2017 Functional Genomics Overview RORY STARK PRINCIPAL BIOINFORMATICS ANALYST CRUK CAMBRIDGE INSTITUTE 18 SEPTEMBER 2017 Agenda What is Functional Genomics? RNA Transcription/Gene Expression Measuring Gene

More information

BIOINFORMATICS 1 SEQUENCING TECHNOLOGY. DNA story. DNA story. Sequencing: infancy. Sequencing: beginnings 26/10/16. bioinformatic challenges

BIOINFORMATICS 1 SEQUENCING TECHNOLOGY. DNA story. DNA story. Sequencing: infancy. Sequencing: beginnings 26/10/16. bioinformatic challenges BIOINFORMATICS 1 or why biologists need computers SEQUENCING TECHNOLOGY bioinformatic challenges http://www.bioinformatics.uni-muenster.de/teaching/courses-2012/bioinf1/index.hbi Prof. Dr. Wojciech Makałowski"

More information

SCIENCE CHINA Life Sciences

SCIENCE CHINA Life Sciences SCIENCE CHINA Life Sciences SPECIAL TOPIC February 2013 Vol.56 No.2: 143 155 RESEARCH PAPER doi: 10.1007/s11427-013-4442-z Comparative study of de novo assembly and genome-guided assembly strategies for

More information

Single Cell Genomics

Single Cell Genomics Single Cell Genomics Application Cost Platform/Protoc ol Note Single cell 3 mrna-seq cell lysis/rt/library prep $2460/Sample 10X Genomics Chromium 500-10,000 cells/sample Single cell 5 V(D)J mrna-seq cell

More information

Sequence assembly. Jose Blanca COMAV institute bioinf.comav.upv.es

Sequence assembly. Jose Blanca COMAV institute bioinf.comav.upv.es Sequence assembly Jose Blanca COMAV institute bioinf.comav.upv.es Sequencing project Unknown sequence { experimental evidence result read 1 read 4 read 2 read 5 read 3 read 6 read 7 Computational requirements

More information

INTRODUCCIÓ A LES TECNOLOGIES DE 'NEXT GENERATION SEQUENCING'

INTRODUCCIÓ A LES TECNOLOGIES DE 'NEXT GENERATION SEQUENCING' INTRODUCCIÓ A LES TECNOLOGIES DE 'NEXT GENERATION SEQUENCING' Bioinformàtica per a la Recerca Biomèdica Ricardo Gonzalo Sanz ricardo.gonzalo@vhir.org 14/12/2016 1. Introduction to NGS 2. First Generation

More information

QIAGEN s NGS Solutions for Biomarkers NGS & Bioinformatics team QIAGEN (Suzhou) Translational Medicine Co.,Ltd

QIAGEN s NGS Solutions for Biomarkers NGS & Bioinformatics team QIAGEN (Suzhou) Translational Medicine Co.,Ltd QIAGEN s NGS Solutions for Biomarkers NGS & Bioinformatics team QIAGEN (Suzhou) Translational Medicine Co.,Ltd 1 Our current NGS & Bioinformatics Platform 2 Our NGS workflow and applications 3 QIAGEN s

More information

Microarray Gene Expression Analysis at CNIO

Microarray Gene Expression Analysis at CNIO Microarray Gene Expression Analysis at CNIO Orlando Domínguez Genomics Unit Biotechnology Program, CNIO 8 May 2013 Workflow, from samples to Gene Expression data Experimental design user/gu/ubio Samples

More information

HLA and Next Generation Sequencing it s all about the Data

HLA and Next Generation Sequencing it s all about the Data HLA and Next Generation Sequencing it s all about the Data John Ord, NHSBT Colindale and University of Cambridge BSHI Annual Conference Manchester September 2014 Introduction In 2003 the first full public

More information

Nature Biotechnology: doi: /nbt Supplementary Figure 1. Number and length distributions of the inferred fosmids.

Nature Biotechnology: doi: /nbt Supplementary Figure 1. Number and length distributions of the inferred fosmids. Supplementary Figure 1 Number and length distributions of the inferred fosmids. Fosmid were inferred by mapping each pool s sequence reads to hg19. We retained only those reads that mapped to within a

More information

How much sequencing do I need? Emily Crisovan Genomics Core

How much sequencing do I need? Emily Crisovan Genomics Core How much sequencing do I need? Emily Crisovan Genomics Core How much sequencing? Three questions: 1. How much sequence is required for good experimental design? 2. What type of sequencing run is best?

More information

Sequence Based Function Annotation. Qi Sun Bioinformatics Facility Biotechnology Resource Center Cornell University

Sequence Based Function Annotation. Qi Sun Bioinformatics Facility Biotechnology Resource Center Cornell University Sequence Based Function Annotation Qi Sun Bioinformatics Facility Biotechnology Resource Center Cornell University Usage scenarios for sequence based function annotation Function prediction of newly cloned

More information

CNV and variant detection for human genome resequencing data - for biomedical researchers (II)

CNV and variant detection for human genome resequencing data - for biomedical researchers (II) CNV and variant detection for human genome resequencing data - for biomedical researchers (II) Chuan-Kun Liu 劉傳崑 Senior Maneger National Center for Genome Medican bioit@ncgm.sinica.edu.tw Abstract Common

More information

Supporting Information

Supporting Information Supporting Information Yuan et al. 10.1073/pnas.0906869106 Fig. S1. Heat map showing that Populus ICS is coregulated with orthologs of Arabidopsis genes involved in PhQ biosynthesis and PSI function, but

More information

Gene Regulation Solutions. Microarrays and Next-Generation Sequencing

Gene Regulation Solutions. Microarrays and Next-Generation Sequencing Gene Regulation Solutions Microarrays and Next-Generation Sequencing Gene Regulation Solutions The Microarrays Advantage Microarrays Lead the Industry in: Comprehensive Content SurePrint G3 Human Gene

More information

Agenda. Web Databases for Drosophila. Gene annotation workflow. GEP Drosophila annotation projects 01/01/2018. Annotation adding labels to a sequence

Agenda. Web Databases for Drosophila. Gene annotation workflow. GEP Drosophila annotation projects 01/01/2018. Annotation adding labels to a sequence Agenda GEP annotation project overview Web Databases for Drosophila An introduction to web tools, databases and NCBI BLAST Web databases for Drosophila annotation UCSC Genome Browser NCBI / BLAST FlyBase

More information

L3: Short Read Alignment to a Reference Genome

L3: Short Read Alignment to a Reference Genome L3: Short Read Alignment to a Reference Genome Shamith Samarajiwa CRUK Autumn School in Bioinformatics Cambridge, September 2017 Where to get help! http://seqanswers.com http://www.biostars.org http://www.bioconductor.org/help/mailing-list

More information

Leonardo Mariño-Ramírez, PhD NCBI / NLM / NIH. BIOL 7210 A Computational Genomics 2/18/2015

Leonardo Mariño-Ramírez, PhD NCBI / NLM / NIH. BIOL 7210 A Computational Genomics 2/18/2015 Leonardo Mariño-Ramírez, PhD NCBI / NLM / NIH BIOL 7210 A Computational Genomics 2/18/2015 The $1,000 genome is here! http://www.illumina.com/systems/hiseq-x-sequencing-system.ilmn Bioinformatics bottleneck

More information

Corset: enabling differential gene expression analysis for de novo assembled transcriptomes

Corset: enabling differential gene expression analysis for de novo assembled transcriptomes Davidson and Oshlack Genome Biology 2014, 15:410 METHOD Open Access : enabling differential gene expression analysis for de novo assembled transcriptomes Nadia M Davidson 1 and Alicia Oshlack 1,2* Abstract

More information

Genomic Data Analysis Services Available for PL-Grid Users

Genomic Data Analysis Services Available for PL-Grid Users Domain-oriented services and resources of Polish Infrastructure for Supporting Computational Science in the European Research Space PLGrid Plus Domain-oriented services and resources of Polish Infrastructure

More information

Variant detection analysis in the BRCA1/2 genes from Ion torrent PGM data

Variant detection analysis in the BRCA1/2 genes from Ion torrent PGM data Variant detection analysis in the BRCA1/2 genes from Ion torrent PGM data Bruno Zeitouni Bionformatics department of the Institut Curie Inserm U900 Mines ParisTech Ion Torrent User Meeting 2012, October

More information

De novo genome assembly. Dr Torsten Seemann

De novo genome assembly. Dr Torsten Seemann De novo genome assembly Dr Torsten Seemann IMB Winter School - Brisbane Mon 1 July 2013 Introduction Ideal world I would not need to give this talk! Human DNA Non-existent USB3 device AGTCTAGGATTCGCTA

More information

CMPS 3110 : Bioinformatics. High-Throughput Sequencing and Applications

CMPS 3110 : Bioinformatics. High-Throughput Sequencing and Applications CMPS 3110 : Bioinformatics High-Throughput Sequencing and Applications Sanger (1982) introduced chaintermination sequencing. Main idea: Obtain fragments of all possible lengths, ending in A, C, T, G. Using

More information

2/5/16. Honeypot Ants. DNA sequencing, Transcriptomics and Genomics. Gene sequence changes? And/or gene expression changes?

2/5/16. Honeypot Ants. DNA sequencing, Transcriptomics and Genomics. Gene sequence changes? And/or gene expression changes? 2/5/16 DNA sequencing, Transcriptomics and Genomics Honeypot Ants "nequacatl" BY2208, Mani Lecture 3 Gene sequence changes? And/or gene expression changes? gene expression differences DNA sequencing, Transcriptomics

More information

3. human genomics clone genes associated with genetic disorders. 4. many projects generate ordered clones that cover genome

3. human genomics clone genes associated with genetic disorders. 4. many projects generate ordered clones that cover genome Lectures 30 and 31 Genome analysis I. Genome analysis A. two general areas 1. structural 2. functional B. genome projects a status report 1. 1 st sequenced: several viral genomes 2. mitochondria and chloroplasts

More information

Local assembly and pre-mrna splicing analyses by high-throughput sequencing data

Local assembly and pre-mrna splicing analyses by high-throughput sequencing data Graduate Theses and Dissertations Graduate College 2012 Local assembly and pre-mrna splicing analyses by high-throughput sequencing data Hsien-chao Chou Iowa State University Follow this and additional

More information

Ultrasequencing: Methods and Applications of the New Generation Sequencing Platforms

Ultrasequencing: Methods and Applications of the New Generation Sequencing Platforms Ultrasequencing: Methods and Applications of the New Generation Sequencing Platforms Laura Moya Andérico Master in Advanced Genetics Genomics Class December 16 th, 2015 Brief Overview First-generation

More information