Skip to content

About

Nextflow pipelines for mapping and differential expression analysis

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Latest commit

 

History

13 Commits

Folders and files

Repository files navigation

banner

Nextflow RNA-Seq Pipelines

For PE reads -/+ UMI barcodes

Overview:

Two RNA-Seq pipelines options are available. They both take pre-cleaned, paired-end fastq files, with or without UMI barcodes as inputs. They start with read mapping and proceed through differential expression analysis. I modeled them after nf-core's RNA-seq pipeline, but trimmed them down to function as a specific workflows rather than combining them into one all-purpose toolkit. This helps reduce errors due to runtime setting mistakes (IMHO). The included Docker container can run either pipeline. All code was written entirely by me.

Requirements:

  • Your reads are paired-end and pre-cleaned
  • Your reads can have UMIs (or not). Note: cell barcode splitting is not supported.
  • There are two DE groups (i.e. control vs. test)

Processing Overview:

  • Mapping - RNA-STAR
  • Duplicate Read Removal - UMI Tools (if you have UMIs)
  • Read Counts - FeatureCounts
  • Differential Expression Analysis - DESeq2

Pipeline Selection:

  • RNA-Seq PE: for PE reads without cell or UMI barcodes. It does not perform a deduplication step.
  • RNA-Seq PE UMI: for PE reads with UMI barcodes. It de-duplicates alignments using UMI-tools dedup.

Pre-Run Overview:

  • You must start with pre-cleaned paired-end fastq files, -/+ UMIs. These pipelines will not work with SE fastq files or reads with cell barcodes.
  • See my fastq cleanup scripts if needed.
  • You will be running either RnaSeq_PE.nf or RnaSeq_PE_UMI.nf in the Docker container cbreuer/rnaseq:latest. Make sure you have the Docker container pulled and working before you start.

Inputs: (4)

  1. Metadata file - file names, locations, and control/test label for DE analysis. (See the provided example.)
  2. Fastq files - Default format is "_R1.fastq.gz" "_R2.fastq.gz"
  3. STAR genome index (see below)
  4. Transcripts.gtf file (Example)

Pre-Run Setup:

1) Edit nextflow.config:

  • Indicate the strandedness of your library. Parameter: fc_strand. 0 = unstranded, 1 = stranded and 2 = reversely stranded
  • Set up your output folder tree as you like

2) Ensure the fastq naming convention matches your files.

  • Default is "_R1.fastq.gz" and "R2.fastq.gz"
  • File names can only have one "_". They should look like "sample1_R1.fastq.gz" and "sample1_R2.fastq.gz".

3) Build a STAR genome or download one from iGenomes.

4) Download the transcripts.fasta file for your genome.

5) Use Docker pull to download cbreuer/rnaseq:latest

Run

  • Place your customized metadata file, fastq files, scripts, and STAR genome in local folders - I like to use WSL.
  • Double check that file locations correct in the nextflow.config file.
  • Run "nextflow run .nf".
  • The pipeline will launch in docker container cbreuer/rnaseq.
  • Results will be published to your local output folder.

Outputs

A file tree with outputs from each process:

filetree

Differential expression analysis table

detable

Differential expression results with Log2FC, p-value, and adjusted p-value (test vs control)

  • A filtered table of significant genes. Note that the adjusted p-value significance cutoff can be set at the top of the DESeq2.R script if needed. Default is 0.05.
  • A basic Volcano plot (-log10 p-value vs. Log2 FC).

    volcanoplot

Principal Component Analysis

  • Shows sample clustering, allows for rapid outlier detection pci

Correlation Matrix

  • Pair-wise comparisons with Pearson Correlation Coefficients
    pci

All aligned (.bam) files

  • Before and after de-duplication (using UMIs) bam files

All other intermdediate files

  • countsMatrix
  • bam indices
  • etc...

About

Nextflow pipelines for mapping and differential expression analysis

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages