For PE reads -/+ UMI barcodes
Two RNA-Seq pipelines options are available. They both take pre-cleaned, paired-end fastq files, with or without UMI barcodes as inputs. They start with read mapping and proceed through differential expression analysis. I modeled them after nf-core's RNA-seq pipeline, but trimmed them down to function as a specific workflows rather than combining them into one all-purpose toolkit. This helps reduce errors due to runtime setting mistakes (IMHO). The included Docker container can run either pipeline. All code was written entirely by me.
- Your reads are paired-end and pre-cleaned
- Your reads can have UMIs (or not). Note: cell barcode splitting is not supported.
- There are two DE groups (i.e. control vs. test)
- Mapping - RNA-STAR
- Duplicate Read Removal - UMI Tools (if you have UMIs)
- Read Counts - FeatureCounts
- Differential Expression Analysis - DESeq2
- RNA-Seq PE: for PE reads without cell or UMI barcodes. It does not perform a deduplication step.
- RNA-Seq PE UMI: for PE reads with UMI barcodes. It de-duplicates alignments using UMI-tools dedup.
- You must start with pre-cleaned paired-end fastq files, -/+ UMIs. These pipelines will not work with SE fastq files or reads with cell barcodes.
- See my fastq cleanup scripts if needed.
- You will be running either RnaSeq_PE.nf or RnaSeq_PE_UMI.nf in the Docker container cbreuer/rnaseq:latest. Make sure you have the Docker container pulled and working before you start.
- Metadata file - file names, locations, and control/test label for DE analysis. (See the provided example.)
- Fastq files - Default format is "_R1.fastq.gz" "_R2.fastq.gz"
- STAR genome index (see below)
- Transcripts.gtf file (Example)
- Indicate the strandedness of your library. Parameter: fc_strand. 0 = unstranded, 1 = stranded and 2 = reversely stranded
- Set up your output folder tree as you like
- Default is "_R1.fastq.gz" and "R2.fastq.gz"
- File names can only have one "_". They should look like "sample1_R1.fastq.gz" and "sample1_R2.fastq.gz".
- Example: Human HG38 Release48 from GenCode.
- Place your customized metadata file, fastq files, scripts, and STAR genome in local folders - I like to use WSL.
- Double check that file locations correct in the nextflow.config file.
- Run "nextflow run .nf".
- The pipeline will launch in docker container cbreuer/rnaseq.
- Results will be published to your local output folder.
- A filtered table of significant genes. Note that the adjusted p-value significance cutoff can be set at the top of the DESeq2.R script if needed. Default is 0.05.
- A basic Volcano plot (-log10 p-value vs. Log2 FC).

- Before and after de-duplication (using UMIs) bam files
- countsMatrix
- bam indices
- etc...




