Skip to content

Latest commit

 

History

5 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 

Repository files navigation

COVID_Data_Analysis

Overview

This repository contains data analysis scripts for processing and visualizing fluorescence time-series data from an automated diagnostic platform for SARS-CoV-2 detection, as described in:

A portable smartphone-based nucleic acid amplification test (Science Advances, abj1281)

The scripts parse raw device output files and generate plots of Relative Fluorescence Units (RFU) vs. time for each sample and its corresponding internal amplification control (IAC). These visualizations are used to assess assay performance, amplification kinetics, and diagnostic calls.


Features

  • Parses raw fluorescence data from multiple device runs
  • Automatically maps metadata (sample, replicate, device, well position)
  • Handles multiple data streams and device formats
  • Performs preprocessing:
    • Baseline alignment
    • Removal of initial signal artifacts
    • Time normalization
  • Generates high-density subplot figures (4×4 grid per page)
  • Overlays:
    • Sample signal (Red channel)
    • Internal Amplification Control (Green channel)
  • Annotates plots with:
    • Detection times
    • Software calls (POSITIVE / NEGATIVE)
    • Sample metadata

Repository Structure

COVID_Data_Analysis/ │ ├── new_data_analysis.py # Analysis for updated dataset (multi-run, replicate-aware) ├── old_data_analysis.py # Legacy dataset processing ├── data/ │ ├── 210701_New_Data/ # Raw data (new format) │ └── 210701_Old_Data/ # Raw data (old format) ├── metadata/ │ └── ForShane.xlsx # Sample metadata and run mapping └── output/ └── Figure*.png # Generated plots


Input Data

1. Metadata File

Excel file (ForShane.xlsx) containing:

  • Sample identifiers
  • Device IDs
  • Run numbers
  • Well positions
  • Replicate information
  • Detection times (sample + IAC)
  • Ground truth / expected values

2. Raw Data Files

Tab-delimited .txt files containing:

  • Time (Seconds)
  • Fluorescence values per well/channel:
    • WellXRed (Sample)
    • WellXGreen (IAC)

The scripts automatically locate the correct files using:

  • Device → folder mapping
  • Run number → file naming convention
  • Stream detection logic

How It Works

1. Data Mapping

Each sample is linked to:

  • Device folder
  • Run number
  • Well position
  • Replicate measurements

2. File Resolution

The script searches across multiple streams to identify the correct data file:

testrun{run}_stream{i}_temp_corrected.txt

and stops when a valid file is found.

3. Preprocessing

  • Converts time from seconds → minutes
  • Normalizes to elapsed time
  • Removes early signal artifacts (baseline dips/spikes)
  • Handles special-case samples with custom corrections

4. Signal Extraction

  • Sample signal → Red channel
  • IAC signal → Green channel

5. Visualization

  • Generates paired plots for replicates
  • Displays:
    • RFU vs time
    • Detection times (table overlay)
    • Diagnostic call (POSITIVE / NEGATIVE)
  • Outputs figures in batches (4×4 subplot grids)

Output

  • High-resolution PNG figures:
    • Figure1.png, Figure2.png, ...

Each figure contains:

  • 16 subplots (8 samples × 2 replicates)
  • Sample + IAC traces
  • Detection time annotations

Dependencies

Install required Python packages:

pip install pandas numpy matplotlib openpyxl
Usage
Update file paths in the script:
new_data = pd.read_excel("path/to/ForShane.xlsx", ...)
new_data_folder = "path/to/210701_New_Data"
Run the script:
python new_data_analysis.py
Output figures will be saved in the working directory.
Notes & Customization
The scripts include hard-coded corrections for specific samples to handle known artifacts (e.g., signal dips, missing streams, anomalous runs).
Device-to-folder mappings must match your local directory structure.
Plot formatting is optimized for dense visualization (small fonts, thin lines).
Context

This analysis supports validation of a portable, automated molecular diagnostic system integrating:

Sample preparation
Isothermal amplification
Optical fluorescence detection
Smartphone-based readout

The RFU vs. time curves correspond to amplification kinetics used to determine:

Presence/absence of viral RNA
Time-to-detection (proxy for viral load)
Future Improvements
Modularize preprocessing steps
Replace hard-coded sample corrections with automated heuristics
Add quantitative curve fitting (e.g., Ct extraction)
Export results as structured datasets (CSV/JSON)
Integrate statistical performance analysis (sensitivity/specificity)

About

Data analysis for post-processing of Harmony Covid-19 data

Resources

Stars

0 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages