A self-paced bioinformatics course built from materials by the Kodomo Program at Moscow State University, the IAB textbook by J. Gregory Caporaso, and the Summer School of Bioinformatics.
197 notebooks · 6 tiers · 30 interactive visualizations · 108 glossary terms · Learn-with-AI mode · 213 Claude Code skills
Tier 0 Computational Foundations 12 notebooks
Linux · Git · Bash · Encodings · R · Biostatistics ·
Probability & Statistics (Python) · Advanced R Statistics
Tier 1 Python for Bioinformatics 38 notebooks
Variables → Strings → Control Flow → Functions → Files →
Data Structures → Iterators → Regex → OOP → Decorators →
NumPy/Pandas → Visualization → SQL
Tier 2 Core Bioinformatics 21 notebooks
Databases · BioPython · Alignment · BLAST · MSA ·
Phylogenetics · Protein Structure · Nucleic Acids ·
Chromatograms · Motifs · GO/Pathways · Comparative Genomics ·
Computational Genetics · Hi-C Analysis · Motif Discovery
Tier 3 Applied Bioinformatics 83 notebooks
NGS · Variant Calling · RNA-seq · Microbial Diversity ·
Promoters · Statistics · Machine Learning · Deep Learning ·
Molecular Modeling · Clinical Genomics · Capstone Project ·
Biochemistry & Enzyme Kinetics · Genetic Engineering ·
Population Genetics · Numerical Methods ·
Genome Assembly · Proteomics & Structural Methods ·
GWAS · Spatial Transcriptomics · Copy Number Analysis ·
Bayesian Statistics · TF Footprinting · Cancer Transcriptomics ·
ChIP-seq & Epigenomics · Long-Read Sequencing ·
Shotgun Metagenomics · Multi-Omics Integration ·
Network Biology · Cheminformatics & Drug Discovery
Tier 4 Algorithms & Data Structures 30 notebooks + 30 interactive visualizations
Complexity · Sorting · Searching · Linked Lists · Stacks/Queues ·
BST · AVL · Red-Black Trees · Hash Tables · Bloom Filters ·
KMP · Rabin-Karp · Tries · Suffix Trees · Graphs · DP
Tier 5 Modern AI for Science 13 notebooks
LLM Fine-tuning · Vision RAG · Diffusion & Generative Models ·
AlphaFold & Protein Design · Genomic Foundation Models ·
Protein Language Models · Foundation Models for Single Cell
Each tier starts with a Skills Check — score above 80% and skip ahead.
Tier 4 runs in parallel with Tiers 2-3 — it provides the CS theory behind bioinformatics tools (DP = sequence alignment, string matching = BLAST, graphs = pathways).
See the full table of contents in Course/README.md, Tier 4 README, and Tier 5 README.
This course is built to be studied with an AI tutor, not just next to one. Start with the methodology guide: Course/LEARNING_WITH_AI.md.
The one rule: the notebook is ground truth, the AI is a tutor. The guide shows how to use an LLM to explain dense topics, teach them back, debug your code, and quiz you — and how to catch it when it hallucinates a function, an outdated API, or a plausible-but-wrong fact (with the failure modes specific to bioinformatics).
Every module's README carries this through with a 🤖 Learning with AI section (prime → teach-back → go-deeper prompts, tuned to that topic) and a ✅ Check Your Understanding self-test. And every topic now ships real, worked homework in Assignments/ — tiered 🟢 warm-up → 🟡 core → 🔴 stretch — with full solutions in Solutions/.
Each biology module also links a real-world case study (CASE_STUDY.md) built on a verified public dataset — e.g. the airway/dexamethasone RNA-seq set (GSE52778), the SARS-CoV-2 spike structure (6VSB), the 2014 Ebola outbreak genomes (PRJNA257197) — so you see what real analyses and real output look like. Every module in Tiers 2/3/5 passed a domain-expert accuracy review (see reports/).
- GC content calculator with sliding window
- DNA to protein translator
- FASTA file parser and analyzer
- Restriction site finder and Open Reading Frame (ORF) detector
- Gene expression heatmaps and genome GC landscape visualizer
- Sequence alignment tools with BLOSUM scoring
- BLAST result analyzer with homology assessment
- Molecular visualization scripts (Jmol/PyMol)
- DNA structure models (A/B/Z forms)
- Enzyme kinetics curve fitter (Michaelis-Menten, inhibition models)
- CRISPR guide RNA designer with off-target scoring
- Genetic drift and selection simulator
- Codon optimizer for heterologous expression
- De novo genome assembler using de Bruijn graphs
- Proteomics mass spectrum analyzer with peptide identification
- Numerical curve fitter (interpolation, FFT, least squares)
git clone https://github.com/Pavel-Kravchenko/Bioinformatics.git
cd Bioinformatics/Course
pip install jupyter numpy pandas matplotlib seaborn biopython scikit-learn scipy
jupyter notebookNot sure where to begin? Open Tier_1_Python_for_Bioinformatics/00_Skills_Check/00_skills_check.ipynb.
The Course/Assets/data/ directory contains real biological files for hands-on practice:
FASTA sequences · PDB protein structures · Sanger chromatograms (.ab1) · VCF variant calls · GenBank records · BLOSUM62 matrix
Algorithm visualizations: QuickSort partitioning · Binary search halving
The course is extracted into 213 curated skills for Claude Code — each with version compatibility, code patterns, and common pitfalls. All rewritten for scanner-discoverability and scored on a 5-dimension quality rubric (213/213 grade A/B; run python3 scripts/score_skills.py).
| Category | Count | Topics |
|---|---|---|
| Foundations | 10 | Linux, Git, Bash, R, biostatistics, probability & statistics |
| Python for Bio | 28 | Variables → OOP → NumPy/Pandas → SQL, all with bioinformatics examples |
| Core Bioinformatics | 17 | Databases, BioPython, alignment, BLAST, phylogenetics, protein structure, GO, Hi-C |
| Applied Bioinformatics | 81 | NGS, variant calling & annotation, RNA-seq, scRNA-seq, ChIP-seq, coverage tracks, GWAS, metagenomics, methylation, CRISPR screens, immunogenomics, immune repertoire, ribo-seq, flow cytometry, primer design, metabolomics, virology, network biology, cheminformatics |
| Algorithms & DS | 29 | Sorting, searching, trees, hash tables, string matching (KMP, Aho-Corasick, suffix trees), graphs, dynamic programming |
| Modern AI | 13 | LLM finetuning, diffusion models, AlphaFold, ESM2, Enformer, Geneformer/scGPT |
| Legacy | 35 | Consolidated skills from earlier course versions |
Usage: Reference any skill by name — Claude activates it automatically. See the full index in Skills/README.md.
Quality tools: python3 scripts/score_skills.py scores all skills. python3 scripts/curate_skills.py runs the full curation pipeline.
This course would not exist without the work of the original authors:
Kodomo Bioinformatics Program — Faculty of Bioengineering and Bioinformatics, Lomonosov Moscow State University. A 10-semester curriculum developed by A.V. Golovin, S.A. Spirin, A.V. Alekseevsky, A. Zalevsky, A.S. Zlobin, D. Penzar, Z. Chervontseva, I. Rusinov, A. Zharikova, V.E. Ramensky, V.Yu. Lunin (IMPB RAS), K.S. Mineev (IBCh RAS), O.S. Sokolova, V.D. Maslova, M. Khachaturyan, D. Dibrova, R. Kudrin, I. Diankin, E. Ocheredko, A. Demkiv, A. Ershova, and other faculty members.
An Introduction to Applied Bioinformatics — by J. Gregory Caporaso and collaborators, Caporaso Lab, Northern Arizona University.
Summer School of Bioinformatics — statistical methods, NGS analysis, and promoter research materials.
FBB Semester Materials — Faculty of Bioengineering and Bioinformatics archive covering advanced R biostatistics (Pervushin/Muromskaya), numerical methods, genome assembly and NGS (Logacheva et al.), proteomics and physical-chemical methods, and protein engineering (Suplatov).
Full attribution details in Course/CREDITS.md.
This repository is a personal study compilation for private, non-commercial educational use only. All intellectual property rights for the original materials remain with their respective authors and institutions listed above. Materials have been translated from Russian to English and adapted solely for the purpose of personal learning.
This is not intended for redistribution, resale, or commercial use. If you are a rights holder and wish to have content removed, please open an issue and it will be addressed promptly.
Pavel Kravchenko




