Skip to content

Repository files navigation

MitoAggNet

A research framework for testing multimodal predictors of explicitly defined protein phenotypes.

MitoAggNet is local-GPU research code for asking whether sequence, structure and experimental context predict measured protein phenotypes. The framework contains historical aggregation and proteostasis hypotheses, but its current real-data results do not validate a universal mitochondrial-vulnerability score or proteome-wide ranking.

Status: research prototype and negative target-validation result. PXD075952 fraction-partition measurements were internally reproducible, but tested sequence/structure representations did not predict the homology-held-out non-membrane cohort robustly. The historical v0.5 atlas remains scientifically_unvalidated; v1.3 produced neither an atlas nor a candidate ranking. Not for clinical diagnosis.

Scientific hypothesis

HYPOTHESIS: Native-state structure alone may be insufficient to explain proteinopathy. MitoAggNet tests whether pathological outcomes are better predicted by combining:

sequence + variant + structure + disorder/APR + PTM/state + mitochondrial compartment + proteostasis context

than by sequence or native structure alone.

Primary disease axes:

  • Parkinson's disease: SNCA / alpha-synuclein
  • ALS/FTD: SOD1, TARDBP/TDP-43, FUS
  • Alzheimer's disease: MAPT/tau, APP/Aβ
  • Extension: HTT, C9orf72 DPRs, mitochondrial complex-I proteostasis

v1 tasks

  1. folding_abundance — continuous stability/folding proxy.
  2. dimerization — continuous interaction/stability proxy.
  3. aggregation — binary/continuous aggregation phenotype.
  4. mito_association — mitochondrial association/mislocalization probability.
  5. proteostasis_vulnerability — sensitivity to mitochondrial quality-control perturbation.
  6. pathogenicity — optional disease-associated variant classification.
  7. residue_aggregation — optional residue-level aggregation/APR logits.

Labels may be missing per sample. Training uses masked multi-task losses, so heterogeneous public datasets can be combined without fabricating labels.

Why this is not another AlphaFold

AlphaFold/OpenFold/Boltz outputs are used as structural priors/features, not as the scientific endpoint. MitoAggNet is aimed at pathological state transitions and context dependence: disorder, aggregation, oligomer/fibril state, PTMs, mitochondrial localization and proteostasis.

Core public data registry

The repository includes configs/data_registry.yaml. High-value sources include:

  • MitoCarta3.0 — 1,136 human mitochondrial genes, subcompartments and pathways.
  • GSE328470 — genome-wide CRISPR screen underlying the 2026 mitochondrial “holdout compartment” study.
  • PXD078690 — 2026 alpha-synuclein fibril / mitochondrial-derived-vesicle structural proteomics.
  • PXD063500 — alpha-synuclein mitochondrial LiP-MS / ATP-homeostasis interaction dataset.
  • MSV000101384 / MSV000101385 + GSE327486 — HTRA2–CLPB mitochondrial proteostasis, complex-I aggregation and striatal RNA-seq.
  • SOD1-ddPCA / bioRxiv 2026.07.15.738613 — ~6,000 SOD1 variants with folded-monomer abundance and heterodimerization phenotypes.
  • PXD066868 — human Alzheimer brain proteomics with multiple PTM classes.
  • GSE274283 / GSE274289 — UPRmt in human microglia / neuronal-glial communication.
  • GSE245929 — mitochondrial plaque-related Alzheimer dataset.
  • DisProt — experimentally curated intrinsic disorder.
  • ProteinGym / MaveDB — mutation-effect measurements.
  • AlphaFold DB / PDB — structural priors.
  • Amyloid Atlas / AmyPro — experimentally supported amyloid structures/regions.

Large/raw datasets are deliberately excluded from Git. Download them from the original repositories and keep provenance in data/manifests/.

Recommended v1 architecture

sequence ──> frozen protein LM embeddings ──┐
                                           │
structure/PDB/AF ──> residue graph GNN ────┼──> gated multimodal fusion ──> task heads
                                           │
variant + PTM + disease + mito context ────┘

heads: folding | dimerization | aggregation | mito association | proteostasis | pathogenicity

The included implementation consumes precomputed residue embeddings. This is intentional: embedding generation can be separated from model training, allowing the trainable MitoAggNet component to run on a single consumer GPU.

GPU plan

Approximate practical tiers for the MitoAggNet-specific training component:

VRAM Recommended mode
8 GB precompute PLM embeddings; small GNN/fusion model; short proteins; batch 1–8
12–16 GB comfortable v1 training with 150M PLM embeddings precomputed; mixed precision
24 GB larger batches, longer proteins, optional partial PLM fine-tuning/LoRA
32–48+ GB larger PLMs and more direct structure-model inference; still not required for core v1

Do not make Boltz/OpenFold inference a mandatory training dependency. Cache structural predictions offline.

Installation

python -m venv .venv
# Windows PowerShell:
.venv\\Scripts\\Activate.ps1
pip install -e .[dev]

Smoke test

python scripts/make_mock_batch.py
python -m mitoaggnet.train --config configs/base.yaml --data data/processed/mock_batch.pt
pytest -q

Dataset v0.1: executable real-data workflow

The repository now contains an executable builder for the first real public-data layer:

pip install -e ".[dev]"
mitoaggnet-build-data --config configs/dataset_v01.yaml

It currently integrates MitoCarta3.0 + reviewed human UniProt + AlphaFold DB + DisProt + AmyPro + the pinned 2026 SOD1 ddPCA dataset and creates:

data/processed/v0.1/proteins.csv
data/processed/v0.1/sod1_ddpca.csv
data/processed/v0.1/sod1_variant_tensor.pt
data/manifests/v0.1/downloads.json
data/manifests/v0.1/dataset_summary.json

The SOD1 tensor contains measured folding-abundance and heterodimerization labels plus sparse aggregation/pathogenicity labels, finite-label masks and a reproducible sequence/variant baseline feature matrix. Use it immediately with:

python scripts/train_sod1_baseline.py \
  --tensor data/processed/v0.1/sod1_variant_tensor.pt \
  --target folding_abundance

The baseline split is grouped by residue position to prevent substitutions at the same position leaking across train/test. The final MitoAggNet graph pipeline will subsequently add PLM embeddings, PDB/AlphaFold residue graphs, PTM state and mitochondrial context. See docs/dataset_v01.md.

v0.2: sequence + structure pipeline

MitoAggNet v0.2 adds cached protein-LM residue embeddings, PDB/mmCIF residue graphs, reference-to-structure alignment, variant-local structural features, and a grouped-by-residue SOD1 benchmark.

pip install -e ".[dev,plm]"
mitoaggnet-build-features --config configs/features_v02.yaml
mitoaggnet-train-structural --bundle data/processed/v0.2/sod1_structural_bundle.pt
mitoaggnet-mavs --proteins data/processed/v0.1/proteins.csv --out data/processed/v0.2/mavs_prior.csv

The default PLM is the small ESM-2 8M checkpoint to keep the workflow practical on consumer GPUs. The model adapter is generic Hugging Face AutoModel, so larger checkpoints can be configured without changing MitoAggNet code. MAVS-prior is a transparent, unvalidated sequence-composition heuristic retained for historical benchmarking. It is not an aggregation assay or calibrated aggregation probability. See docs/v02.md.

Critical split rules

Random mutation-level 80/20 splits are not acceptable as the primary validation because variants from the same protein leak sequence/structure information across splits.

Required evaluations:

  • Group split by protein/gene.
  • Leave-one-protein-out (LOPO).
  • Disease-family holdout when possible.
  • Temporal benchmark: train on data available before a cutoff and evaluate on later datasets.
  • Dedicated 2026 challenge sets: SOD1 DMS, HTRA2–CLPB, mitochondrial holdout, alpha-syn mitochondrial proteomics.

Historical research directions (hypotheses, not validated findings)

Historical Paper A proposal — Mitochondrial Vulnerability Atlas

The former plan was to score 1,136 MitoCarta proteins. Current claim gates do not permit that ranking: the v0.5 output is retained for provenance only and is scientifically unvalidated.

Paper B hypothesis — Variant → folding → aggregation transfer

Test, rather than assume, transfer from broad DMS/ProteinGym data to an independently measured SOD1 cohort with phenotype-specific labels.

Paper C hypothesis — State-dependent mitochondrial interactome

Evaluate explicitly measured differences between monomeric, truncated and oligomeric/fibrillar alpha-synuclein mitochondrial interactions.

Reproducibility

  • Python package under src/.
  • YAML configs under configs/.
  • No raw data committed.
  • GitHub Actions CPU smoke tests.
  • Seeded training.
  • Dataset provenance manifest.
  • CITATION.cff for later paper/release citation.

License

Code: MIT. External datasets and model weights remain subject to their original licenses/terms. See DATA_LICENSES.md.

v0.3 — mitochondrial context + proteostasis

v0.3 adds a heterogeneous mitochondrial evidence layer on top of the sequence/structure work in v0.2. It intentionally keeps mechanistically different assays separate rather than inventing a single target.

Evidence channels:

  • GSE328470: genome-wide CRISPR-Cas9 screen of degradation of the aggregation-prone YFP-CL1 reporter. YFP-high and YFP-low populations remain distinct endpoints.
  • MSV000101384 / MSV000101385: HTRA2/CLPB IP, HEK, K562 and mouse-striatum proteomics for mitochondrial intermembrane-space proteostasis and complex-I stability.
  • PXD063500: conformation-dependent alpha-synuclein mitochondrial LiP-MS / structural proteomics.
  • PXD078690: alpha-synuclein fibril-induced mitochondrial-derived-vesicle proteomics.
  • GSE327486: Htra2 mutant/control striatal RNA-seq under normoxia/hypoxia, retained as explicit transcriptomic context.

Critical missing-label rule

A protein not reported as a hit is not automatically a negative. Curated mechanistic observations are stored as positive evidence only. The learned MAVS component is enabled only when quantitative processed tables provide enough measured values.

Quick start v0.3

# Optional: downloads small processed/OA supplementary packages; avoids multi-GB raw MS/SRA files.
mitoaggnet-fetch-context --out data/raw/v0.3

# Build sparse heterogeneous evidence table + tensors.
mitoaggnet-build-context --config configs/context_v03.yaml

# Training intentionally fails if there are too few quantitative labels.
mitoaggnet-train-context \
  --bundle data/processed/v0.3/context_bundle.pt \
  --out outputs/v0.3

# Historical proteome-wide model output retained for provenance only.
mitoaggnet-predict-mavs \
  --bundle data/processed/v0.3/context_bundle.pt \
  --model outputs/v0.3/context_model.pt \
  --table data/processed/v0.3/protein_context.csv \
  --out data/processed/v0.3/mavs_learned.csv

mavs_learned_v03 is a historical, scientifically unvalidated computational output trained against sparse heterogeneous perturbation evidence. It is retained for provenance and method development, not as a candidate-prioritization score, clinical probability, diagnostic score or experimentally validated aggregation measurement.

v0.4 — MitoCarta-wide homology-aware ablation benchmark

v0.4 converts the v0.3 protein-context table into a single proteome-wide benchmark with explicit feature groups and a homology-clustered train/validation/test split.

# Summarize cached AlphaFold/PDB structures.
mitoaggnet-structure-summary \
  --proteins data/processed/v0.3/protein_context.csv \
  --structure-dir data/features/v0.2/structures \
  --out data/processed/v0.4/structure_summary.csv

# Requires MMseqs2 for the full MitoCarta proteome.
mitoaggnet-build-atlas --config configs/atlas_v04.yaml

# Pre-registered ablations on the same held-out structural cohort.
mitoaggnet-benchmark \
  --bundle data/processed/v0.4/mitocarta_atlas_bundle.pt \
  --target mavs_weak_target \
  --out outputs/v0.4 \
  --repeats 3

Ablations: sequence_only, structure_only, sequence_structure, and sequence_structure_mitochondria, plus a shuffled-training-label negative control. Homologous proteins are assigned to the same split. Quantitative ev_* columns and mavs_weak_target are forbidden as input features to prevent target leakage. See docs/v04.md.

v0.5 — Mitochondrial Vulnerability Atlas

Historical status: scientifically_unvalidated. The v0.5 files are retained for provenance, but their rankings are not validated predictions and must not be used to prioritize a proteome-wide screen.

v0.5 originally turned the locked v0.4 homology-aware matrix into a discovery atlas with uncertainty and novelty kept separate from the prediction itself.

mitoaggnet-discovery-atlas --config configs/atlas_v05.yaml

Primary outputs:

  • mitochondrial_vulnerability_atlas.csv: ranked MitoCarta proteins with ensemble MAVS-L, split-conformal interval, ensemble uncertainty, PCA+kNN OOD score and confidence;
  • novel_candidates.csv: high-confidence, high-MAVS proteins with no quantitative target label, no direct mechanistic evidence in the v0.3 table and low prior neurodegeneration evidence;
  • pathway_enrichment.csv: MitoPathway/subcompartment enrichment among top predictions with Benjamini-Hochberg correction;
  • feature_group_importance.csv: sequence/structure/mitochondrial group permutation importance on the untouched homology-held-out test set;
  • atlas_report.md: compact research report.

Known disease associations are post-hoc only. They are never predictive features. The bundled curated seed can be replaced with a versioned Open Targets snapshot:

mitoaggnet-fetch-known-evidence \
  --disease EFO_OR_MONDO_ID:"Disease label" \
  --out data/manifests/v0.5/opentargets_neurodegeneration.csv

The resulting historical score is scientifically unvalidated and retained for provenance only; it is not a research-prioritization signal, disease probability or clinical risk. See docs/v05.md.

v0.6 — Mechanistic candidate dossiers

Historical scientifically unvalidated output — provenance only. These dossiers are not validated candidates or biological discoveries. GPT2 is a “legacy exploratory pilot hypothesis derived from an unvalidated historical model.” MMUT is a “legacy high-OOD exploratory example; not prioritized.”

v0.6 converted the post-hoc v0.5 model shortlist into hypothesis dossiers. It preserves the locked historical MAVS-L output and keeps all literature, variant, interaction and disease evidence post-hoc.

Optional evidence cache:

mitoaggnet-fetch-dossier-evidence \
  --candidates outputs/v0.5/novel_candidates.csv \
  --proteins data/processed/v0.4/mitocarta_training_matrix.csv \
  --out data/raw/v0.6/dossiers

Build the dossiers:

mitoaggnet-candidate-dossiers --config configs/dossier_v06.yaml

Each dossier contains the locked v0.5 model signal, MitoCarta context, transparent residue-window prioritization, available v0.3 experimental evidence, a versioned literature/feature snapshot, an evidence-gap score, a mechanism hypothesis, and an explicit falsification-oriented validation plan. Primary outputs are candidate_dossier_index.csv, evidence_gap_ranking.csv, mechanistic_hypotheses.csv, falsification_plans.csv, candidate_report.md, and one directory per gene under outputs/v0.6/candidate_dossiers/.

The evidence-gap score describes gaps in the supplied/versioned evidence sources; it must not be interpreted as proof that a gene has never been studied. Region-level scores are research proxies, not calibrated amyloid probabilities. See docs/v06.md.

v0.7 — evidence-typed mechanistic hypothesis graphs

v0.7 connects the v0.6 candidate dossiers to positional variant/PTM annotations and experimentally curated molecular interactions.

mitoaggnet-fetch-graph-evidence \
  --candidates outputs/v0.6/candidate_dossier_index.csv \
  --proteins data/processed/v0.4/mitocarta_training_matrix.csv \
  --out data/raw/v0.7/graphs

mitoaggnet-mechanistic-graphs --config configs/graph_v07.yaml

Each candidate receives JSON and GraphML representations plus ranked mechanistic paths and variant/PTM-to-risk-region overlaps. Edges are explicitly typed as observed_direct, observed_interaction, curated_annotation, model_prediction, or hypothesis. Direction encodes a falsifiable mechanistic narrative; v0.7 does not claim causal identification from observational database integration. See docs/v07.md.

v0.8 — historical intervention-response records

v0.8 adds a distinct experimental-intervention layer on top of the evidence-typed v0.7 graphs. It can normalize CRISPR/MAGeCK screens, processed perturbation proteomics and factorial mutant/rescue expression designs without treating unmeasured genes as negatives.

mitoaggnet-fetch-interventions --out data/raw/v0.8
mitoaggnet-intervention-layer --config configs/interventions_v08.yaml

The historical Intervention Evidence Score combines recorded intervention-design strength, readout directness, statistical confidence, descriptive movement toward reference and supplied source/modality diversity. Graph edges distinguish intervention_design from observed_intervention_response, but the score is provenance-only: it is not validation of a candidate, mechanism, rescue effect or causal mediation estimate. See docs/v08.md.

v0.9 — cross-intervention triangulation

v0.9 historically tested whether intervention-response records formed convergent, paired or sequential patterns. The current artifacts contain 24,940 paired readout rows from one genuinely independent paired design, but zero convergent mediators, zero sequential chains and zero formal mediation analyses. The retained Triangulation Score is a historical descriptive heuristic, not mechanism-level evidence or candidate validation.

mitoaggnet-triangulate --config configs/triangulation_v09.yaml

Primary outputs are written under outputs/v0.9/, including convergent_mediators.csv, rescue_pairs.csv, sequential_chains.csv, triangulation_scores.csv, audited optional mediation results, and updated JSON/GraphML graphs.

Formal mediation is deliberately separated from summary-level triangulation. The optional mitoaggnet-mediation-check command requires replicate-level data and returns an assumption-dependent linear indirect-effect estimate; causal interpretation still depends on the experimental design and mediator–outcome confounding assumptions.

v1.0 — publication/reproducibility release controller

v1.0 closes the v0.1→v0.9 research pipeline with a fail-closed release audit and deterministic paper-asset builder. It does not convert a smoke test into a biological result.

# audit an already completed real-data run
mitoaggnet-release --config configs/release_v10.yaml

# execute the configured pipeline, log every stage, then audit
mitoaggnet-release --config configs/release_v10.yaml --run-pipeline

# create paper figures/tables only from locked outputs
mitoaggnet-paper-assets --config configs/release_v10.yaml

The publication gate requires full 1,136-protein MitoCarta coverage, MMseqs2 homology splitting, no cross-cluster leakage, no mock/synthetic release artifacts, no declared target/evidence features, frozen checksums and the intervention/triangulation layers. See docs/v10.md.

v1.2 — target redesign and validation repair

The first real-data execution established engineering reproducibility but produced a negative scientific result: on the held-out cohort, the old weak MAVS model had MAE 0.03279, R² -0.09898 and Spearman -0.01376, while the train-median baseline had MAE 0.02779 and a label-shuffle control also outperformed the learned model. v1.2 preserves this result and corrects the target semantics instead of tuning against the test set.

The target matrix now distinguishes direct intrinsic readouts, perturbation targets, context responses, rescue responses, interaction evidence and mechanistic annotation. Missing measurements remain missing rather than becoming zero. Every defensible endpoint is evaluated against mean, median, source-aware median and ridge baselines using homology-held-out and source-aware regimes. Repeated permutations shuffle training labels only; test labels and the evaluation cohort remain fixed.

An endpoint may support atlas generation only when all preregistered learnability conditions pass, including improvement over the median baseline, positive held-out rank signal, permutation significance, source generalization and adequate independent protein/homology-cluster counts. Failure is recorded as target_not_learnable_do_not_rank. The historical v0.5 atlas remains available only as a provenance artifact with status scientifically_unvalidated.

mitoaggnet-target-redesign --config configs/target_redesign_v12.yaml
mitoaggnet-target-validate --config configs/target_validation_v12.yaml
mitoaggnet-v12-report --root . --out-dir outputs/v1.2

See docs/v12.md and RELEASE_NOTES_v1.2.md. If no endpoint passes the gate, the required conclusion is:

No current endpoint supports proteome-wide ranking. Further direct experimental labels are required.

v1.3 — direct mitochondrial fraction-partition benchmark

v1.3 asks whether sequence and/or structure predict a directly measured mitochondrial soluble/insoluble partition phenotype without relying on the trivial shortcut that membrane proteins partition into the insoluble fraction. The conservative endpoint is fraction_partition_score; it is not called aggregation, misfolding, amyloid or pathogenicity.

PXD075952 CTRL HeLa samples define the primary DIRECT_INTRINSIC_READOUT. DNAJA3-loss measurements remain a separate CONTEXT_RESPONSE. PXD052817 is a cross_protocol_domain_shift_evaluation, not a matched external validation: it uses K-562 cells, DSSO-crosslinked mitochondria and a different mass-spectrometry workflow. Missing fractions are censored observations, not zeros or pseudocount-derived continuous labels.

mitoaggnet-direct-benchmark --config configs/direct_solubility_v13.yaml
mitoaggnet-credibility-audit --config configs/scientific_credibility.yaml

The benchmark performs proteomics QC, normalization sensitivity, membrane-confound analysis, label-reproducibility checks, homology-aware baseline evaluation, 100 train-label permutations and a frozen cross-protocol domain-shift stress test. The credibility audit adds missingness/selection bias, transparent topology variance decomposition and a machine-readable claim gate.

Current conclusion: Direct mitochondrial fraction-partition measurements were highly reproducible within PXD075952, but the tested sequence- and structure-derived representations did not show robust predictive performance in homology-held-out non-membrane proteins. Membrane/topology features contributed a detectable signal in the full cohort, while transfer to the substantially shifted PXD052817 experimental domain was poor. These findings do not support a validated universal proteome-wide mitochondrial vulnerability ranking from the currently tested data and representations.

Regardless of the result, v1.3 creates no atlas and no candidate ranking. A hypothetical positive result would not retroactively validate the old MAVS target or historical v0.5 atlas. See docs/v13.md and RELEASE_NOTES_v1.3.md.

About

Research framework for testing multimodal predictors of explicitly defined protein phenotypes

Resources

Contributing

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages