Skip to content

[M5][DRUG-DISCOVERY] Integrate Multiple Myeloma pharmacogenomics and dependency data #57

Description

@andreazedda

Objective

Connect bmyCure4MM drug-discovery capabilities to Multiple Myeloma-specific biological dependencies, drug-response data, molecular states, and evidence rather than relying on generic molecular similarity or drug-likeness alone.

This issue owns the MM-specific evidence/data layer consumed by the end-to-end mechanism-to-molecule pipeline in #77.

Initial data sources

  • DepMap CRISPR dependencies and omics;
  • GDSC drug-response data;
  • PRISM drug-repurposing data;
  • CoMMpass genomics, transcriptomics, treatment, and outcomes;
  • curated MM cell-line metadata and relevant public functional screens;
  • ChEMBL, PubChem, PDB, and other compound/target resources as supporting layers.

Scope

  • resolve MM cell-line and sample identities across sources;
  • construct gene-dependency, drug-sensitivity, target, pathway, and biomarker matrices;
  • record tissue, subtype, treatment context, assay, concentration, and quality;
  • define train/test splits that avoid cell-line, compound, or scaffold leakage where applicable;
  • establish transparent RF/linear/elastic-net baselines before deep models;
  • add interpretable genomic/transcriptomic response-prediction baselines;
  • connect candidates to the resistance and evidence graphs;
  • separate target validation, compound activity, ADME, and clinical plausibility.

Relationship to #77

#57 = MM-specific evidence and dependency substrate
#77 = mechanism → target → modality → molecule → developability → PK/PD → virtual-patient counterfactual pipeline

#57 therefore provides governed evidence for the Mechanism & Pathway Atlas, patient/subtype mechanistic profiles, target prioritization, MM-specific candidate relevance and drug-response validation in #77.

A generic molecular score cannot satisfy #57 or #77. Candidate-level outputs must keep distinct:

binding / docking
molecular similarity
MM dependency
functional response
selectivity
off-target liability
clinical / translational evidence

The three discovery lanes in #77—REPURPOSING, OPTIMIZATION, DE_NOVO_DESIGN—may consume this evidence differently, but all MM-relevance claims require an explicit #57 evidence record or an explicit INSUFFICIENT_MM_SPECIFIC_EVIDENCE state.

Acceptance criteria

  • MM-specific evidence is required before a candidate is presented as MM-relevant.
  • Cell-line identity and cross-dataset mapping are validated.
  • Drug-response units, assay conditions, and censored values are explicit.
  • Baseline models and leakage-resistant splits are implemented.
  • External validation or held-out domain evaluation is reported.
  • Molecular similarity, binding, dependency and clinical evidence remain separate scores.
  • Candidate ranking exposes evidence, uncertainty, contradictions and failure reasons.
  • GNN or generative models are added only after strong baselines and task definition.
  • [M5][DRUG-DISCOVERY] Build mechanism-to-molecule discovery and virtual-patient counterfactual validation pipeline #77 can consume [M5][DRUG-DISCOVERY] Integrate Multiple Myeloma pharmacogenomics and dependency data #57 outputs through versioned, provenance-preserving contracts without duplicating dataset authority.
  • A candidate with no supported MM-specific dependency remains explicitly unsupported even when docking, ADME or generic drug-likeness scores are favorable.

Non-goals

This issue does not claim that an in-vitro response or computational score is a clinically effective treatment. It does not own molecule generation, retrosynthesis, developability orchestration, candidate PK/PD or virtual-patient counterfactual comparison; those are integrated under #77.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Projects

No projects

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions