Skip to content
drjoykarmakarPublic

About

An evidence-first, endpoint-specific ADMET prediction and medicinal-chemistry decision-support system.

Topics

Resources

Contributing

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

ADMET Platform

An evidence-first, endpoint-specific ADMET prediction and medicinal-chemistry decision-support system.

CI Python API License

ADMET Platform is a production-oriented starter repository for building models and workflows that estimate absorption, distribution, metabolism, excretion, and toxicity properties of small molecules. It is designed for medicinal chemists, DMPK scientists, computational chemists, machine-learning engineers, and platform teams who need more than a black-box number.

Every prediction is intended to answer five questions:

  1. What is the estimate?
  2. How uncertain is it?
  3. Is the molecule inside the model’s supported chemical domain?
  4. Which exact model and data snapshot produced it?
  5. What should a scientist do with the result?

The repository is deliberately opinionated: an ADMET model is not a product until its assay context, data lineage, validation design, uncertainty, applicability domain, security, and review workflow are treated as first-class engineering objects.

Caution

This repository is a research and engineering starter. The bundled demonstration model is trained on synthetic example labels and is not scientifically validated. It must not be used for clinical, regulatory, patient-level, or real drug-development decisions.


Table of contents


Why this project exists

A large portion of drug-discovery effort is spent discovering that promising molecules have unacceptable properties. The difficulty is not merely predicting a single endpoint. Real projects must reconcile multiple, frequently conflicting objectives:

  • improve target potency without increasing lipophilicity;
  • improve permeability without creating efflux or solubility problems;
  • reduce clearance without increasing CYP inhibition;
  • lower hERG risk while preserving exposure;
  • interpret measurements produced under different protocols;
  • decide whether an apparently precise prediction is actually supported by nearby chemistry.

Many modeling demonstrations stop after producing a test-set metric. A medicinal-chemistry organization needs a system that also handles:

  • chemical structures, salts, tautomers and stereochemistry;
  • endpoint and assay definitions;
  • units, censored measurements and replicates;
  • data lineage and rights;
  • temporal and chemical-series validation;
  • model calibration;
  • applicability-domain detection;
  • model versioning and monitoring;
  • scientist review and override;
  • confidential customer data.

This repository establishes those boundaries early so that a successful prototype can evolve into a trustworthy product rather than being rewritten after the first enterprise pilot.


Product thesis

The highest-probability entry point is not “predict every ADMET property.” It is:

Own one expensive, frequent medicinal-chemistry decision and become the most trustworthy place to make it.

A practical initial wedge is hERG liability decision support for lead optimization. The first product should accept a compound series, produce calibrated risk estimates with uncertainty and domain evidence, identify the nearest supporting chemistry, and help a chemist prioritize which compounds to synthesize or test next.

The platform then expands into a coherent DMPK and safety panel:

  • hERG inhibition;
  • CYP3A4, CYP2D6 and CYP2C9 inhibition;
  • microsomal and hepatocyte stability;
  • aqueous solubility;
  • permeability and efflux;
  • plasma protein binding;
  • clearance and exposure-related endpoints;
  • CNS-specific properties when the chosen market requires them.

The defensible asset is not an individual algorithm. It is the accumulated system of:

  • assay-normalized measurements;
  • endpoint-specific models;
  • customer-private calibration;
  • prospective outcomes;
  • model-monitoring history;
  • medicinal-chemistry decisions and rejection reasons;
  • workflow integration.

What is included

Runnable backend

  • FastAPI application with OpenAPI documentation;
  • health and endpoint-discovery routes;
  • batch prediction contract;
  • API-key guard outside local mode;
  • structure validation and standardization;
  • endpoint model loading;
  • uncertainty intervals;
  • a basic applicability-domain score;
  • explicit warnings and model version in every response.

Scientific core

  • molecule and endpoint domain objects;
  • RDKit-enabled standardization with a minimal CI fallback;
  • descriptor calculation;
  • measurement-table validation;
  • preservation points for assay and provenance metadata.

Machine-learning package

  • deterministic feature pipeline;
  • random-forest classification and regression baselines;
  • ensemble-derived uncertainty intervals;
  • feature-space domain scoring;
  • classification and regression metrics;
  • deterministic group split utility;
  • serializable model bundles.

Product workspace

  • a clean Next.js scientific dashboard concept;
  • batch status and endpoint overview;
  • risk, confidence and domain presentation;
  • design language suitable for an enterprise scientific application.

Platform engineering

  • Docker Compose for API, PostgreSQL, Redis and MinIO;
  • API container definition;
  • environment template;
  • GitHub Actions CI;
  • linting, unit tests and repository checks;
  • model and report artifact directories.

Documentation

  • architecture and bounded contexts;
  • scientific data model;
  • model-validation protocol;
  • model-card template;
  • security and tenant-isolation baseline;
  • product roadmap;
  • contribution expectations.

System architecture

flowchart TB
    subgraph Experience
        WEB[Scientific web workspace]
        SDK[Python / REST clients]
    end

    subgraph Application
        API[FastAPI gateway]
        AUTH[Identity and tenant policy]
        JOBS[Batch job coordinator]
        AUDIT[Audit event service]
    end

    subgraph Scientific Services
        REGISTRY[Compound and assay registry]
        STANDARDIZE[Structure standardization]
        FEATURES[Feature service]
        INFERENCE[Endpoint inference]
        DOMAIN[Applicability domain]
        EXPLAIN[Evidence and explanation]
    end

    subgraph Data and MLOps
        PG[(PostgreSQL)]
        OBJECTS[(Object storage)]
        REDIS[(Redis / queue)]
        MODELS[(Model registry)]
        MONITOR[Monitoring and drift]
    end

    subgraph Training
        INGEST[Curated data ingestion]
        QC[Scientific quality control]
        SPLIT[Temporal / scaffold splits]
        TRAIN[Training and calibration]
        VALIDATE[Validation and release gate]
    end

    WEB --> API
    SDK --> API
    API --> AUTH
    API --> JOBS
    API --> STANDARDIZE
    STANDARDIZE --> REGISTRY
    STANDARDIZE --> FEATURES
    FEATURES --> INFERENCE
    INFERENCE --> MODELS
    INFERENCE --> DOMAIN
    DOMAIN --> EXPLAIN
    EXPLAIN --> API
    API --> AUDIT
    AUDIT --> PG
    JOBS --> REDIS
    REGISTRY --> PG
    API --> OBJECTS

    INGEST --> QC --> SPLIT --> TRAIN --> VALIDATE --> MODELS
    MODELS --> MONITOR
    API --> MONITOR
Loading

Architectural boundaries

Chemical truth is deterministic. Structure parsing, standardization, descriptor calculation and identifier generation are implemented as versioned chemistry functions. A language model must never fabricate or silently modify chemical structures.

Assay context is explicit. “hERG” is not a sufficient label by itself. Production data should retain protocol, technology, species, conditions, laboratory, quantification limits and mapping policy.

Models are endpoint-specific artifacts. Each model bundle contains the task, features, training-domain summary and version. Production bundles should also include dataset IDs, code commit, environment digest, calibration object, acceptance thresholds and model-card location.

Prediction and explanation are separate. The estimator generates numerical outputs. An explanation service may summarize them, but cannot change values or claim evidence that is absent.

Every action is auditable. Inputs, normalized structures, model version, outputs, user, tenant, timestamp and export events should be recoverable.

See docs/architecture.md for the concise architecture specification.


Scientific principles

1. Model the assay, not the marketing label

Different protocols can produce systematically different values. Data should be pooled only when the mapping rule is scientifically justified and validated. Endpoint definitions are versioned product objects.

2. Preserve source structures

Store the submitted structure and the standardized modeling parent. Never destroy salts, stereochemistry, charge state or original identifiers simply because the model consumes a normalized representation.

3. Preserve censored measurements

Values such as <0.1, >30 and “not detected” carry information. They should not be silently converted to exact numbers. A production pipeline may use censored-regression methods or clearly documented imputation rules.

4. Separate retrospective fit from prospective utility

Random-split performance can be encouraging while failing on novel chemical series. Temporal, scaffold, project and external holdouts are required. The strongest proof is a locked prospective prediction evaluated after experimental results arrive.

5. Report uncertainty

A point estimate without uncertainty creates false precision. Prediction intervals, ensemble disagreement, calibration and chemical-domain evidence should be visible in the product.

6. Make abstention a feature

The system should be allowed to say: “insufficient supporting chemistry; test this compound experimentally.” Low-confidence abstention is often more valuable than forced ranking.

7. Keep human scientific authority

The product supports prioritization. Scientists approve the endpoint mapping, interpret trade-offs and decide what to synthesize or test. Override reasons become valuable decision data.

8. Prefer interpretable baselines before complexity

Descriptor models, fingerprints, nearest neighbors, matched molecular pairs and tree ensembles create strong scientific baselines. More complex graph or foundation models should earn their operational complexity through reproducible, prospective improvement.


Initial endpoint strategy

Wedge: hERG inhibition

Why begin here:

  • the liability is widely recognized and decision-relevant;
  • medicinal chemists often need to address it during optimization;
  • both classification and continuous formulations are possible;
  • chemical-series context and local analogues matter;
  • uncertainty and assay heterogeneity are highly visible, making trust features valuable.

The product should not reduce hERG to a universal binary cutoff. A production implementation should define one or more assay-specific endpoints, for example:

  • probability that measured IC50 falls below a configured threshold;
  • regression on log-transformed IC50 for a specified protocol;
  • categorical risk with an explicit gray zone;
  • project-calibrated ranking using both model output and nearest-neighbor evidence.

Expansion endpoints

Endpoint Typical task Key difficulties Product value
CYP inhibition Classification or regression Multiple isoforms, protocol variation, class imbalance Drug–drug interaction risk triage
Microsomal stability Regression or ordinal Species and matrix differences, censoring Clearance optimization
Solubility Regression Method, pH, kinetic vs thermodynamic values Formulation and exposure risk
Permeability Regression or classification Cell line and protocol dependence Oral/CNS exposure triage
Efflux ratio Regression Transporter and cell-system variation CNS and oral design support
Plasma protein binding Regression High-bound region is difficult; species effects Exposure interpretation
Hepatotoxicity panels Multi-task classification Endpoint ambiguity, class imbalance, causal uncertainty Early safety screening

The endpoint definitions in configs/endpoints.yaml are intentionally small and readable. They should eventually be stored in a governed endpoint registry.


Repository structure

admet-platform/
├── apps/
│   ├── api/                         # FastAPI prediction service
│   │   ├── src/admet_api/
│   │   │   ├── main.py              # routes and service metadata
│   │   │   ├── schemas.py           # versioned API contracts
│   │   │   ├── service.py           # inference orchestration
│   │   │   └── settings.py          # environment configuration
│   │   └── tests/
│   └── web/                         # Next.js scientific workspace concept
├── packages/
│   ├── admet_core/
│   │   ├── src/admet_core/
│   │   │   ├── chemistry.py         # standardization and descriptors
│   │   │   ├── domain.py            # scientific domain objects
│   │   │   └── validation.py        # input data validation
│   │   └── tests/
│   └── admet_ml/
│       ├── src/admet_ml/
│       │   ├── features.py           # feature construction
│       │   ├── metrics.py            # endpoint metrics
│       │   ├── modeling.py           # training, inference, domain score
│       │   └── splits.py             # group-aware splitting
│       └── tests/
├── configs/
│   ├── data_contract.yaml           # input measurement contract
│   └── endpoints.yaml               # initial endpoint definitions
├── data/examples/                   # synthetic demonstration data only
├── docs/
│   ├── architecture.md
│   ├── data-model.md
│   ├── model-card-template.md
│   ├── model-validation.md
│   ├── product-roadmap.md
│   └── security.md
├── infra/docker/api.Dockerfile
├── scripts/
│   ├── check_repo.py
│   ├── demo_prediction.py
│   └── train_baselines.py
├── artifacts/                       # ignored model and report outputs
├── .github/workflows/ci.yml
├── docker-compose.yml
├── Makefile
├── pyproject.toml
└── README.md

Quick start

Prerequisites

  • Python 3.11 or newer;
  • Git;
  • Docker and Docker Compose for the full local stack;
  • Node.js 20 or newer only when running the web workspace;
  • RDKit strongly recommended for chemistry-grade parsing and descriptors.

1. Create a virtual environment

python -m venv .venv
source .venv/bin/activate

Windows PowerShell:

py -m venv .venv
.venv\Scripts\Activate.ps1

2. Install the Python project

Minimal development environment:

python -m pip install --upgrade pip
python -m pip install -e '.[dev]'

With RDKit support:

python -m pip install -e '.[dev,chem]'

With experiment tracking and optimization tools:

python -m pip install -e '.[dev,chem,mlops]'

3. Run the test suite

make test

4. Start the API

make api

Open:

  • API documentation: http://localhost:8000/docs
  • health check: http://localhost:8000/health
  • endpoint registry: http://localhost:8000/v1/endpoints

5. Run the bundled demonstration

make demo

The output includes:

  • point estimate;
  • lower and upper ensemble quantiles;
  • confidence category;
  • applicability-domain score;
  • model version;
  • warnings.

Warning

When no trained artifact exists, the API loads a deterministic demonstration classifier trained on synthetic labels. This behavior is useful for contract testing only. A production deployment should fail closed when an approved endpoint model is unavailable.

6. Train demonstration baseline artifacts

make train

This reads data/examples/admet_training_demo.csv and writes endpoint model bundles to artifacts/models/.

7. Start the infrastructure stack

cp .env.example .env
docker compose up --build

Services:

Service Address Purpose
API localhost:8000 prediction and metadata API
PostgreSQL localhost:5432 governed metadata and audit records
Redis localhost:6379 queues and caching
MinIO localhost:9000 datasets, model artifacts and reports
MinIO console localhost:9001 local object-store administration

8. Run the web workspace

cd apps/web
npm install
npm run dev

Open http://localhost:3000.


API usage

Health

curl http://localhost:8000/health

Discover endpoint status

curl http://localhost:8000/v1/endpoints

Predict a batch

curl -X POST http://localhost:8000/v1/predictions \
  -H 'Content-Type: application/json' \
  -d '{
    "endpoint": "herg",
    "molecules": [
      {
        "compound_id": "CMP-001",
        "smiles": "CCOc1ccc2nc(S(N)(=O)=O)sc2c1"
      },
      {
        "compound_id": "CMP-002",
        "smiles": "CN1CCC(CC1)Oc2ccc(C#N)cc2"
      }
    ]
  }'

Example response shape:

{
  "request_id": "0faad95c-d285-4e23-a1b9-2f1d11ec48ce",
  "predictions": [
    {
      "compound_id": "CMP-001",
      "endpoint": "herg",
      "value": 0.27,
      "lower": 0.05,
      "upper": 0.61,
      "confidence": "medium",
      "in_domain": true,
      "domain_score": 0.58,
      "model_version": "baseline-0.1.0",
      "warnings": []
    }
  ]
}

Production contract recommendations

Before external release, add:

  • tenant and project identifiers;
  • request idempotency key;
  • standardization-policy version;
  • model selection policy;
  • endpoint version;
  • synchronous and asynchronous batch routes;
  • source-structure echo and standardized structure;
  • structured warning codes;
  • nearest-neighbor evidence;
  • prediction timestamp and artifact digest;
  • audit-event identifier;
  • explicit units for regression outputs;
  • explainability object;
  • error object with row-level validation failures.

A useful production route family could be:

POST /v1/projects/{project_id}/datasets
POST /v1/projects/{project_id}/prediction-jobs
GET  /v1/prediction-jobs/{job_id}
GET  /v1/prediction-jobs/{job_id}/results
GET  /v1/models/{endpoint}/{version}/card
POST /v1/predictions/{prediction_id}/review

Data contract

The minimal normalized training input is:

Field Required Meaning
compound_id yes stable source identifier
smiles yes source molecular representation
endpoint yes governed endpoint identifier
value yes measured value before endpoint transform
unit yes source unit
relation recommended =, <, <=, >, >=
assay_id recommended immutable assay/protocol identifier
assay_type recommended technology or format
species endpoint-dependent biological species
matrix endpoint-dependent microsomes, hepatocytes, plasma, cells, etc.
pH endpoint-dependent relevant experimental condition
temperature_c endpoint-dependent experimental temperature
source yes in production source and licensing record
measured_at recommended enables temporal validation

The declarative starter contract is in configs/data_contract.yaml.

Data ingestion stages

  1. Land raw data immutably. Save the original file, checksum, source, license, uploader and timestamp.
  2. Parse without scientific assumptions. Preserve all values, units and relation operators.
  3. Standardize structures. Generate the modeling parent while retaining source structures.
  4. Map units. Convert only through governed, endpoint-specific rules.
  5. Map assays to endpoint versions. Do not pool merely because labels look similar.
  6. Detect duplicates and conflicts. Preserve provenance for every replicate.
  7. Apply quality flags. Invalid structures, ambiguous stereochemistry, impossible units, outliers and protocol gaps.
  8. Create an immutable dataset snapshot. Assign a dataset ID and content digest.
  9. Generate a quality report. Document exclusions and transformation counts.
  10. Approve for training. Scientific review is separate from pipeline success.

Duplicate policy

Possible duplicate strategies include:

  • retain every replicate and model replicate structure;
  • aggregate with median or geometric mean after protocol matching;
  • create separate labels by laboratory or assay version;
  • select the most recent or highest-quality measurement;
  • exclude conflicting records above a prespecified variance threshold.

The choice must be endpoint-specific and documented in the model card.

Data rights

Before using any dataset, record:

  • origin;
  • license or contract;
  • permitted purposes;
  • redistribution restrictions;
  • whether model training is allowed;
  • whether derived model artifacts may be commercialized;
  • retention requirements;
  • customer-private status.

Public accessibility does not remove the need for a rights review.


Chemical standardization

The core function is standardize_smiles in packages/admet_core/src/admet_core/chemistry.py.

When RDKit is installed, the starter pipeline performs:

  1. SMILES parsing;
  2. RDKit cleanup;
  3. fragment-parent selection;
  4. charge neutralization where supported;
  5. canonical isomeric SMILES generation;
  6. SHA-256 structure digest generation.

A production policy should explicitly address:

  • salts and solvates;
  • mixtures and disconnected fragments;
  • organometallic compounds;
  • isotope labels;
  • undefined stereocenters;
  • atropisomerism;
  • tautomer canonicalization;
  • protonation state;
  • valence and aromaticity errors;
  • covalent warheads;
  • duplicate detection across representations.

Version standardization policies

Changing structure standardization can change fingerprints, descriptors, duplicates and model labels. Therefore the output should include a standardization policy ID, for example:

small-molecule-parent/2026-01

Datasets and models should reference the exact policy version. Reprocessing an old dataset should produce a new snapshot rather than mutating historical records.

Fallback behavior

The repository includes a very small fallback parser so unit tests can run without RDKit. It is not chemistry-grade and returns a warning:

rdkit_not_installed_fallback_used

Production startup should refuse to serve predictions when the approved chemistry runtime is absent.


Feature engineering

The starter uses a compact descriptor vector:

  • molecular weight;
  • calculated logP;
  • topological polar surface area;
  • hydrogen-bond donors;
  • hydrogen-bond acceptors;
  • rotatable bonds;
  • heavy-atom count.

This is intentionally a baseline. A mature feature registry may include:

Two-dimensional molecular features

  • Morgan/ECFP fingerprints;
  • MACCS or other structural keys;
  • topological descriptors;
  • fragment counts;
  • charge and ionization features;
  • aromatic ring and heterocycle features;
  • structural alerts;
  • matched molecular-pair context.

Three-dimensional features

  • conformer ensembles;
  • shape and pharmacophore descriptors;
  • electrostatic features;
  • solvent-accessible surface estimates;
  • quantum-derived descriptors when justified.

Learned representations

  • graph neural-network embeddings;
  • SMILES or molecular-language embeddings;
  • pre-trained chemical foundation-model embeddings;
  • multi-task endpoint representations.

Context features

  • assay protocol;
  • species;
  • matrix;
  • laboratory/vendor;
  • project/series identity for hierarchical models;
  • measurement date;
  • batch quality information.

Every feature set should have:

  • a version;
  • deterministic implementation;
  • test vectors;
  • missing-value policy;
  • computation environment;
  • performance and cost profile;
  • explanation of chemical invariances.

Model development

Baseline ladder

Do not jump directly to the most sophisticated architecture. Establish a ladder:

  1. endpoint prevalence or mean baseline;
  2. nearest-neighbor baseline;
  3. descriptor linear/logistic model;
  4. fingerprint similarity or kernel method;
  5. random forest or gradient boosting;
  6. graph neural network;
  7. multi-task model;
  8. project-specific fine-tuning or calibration;
  9. ensembles.

A model is valuable only if it beats simpler baselines under the split that represents actual deployment.

Classification

For liability classification, the starter uses a class-balanced random forest. Production work should evaluate:

  • class-weighted gradient boosting;
  • calibrated tree ensembles;
  • fingerprint-based models;
  • graph models;
  • conformal classification;
  • endpoint-specific gray zones;
  • cost-sensitive thresholds.

The probability must be calibrated. A model that labels compounds correctly but produces unreliable probabilities cannot support rational risk thresholds.

Regression

For continuous endpoints, the starter provides a random-forest regression path. Production models should evaluate:

  • log transforms aligned with assay interpretation;
  • censored regression;
  • heteroscedastic uncertainty;
  • robust losses;
  • rank-aware objectives;
  • hierarchical assay effects;
  • quantile regression;
  • conformal prediction intervals.

Multi-task learning

Multi-task models can transfer information among related endpoints, but they create governance and debugging complexity. Use them after single-endpoint baselines are stable. Validate whether transfer helps each endpoint and chemical domain; aggregate improvement can hide endpoint regressions.

Project-specific adaptation

Customer projects frequently occupy a narrow chemical series. A strong platform should support:

  • recalibration using local measurements;
  • local nearest-neighbor models;
  • hierarchical global-plus-project models;
  • uncertainty reduction as project data accumulate;
  • private model artifacts;
  • explicit fallbacks when local data are too sparse.

The shared global model offers an initial estimate; the private project layer creates increasing customer value and defensibility.


Uncertainty and applicability domain

The API reports ensemble quantiles and a simple feature-space domain score. This is a scaffold, not a final scientific method.

Production uncertainty stack

Use several independent signals:

  • ensemble variance;
  • calibrated prediction intervals;
  • conformal methods;
  • nearest-neighbor similarity;
  • feature-space distance;
  • disagreement among model families;
  • training-data density;
  • assay noise floor;
  • stereochemistry and structure-quality flags.

Domain classification

A useful user-facing result is:

Domain status Meaning Recommended action
In domain supported by nearby, consistent training chemistry use as decision support
Edge of domain partial support or moderate model disagreement inspect analogues and consider testing
Out of domain little relevant training support abstain or treat as hypothesis only

The threshold should be endpoint-specific and validated against empirical error. Do not choose it solely because a score “looks reasonable.”

Nearest-neighbor evidence

For every prediction, the product should show several relevant training analogues with:

  • structure;
  • similarity;
  • assay value;
  • assay/protocol match;
  • source;
  • model inclusion status;
  • important structural differences.

This allows medicinal chemists to combine model output with recognizable SAR evidence.

Confidence labels

Avoid presenting “high,” “medium” and “low” without a definition. A model card should define each label using measurable criteria such as interval width, calibration bucket, minimum similarity and ensemble disagreement.


Evaluation and release gates

Detailed requirements are in docs/model-validation.md.

Split strategy

Use multiple evaluation views:

  • random split: debugging and continuity only;
  • scaffold split: tests structural generalization;
  • chemical-series/project split: tests transfer to unseen programs;
  • temporal split: simulates future prediction;
  • source/laboratory split: tests protocol and organization transfer;
  • external set: tests independent reproducibility;
  • prospective set: tests actual future utility.

Classification metrics

At minimum:

  • ROC-AUC;
  • PR-AUC;
  • balanced accuracy;
  • sensitivity and specificity;
  • positive and negative predictive value;
  • Matthews correlation coefficient;
  • Brier score;
  • expected calibration error;
  • calibration plot;
  • confusion matrices at decision thresholds.

Evaluate by:

  • scaffold;
  • source;
  • value range;
  • molecular-weight and lipophilicity bands;
  • ionization class;
  • nearest-neighbor similarity;
  • domain status.

Regression metrics

At minimum:

  • MAE;
  • RMSE;
  • R²;
  • Spearman rank correlation;
  • median absolute error;
  • error by assay range;
  • interval coverage and width;
  • residual plots;
  • error by chemical series and domain status.

Prospective protocol

  1. Freeze model, code, data and thresholds.
  2. Timestamp and sign predictions before experiments.
  3. Record all proposed compounds, not just successful selections.
  4. Receive assay results without retroactive endpoint remapping.
  5. Evaluate predeclared metrics.
  6. Investigate failures and domain behavior.
  7. Publish an internal prospective report.
  8. Decide whether to promote, revise or reject the model.

Release gate

A model can be promoted only after:

  • data lineage is complete;
  • rights review is complete;
  • scientific mapping is approved;
  • required split metrics pass;
  • calibration is acceptable;
  • uncertainty intervals have empirical support;
  • domain rules are evaluated;
  • failure modes are documented;
  • security scan passes;
  • model card is approved;
  • rollback artifact exists.

MLOps and reproducibility

Model bundle

The starter ModelBundle contains:

  • endpoint;
  • task;
  • fitted estimator;
  • feature columns;
  • version;
  • training-domain center and scale.

A production artifact should additionally contain:

  • endpoint version;
  • dataset snapshot IDs;
  • model code commit;
  • standardization version;
  • feature version;
  • training configuration;
  • seed;
  • dependency lock digest;
  • calibration object;
  • decision thresholds;
  • metric report location;
  • model-card digest;
  • artifact checksum;
  • reviewer approvals.

Experiment tracking

The optional mlops dependency group includes MLflow and Optuna. A practical experiment record should log:

  • full configuration;
  • dataset IDs, not local filenames;
  • split assignments;
  • metrics and confidence intervals;
  • plots;
  • environment;
  • trained artifact;
  • feature importance or explanation summaries;
  • subgroup metrics;
  • failure examples.

Promotion environments

Suggested lifecycle:

candidate -> scientifically reviewed -> staging -> prospective validation -> production -> retired

Production must reference immutable artifacts. Never overwrite a model file in place.

Monitoring

Monitor:

  • request volume and latency;
  • invalid structure rate;
  • domain-score distribution;
  • feature drift;
  • prediction distribution;
  • calibration after labels arrive;
  • interval coverage;
  • error by chemical series;
  • model and data-version adoption;
  • user overrides and rejection reasons;
  • assay turnaround and label delay.

Drift alone should not automatically retrain a model. It should trigger review, because drift may represent a valuable new project domain requiring dedicated data and validation.


Product experience

The web concept lives in apps/web. The intended product is a scientific workspace, not a chatbot-first interface.

Primary screens

1. Project dashboard

  • project objectives and endpoint thresholds;
  • current compound-series distribution;
  • outstanding predictions and experiments;
  • model coverage and warnings;
  • recent decisions.

2. Batch prediction table

Columns should include:

  • compound ID and structure;
  • source structure status;
  • endpoint value/probability;
  • interval;
  • confidence;
  • domain status;
  • nearest-neighbor similarity;
  • important alerts;
  • model version;
  • scientist review state.

3. Compound detail

  • 2D structure;
  • all endpoint predictions;
  • confidence and domain evidence;
  • nearest measured analogues;
  • property radar or parallel coordinates;
  • structural alerts;
  • project comparison;
  • model history;
  • reviewer comments.

4. Model card

  • intended use;
  • assay definition;
  • training-data coverage;
  • validation results;
  • limitations;
  • current monitoring status;
  • release owner.

5. Experimental feedback

  • upload results;
  • map assays;
  • confirm structures and batches;
  • compare predicted vs measured;
  • record design-hypothesis outcome;
  • approve data for private retraining.

Decision language

Prefer:

“Predicted hERG risk is elevated, with moderate confidence. The molecule lies near the edge of the supported domain. Two close analogues show lower risk after removal of the basic side chain. Experimental confirmation is recommended.”

Avoid:

“This molecule is safe.”

The software predicts assay outcomes; it does not establish whole-organism or clinical safety.


Security and enterprise deployment

See docs/security.md.

Required controls

  • tenant-scoped authorization in every service;
  • encryption at rest and in transit;
  • private object prefixes and database row-level policies;
  • secrets management;
  • immutable audit events;
  • malware and content scanning on upload;
  • sandboxed parsing workers;
  • rate limits and quotas;
  • artifact signature or checksum validation;
  • backup and disaster-recovery testing;
  • incident response;
  • no-training flags for customer data;
  • retention and deletion policies;
  • administrative access review;
  • dependency and container scanning.

Deployment patterns

Shared SaaS

Best economics and iteration speed. Requires strong logical isolation and customer acceptance.

Dedicated tenant

Separate database, object storage and workers for each customer. Higher cost, easier boundary explanation.

Customer VPC

Application and model runtime deployed into the customer’s cloud account. Useful for highly sensitive programs.

On-premises or disconnected

Highest operational burden. Requires artifact distribution, upgrade procedures, offline license and support strategy.

Customer-data training policy

The default should be:

Customer data are not used to train shared models unless the customer has explicitly agreed to that use.

Private project adaptation can still occur inside the tenant boundary. The product should visibly distinguish:

  • global public/licensed model;
  • organization-private model;
  • project-private calibration;
  • experimental or unapproved model.

Implementation plan

Phase 0 — Customer and assay discovery: weeks 1–4

Goal: define one narrow endpoint product with real demand.

Deliverables:

  • 25–40 interviews with medicinal chemists, DMPK scientists and project leaders;
  • five observed design or triage workflows;
  • endpoint decision map;
  • current tools and failure modes;
  • willingness-to-pay evidence;
  • first pilot-design-partner criteria;
  • hERG assay taxonomy and data-contract draft;
  • success metrics.

Exit criteria:

  • one decision is selected;
  • at least three teams will provide historical evaluation data or a paid pilot;
  • assay context can be defined tightly enough for modeling.

Phase 1 — Data foundation: weeks 3–10

Deliverables:

  • source and rights registry;
  • immutable raw-data landing;
  • structure-standardization policy;
  • assay and unit mapping;
  • duplicate/replicate policy;
  • quality dashboard;
  • versioned dataset snapshots;
  • 10–20 historical project sets where possible.

Exit criteria:

  • every modeled row is traceable to source;
  • exclusions are explainable;
  • endpoint class/value distribution is understood;
  • scientific reviewer approves mapping rules.

Phase 2 — Baselines and benchmark: weeks 7–14

Deliverables:

  • prevalence/mean baseline;
  • nearest-neighbor baseline;
  • descriptor and fingerprint models;
  • calibrated tree ensemble;
  • scaffold and temporal benchmarks;
  • subgroup error report;
  • domain prototype;
  • draft model card.

Exit criteria:

  • model beats relevant simple baselines;
  • probabilities are sufficiently calibrated for pilot thresholds;
  • unsupported chemistry can be detected better than chance;
  • limitations are acceptable for a controlled historical test.

Phase 3 — Pilot product: weeks 11–20

Deliverables:

  • secure project workspace;
  • CSV/SDF upload;
  • structure and assay validation report;
  • batch predictions;
  • analogue evidence;
  • reviewer workflow;
  • audit trail;
  • exportable report;
  • pilot administration tools.

Exit criteria:

  • scientists can complete the workflow without developer intervention;
  • prediction provenance is visible;
  • security review passes for pilot data;
  • historical evaluation demonstrates useful time savings or prioritization value.

Phase 4 — Prospective study: months 5–9

Deliverables:

  • locked predictions before synthesis/testing;
  • experimental-result ingestion;
  • prospective analysis plan;
  • recommendation acceptance and rejection data;
  • calibration and error analysis;
  • customer outcome report.

Exit criteria:

  • at least one customer synthesizes or tests model-supported compounds;
  • prospective value is measurable;
  • failures produce actionable model or workflow improvements;
  • customer agrees to continue or expand.

Phase 5 — DMPK panel: months 8–15

Deliverables:

  • two to five additional endpoints;
  • multi-endpoint table;
  • project-defined constraints;
  • Pareto-front ranking;
  • customer-private recalibration;
  • endpoint-specific model monitoring;
  • expanded integrations.

Phase 6 — Design guidance: months 12–24

Deliverables:

  • matched molecular-pair explanations;
  • liability-reducing transforms;
  • constrained analogue suggestions;
  • synthetic-feasibility integration;
  • active-learning experiment selection;
  • closed-loop project learning.

Team and operating model

Minimum founding team

Medicinal chemistry / product founder

Owns:

  • customer discovery;
  • endpoint choice;
  • scientific requirements;
  • output review;
  • pilot sales;
  • design-partner relationships;
  • business narrative.

Cheminformatics engineer

Owns:

  • structure standardization;
  • fingerprints and descriptors;
  • chemical database design;
  • similarity and substructure search;
  • data ingestion;
  • scientific performance.

ADMET machine-learning scientist

Owns:

  • endpoint formulation;
  • model baselines;
  • uncertainty and calibration;
  • split design;
  • validation;
  • model cards;
  • prospective study analysis.

Full-stack/platform engineer

Owns:

  • APIs;
  • project workspace;
  • authentication;
  • jobs and storage;
  • deployment;
  • observability;
  • integrations.

DMPK scientific lead or adviser

Owns:

  • assay taxonomy;
  • data-combination policy;
  • interpretation;
  • experimental-validation design;
  • customer scientific credibility.

First 10 hires

A sensible sequence:

  1. medicinal-chemistry/product founder;
  2. senior cheminformatics engineer;
  3. ML scientist;
  4. product-minded full-stack engineer;
  5. DMPK scientist;
  6. data engineer;
  7. enterprise/backend engineer;
  8. product designer;
  9. customer scientific success lead;
  10. security/infra engineer.

Avoid building a large pure-research team before securing real project data and workflow adoption.

ADHD-aware founder operating system

For a founder with ADHD, design the company so that critical execution does not depend on memory or sustained administrative attention:

  • one weekly company scorecard;
  • three quarterly priorities;
  • daily written top-three tasks;
  • operations-oriented counterpart;
  • explicit owner and due date for every decision;
  • short design-partner feedback cycles;
  • automated notes and follow-ups;
  • protected scientific deep-work blocks;
  • no more than one new endpoint entering development at a time;
  • decision log for ideas that are intentionally deferred.

Hyperfocus belongs on product insight, scientific synthesis, recruiting and major negotiations. Repetitive release, security and project-management processes should be systematized and delegated.


Budget scenarios

These are planning ranges, not quotes.

Technical proof: approximately $75,000–$250,000

Suitable for:

  • two or three founders;
  • low founder salaries;
  • one endpoint;
  • public or already licensed data;
  • historical benchmark;
  • basic API and local workspace;
  • no enterprise compliance program.

Primary spending:

  • personnel;
  • data cleaning;
  • compute;
  • legal/data-rights review;
  • customer discovery.

Pilot-ready platform: approximately $900,000–$2.5 million

Suitable for:

  • five to eight people;
  • 12–18 months;
  • secure design-partner pilots;
  • production ingestion and lineage;
  • endpoint validation;
  • scientific UI;
  • audit and tenant controls;
  • prospective study.

Enterprise-ready platform: approximately $3 million–$8 million

Suitable for:

  • multiple endpoint teams;
  • VPC or dedicated deployments;
  • formal security program;
  • data integrations;
  • model monitoring;
  • customer scientific success;
  • private adaptation;
  • larger prospective evidence base.

Laboratory-connected platform: $10 million and above

Adding synthesis and assays creates a different business:

  • compound procurement;
  • CRO contracts or facilities;
  • sample logistics;
  • analytical chemistry;
  • assay development;
  • laboratory informatics;
  • quality systems;
  • longer working-capital cycles.

Begin with software and carefully selected external experiments unless laboratory ownership is central to a differentiated data strategy.


Business model and defensibility

Possible commercial models

Annual enterprise subscription

Access to platform, endpoint panel, seats and a prediction volume allowance.

Project or program license

Useful for smaller biotechnology companies with one or two active programs.

Private model package

Additional fee for customer-specific training, validation and deployment.

API usage

Suitable for integration into internal design systems, registration tools and ELNs.

Scientific services

Data mapping, historical benchmark and model validation can accelerate early adoption, but should feed a scalable product rather than becoming unlimited consulting.

Value metrics

Price against measurable value:

  • fewer unnecessary compounds synthesized;
  • faster design cycles;
  • improved probability of selecting acceptable compounds;
  • reduced time preparing DMPK reviews;
  • increased experimental learning per cycle;
  • earlier termination of poor series;
  • improved traceability of model-assisted decisions.

Defensibility ladder

Weakest to strongest:

  1. generic model served by API;
  2. better endpoint benchmarks;
  3. unique assay-normalized datasets;
  4. workflow integration;
  5. customer-private calibration;
  6. prospective experimental feedback;
  7. medicinal-chemistry decision and rejection data;
  8. cross-endpoint design optimization;
  9. closed design–make–test learning loop.

The long-term objective is to become the system where project data return after every experiment and the models become increasingly specific to the customer’s chemistry and assays.


Testing

Run all checks:

make check

Run tests directly:

pytest

Run linting:

ruff check .

Included tests

  • API health contract;
  • prediction response contract;
  • invalid-SMILES handling;
  • deterministic structure hashing;
  • descriptor-schema check;
  • model prediction interval shape;
  • applicability-domain score bounds.

Testing that a production system still needs

Chemistry golden tests

Maintain structures that cover:

  • salts;
  • zwitterions;
  • charged compounds;
  • stereocenters;
  • isotopes;
  • aromaticity edge cases;
  • mixtures;
  • invalid valence;
  • tautomer-sensitive cases.

Expected standardized outputs should be versioned and reviewed by chemists.

Data-pipeline tests

  • unit conversions;
  • censored values;
  • duplicate aggregation;
  • assay mapping;
  • provenance preservation;
  • immutable snapshot hashing;
  • schema evolution.

Model tests

  • split leakage detection;
  • deterministic training within tolerance;
  • metric regression tests;
  • calibration tests;
  • interval coverage;
  • domain-error relationship;
  • artifact load and signature validation.

Product tests

  • tenant isolation;
  • permissions;
  • upload abuse;
  • large batches;
  • retries and idempotency;
  • audit completeness;
  • export reproducibility;
  • rollback.

Known limitations

  • The bundled data are synthetic and have no scientific meaning.
  • The demonstration model is not approved for any real endpoint decision.
  • The fallback SMILES validator is intentionally primitive.
  • The descriptor set is too small for competitive endpoint modeling.
  • The domain score is a feature-distance heuristic, not an empirically calibrated applicability-domain method.
  • Ensemble quantiles from random-forest trees are not guaranteed calibrated intervals.
  • The API currently performs synchronous in-process inference.
  • Database models, migrations, audit persistence and tenant authorization are architectural placeholders rather than complete implementations.
  • The web application is a visual product concept, not a connected production client.
  • Regression endpoint training needs endpoint transforms and unit normalization.
  • Censored data are documented but not yet modeled.
  • Nearest-neighbor evidence and chemical drawings are not implemented in the API.
  • No model-signing or artifact-integrity layer is implemented.
  • No formal regulated-software claim is made.

These limitations are explicit so the repository is not mistaken for a validated commercial product.


Roadmap

The detailed roadmap is in docs/product-roadmap.md.

Near term

  • real RDKit-only production chemistry runtime;
  • Pydantic endpoint registry;
  • assay-aware ingestion;
  • PostgreSQL schema and migrations;
  • fingerprint and nearest-neighbor service;
  • calibrated hERG baseline;
  • scaffold and temporal benchmark reports;
  • batch file jobs;
  • model cards served through API;
  • connected web workspace.

Medium term

  • additional CYP and DMPK endpoints;
  • MLflow registry integration;
  • conformal intervals;
  • project-private recalibration;
  • SSO and tenant policy;
  • audit persistence;
  • drift monitoring;
  • historical customer benchmark workflow;
  • prospective-study module.

Long term

  • multi-endpoint optimization;
  • matched molecular-pair recommendations;
  • experiment selection;
  • ELN and compound-registration integrations;
  • customer VPC deployment;
  • private model federation where appropriate;
  • closed-loop design–make–test orchestration.

Contributing

Read CONTRIBUTING.md.

Scientific changes require more than passing unit tests. A pull request that changes any of the following must document the expected impact and migration path:

  • structure standardization;
  • endpoint definition;
  • unit mapping;
  • duplicate policy;
  • dataset inclusion/exclusion;
  • features;
  • split strategy;
  • calibration;
  • model threshold;
  • applicability-domain logic;
  • prediction contract.

Never commit proprietary compounds, customer data, assay results, credentials or model artifacts derived from restricted data.


Final perspective

A credible ADMET company is built by combining medicinal chemistry, DMPK science, data engineering, machine learning, software product design, security, and prospective experimental evidence.

The winning system will not be the one that produces the most predictions. It will be the one that scientists trust to distinguish:

  • supported prediction from speculation;
  • assay signal from data artifact;
  • local evidence from global extrapolation;
  • useful optimization direction from confident nonsense.

Build one endpoint deeply. Preserve every source and decision. Validate prospectively. Then expand the panel and learning loop.

About

An evidence-first, endpoint-specific ADMET prediction and medicinal-chemistry decision-support system.

Topics

Resources

Contributing

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages