An evidence-first, endpoint-specific ADMET prediction and medicinal-chemistry decision-support system.
ADMET Platform is a production-oriented starter repository for building models and workflows that estimate absorption, distribution, metabolism, excretion, and toxicity properties of small molecules. It is designed for medicinal chemists, DMPK scientists, computational chemists, machine-learning engineers, and platform teams who need more than a black-box number.
Every prediction is intended to answer five questions:
- What is the estimate?
- How uncertain is it?
- Is the molecule inside the model’s supported chemical domain?
- Which exact model and data snapshot produced it?
- What should a scientist do with the result?
The repository is deliberately opinionated: an ADMET model is not a product until its assay context, data lineage, validation design, uncertainty, applicability domain, security, and review workflow are treated as first-class engineering objects.
Caution
This repository is a research and engineering starter. The bundled demonstration model is trained on synthetic example labels and is not scientifically validated. It must not be used for clinical, regulatory, patient-level, or real drug-development decisions.
- Why this project exists
- Product thesis
- What is included
- System architecture
- Scientific principles
- Initial endpoint strategy
- Repository structure
- Quick start
- API usage
- Data contract
- Chemical standardization
- Feature engineering
- Model development
- Uncertainty and applicability domain
- Evaluation and release gates
- MLOps and reproducibility
- Product experience
- Security and enterprise deployment
- Implementation plan
- Team and operating model
- Budget scenarios
- Business model and defensibility
- Testing
- Known limitations
- Roadmap
- Contributing
A large portion of drug-discovery effort is spent discovering that promising molecules have unacceptable properties. The difficulty is not merely predicting a single endpoint. Real projects must reconcile multiple, frequently conflicting objectives:
- improve target potency without increasing lipophilicity;
- improve permeability without creating efflux or solubility problems;
- reduce clearance without increasing CYP inhibition;
- lower hERG risk while preserving exposure;
- interpret measurements produced under different protocols;
- decide whether an apparently precise prediction is actually supported by nearby chemistry.
Many modeling demonstrations stop after producing a test-set metric. A medicinal-chemistry organization needs a system that also handles:
- chemical structures, salts, tautomers and stereochemistry;
- endpoint and assay definitions;
- units, censored measurements and replicates;
- data lineage and rights;
- temporal and chemical-series validation;
- model calibration;
- applicability-domain detection;
- model versioning and monitoring;
- scientist review and override;
- confidential customer data.
This repository establishes those boundaries early so that a successful prototype can evolve into a trustworthy product rather than being rewritten after the first enterprise pilot.
The highest-probability entry point is not “predict every ADMET property.” It is:
Own one expensive, frequent medicinal-chemistry decision and become the most trustworthy place to make it.
A practical initial wedge is hERG liability decision support for lead optimization. The first product should accept a compound series, produce calibrated risk estimates with uncertainty and domain evidence, identify the nearest supporting chemistry, and help a chemist prioritize which compounds to synthesize or test next.
The platform then expands into a coherent DMPK and safety panel:
- hERG inhibition;
- CYP3A4, CYP2D6 and CYP2C9 inhibition;
- microsomal and hepatocyte stability;
- aqueous solubility;
- permeability and efflux;
- plasma protein binding;
- clearance and exposure-related endpoints;
- CNS-specific properties when the chosen market requires them.
The defensible asset is not an individual algorithm. It is the accumulated system of:
- assay-normalized measurements;
- endpoint-specific models;
- customer-private calibration;
- prospective outcomes;
- model-monitoring history;
- medicinal-chemistry decisions and rejection reasons;
- workflow integration.
- FastAPI application with OpenAPI documentation;
- health and endpoint-discovery routes;
- batch prediction contract;
- API-key guard outside local mode;
- structure validation and standardization;
- endpoint model loading;
- uncertainty intervals;
- a basic applicability-domain score;
- explicit warnings and model version in every response.
- molecule and endpoint domain objects;
- RDKit-enabled standardization with a minimal CI fallback;
- descriptor calculation;
- measurement-table validation;
- preservation points for assay and provenance metadata.
- deterministic feature pipeline;
- random-forest classification and regression baselines;
- ensemble-derived uncertainty intervals;
- feature-space domain scoring;
- classification and regression metrics;
- deterministic group split utility;
- serializable model bundles.
- a clean Next.js scientific dashboard concept;
- batch status and endpoint overview;
- risk, confidence and domain presentation;
- design language suitable for an enterprise scientific application.
- Docker Compose for API, PostgreSQL, Redis and MinIO;
- API container definition;
- environment template;
- GitHub Actions CI;
- linting, unit tests and repository checks;
- model and report artifact directories.
- architecture and bounded contexts;
- scientific data model;
- model-validation protocol;
- model-card template;
- security and tenant-isolation baseline;
- product roadmap;
- contribution expectations.
flowchart TB
subgraph Experience
WEB[Scientific web workspace]
SDK[Python / REST clients]
end
subgraph Application
API[FastAPI gateway]
AUTH[Identity and tenant policy]
JOBS[Batch job coordinator]
AUDIT[Audit event service]
end
subgraph Scientific Services
REGISTRY[Compound and assay registry]
STANDARDIZE[Structure standardization]
FEATURES[Feature service]
INFERENCE[Endpoint inference]
DOMAIN[Applicability domain]
EXPLAIN[Evidence and explanation]
end
subgraph Data and MLOps
PG[(PostgreSQL)]
OBJECTS[(Object storage)]
REDIS[(Redis / queue)]
MODELS[(Model registry)]
MONITOR[Monitoring and drift]
end
subgraph Training
INGEST[Curated data ingestion]
QC[Scientific quality control]
SPLIT[Temporal / scaffold splits]
TRAIN[Training and calibration]
VALIDATE[Validation and release gate]
end
WEB --> API
SDK --> API
API --> AUTH
API --> JOBS
API --> STANDARDIZE
STANDARDIZE --> REGISTRY
STANDARDIZE --> FEATURES
FEATURES --> INFERENCE
INFERENCE --> MODELS
INFERENCE --> DOMAIN
DOMAIN --> EXPLAIN
EXPLAIN --> API
API --> AUDIT
AUDIT --> PG
JOBS --> REDIS
REGISTRY --> PG
API --> OBJECTS
INGEST --> QC --> SPLIT --> TRAIN --> VALIDATE --> MODELS
MODELS --> MONITOR
API --> MONITOR
Chemical truth is deterministic. Structure parsing, standardization, descriptor calculation and identifier generation are implemented as versioned chemistry functions. A language model must never fabricate or silently modify chemical structures.
Assay context is explicit. “hERG” is not a sufficient label by itself. Production data should retain protocol, technology, species, conditions, laboratory, quantification limits and mapping policy.
Models are endpoint-specific artifacts. Each model bundle contains the task, features, training-domain summary and version. Production bundles should also include dataset IDs, code commit, environment digest, calibration object, acceptance thresholds and model-card location.
Prediction and explanation are separate. The estimator generates numerical outputs. An explanation service may summarize them, but cannot change values or claim evidence that is absent.
Every action is auditable. Inputs, normalized structures, model version, outputs, user, tenant, timestamp and export events should be recoverable.
See docs/architecture.md for the concise architecture specification.
Different protocols can produce systematically different values. Data should be pooled only when the mapping rule is scientifically justified and validated. Endpoint definitions are versioned product objects.
Store the submitted structure and the standardized modeling parent. Never destroy salts, stereochemistry, charge state or original identifiers simply because the model consumes a normalized representation.
Values such as <0.1, >30 and “not detected” carry information. They should not be silently converted to exact numbers. A production pipeline may use censored-regression methods or clearly documented imputation rules.
Random-split performance can be encouraging while failing on novel chemical series. Temporal, scaffold, project and external holdouts are required. The strongest proof is a locked prospective prediction evaluated after experimental results arrive.
A point estimate without uncertainty creates false precision. Prediction intervals, ensemble disagreement, calibration and chemical-domain evidence should be visible in the product.
The system should be allowed to say: “insufficient supporting chemistry; test this compound experimentally.” Low-confidence abstention is often more valuable than forced ranking.
The product supports prioritization. Scientists approve the endpoint mapping, interpret trade-offs and decide what to synthesize or test. Override reasons become valuable decision data.
Descriptor models, fingerprints, nearest neighbors, matched molecular pairs and tree ensembles create strong scientific baselines. More complex graph or foundation models should earn their operational complexity through reproducible, prospective improvement.
Why begin here:
- the liability is widely recognized and decision-relevant;
- medicinal chemists often need to address it during optimization;
- both classification and continuous formulations are possible;
- chemical-series context and local analogues matter;
- uncertainty and assay heterogeneity are highly visible, making trust features valuable.
The product should not reduce hERG to a universal binary cutoff. A production implementation should define one or more assay-specific endpoints, for example:
- probability that measured IC50 falls below a configured threshold;
- regression on log-transformed IC50 for a specified protocol;
- categorical risk with an explicit gray zone;
- project-calibrated ranking using both model output and nearest-neighbor evidence.
| Endpoint | Typical task | Key difficulties | Product value |
|---|---|---|---|
| CYP inhibition | Classification or regression | Multiple isoforms, protocol variation, class imbalance | Drug–drug interaction risk triage |
| Microsomal stability | Regression or ordinal | Species and matrix differences, censoring | Clearance optimization |
| Solubility | Regression | Method, pH, kinetic vs thermodynamic values | Formulation and exposure risk |
| Permeability | Regression or classification | Cell line and protocol dependence | Oral/CNS exposure triage |
| Efflux ratio | Regression | Transporter and cell-system variation | CNS and oral design support |
| Plasma protein binding | Regression | High-bound region is difficult; species effects | Exposure interpretation |
| Hepatotoxicity panels | Multi-task classification | Endpoint ambiguity, class imbalance, causal uncertainty | Early safety screening |
The endpoint definitions in configs/endpoints.yaml are intentionally small and readable. They should eventually be stored in a governed endpoint registry.
admet-platform/
├── apps/
│ ├── api/ # FastAPI prediction service
│ │ ├── src/admet_api/
│ │ │ ├── main.py # routes and service metadata
│ │ │ ├── schemas.py # versioned API contracts
│ │ │ ├── service.py # inference orchestration
│ │ │ └── settings.py # environment configuration
│ │ └── tests/
│ └── web/ # Next.js scientific workspace concept
├── packages/
│ ├── admet_core/
│ │ ├── src/admet_core/
│ │ │ ├── chemistry.py # standardization and descriptors
│ │ │ ├── domain.py # scientific domain objects
│ │ │ └── validation.py # input data validation
│ │ └── tests/
│ └── admet_ml/
│ ├── src/admet_ml/
│ │ ├── features.py # feature construction
│ │ ├── metrics.py # endpoint metrics
│ │ ├── modeling.py # training, inference, domain score
│ │ └── splits.py # group-aware splitting
│ └── tests/
├── configs/
│ ├── data_contract.yaml # input measurement contract
│ └── endpoints.yaml # initial endpoint definitions
├── data/examples/ # synthetic demonstration data only
├── docs/
│ ├── architecture.md
│ ├── data-model.md
│ ├── model-card-template.md
│ ├── model-validation.md
│ ├── product-roadmap.md
│ └── security.md
├── infra/docker/api.Dockerfile
├── scripts/
│ ├── check_repo.py
│ ├── demo_prediction.py
│ └── train_baselines.py
├── artifacts/ # ignored model and report outputs
├── .github/workflows/ci.yml
├── docker-compose.yml
├── Makefile
├── pyproject.toml
└── README.md
- Python 3.11 or newer;
- Git;
- Docker and Docker Compose for the full local stack;
- Node.js 20 or newer only when running the web workspace;
- RDKit strongly recommended for chemistry-grade parsing and descriptors.
python -m venv .venv
source .venv/bin/activateWindows PowerShell:
py -m venv .venv
.venv\Scripts\Activate.ps1Minimal development environment:
python -m pip install --upgrade pip
python -m pip install -e '.[dev]'With RDKit support:
python -m pip install -e '.[dev,chem]'With experiment tracking and optimization tools:
python -m pip install -e '.[dev,chem,mlops]'make testmake apiOpen:
- API documentation:
http://localhost:8000/docs - health check:
http://localhost:8000/health - endpoint registry:
http://localhost:8000/v1/endpoints
make demoThe output includes:
- point estimate;
- lower and upper ensemble quantiles;
- confidence category;
- applicability-domain score;
- model version;
- warnings.
Warning
When no trained artifact exists, the API loads a deterministic demonstration classifier trained on synthetic labels. This behavior is useful for contract testing only. A production deployment should fail closed when an approved endpoint model is unavailable.
make trainThis reads data/examples/admet_training_demo.csv and writes endpoint model bundles to artifacts/models/.
cp .env.example .env
docker compose up --buildServices:
| Service | Address | Purpose |
|---|---|---|
| API | localhost:8000 |
prediction and metadata API |
| PostgreSQL | localhost:5432 |
governed metadata and audit records |
| Redis | localhost:6379 |
queues and caching |
| MinIO | localhost:9000 |
datasets, model artifacts and reports |
| MinIO console | localhost:9001 |
local object-store administration |
cd apps/web
npm install
npm run devOpen http://localhost:3000.
curl http://localhost:8000/healthcurl http://localhost:8000/v1/endpointscurl -X POST http://localhost:8000/v1/predictions \
-H 'Content-Type: application/json' \
-d '{
"endpoint": "herg",
"molecules": [
{
"compound_id": "CMP-001",
"smiles": "CCOc1ccc2nc(S(N)(=O)=O)sc2c1"
},
{
"compound_id": "CMP-002",
"smiles": "CN1CCC(CC1)Oc2ccc(C#N)cc2"
}
]
}'Example response shape:
{
"request_id": "0faad95c-d285-4e23-a1b9-2f1d11ec48ce",
"predictions": [
{
"compound_id": "CMP-001",
"endpoint": "herg",
"value": 0.27,
"lower": 0.05,
"upper": 0.61,
"confidence": "medium",
"in_domain": true,
"domain_score": 0.58,
"model_version": "baseline-0.1.0",
"warnings": []
}
]
}Before external release, add:
- tenant and project identifiers;
- request idempotency key;
- standardization-policy version;
- model selection policy;
- endpoint version;
- synchronous and asynchronous batch routes;
- source-structure echo and standardized structure;
- structured warning codes;
- nearest-neighbor evidence;
- prediction timestamp and artifact digest;
- audit-event identifier;
- explicit units for regression outputs;
- explainability object;
- error object with row-level validation failures.
A useful production route family could be:
POST /v1/projects/{project_id}/datasets
POST /v1/projects/{project_id}/prediction-jobs
GET /v1/prediction-jobs/{job_id}
GET /v1/prediction-jobs/{job_id}/results
GET /v1/models/{endpoint}/{version}/card
POST /v1/predictions/{prediction_id}/review
The minimal normalized training input is:
| Field | Required | Meaning |
|---|---|---|
compound_id |
yes | stable source identifier |
smiles |
yes | source molecular representation |
endpoint |
yes | governed endpoint identifier |
value |
yes | measured value before endpoint transform |
unit |
yes | source unit |
relation |
recommended | =, <, <=, >, >= |
assay_id |
recommended | immutable assay/protocol identifier |
assay_type |
recommended | technology or format |
species |
endpoint-dependent | biological species |
matrix |
endpoint-dependent | microsomes, hepatocytes, plasma, cells, etc. |
pH |
endpoint-dependent | relevant experimental condition |
temperature_c |
endpoint-dependent | experimental temperature |
source |
yes in production | source and licensing record |
measured_at |
recommended | enables temporal validation |
The declarative starter contract is in configs/data_contract.yaml.
- Land raw data immutably. Save the original file, checksum, source, license, uploader and timestamp.
- Parse without scientific assumptions. Preserve all values, units and relation operators.
- Standardize structures. Generate the modeling parent while retaining source structures.
- Map units. Convert only through governed, endpoint-specific rules.
- Map assays to endpoint versions. Do not pool merely because labels look similar.
- Detect duplicates and conflicts. Preserve provenance for every replicate.
- Apply quality flags. Invalid structures, ambiguous stereochemistry, impossible units, outliers and protocol gaps.
- Create an immutable dataset snapshot. Assign a dataset ID and content digest.
- Generate a quality report. Document exclusions and transformation counts.
- Approve for training. Scientific review is separate from pipeline success.
Possible duplicate strategies include:
- retain every replicate and model replicate structure;
- aggregate with median or geometric mean after protocol matching;
- create separate labels by laboratory or assay version;
- select the most recent or highest-quality measurement;
- exclude conflicting records above a prespecified variance threshold.
The choice must be endpoint-specific and documented in the model card.
Before using any dataset, record:
- origin;
- license or contract;
- permitted purposes;
- redistribution restrictions;
- whether model training is allowed;
- whether derived model artifacts may be commercialized;
- retention requirements;
- customer-private status.
Public accessibility does not remove the need for a rights review.
The core function is standardize_smiles in packages/admet_core/src/admet_core/chemistry.py.
When RDKit is installed, the starter pipeline performs:
- SMILES parsing;
- RDKit cleanup;
- fragment-parent selection;
- charge neutralization where supported;
- canonical isomeric SMILES generation;
- SHA-256 structure digest generation.
A production policy should explicitly address:
- salts and solvates;
- mixtures and disconnected fragments;
- organometallic compounds;
- isotope labels;
- undefined stereocenters;
- atropisomerism;
- tautomer canonicalization;
- protonation state;
- valence and aromaticity errors;
- covalent warheads;
- duplicate detection across representations.
Changing structure standardization can change fingerprints, descriptors, duplicates and model labels. Therefore the output should include a standardization policy ID, for example:
small-molecule-parent/2026-01
Datasets and models should reference the exact policy version. Reprocessing an old dataset should produce a new snapshot rather than mutating historical records.
The repository includes a very small fallback parser so unit tests can run without RDKit. It is not chemistry-grade and returns a warning:
rdkit_not_installed_fallback_used
Production startup should refuse to serve predictions when the approved chemistry runtime is absent.
The starter uses a compact descriptor vector:
- molecular weight;
- calculated logP;
- topological polar surface area;
- hydrogen-bond donors;
- hydrogen-bond acceptors;
- rotatable bonds;
- heavy-atom count.
This is intentionally a baseline. A mature feature registry may include:
- Morgan/ECFP fingerprints;
- MACCS or other structural keys;
- topological descriptors;
- fragment counts;
- charge and ionization features;
- aromatic ring and heterocycle features;
- structural alerts;
- matched molecular-pair context.
- conformer ensembles;
- shape and pharmacophore descriptors;
- electrostatic features;
- solvent-accessible surface estimates;
- quantum-derived descriptors when justified.
- graph neural-network embeddings;
- SMILES or molecular-language embeddings;
- pre-trained chemical foundation-model embeddings;
- multi-task endpoint representations.
- assay protocol;
- species;
- matrix;
- laboratory/vendor;
- project/series identity for hierarchical models;
- measurement date;
- batch quality information.
Every feature set should have:
- a version;
- deterministic implementation;
- test vectors;
- missing-value policy;
- computation environment;
- performance and cost profile;
- explanation of chemical invariances.
Do not jump directly to the most sophisticated architecture. Establish a ladder:
- endpoint prevalence or mean baseline;
- nearest-neighbor baseline;
- descriptor linear/logistic model;
- fingerprint similarity or kernel method;
- random forest or gradient boosting;
- graph neural network;
- multi-task model;
- project-specific fine-tuning or calibration;
- ensembles.
A model is valuable only if it beats simpler baselines under the split that represents actual deployment.
For liability classification, the starter uses a class-balanced random forest. Production work should evaluate:
- class-weighted gradient boosting;
- calibrated tree ensembles;
- fingerprint-based models;
- graph models;
- conformal classification;
- endpoint-specific gray zones;
- cost-sensitive thresholds.
The probability must be calibrated. A model that labels compounds correctly but produces unreliable probabilities cannot support rational risk thresholds.
For continuous endpoints, the starter provides a random-forest regression path. Production models should evaluate:
- log transforms aligned with assay interpretation;
- censored regression;
- heteroscedastic uncertainty;
- robust losses;
- rank-aware objectives;
- hierarchical assay effects;
- quantile regression;
- conformal prediction intervals.
Multi-task models can transfer information among related endpoints, but they create governance and debugging complexity. Use them after single-endpoint baselines are stable. Validate whether transfer helps each endpoint and chemical domain; aggregate improvement can hide endpoint regressions.
Customer projects frequently occupy a narrow chemical series. A strong platform should support:
- recalibration using local measurements;
- local nearest-neighbor models;
- hierarchical global-plus-project models;
- uncertainty reduction as project data accumulate;
- private model artifacts;
- explicit fallbacks when local data are too sparse.
The shared global model offers an initial estimate; the private project layer creates increasing customer value and defensibility.
The API reports ensemble quantiles and a simple feature-space domain score. This is a scaffold, not a final scientific method.
Use several independent signals:
- ensemble variance;
- calibrated prediction intervals;
- conformal methods;
- nearest-neighbor similarity;
- feature-space distance;
- disagreement among model families;
- training-data density;
- assay noise floor;
- stereochemistry and structure-quality flags.
A useful user-facing result is:
| Domain status | Meaning | Recommended action |
|---|---|---|
| In domain | supported by nearby, consistent training chemistry | use as decision support |
| Edge of domain | partial support or moderate model disagreement | inspect analogues and consider testing |
| Out of domain | little relevant training support | abstain or treat as hypothesis only |
The threshold should be endpoint-specific and validated against empirical error. Do not choose it solely because a score “looks reasonable.”
For every prediction, the product should show several relevant training analogues with:
- structure;
- similarity;
- assay value;
- assay/protocol match;
- source;
- model inclusion status;
- important structural differences.
This allows medicinal chemists to combine model output with recognizable SAR evidence.
Avoid presenting “high,” “medium” and “low” without a definition. A model card should define each label using measurable criteria such as interval width, calibration bucket, minimum similarity and ensemble disagreement.
Detailed requirements are in docs/model-validation.md.
Use multiple evaluation views:
- random split: debugging and continuity only;
- scaffold split: tests structural generalization;
- chemical-series/project split: tests transfer to unseen programs;
- temporal split: simulates future prediction;
- source/laboratory split: tests protocol and organization transfer;
- external set: tests independent reproducibility;
- prospective set: tests actual future utility.
At minimum:
- ROC-AUC;
- PR-AUC;
- balanced accuracy;
- sensitivity and specificity;
- positive and negative predictive value;
- Matthews correlation coefficient;
- Brier score;
- expected calibration error;
- calibration plot;
- confusion matrices at decision thresholds.
Evaluate by:
- scaffold;
- source;
- value range;
- molecular-weight and lipophilicity bands;
- ionization class;
- nearest-neighbor similarity;
- domain status.
At minimum:
- MAE;
- RMSE;
- R²;
- Spearman rank correlation;
- median absolute error;
- error by assay range;
- interval coverage and width;
- residual plots;
- error by chemical series and domain status.
- Freeze model, code, data and thresholds.
- Timestamp and sign predictions before experiments.
- Record all proposed compounds, not just successful selections.
- Receive assay results without retroactive endpoint remapping.
- Evaluate predeclared metrics.
- Investigate failures and domain behavior.
- Publish an internal prospective report.
- Decide whether to promote, revise or reject the model.
A model can be promoted only after:
- data lineage is complete;
- rights review is complete;
- scientific mapping is approved;
- required split metrics pass;
- calibration is acceptable;
- uncertainty intervals have empirical support;
- domain rules are evaluated;
- failure modes are documented;
- security scan passes;
- model card is approved;
- rollback artifact exists.
The starter ModelBundle contains:
- endpoint;
- task;
- fitted estimator;
- feature columns;
- version;
- training-domain center and scale.
A production artifact should additionally contain:
- endpoint version;
- dataset snapshot IDs;
- model code commit;
- standardization version;
- feature version;
- training configuration;
- seed;
- dependency lock digest;
- calibration object;
- decision thresholds;
- metric report location;
- model-card digest;
- artifact checksum;
- reviewer approvals.
The optional mlops dependency group includes MLflow and Optuna. A practical experiment record should log:
- full configuration;
- dataset IDs, not local filenames;
- split assignments;
- metrics and confidence intervals;
- plots;
- environment;
- trained artifact;
- feature importance or explanation summaries;
- subgroup metrics;
- failure examples.
Suggested lifecycle:
candidate -> scientifically reviewed -> staging -> prospective validation -> production -> retired
Production must reference immutable artifacts. Never overwrite a model file in place.
Monitor:
- request volume and latency;
- invalid structure rate;
- domain-score distribution;
- feature drift;
- prediction distribution;
- calibration after labels arrive;
- interval coverage;
- error by chemical series;
- model and data-version adoption;
- user overrides and rejection reasons;
- assay turnaround and label delay.
Drift alone should not automatically retrain a model. It should trigger review, because drift may represent a valuable new project domain requiring dedicated data and validation.
The web concept lives in apps/web. The intended product is a scientific workspace, not a chatbot-first interface.
- project objectives and endpoint thresholds;
- current compound-series distribution;
- outstanding predictions and experiments;
- model coverage and warnings;
- recent decisions.
Columns should include:
- compound ID and structure;
- source structure status;
- endpoint value/probability;
- interval;
- confidence;
- domain status;
- nearest-neighbor similarity;
- important alerts;
- model version;
- scientist review state.
- 2D structure;
- all endpoint predictions;
- confidence and domain evidence;
- nearest measured analogues;
- property radar or parallel coordinates;
- structural alerts;
- project comparison;
- model history;
- reviewer comments.
- intended use;
- assay definition;
- training-data coverage;
- validation results;
- limitations;
- current monitoring status;
- release owner.
- upload results;
- map assays;
- confirm structures and batches;
- compare predicted vs measured;
- record design-hypothesis outcome;
- approve data for private retraining.
Prefer:
“Predicted hERG risk is elevated, with moderate confidence. The molecule lies near the edge of the supported domain. Two close analogues show lower risk after removal of the basic side chain. Experimental confirmation is recommended.”
Avoid:
“This molecule is safe.”
The software predicts assay outcomes; it does not establish whole-organism or clinical safety.
See docs/security.md.
- tenant-scoped authorization in every service;
- encryption at rest and in transit;
- private object prefixes and database row-level policies;
- secrets management;
- immutable audit events;
- malware and content scanning on upload;
- sandboxed parsing workers;
- rate limits and quotas;
- artifact signature or checksum validation;
- backup and disaster-recovery testing;
- incident response;
- no-training flags for customer data;
- retention and deletion policies;
- administrative access review;
- dependency and container scanning.
Best economics and iteration speed. Requires strong logical isolation and customer acceptance.
Separate database, object storage and workers for each customer. Higher cost, easier boundary explanation.
Application and model runtime deployed into the customer’s cloud account. Useful for highly sensitive programs.
Highest operational burden. Requires artifact distribution, upgrade procedures, offline license and support strategy.
The default should be:
Customer data are not used to train shared models unless the customer has explicitly agreed to that use.
Private project adaptation can still occur inside the tenant boundary. The product should visibly distinguish:
- global public/licensed model;
- organization-private model;
- project-private calibration;
- experimental or unapproved model.
Goal: define one narrow endpoint product with real demand.
Deliverables:
- 25–40 interviews with medicinal chemists, DMPK scientists and project leaders;
- five observed design or triage workflows;
- endpoint decision map;
- current tools and failure modes;
- willingness-to-pay evidence;
- first pilot-design-partner criteria;
- hERG assay taxonomy and data-contract draft;
- success metrics.
Exit criteria:
- one decision is selected;
- at least three teams will provide historical evaluation data or a paid pilot;
- assay context can be defined tightly enough for modeling.
Deliverables:
- source and rights registry;
- immutable raw-data landing;
- structure-standardization policy;
- assay and unit mapping;
- duplicate/replicate policy;
- quality dashboard;
- versioned dataset snapshots;
- 10–20 historical project sets where possible.
Exit criteria:
- every modeled row is traceable to source;
- exclusions are explainable;
- endpoint class/value distribution is understood;
- scientific reviewer approves mapping rules.
Deliverables:
- prevalence/mean baseline;
- nearest-neighbor baseline;
- descriptor and fingerprint models;
- calibrated tree ensemble;
- scaffold and temporal benchmarks;
- subgroup error report;
- domain prototype;
- draft model card.
Exit criteria:
- model beats relevant simple baselines;
- probabilities are sufficiently calibrated for pilot thresholds;
- unsupported chemistry can be detected better than chance;
- limitations are acceptable for a controlled historical test.
Deliverables:
- secure project workspace;
- CSV/SDF upload;
- structure and assay validation report;
- batch predictions;
- analogue evidence;
- reviewer workflow;
- audit trail;
- exportable report;
- pilot administration tools.
Exit criteria:
- scientists can complete the workflow without developer intervention;
- prediction provenance is visible;
- security review passes for pilot data;
- historical evaluation demonstrates useful time savings or prioritization value.
Deliverables:
- locked predictions before synthesis/testing;
- experimental-result ingestion;
- prospective analysis plan;
- recommendation acceptance and rejection data;
- calibration and error analysis;
- customer outcome report.
Exit criteria:
- at least one customer synthesizes or tests model-supported compounds;
- prospective value is measurable;
- failures produce actionable model or workflow improvements;
- customer agrees to continue or expand.
Deliverables:
- two to five additional endpoints;
- multi-endpoint table;
- project-defined constraints;
- Pareto-front ranking;
- customer-private recalibration;
- endpoint-specific model monitoring;
- expanded integrations.
Deliverables:
- matched molecular-pair explanations;
- liability-reducing transforms;
- constrained analogue suggestions;
- synthetic-feasibility integration;
- active-learning experiment selection;
- closed-loop project learning.
Owns:
- customer discovery;
- endpoint choice;
- scientific requirements;
- output review;
- pilot sales;
- design-partner relationships;
- business narrative.
Owns:
- structure standardization;
- fingerprints and descriptors;
- chemical database design;
- similarity and substructure search;
- data ingestion;
- scientific performance.
Owns:
- endpoint formulation;
- model baselines;
- uncertainty and calibration;
- split design;
- validation;
- model cards;
- prospective study analysis.
Owns:
- APIs;
- project workspace;
- authentication;
- jobs and storage;
- deployment;
- observability;
- integrations.
Owns:
- assay taxonomy;
- data-combination policy;
- interpretation;
- experimental-validation design;
- customer scientific credibility.
A sensible sequence:
- medicinal-chemistry/product founder;
- senior cheminformatics engineer;
- ML scientist;
- product-minded full-stack engineer;
- DMPK scientist;
- data engineer;
- enterprise/backend engineer;
- product designer;
- customer scientific success lead;
- security/infra engineer.
Avoid building a large pure-research team before securing real project data and workflow adoption.
For a founder with ADHD, design the company so that critical execution does not depend on memory or sustained administrative attention:
- one weekly company scorecard;
- three quarterly priorities;
- daily written top-three tasks;
- operations-oriented counterpart;
- explicit owner and due date for every decision;
- short design-partner feedback cycles;
- automated notes and follow-ups;
- protected scientific deep-work blocks;
- no more than one new endpoint entering development at a time;
- decision log for ideas that are intentionally deferred.
Hyperfocus belongs on product insight, scientific synthesis, recruiting and major negotiations. Repetitive release, security and project-management processes should be systematized and delegated.
These are planning ranges, not quotes.
Suitable for:
- two or three founders;
- low founder salaries;
- one endpoint;
- public or already licensed data;
- historical benchmark;
- basic API and local workspace;
- no enterprise compliance program.
Primary spending:
- personnel;
- data cleaning;
- compute;
- legal/data-rights review;
- customer discovery.
Suitable for:
- five to eight people;
- 12–18 months;
- secure design-partner pilots;
- production ingestion and lineage;
- endpoint validation;
- scientific UI;
- audit and tenant controls;
- prospective study.
Suitable for:
- multiple endpoint teams;
- VPC or dedicated deployments;
- formal security program;
- data integrations;
- model monitoring;
- customer scientific success;
- private adaptation;
- larger prospective evidence base.
Adding synthesis and assays creates a different business:
- compound procurement;
- CRO contracts or facilities;
- sample logistics;
- analytical chemistry;
- assay development;
- laboratory informatics;
- quality systems;
- longer working-capital cycles.
Begin with software and carefully selected external experiments unless laboratory ownership is central to a differentiated data strategy.
Access to platform, endpoint panel, seats and a prediction volume allowance.
Useful for smaller biotechnology companies with one or two active programs.
Additional fee for customer-specific training, validation and deployment.
Suitable for integration into internal design systems, registration tools and ELNs.
Data mapping, historical benchmark and model validation can accelerate early adoption, but should feed a scalable product rather than becoming unlimited consulting.
Price against measurable value:
- fewer unnecessary compounds synthesized;
- faster design cycles;
- improved probability of selecting acceptable compounds;
- reduced time preparing DMPK reviews;
- increased experimental learning per cycle;
- earlier termination of poor series;
- improved traceability of model-assisted decisions.
Weakest to strongest:
- generic model served by API;
- better endpoint benchmarks;
- unique assay-normalized datasets;
- workflow integration;
- customer-private calibration;
- prospective experimental feedback;
- medicinal-chemistry decision and rejection data;
- cross-endpoint design optimization;
- closed design–make–test learning loop.
The long-term objective is to become the system where project data return after every experiment and the models become increasingly specific to the customer’s chemistry and assays.
Run all checks:
make checkRun tests directly:
pytestRun linting:
ruff check .- API health contract;
- prediction response contract;
- invalid-SMILES handling;
- deterministic structure hashing;
- descriptor-schema check;
- model prediction interval shape;
- applicability-domain score bounds.
Maintain structures that cover:
- salts;
- zwitterions;
- charged compounds;
- stereocenters;
- isotopes;
- aromaticity edge cases;
- mixtures;
- invalid valence;
- tautomer-sensitive cases.
Expected standardized outputs should be versioned and reviewed by chemists.
- unit conversions;
- censored values;
- duplicate aggregation;
- assay mapping;
- provenance preservation;
- immutable snapshot hashing;
- schema evolution.
- split leakage detection;
- deterministic training within tolerance;
- metric regression tests;
- calibration tests;
- interval coverage;
- domain-error relationship;
- artifact load and signature validation.
- tenant isolation;
- permissions;
- upload abuse;
- large batches;
- retries and idempotency;
- audit completeness;
- export reproducibility;
- rollback.
- The bundled data are synthetic and have no scientific meaning.
- The demonstration model is not approved for any real endpoint decision.
- The fallback SMILES validator is intentionally primitive.
- The descriptor set is too small for competitive endpoint modeling.
- The domain score is a feature-distance heuristic, not an empirically calibrated applicability-domain method.
- Ensemble quantiles from random-forest trees are not guaranteed calibrated intervals.
- The API currently performs synchronous in-process inference.
- Database models, migrations, audit persistence and tenant authorization are architectural placeholders rather than complete implementations.
- The web application is a visual product concept, not a connected production client.
- Regression endpoint training needs endpoint transforms and unit normalization.
- Censored data are documented but not yet modeled.
- Nearest-neighbor evidence and chemical drawings are not implemented in the API.
- No model-signing or artifact-integrity layer is implemented.
- No formal regulated-software claim is made.
These limitations are explicit so the repository is not mistaken for a validated commercial product.
The detailed roadmap is in docs/product-roadmap.md.
- real RDKit-only production chemistry runtime;
- Pydantic endpoint registry;
- assay-aware ingestion;
- PostgreSQL schema and migrations;
- fingerprint and nearest-neighbor service;
- calibrated hERG baseline;
- scaffold and temporal benchmark reports;
- batch file jobs;
- model cards served through API;
- connected web workspace.
- additional CYP and DMPK endpoints;
- MLflow registry integration;
- conformal intervals;
- project-private recalibration;
- SSO and tenant policy;
- audit persistence;
- drift monitoring;
- historical customer benchmark workflow;
- prospective-study module.
- multi-endpoint optimization;
- matched molecular-pair recommendations;
- experiment selection;
- ELN and compound-registration integrations;
- customer VPC deployment;
- private model federation where appropriate;
- closed-loop design–make–test orchestration.
Read CONTRIBUTING.md.
Scientific changes require more than passing unit tests. A pull request that changes any of the following must document the expected impact and migration path:
- structure standardization;
- endpoint definition;
- unit mapping;
- duplicate policy;
- dataset inclusion/exclusion;
- features;
- split strategy;
- calibration;
- model threshold;
- applicability-domain logic;
- prediction contract.
Never commit proprietary compounds, customer data, assay results, credentials or model artifacts derived from restricted data.
A credible ADMET company is built by combining medicinal chemistry, DMPK science, data engineering, machine learning, software product design, security, and prospective experimental evidence.
The winning system will not be the one that produces the most predictions. It will be the one that scientists trust to distinguish:
- supported prediction from speculation;
- assay signal from data artifact;
- local evidence from global extrapolation;
- useful optimization direction from confident nonsense.
Build one endpoint deeply. Preserve every source and decision. Validate prospectively. Then expand the panel and learning loop.