Autonomous Scientific Agent for Target Inhibitor Discovery
Maintainer: @drJoyKarmakar | Version: 0.1.0 | License: MIT
Initial research preview. The GitHub publishing workflow marks this version as a pre-release. No trained binding-affinity model or experimental validation is included. This package was prepared locally; see the repository's actual Actions and Releases pages for remote validation and publication status.
An MIT-licensed, modular research workflow for target-associated compound retrieval, RDKit screening, supervised property-model adaptation, template retrosynthesis, and evidence-linked Markdown reports.
This is research software, not a validated drug-discovery product. No pretrained EGFR affinity model, complete ADMET model, experimental synthesis validation, or clinical recommendation is bundled. Missing predictions remain missing.
Python 3.11 or newer. From a published checkout:
git clone https://github.com/drJoyKarmakar/BioAgent-X.git
cd BioAgent-XFrom the prepared archive, extract it and enter its BioAgent-X directory instead.
python -m venv .venv
source .venv/bin/activate
python -m pip install -e '.[dev]'
# Optional supervised model training and inference:
pip install -e '.[ml,dev]'
python -m pytest -q
python -m bioagent_x --versionOn Windows, activate .venv\Scripts\activate instead. Core functionality does not
import Torch, Transformers or PEFT unless a model is trained or loaded. Dependency
ranges are compatibility bounds, not a fully tested lockfile. Freeze your verified
local environment before a scientific production run.
Use an API/model that supports strict Responses API function calling. This is
not a Chat Completions-only interface. Environment variables are read directly;
.env.example is documentation and is not loaded automatically.
export BIOAGENT_LLM_API_KEY='your-key'
export BIOAGENT_LLM_MODEL='your-responses-function-calling-model-id'
# Default base URL: https://api.openai.com/v1
python main.py run \
'Find potential small-molecule inhibitors for human EGFR, screen molecular properties, and propose synthesis routes' \
--target-id CHEMBL203 --output runs/egfrPlanner, selector and reviewer are separate role invocations using the same configured model. They never fabricate chemistry-tool results. Python owns verified structures, model paths, target identity, dependency gating and budgets. Prompts and compact observations are sent to the configured LLM service; review your data-sharing policy before using confidential research data.
This still requires network access to ChEMBL. It is NOT a fake-data demo.
python main.py run 'Screen known human EGFR-associated molecules' \
--no-llm --target EGFR --target-id CHEMBL203 --endpoint pIC50 \
--limit 30 --output runs/egfr-no-llmUse a new or empty output directory. Exit code 0 means available stages finished; exit code 2 indicates a failure or an incomplete workflow. Missing optional predictions are reported explicitly but do not fail a descriptor-only run.
Supply a genuinely curated CSV with the following columns:
smiles,label,endpoint,target_chembl_id,assay_context
For binding models, endpoint must be pKi or pIC50; label is the negative
base-10 logarithm of the molar Ki or IC50. For nM concentrations,
pActivity = 9 - log10(value_nm). Do not mix endpoint labels, targets, wild-type
and mutant proteins, incompatible assay conditions, or censored observations.
For LogP models use endpoint logP and an empty target field unless intentionally
building a target-associated property dataset. assay_context must exactly match
one declared, user-curated context for all rows; an identical string does not
independently verify experimental comparability.
No fabricated training dataset is included. At least 20 unique molecules and three scaffolds are required, with at least three molecules in every split. These small minimums are software guardrails, NOT statistical adequacy criteria.
python main.py train \
--data data/egfr_curated.csv --output models/egfr-pic50 \
--endpoint pIC50 --target-id CHEMBL203 \
--assay-context 'curated-wt-egfr-biochemical-v1' \
--epochs 10 --batch-size 16 --device cpu
python main.py run 'Screen known human EGFR-associated molecules' \
--no-llm --target-id CHEMBL203 --endpoint pIC50 \
--checkpoint models/egfr-pic50 --output runs/egfr-modelTo train inside the pipeline, supply --training-data, --target-id, and
--assay-context to run. The model is never fine-tuned from LLM-generated
pseudo-labels. A checkpoint can predict only its declared endpoint and target.
Exact training-set molecules receive an abstention rather than an in-sample
prediction presented as evidence. Fingerprint-domain warnings are heuristic and
not confidence intervals. The report compares held-out RMSE against a mean-label
baseline and does not infer scientific validity from a successful training run.
The trainer supports RoBERTa/ChemBERTa architectures, LoRA query/value projections,
and a fully trained/saved classifier head. Its default backbone is
DeepChem/ChemBERTa-77M-MLM, a masked language model, NOT a pretrained affinity
predictor. Train/validation/test partitions group Murcko scaffolds. Validation
chooses the epoch; the test partition is evaluated once. Data hashes, local
backbone hashes or a remote revision, split membership and dependency versions
are persisted. GPU numerical reproducibility is not guaranteed merely by a seed.
Treat model directories as trusted artifacts: hashes detect accidental changes,
not maliciously forged manifests. Audit upstream model/data licenses separately.
The requested pubchem_search(target_name) API name is retained, but the actual
lookup is ChEMBL target -> assay/activity -> molecule. A protein name must not be
looked up as though it were a PubChem compound name. Exact target synonyms or an
explicit ChEMBL ID are required. Only single-protein, organism-matched targets
and direct, confidence-9 binding assays are accepted. Annotated variants, censored
measurements, duplicate flags and data validity comments are excluded; missing
annotations do not establish wild type. Discovery is capped at 2,000 activity
rows ordered by ID, not claimed exhaustive or potency-ranked.
The chemistry layer performs actual MolWt, MolLogP, Lipinski donor/acceptor, TPSA, rotatable-bond, QED and PAINS calculations. It attempts normalization and full sanitization, then skips invalid structures. Salts/mixtures are rejected rather than silently converted into a different compound. The filter is <=1 Rule-of-Five violation and no PAINS alert. Reports order candidates by QED, not by supposed efficacy. A candidate remains merely target-associated until the specific activity and mechanism evidence is reviewed.
Retrosynthesis is an intentionally small, bounded reaction-SMARTS engine. It
attempts up to three connected disconnections and verifies each forward graph
round trip. It does not invent extra steps when coverage is insufficient. It
returns partial or unsupported, and provides no fabricated conditions,
selectivities, yields, costs or stock availability. Expert synthetic review is
required. This is not an AiZynthFinder-quality route search.
ADMET output is limited to physicochemical and structural-alert screening. No absorption, clearance, CYP, hERG, toxicity, metabolic-stability, exposure or selectivity model is bundled. No docking or wet-lab experiment is performed.
Each successful report export contains report.md, results.json, a
run_manifest.json of hashes/versions, and locally rendered PNG/SVG files in
structures/. Successful retrieval also saves api_snapshot.json with fetched
responses. Full activity records retain ChEMBL activity/assay/document IDs.
Mechanism text is taken from curated molecule-specific records, not extrapolated
to all candidates. The LLM's reviewer notes are labeled non-evidentiary.
See TEST_RESULTS.md for the tests actually executed during development. Tests
use explicitly synthetic ChEMBL/LLM fixtures, never claim those are experimental
results, and do not need network access. The optional ML test constructs a tiny
local RoBERTa backbone and tokenizer to exercise real train/save/load behavior.
That test is skipped when Transformers/PEFT are absent.
Before production deployment: execute live API contract tests, run the optional ML test in your environment, verify a dependency lock, curate assay-compatible data, validate externally on appropriate chemistry/assays, calibrate uncertainty, and add domain-specific ADMET/route-validation services as needed.
- ChEMBL REST API: https://www.ebi.ac.uk/chembl/api/data/docs
- ChEMBL web services: https://chembl.gitbook.io/chembl-interface-documentation/web-services/chembl-data-web-services
- RDKit QED: https://www.rdkit.org/docs/source/rdkit.Chem.QED.html
- RDKit filter catalogs: https://www.rdkit.org/docs/source/rdkit.Chem.rdfiltercatalog.html
- RDKit reactions: https://www.rdkit.org/docs/source/rdkit.Chem.rdChemReactions.html
- PEFT LoRA: https://huggingface.co/docs/peft/en/package_reference/lora
- Native function calling: https://developers.openai.com/api/docs/guides/function-calling
- Default backbone: https://huggingface.co/DeepChem/ChemBERTa-77M-MLM
The included publisher is configured for drJoyKarmakar/BioAgent-X. It defaults
to a preview and performs no remote writes until --execute is passed:
# Install Git and GitHub CLI separately, then authenticate as drJoyKarmakar.
gh auth login --hostname github.com --git-protocol https --web
python scripts/publish_github.py
python scripts/publish_github.py --executeExecution creates a public repository, pushes main, and pushes v0.1.0.
The tag workflow runs core, packaging, and real ML smoke checks before publishing
an explicitly labeled research pre-release with wheel, source archive, and hashes.
Failed checks block publication. No PyPI upload is configured.
Use --visibility private before first creation to keep the repository private.
See release instructions, release notes,
contribution guidance, and security guidance.