Skip to content

About

Autonomous LLM-powered scientific agent for computational drug discovery, integrating RDKit, ChEMBL/PubChem, LoRA property prediction, ADMET screening, retrosynthesis, and automated scientific reporting.

Resources

Contributing

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

BioAgent-X

Autonomous Scientific Agent for Target Inhibitor Discovery

Maintainer: @drJoyKarmakar | Version: 0.1.0 | License: MIT

Initial research preview. The GitHub publishing workflow marks this version as a pre-release. No trained binding-affinity model or experimental validation is included. This package was prepared locally; see the repository's actual Actions and Releases pages for remote validation and publication status.

An MIT-licensed, modular research workflow for target-associated compound retrieval, RDKit screening, supervised property-model adaptation, template retrosynthesis, and evidence-linked Markdown reports.

This is research software, not a validated drug-discovery product. No pretrained EGFR affinity model, complete ADMET model, experimental synthesis validation, or clinical recommendation is bundled. Missing predictions remain missing.

Install

Python 3.11 or newer. From a published checkout:

git clone https://github.com/drJoyKarmakar/BioAgent-X.git
cd BioAgent-X

From the prepared archive, extract it and enter its BioAgent-X directory instead.

python -m venv .venv
source .venv/bin/activate
python -m pip install -e '.[dev]'
# Optional supervised model training and inference:
pip install -e '.[ml,dev]'
python -m pytest -q
python -m bioagent_x --version

On Windows, activate .venv\Scripts\activate instead. Core functionality does not import Torch, Transformers or PEFT unless a model is trained or loaded. Dependency ranges are compatibility bounds, not a fully tested lockfile. Freeze your verified local environment before a scientific production run.

Natural-language, LLM-controlled run

Use an API/model that supports strict Responses API function calling. This is not a Chat Completions-only interface. Environment variables are read directly; .env.example is documentation and is not loaded automatically.

export BIOAGENT_LLM_API_KEY='your-key'
export BIOAGENT_LLM_MODEL='your-responses-function-calling-model-id'
# Default base URL: https://api.openai.com/v1
python main.py run \
  'Find potential small-molecule inhibitors for human EGFR, screen molecular properties, and propose synthesis routes' \
  --target-id CHEMBL203 --output runs/egfr

Planner, selector and reviewer are separate role invocations using the same configured model. They never fabricate chemistry-tool results. Python owns verified structures, model paths, target identity, dependency gating and budgets. Prompts and compact observations are sent to the configured LLM service; review your data-sharing policy before using confidential research data.

Deterministic run without an LLM

This still requires network access to ChEMBL. It is NOT a fake-data demo.

python main.py run 'Screen known human EGFR-associated molecules' \
  --no-llm --target EGFR --target-id CHEMBL203 --endpoint pIC50 \
  --limit 30 --output runs/egfr-no-llm

Use a new or empty output directory. Exit code 0 means available stages finished; exit code 2 indicates a failure or an incomplete workflow. Missing optional predictions are reported explicitly but do not fail a descriptor-only run.

Supervised LoRA training

Supply a genuinely curated CSV with the following columns:

smiles,label,endpoint,target_chembl_id,assay_context

For binding models, endpoint must be pKi or pIC50; label is the negative base-10 logarithm of the molar Ki or IC50. For nM concentrations, pActivity = 9 - log10(value_nm). Do not mix endpoint labels, targets, wild-type and mutant proteins, incompatible assay conditions, or censored observations. For LogP models use endpoint logP and an empty target field unless intentionally building a target-associated property dataset. assay_context must exactly match one declared, user-curated context for all rows; an identical string does not independently verify experimental comparability.

No fabricated training dataset is included. At least 20 unique molecules and three scaffolds are required, with at least three molecules in every split. These small minimums are software guardrails, NOT statistical adequacy criteria.

python main.py train \
  --data data/egfr_curated.csv --output models/egfr-pic50 \
  --endpoint pIC50 --target-id CHEMBL203 \
  --assay-context 'curated-wt-egfr-biochemical-v1' \
  --epochs 10 --batch-size 16 --device cpu

python main.py run 'Screen known human EGFR-associated molecules' \
  --no-llm --target-id CHEMBL203 --endpoint pIC50 \
  --checkpoint models/egfr-pic50 --output runs/egfr-model

To train inside the pipeline, supply --training-data, --target-id, and --assay-context to run. The model is never fine-tuned from LLM-generated pseudo-labels. A checkpoint can predict only its declared endpoint and target. Exact training-set molecules receive an abstention rather than an in-sample prediction presented as evidence. Fingerprint-domain warnings are heuristic and not confidence intervals. The report compares held-out RMSE against a mean-label baseline and does not infer scientific validity from a successful training run.

The trainer supports RoBERTa/ChemBERTa architectures, LoRA query/value projections, and a fully trained/saved classifier head. Its default backbone is DeepChem/ChemBERTa-77M-MLM, a masked language model, NOT a pretrained affinity predictor. Train/validation/test partitions group Murcko scaffolds. Validation chooses the epoch; the test partition is evaluated once. Data hashes, local backbone hashes or a remote revision, split membership and dependency versions are persisted. GPU numerical reproducibility is not guaranteed merely by a seed. Treat model directories as trusted artifacts: hashes detect accidental changes, not maliciously forged manifests. Audit upstream model/data licenses separately.

Scientific policy

The requested pubchem_search(target_name) API name is retained, but the actual lookup is ChEMBL target -> assay/activity -> molecule. A protein name must not be looked up as though it were a PubChem compound name. Exact target synonyms or an explicit ChEMBL ID are required. Only single-protein, organism-matched targets and direct, confidence-9 binding assays are accepted. Annotated variants, censored measurements, duplicate flags and data validity comments are excluded; missing annotations do not establish wild type. Discovery is capped at 2,000 activity rows ordered by ID, not claimed exhaustive or potency-ranked.

The chemistry layer performs actual MolWt, MolLogP, Lipinski donor/acceptor, TPSA, rotatable-bond, QED and PAINS calculations. It attempts normalization and full sanitization, then skips invalid structures. Salts/mixtures are rejected rather than silently converted into a different compound. The filter is <=1 Rule-of-Five violation and no PAINS alert. Reports order candidates by QED, not by supposed efficacy. A candidate remains merely target-associated until the specific activity and mechanism evidence is reviewed.

Retrosynthesis is an intentionally small, bounded reaction-SMARTS engine. It attempts up to three connected disconnections and verifies each forward graph round trip. It does not invent extra steps when coverage is insufficient. It returns partial or unsupported, and provides no fabricated conditions, selectivities, yields, costs or stock availability. Expert synthetic review is required. This is not an AiZynthFinder-quality route search.

ADMET output is limited to physicochemical and structural-alert screening. No absorption, clearance, CYP, hERG, toxicity, metabolic-stability, exposure or selectivity model is bundled. No docking or wet-lab experiment is performed.

Outputs

Each successful report export contains report.md, results.json, a run_manifest.json of hashes/versions, and locally rendered PNG/SVG files in structures/. Successful retrieval also saves api_snapshot.json with fetched responses. Full activity records retain ChEMBL activity/assay/document IDs. Mechanism text is taken from curated molecule-specific records, not extrapolated to all candidates. The LLM's reviewer notes are labeled non-evidentiary.

Tests and deployment status

See TEST_RESULTS.md for the tests actually executed during development. Tests use explicitly synthetic ChEMBL/LLM fixtures, never claim those are experimental results, and do not need network access. The optional ML test constructs a tiny local RoBERTa backbone and tokenizer to exercise real train/save/load behavior. That test is skipped when Transformers/PEFT are absent.

Before production deployment: execute live API contract tests, run the optional ML test in your environment, verify a dependency lock, curate assay-compatible data, validate externally on appropriate chemistry/assays, calibrate uncertainty, and add domain-specific ADMET/route-validation services as needed.

Primary technical references

Publish this repository

The included publisher is configured for drJoyKarmakar/BioAgent-X. It defaults to a preview and performs no remote writes until --execute is passed:

# Install Git and GitHub CLI separately, then authenticate as drJoyKarmakar.
gh auth login --hostname github.com --git-protocol https --web
python scripts/publish_github.py
python scripts/publish_github.py --execute

Execution creates a public repository, pushes main, and pushes v0.1.0. The tag workflow runs core, packaging, and real ML smoke checks before publishing an explicitly labeled research pre-release with wheel, source archive, and hashes. Failed checks block publication. No PyPI upload is configured. Use --visibility private before first creation to keep the repository private. See release instructions, release notes, contribution guidance, and security guidance.

About

Autonomous LLM-powered scientific agent for computational drug discovery, integrating RDKit, ChEMBL/PubChem, LoRA property prediction, ADMET screening, retrosynthesis, and automated scientific reporting.

Resources

Contributing

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages