This repository represents an extension of BacterAI produced by Pacific Northwest National Laboratory. This repository allows for a CLI to control BacterAI and allows users to set their own lists of experimental conditions. BacterAI was first developed by the Jensen Lab at the University of Michigan. Those who want more information about BacterAI are encouraged to contact the original authors at manager@jensenlab.net, read through the original repository (https://github.com/jensenlab/BacterAI), as well as the canonical paper:
Adam C. Dama, Kevin S. Kim, Danielle M. Leyva, Annamarie P. Lunkes, Noah S. Schmid, Kenan Jijakli & Paul A. Jensen. BacterAI maps microbial metabolism without prior knowledge. Nat Microbiol 8, 1018–1025 (2023). https://doi.org/10.1038/s41564-023-01376-0
See example.md for short vignette of our workflow.
This document summarizes:
- How to clone the repo and run the CLI locally
- What the current CLI does
- Command reference and usage examples
git- Python
3.12+ - One environment manager (
condaorvenv)
git clone https://github.com/PNNL-Predictive-Phenomics/BacterAI_pnnl.git
cd BacterAI_pnnl
conda env create --name bacterai --file bacterai_env.yml
conda activate bacterai
pip install -e .
bacterai -hgit clone https://github.com/PNNL-Predictive-Phenomics/BacterAI_pnnl.git
cd BacterAI_pnnl
python3.12 -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install -r requirements.txt
python -m pip install -e .
bacterai -hbacterai -h
bacterai experiment -hIf your shell does not find bacterai, use:
python -m bacterai.main -hThe bacterai CLI is a unified interface for preparing, running, and post-processing BacterAI experiments.
At a high level it helps you:
- Convert experiment spreadsheets into
config.jsonfiles - Convert ingredient spreadsheets into
ingredients.json - Interactively scaffold experiment folders and write config/ingredient files
- Execute experiment rounds (
Round1,Round2, ...) - Process raw plate reader data (Biotek/Tecan) into
mapped_data_*.csvfor model training
- Console script:
bacterai - Backed by:
src/bacterai/main.py(bacterai.main:maininpyproject.toml)
Use help anytime:
bacterai -h
bacterai <command> -hConvert experiment setup sheets (.csv/.xlsx) into one or more config.json files.
bacterai experiment <input> [--sheet <name_or_index>] [--outfile-name config.json] [--force-format vertical|horizontal] [--verbose]- Prompts for handling existing directories
- Writes one config per experiment directory
--force-formathelps when auto-detection is ambiguous
Example:
bacterai experiment ./inputs/experiment_design.xlsx --sheet 0 --verboseInteractive end-to-end setup for experiment configs and ingredients.
bacterai setup <input> [--sheet <name_or_index>] [--outfile-name config.json] [--force-format vertical|horizontal] [--verbose]- Uses same experiment parsing as
experiment - If
experiment_pathis missing, prompts to use current working directory (then asks for experiment folder name) or provide a full experiment folder path - Creates the experiment folder path if it does not exist
- If target folder already exists and contains files, prompts to overwrite or provide a different full path
- Prompts whether ingredients are shared across experiments or individual
- Writes both config and ingredient files through setup utilities
Example:
bacterai setup ./inputs/experiment_design.xlsx --sheet PlanningConvert ingredient sheets (.csv/.xlsx) to JSON payload.
bacterai ingredients <input> [--sheet <name_or_index>] [--output <path.json>] [--id-map <map.csv>] [--verbose]- If
--outputis omitted, JSON is printed to stdout --id-mapcan map ingredient names to IDs
Example:
bacterai ingredients ./inputs/ingredients.xlsx --sheet 1 --output ./exp1/ingredients.jsonExecute a BacterAI experiment round from an experiment directory.
bacterai run <experiment_path> [--plot-only] [--verbose]- Verifies experiment path exists
- Auto-detects next round number using existing
Round*folders - Executes full pipeline unless
--plot-onlyis set
Example:
bacterai run /path/to/experiment --verboseProcess raw plate reader files and write mapped data CSV for training.
bacterai process_data {biotek|tecan} [path] --round <N> --feature <name> [--date <id>] [--signal <int>] [--output <csv>] [--verbose]- Supports experiment request layouts at either:
<experiment_path>/experiment_request/<experiment_path>/RoundN/experiment_request/
--signalis required forbiotekand ignored fortecan- Default output name (if not provided):
RoundN/mapped_data_<date_or_today>_<reader>_<feature>_data.csv
Example (Biotek):
bacterai process_data biotek /path/to/experiment --date 2026-03-12 --round 1 --signal 600 --feature delta_od --verboseExample (Tecan):
bacterai process_data tecan /path/to/experiment --round 1 --feature delta_odbacterai experiment ...orbacterai setup ...bacterai run <experiment_path>for Round 1- Collect plate reader outputs externally
bacterai process_data ... --round 1 ...bacterai run <experiment_path>for Round 2- Repeat processing and run for subsequent rounds
The config file represents the main way to control the experiment.
{
"experiment_path": "/path/to/experiment",
"ingredients_file": "ingredients.json",
"grow_threshold": 0.25,
"nickname": "PpKT2440",
"batch_size": 100,
"timeout_min": 240,
"model_type": 0,
"direction": 0,
"beyond_frontier": true,
"use_unique": false,
"n_rollouts": 10,
"n_bags": 25,
"transfer_model_folder": null,
"transfer_data_dir": null,
"redo_size": 0,
"redo_threshold": [0, 1],
"aas_only": false,
"separate_redos": false,
"simulation_types": [2],
"random_walk_increment": 10
}- Used in
main()function inrun.py - Default "growth/no-growth" threshold for model training. This is described as the grow threshold but it is more accurate to describe it as the fitness threshold because this represents the values of the fitness (or "y") variable after being normalized by each plate to a generic fitness score
0.25was the original default value;0.1 - 0.15seem good for M9 media- To set this value use data from piloting trials or look at the distribution of values from the first round of data. A bimodal distribution of values indicates clusters of growth or non-growth and can show where a threshold setting is most appropriate. This will depend on the target organism, the growth conditions, and the model of the plate reader (i.e., the default value cannot be relied upon to suit every experiment)
- What to call your experiment
- The name of the file containing the list of ingredients to consider for BacterAI experiments
- BacterAI expects a JSON file
- Default:
"ingredients.json"
- Number of experiments to run during a single BacterAI "round"
- Determined by your plate size, liquid handling capacity, and sample measurement process
- Max time (minutes) to run a "round" in BacterAI before timing out
- Default:
4 * 60(i.e., 4 hours) - Note: There is some indication that this may actually be in seconds
- Used in
main()function inrun.py - Selects the predictive model
- GPR is a form of Bayesian optimization and will work best for continuous variables and where the full kernel of a density distribution can be explored across the possible range of values of a variable
- Bagged neural nets may work best when confronted with binary or ordinal factors or a mix of categorical and continuous variables
0 = GPR (Gaussian process regression),1 = NEURAL_NET- Default:
0
- Used in
main()andmake_batch()functions inrun.py, but mainly inperform_simulations()function insim.py - Indicates whether to put all ingredients in and remove (
DOWN), leave all out and add-in (UP), or split runs to do both 0 = DOWN,1 = UP,2 = BOTH- Default:
0 - Note that it is important to produce a balanced data set where both growth and no-growth can occur. If the predictive models do not experience enough non-growth conditions, it can be difficult to predict experiments at or beyond the growth frontier and as a result, searches fail to resolve new experiments before timing out
- Used in
main()andmake_batch()functions inrun.py; mainly inperform_simulations()function insim.py - Can have length > 1 if multiple types are specified. If so, new batch simulations are divided across types:
randomperforms random take-one-out actionsgreedytakes all leave-one-out actionsrollouttakes all leave-one-out actions, then callsrollout_trajectory()function to perform random walks of available actions and ranks those with the most viable "growth pathways"rollout_propuses a rollout method with a mix of deterministic and stochastic exploration using a softmax distribution, it produces more stochastic searches if the deterministic search pattern becomes stuck by repeatedly arriving at the same set of "new" proposed experiments
0 = RANDOM,1 = GREEDY,2 = ROLLOUT,3 = ROLLOUT_PROB- Default:
2 - It is recommended, but not required, that some combination of random and systematic (deterministic) searching be applied in order to allow for sufficient exploration of the parameter space of the experiment. A good combination is using both
RANDOMandROLLOUT_PROBas seen in Damaet al.
- Used in
perform_simulations()function insim.py - When simulating new runs, determines whether to go beyond the growth frontier
- If
True, new untested conditions supporting predicted growth are included
- If
- Default:
True
- Used in
perform_simulations()function insim.py - Whether to take only unique states for the batch
- Default:
True
- Used in
perform_simulations()insim.pyand within therollout_trajectory()function (called byperform_simulations()whensimulation_types = ROLLOUT*) - Number of random-walk rollouts to assess with leave-one-out approaches
- Default:
1(must be > 0 if using any rollout type of simulation). Values of5-20may be appropriate
- Used in
train_bagged()function innet.py - Number of bagged models to run (relevant for NeuralNet model training)
- Number of bagged models should match the number of transfer learning models, if applicable
- Default:
25
- Location of the transfer learning model
- Default:
None
- Location of the transfer learning data
- Default:
None
- Pre-specifies the number of experiments to redo from the previous round for QC purposes. At present, these intentional redo experiments are not included in any training data
- Used in
process_results()function (run.py); in code, assigned toN_REDOS- Should be
0, rather thanNonewhich may force the redo the entire previous round
- Should be
- Use in conjunction with
redo_threshold - Default:
0
- Threshold range used to determine which experiments to redo
- Used in
process_results()(run.py)- An intentional QA strategy or a mechanism to refine data for borderline growth conditions
- Default:
[0, 1]
- Determines whether ingredients consist only of amino acids
- Default:
false
- Whether or not to separate redos into a dedicated batch of experiments (e.g., a separate plate)
- Default:
false
- Number of increments to break a continuous variable into for simulation runs
- Used in
make_batch(),perform_simulations(), and therollout_trajectory()function (if using a rollout simulation type)- Applied only to continuous variables requiring exploration (
N_STATES = NA) - Larger values increase exploration speed but decrease precision
- Potential enhancement: make dynamic so increments decrease as rounds increase
- Applied only to continuous variables requiring exploration (
- Default:
10