Paste a written lab protocol. Get back a deterministic, feasibility-checked simulation of a robot executing it — rendered as a 3D digital twin with a real run-time estimate.
A scientist names a run, pastes an SOP (or drops in a PDF/DOCX), confirms the bench the platform infers from it, and watches a cobot arm on a linear rail carry the sample between instruments and pipette every step, on a watchable clock. Every plan is saved, replayable, and shareable by link.
1 · Experiment → 2 · Protocol → 3 · Confirm bench → 4 · Breakdown → 5 · Run
name the run paste the SOP edit the inferred four-layer view 3D console
infer the bench machines & labware run-time · issues
Lab-automation onboarding stalls at the same point every time: a scientist has a protocol on paper, an engineer has a robot, and nobody can tell — before anything is built — whether the protocol is even runnable on that bench, how long it will take, or where it will fail. This platform answers those three questions from the protocol text alone, deterministically, before a single tip is picked up.
The engine is split into a probabilistic half and a deterministic half, with a small frozen contract between them.
- Reasoning (probabilistic) — one LLM call at temperature 0 turns
bench + protocol textinto an orderedOperationListdrawn from a closed vocabulary of 9 ops (aspirate,dispense,mix,move,incubate,wash,discard,pick_up_tip,drop_tip). The prompt injects the vocabulary and the declared bench, so the model can only emit known ops against real targets. A modal self-consistency wrapper parses k times and ships the sequence the majority agrees on; the agreement fraction is shown on every run ("Deterministic — 3/3 parses agree" vs "1/3 parses agree — review"). - Grounding (deterministic) — a validator checks every op against the library's geometry and each machine's capability envelope, modelled with PyLabRobot as the resource model. Then the motion expander turns ops into timed robot motions (pick, travel, place, pipette, hold), and the feasibility engine runs stateful checks — tip loaded before aspirating, volumes inside the machine's range, wells that exist, deck slots that fit, waste capacity, uncaptured repetitions — each finding carrying a severity and, where possible, a fix.
The Operation/Bench schemas, the op vocabulary, the parser and the expander/validator were frozen as contract-v1.0. Everything built afterwards — bench inference, the digital twin, the scheduler, the importers, run history — is layered strictly on top and never edits the frozen modules. Same protocol + same bench → identical motion sequence, every time. The backend suite (156 tests) runs fully offline with the LLM injected as a stub.
3D digital twin (frontend/src/three/, Three.js). A single Franka-style cobot arm with 2-link IK on a traversing linear rail carries one sample plate between instruments; pipetting descends to the labware; incubations glow steadily on the instrument. Instruments are procedural meshes built from spec — an open-deck 96-channel liquid handler, thermocycler, cooling block, tube racks, reservoirs, plates, tip racks, waste — with a curated-GLTF tier wired in for hero instruments. A run is hold-dominated (a 1-hour incubate is nearly all the real time), so playback uses per-beat pacing: every motion beat plays over a watchable slot while long holds compress, with 0.5×/1×/2× speed.
![]() |
![]() |
| The arm pipetting on the liquid handler deck | Hand-off: plate placed into the thermocycler before the hold |
Four-layer breakdown (frontend/src/lib/breakdown.ts). One deriveBreakdown() turns the plan's motions into four layers — mechanism, the protocol's own numbered steps, grouped robot actions, and atomic actions colour-coded by the machine that performs them. Atomics are 1:1 with the plan's motions, so the beat under the playhead is the atomic the strip highlights.
Feasibility, not just animation. When the plan has problems it still runs, but every finding is listed with its fix — here an under-specified protocol with missing mix volumes and aspirates with no tip loaded, flagged before anything reaches a robot. The inspection panel shows tracked state at any beat (where the sample is, tip state, liquid moved, wells touched), and the schedule view lays the run out as per-instrument Gantt lanes and lets you ask what happens with 2, 4 or 8 plates in flight — the arm is freed during holds, so plate B's pipetting overlaps plate A's incubation until they contend for the same instrument.
Run history. Every planned run is persisted (SQLite) with its plan, run-time, motion count, issue count and determinism score; open one to replay it in the 3D console or copy a ?run=<id> link to share it.
- Protocol importers (
backend/app/importers/) — Opentrons Protocol API and PyLabRobot Python (AST-walked, never executed) and Autoprotocol JSON are lifted into a rich-op layer, complex commands (transfer,distribute,consolidate, with Opentrons order-of-operations and tip-capacity splitting) are expanded to atomics, and the result is lowered onto the frozen 9-op contract for the same grounding pipeline.POST /api/import-protocol. - Bench inference (
reasoning/infer_bench.py) — protocol text → the machines and labware it needs (LLM extraction, then deterministic catalog matching); unmatched items get a "needs your pick" strip wired to the library picker. - Renderer-agnostic hand-off contract (
backend/app/handoff/) — a saved run can be emitted as a scene inventory + timed semantic trajectory + per-beat state deltas, so a physics renderer such as NVIDIA Isaac Sim (or any other viewer) can consume the same plan the Three.js console draws. See docs/ISAAC-HANDOFF-CONTRACT.md. - SBS labware geometry table — item → well pitch, A1 offset, rows×cols, well depth, tip-pickup height;
GET /api/labware-geometry/{item}/well/{A1}resolves any well to a centre offset in mm. - File ingestion — protocol text is extracted from uploaded PDF / DOC(X) / TXT / MD.
| Layer | Tech |
|---|---|
| Reasoning | Anthropic Claude at temperature 0, structured output, modal self-consistency |
| Grounding | Python · PyLabRobot resource/geometry model · deterministic expander, validator, feasibility engine, scheduler |
| API | FastAPI · Pydantic schemas are the shared contract (/docs for OpenAPI) |
| Frontend | React · TypeScript · Vite · Zustand · Three.js (lazy-loaded so the entry screens ship no 3D) |
| Data | JSON library catalog · SQLite run store · JSON world persistence |
~5,500 lines of Python (156 tests, 26 files) and ~7,300 lines of TypeScript.
Backend (Python 3.10+)
cd backend
python -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
cp .env.example .env # add ANTHROPIC_API_KEY (only needed for the LLM parse / bench inference)
python -m pytest -q # 156 passed — fully offline
uvicorn app.main:app --reload # http://localhost:8000/docsFrontend
cd frontend
npm install
npm run dev # http://localhost:5173 — proxies /api to the backendOpen the app, click load sample IVT protocol on the Protocol screen, infer the bench, confirm it, and run. Without an API key you can still exercise the deterministic layers — the importers (/api/import-protocol), op expansion (/api/expand-ops), feasibility (/api/feasibility-check), scheduling (/api/schedule) and the geometry table — and replay any saved run.
| Endpoint | Purpose |
|---|---|
POST /api/infer-bench |
protocol text → required machines / labware / unmatched |
POST /api/world · GET /api/world/{id} |
save / fetch a confirmed bench (the digital-twin world) |
POST /api/plan-protocol |
pasted protocol on a world → moves, operations, motions, run-time, issues, consensus |
POST /api/step/{id}/plan |
the same, for a step of a library workflow |
POST /api/import-protocol · POST /api/expand-ops |
Opentrons / PyLabRobot / Autoprotocol → rich ops → frozen op list |
POST /api/feasibility-check · POST /api/feasibility |
the feasibility engine on an op list + world; per-step workflow feasibility |
POST /api/schedule |
motions × N plates → per-instrument lanes + wall-clock |
GET /api/runs · GET /api/run/{id} · GET /api/run/{id}/handoff |
run history, replay, renderer hand-off contract |
GET /api/library · GET /api/labware-geometry · GET /api/assay-schemas |
catalog, SBS geometry table, assay schema registry |
POST /api/extract-protocol |
protocol text out of an uploaded file |
backend/
app/
schemas/ Operation · Bench · LibraryItem · WorldModel · Workflow (the contract)
reasoning/ Claude client, parser prompt, parser, consensus wrapper, bench inference
grounding/ validator, capability envelope, motion expander (frozen)
ops/ rich-op layer, complex-command expansion, lowering to frozen ops
importers/ Opentrons / PyLabRobot (AST) and Autoprotocol importers
feasibility/ stateful pre-run checks with severity + fixes
scheduling/ resource-lane discrete-event scheduler (arm + instrument lanes)
workflow/ transport routing, step and protocol planning, step↔motion linking
handoff/ renderer-agnostic scene + trajectory contract
geometry/ SBS labware geometry table and well resolution
storage/ run store (SQLite) and world store (JSON)
api/routes/ FastAPI routes
data/ library catalog, sample protocols, workflows
scripts/ spike.py (end-to-end), score.py (corpus scorer), fixtures/
tests/ 156 offline tests
frontend/
src/steps/ Experiment · PasteProtocol · ConfirmBench · Breakdown · ProtocolRunConsole · History
src/three/ Scene (arm + rail + IK), archetypes (procedural instruments), animation, layout
src/lib/ four-layer breakdown, inspection state, durations
docs/ architecture, state & roadmap, op vocabulary, hand-off contract, design record
- docs/architecture.md — the two halves and the seam between them
- docs/STATE-AND-ROADMAP.md — ground-truth snapshot of the build, candid limitations, next directions
- docs/REVAMP-V2.md — the design record: what is frozen, what was layered on top, build sequence, contracts
- docs/OP-VOCABULARY-V2.md — the rich-op layer and complex-command expansion rules
- docs/ISAAC-HANDOFF-CONTRACT.md — the renderer hand-off contract
- docs/spike-memo.md — the Phase-0 de-risk that froze the contract
Bench inference is LLM-backed and the least deterministic link in the chain. The twin is single-sample, single-arm (the scheduler reasons about N plates; the scene draws one). Persistence is SQLite + flat JSON — no accounts. The scene is procedural meshes; only hero instruments would get curated GLTF. See the roadmap doc for the full list.
Designed and built by Anupama Kozhiyalam — industrial biotechnologist working at the seam between wet-lab process development and lab automation software. MIT licensed.





