Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
92 changes: 91 additions & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -37,6 +37,10 @@ autonomous-lab throughput single_cell_genomics # plates/day, or why that number
autonomous-lab provenance single_cell_genomics # what could be proven about a run afterwards
autonomous-lab lineage single_cell_genomics # which cell did this read come from?
autonomous-lab knowledge # encoded expert judgment and robot benchmarks

autonomous-lab evidence # what the published record shows, and what it does not
autonomous-lab authority # which decisions a model may make, and which need a person
autonomous-lab cadence 500 --endpoint started --hours 16 # what that target actually costs
```

## What it reports today
Expand Down Expand Up @@ -318,6 +322,92 @@ adequacy. The STAR's dry motion is validated; its volumetric accuracy at the wor
has never been measured, and no camera retires that benchmark. It needs a calibration
experiment.

## Is the ambition supported by anything published?

Every layer above asks what *this* lab can prove. Three more ask a question none of them
can, because it is not about this lab: **is the plan supported by anything published, who
is allowed to decide, and what does a stated throughput target actually cost?**

```
autonomous-lab evidence # what the record shows -- and what it does not
autonomous-lab authority # which decisions a model may make
autonomous-lab cadence 500 --endpoint started --hours 16
```

### Evidence, scoped rather than cited

Autonomous-lab claims get justified by pointing at famous papers whose scope is far
narrower than the claim. A materials-synthesis campaign gets cited to justify a cell-assay
production line. So every entry carries what it **does not establish**, and support is
resolved by scope rather than by fame.

Fifteen works, each independently verified against the live record before it was encoded --
**ten confirmed, five partial, and eleven carrying figures that could not be confirmed**,
listed rather than rounded away. An entry whose numbers were checked against a preprint
rather than the version of record is `PARTIAL`, however solid the work is. The verification changed the content:

- The A-Lab paper was **corrected**. Forty-one successes became thirty-six after four were
reclassified as inconclusive, and the "43 new materials" figure in news coverage matches
no version of the paper.
- Assay Guidance Manual chapters carry **no DOI**, so any citation supplying one is wrong.
- PyLabRobot makes an *architectural assertion* about cross-family simulation. It does not
demonstrate it, and the entry says so.

`unsupported_claims()` is the standing list of things **no** cited work establishes. It is
computed from the table, so it shrinks on its own the day real evidence arrives.

### Authority: which decisions a model may make

"AI-native" gets read as "the model drives the instruments". The benchmark evidence is that
**no evaluated model exceeded 70% hazard-identification accuracy**, and fluent output is
not evidence of laboratory safety competence. But the honest position is not that models
are useless -- it is that decision classes differ.

Thirteen classes across four levels: `MODEL` may decide alone, `MODEL_PROPOSES` needs a
deterministic check or a human to ratify, `DETERMINISTIC` needs a coded interlock with no
model in the path, and `HUMAN` needs a named accountable person. **Authority may be
strengthened and never weakened** -- putting a person on a decision a model could make is a
cost, and the reverse is the defect `violations()` exists to catch.

Each class states *why* it sits where it does, against a specific published finding. Triage
sits at `MODEL_PROPOSES` because of a documented case where an unexpected intramolecular
cyclization carried the same molecular weight as the expected product, was indistinguishable
by chromatogram or MS, and passed the automated rule -- until a human read the NMR.

The qualification ladder runs simulation -> device acceptance -> integrated dry run ->
volumetric qualification -> assay transfer -> sustained production. **No rung implies the
one above it.** A passed simulation qualifies nothing physical.

### Cadence: what a throughput target actually costs

A target is the easiest number in a lab to state and the easiest to state meaninglessly.
`cadence` refuses one until it says what it counts, because plates *started*, *completed*,
*QC-passed*, and *released* differ by exactly the failure and rerun rate -- a measurement,
not an assumption. Resolving that silently would pick the loosest reading in the caller's
favour. It refuses twice: a downstream endpoint is a **return** rate, so reaching 500
released means admitting more than 500, and without a measured survival fraction the module
will not convert. Dividing the face value would report the *started* interval under a
*released* label.

Given an endpoint it is pure arithmetic on the caller's inputs, needing no measured
durations: 500 plates a day is one every **172.8 s** round the clock, or **115.2 s** inside
a sixteen-hour window. The ratio between those is the walk-away penalty, and it depends on
the clock alone -- not on the target, protocol, or instruments, which is the whole
quantitative case for closing the steps that need a human.

It then names what the interval forces, and refuses where the input does not exist: a
buffer depth is a makespan divided by an interval, so it inherits every unmeasured second
and is reported as not computable rather than estimated. `headroom()` is where "just buy a
faster robot" dies -- an intervention rate of one per two plates at ten minutes each
consumes **260%** of a 115.2 s interval, and instrument speed is not in that arithmetic.

And `demonstrated_support()` answers the question everybody skips. No cited system
establishes integrated multi-domain operation at a production rate: the longest integrated
campaign in the record is seventeen days in a single domain, and the one system running
continuously at scale across institutions is a cage sensor that observes rather than
actuates. That emptiness is **computed from the evidence table**, so it flips the day
somebody publishes a counterexample rather than staying a stale assertion.

## The RE queue is computed, not argued about

```
Expand Down Expand Up @@ -472,7 +562,7 @@ event receiver and silently steals the first one's callbacks.
pip install -e '.[dev]' && pytest
```

149 device-free tests. The ones that matter most try to make the layer lie:
292 device-free tests. The ones that matter most try to make the layer lie:

- claim a step is automated when its command is undecoded; claim a decoded command is
runnable while its siblings are not; claim a federated leg runs when no run card was ever
Expand Down
72 changes: 72 additions & 0 deletions autonomous_lab/__init__.py
Original file line number Diff line number Diff line change
Expand Up @@ -35,6 +35,20 @@
lineage which cell a read came from -- and whether pooling already destroyed the
answer, which is a different question from whether the plate was tracked.
intelligence the tacit expert judgment and the benchmarks a robot must meet first.

Three further layers ask the question the others cannot, because it is not about this lab
at all: is the ambition supported by anything published, who is allowed to decide, and what
does a stated throughput target actually cost?

literature what the published record demonstrates -- and, entry by entry, what it
explicitly does NOT establish. Scope, not fame, decides whether a paper
supports a claim.
authority which decisions a model may make, which need a deterministic interlock, and
which need a named accountable human. Authority may be strengthened, never
weakened.
cadence a plates-per-day target converted into the admission interval it implies,
and refused until it says whether it counts plates started, completed,
QC-passed, or released.
"""

from .executor import Executor, Handoff, RunReport, StepResult
Expand All @@ -48,7 +62,39 @@
loop_closure,
trusted_for,
)
from .authority import (
Authority,
DecisionClass,
Plan,
QualificationRung,
Rung,
Violation,
qualified_for,
violations,
)
from .cadence import (
Cadence,
Demonstration,
Endpoint,
Target,
cadence_for,
demonstrated_support,
headroom,
launch_interval,
sustained_rate_demonstrations,
)
from .ledger import Ledger, StepVerdict, Unlock, build_ledger, cost_step, rank_unlocks
from .literature import (
EVIDENCE,
Confidence,
Domain,
Evidence,
EvidenceKind,
Scope,
Support,
support_for,
unsupported_claims,
)
from .lineage import (
MISASSIGNMENT,
Cohort,
Expand Down Expand Up @@ -92,13 +138,23 @@ def loop_closure_for(protocol, workcell=None):

__all__ = [
"Artifact",
"Authority",
"Cadence",
"Cohort",
"Confidence",
"MISASSIGNMENT",
"Attestation",
"Basis",
"Benchmark",
"BenchmarkStatus",
"Criterion",
"DecisionClass",
"Demonstration",
"Domain",
"EVIDENCE",
"Endpoint",
"Evidence",
"EvidenceKind",
"CustodyGap",
"Decision",
"Detection",
Expand All @@ -122,44 +178,60 @@ def loop_closure_for(protocol, workcell=None):
"Misassignment",
"LoopClosure",
"Observable",
"Plan",
"Protocol",
"QualificationRung",
"Readiness",
"Role",
"RunRecord",
"Rung",
"Scope",
"Separability",
"RunReport",
"Severity",
"Step",
"StepResult",
"Support",
"StepVerdict",
"Tier",
"Target",
"TimeBasis",
"Traceability",
"Transform",
"UndeclaredTransform",
"Unlock",
"Verdict",
"Violation",
"VisionCapability",
"VisionRequirement",
"VisualCheck",
"Workcell",
"ZeroDecodeOp",
"build_ledger",
"build_lineage",
"cadence_for",
"cost_step",
"declared",
"demonstrated_support",
"estimate",
"evaluate",
"gate_report",
"headroom",
"knowledge_summary",
"launch_interval",
"lineage_report",
"loop_closure",
"loop_closure_for",
"provenance_report",
"qualified_for",
"rank_unlocks",
"recovery_report",
"registry",
"spec",
"support_for",
"sustained_rate_demonstrations",
"trusted_for",
"undeclared_transforms",
"unsupported_claims",
"violations",
]
Loading
Loading