Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
84 changes: 83 additions & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -37,6 +37,11 @@ autonomous-lab throughput single_cell_genomics # plates/day, or why that number
autonomous-lab provenance single_cell_genomics # what could be proven about a run afterwards
autonomous-lab lineage single_cell_genomics # which cell did this read come from?
autonomous-lab knowledge # encoded expert judgment and robot benchmarks

autonomous-lab coverage single_cell_genomics # would anything catch each failure?
autonomous-lab durability # what each instrument may be trusted with
autonomous-lab teaching single_cell_genomics # what an expert demonstrated, what the machine attains
autonomous-lab feedback single_cell_genomics # can a control loop actually close?
```

## What it reports today
Expand Down Expand Up @@ -318,6 +323,83 @@ adequacy. The STAR's dry motion is validated; its volumetric accuracy at the wor
has never been measured, and no camera retires that benchmark. It needs a calibration
experiment.

## Would anything catch it? Vision and gates, composed

Three layers here each hold a third of one answer and never meet. `vision` knows what a
camera can and can never see. `qc` knows which gates exist and whether they can fire.
`recovery` knows which failures destroy material. Nobody joined them, so nobody could answer
the question a lab actually asks: **for every failure that destroys material, is there
anything at all that would catch it?**

```
$ autonomous-lab coverage single_cell_genomics

failure modes 14
covered 0
uncovered 14, of which 9 destroy material
closable as it stands 7 (build the detector, or make the gate fire)
needs a new instrument 7
```

The composition is the product, because vision and gates are **complementary instruments,
not alternatives**. Where the photons are identical either way, no model resolves it and
only an assay does. So the report splits three ways, and each third is a different job:

- **six a camera would catch** -- `bead_pellet_aspirated` among them, the failure the
recovery layer already calls the worst silent one. The check is declared and undeployable:
missing camera, pose, labels, model, validation. That is a CV project with a known payoff.
- **one invisible, but a declared gate already reads it** -- `enzyme_activity_lost` sits
behind `library_quant_before_flow_cell`, which is unsatisfiable because the reader is
broken. No camera is involved. Repair the instrument and the coverage arrives.
- **seven no camera reaches at any capability** -- these are `mandatory_gates()`. Buying a
better model buys nothing against them. They need an assay that does not currently exist
in the protocol.

`sota_lift(current, proposed)` prices a model upgrade honestly. It returns both lists, and
the second is required output: **a report showing only what a capability lifts is a purchase
justification**, and every purchase justification is correct about the failures it lists and
silent about the ones that make the purchase insufficient.

Two things it refuses to count as coverage: a vision check that exists but whose
requirements are unmet, and a gate that exists but cannot be evaluated. Both are the
vacuous-pass failure in a new costume -- a plan that counts unbuilt detectors and unfirable
gates reports a lab as covered while material is quietly destroyed.

## Keeping it running, teaching it, and closing the loop

Three further layers cover what a workcell needs once it has to survive contact with a
calendar.

**`durability`** asks what an instrument is currently *entitled* to be trusted with. It is
deliberately not a failure predictor -- that needs population reliability data this package
does not have, and an invented MTBF becomes a specification nobody measured. What it does
compute is whether a planned campaign **crosses a service or calibration boundary mid-run**,
which is the insurance question: a plate that started before an expiry and finished after it
has an ambiguous provenance, and nothing downstream repairs that. An instrument with no
service history reports as unmeasured rather than healthy, matching how an unbenchmarked
operation is untrusted by default.

**`teaching`** models the transfer that makes any of this worth doing: an expert
demonstrating an operation, and a machine measured against that demonstration. The honest
core is that **a demonstration is data, not authority.** A scientist doing something twice is
not a specification -- it is two observations with a spread, and the spread is the
information. So an envelope refuses to produce a tolerance from a single demonstration, and
a machine with one good run reports as indistinguishable from unmeasured rather than as
meeting the bar. Most operations have no envelope at all, and that list is the real backlog
of an automation programme; `demonstration_queue` ranks what an expert should demonstrate
next by how many operations it unblocks, the same way `unlocks()` ranks decoding work.

**`feedback`** asks whether a control loop can actually close. Each of measure, compare, and
act has its own way of being fake, and the load-bearing one is latency: **a measurement taken
after the material is consumed cannot steer anything, however accurate it is.** It is a
post-mortem wearing the costume of a control loop. So a loop declares where it senses and
where it corrects, and any loop whose sensor sits downstream of its actuator is refused --
pure graph reasoning over the protocol, and it kills most proposed closed-loop designs. What
it will not do is model gain, overshoot, or settling, because that needs a plant model and
measured response data nothing here has. What it *can* say is how many plates are already in
flight between sensor and actuator, since every one of them is committed before the
correction lands.

## The RE queue is computed, not argued about

```
Expand Down Expand Up @@ -472,7 +554,7 @@ event receiver and silently steals the first one's callbacks.
pip install -e '.[dev]' && pytest
```

149 device-free tests. The ones that matter most try to make the layer lie:
351 device-free tests. The ones that matter most try to make the layer lie:

- claim a step is automated when its command is undecoded; claim a decoded command is
runnable while its siblings are not; claim a federated leg runs when no run card was ever
Expand Down
120 changes: 116 additions & 4 deletions autonomous_lab/__init__.py
Original file line number Diff line number Diff line change
Expand Up @@ -35,9 +35,65 @@
lineage which cell a read came from -- and whether pooling already destroyed the
answer, which is a different question from whether the plate was tracked.
intelligence the tacit expert judgment and the benchmarks a robot must meet first.

Four further layers cover what a workcell needs once more than one instrument is involved
and the lab has to keep running:

coverage vision and QC gates composed. For every failure that destroys material, is
there ANYTHING that would catch it? Where a camera is physically incapable,
an assay is the only option left, so an invisible failure with no gate is
uncovered and no CV budget ever changes that.
durability what an instrument is currently entitled to be trusted with, and whether a
planned campaign crosses a service boundary mid-run. Not a failure
predictor -- that needs reliability data this package does not have.
teaching an expert demonstrating an operation, and a machine measured against that
demonstration. A demonstration is data, not authority: one performance is a
value with no tolerance, and saying so is the point.
feedback whether a control loop can actually close. A sensor downstream of its own
actuator is a post-mortem wearing the costume of a control loop.

`Envelope` is defined in both `teaching` (what an expert demonstrated) and `feedback` (what
a loop steers toward). They are related and not the same, so neither is re-exported here --
import from the module that means the one you want.
"""

from .coverage import (
Ceiling,
Cover,
CoverageReport,
CoverageRow,
Demand,
SotaLift,
coverage_report,
mandatory_gates,
sota_lift,
unmet_demands,
)
from .durability import (
Accrued,
BoundaryReport,
Campaign,
Entitlement,
InstrumentHealth,
Interval,
IntervalKind,
ServiceRecord,
crosses_boundary,
entitlement_summary,
untrusted_instruments,
)
from .executor import Executor, Handoff, RunReport, StepResult
from .feedback import (
Closable,
Closure,
Controller,
Correction,
FeedbackReport,
Loop,
can_close,
feedback_report,
in_flight_exposure,
)
from .intelligence import (
Benchmark,
BenchmarkStatus,
Expand Down Expand Up @@ -66,6 +122,20 @@
from .model import Artifact, Protocol, Role, Step, Tier, Transform, Verdict, ZeroDecodeOp
from .provenance import Attestation, CustodyGap, Event, RunRecord, provenance_report
from .qc import Basis, Criterion, Decision, Gate, Readiness, evaluate, gate_report
from .teaching import (
Attainment,
Demonstration,
MachineObservation,
NextDemonstration,
TransferReport,
TransferRow,
attainment,
demonstration_queue,
envelope_for,
taught,
transfer_report,
untaught_operations,
)
from .recovery import Detection, FailureMode, Latency, Severity, recovery_report
from .registry import FEDERATED, FederatedSpec, InstrumentSpec, declared, registry, spec
from .throughput import Duration, TimeBasis, estimate
Expand All @@ -91,50 +161,76 @@ def loop_closure_for(protocol, workcell=None):


__all__ = [
"Accrued",
"Artifact",
"Cohort",
"MISASSIGNMENT",
"Attainment",
"Attestation",
"Basis",
"Benchmark",
"BenchmarkStatus",
"BoundaryReport",
"Campaign",
"Ceiling",
"Closable",
"Closure",
"Cohort",
"Controller",
"Correction",
"Cover",
"CoverageReport",
"CoverageRow",
"Criterion",
"CustodyGap",
"Decision",
"Demand",
"Demonstration",
"Detection",
"Duration",
"Entitlement",
"Event",
"Executor",
"FEDERATED",
"FailureMode",
"FederatedSpec",
"FeedbackReport",
"Gate",
"Handoff",
"InstrumentConfig",
"InstrumentHealth",
"InstrumentSpec",
"Interval",
"IntervalKind",
"Judgment",
"Latency",
"Ledger",
"Leg",
"LineageEdge",
"LineageGraph",
"LineageReport",
"Misassignment",
"Loop",
"LoopClosure",
"MISASSIGNMENT",
"MachineObservation",
"Misassignment",
"NextDemonstration",
"Observable",
"Protocol",
"Readiness",
"Role",
"RunRecord",
"Separability",
"RunReport",
"Separability",
"ServiceRecord",
"Severity",
"SotaLift",
"Step",
"StepResult",
"StepVerdict",
"Tier",
"TimeBasis",
"Traceability",
"TransferReport",
"TransferRow",
"Transform",
"UndeclaredTransform",
"Unlock",
Expand All @@ -144,22 +240,38 @@ def loop_closure_for(protocol, workcell=None):
"VisualCheck",
"Workcell",
"ZeroDecodeOp",
"attainment",
"build_ledger",
"build_lineage",
"can_close",
"cost_step",
"coverage_report",
"crosses_boundary",
"declared",
"demonstration_queue",
"entitlement_summary",
"envelope_for",
"estimate",
"evaluate",
"feedback_report",
"gate_report",
"in_flight_exposure",
"knowledge_summary",
"lineage_report",
"loop_closure",
"loop_closure_for",
"mandatory_gates",
"provenance_report",
"rank_unlocks",
"recovery_report",
"registry",
"sota_lift",
"spec",
"taught",
"transfer_report",
"trusted_for",
"undeclared_transforms",
"unmet_demands",
"untaught_operations",
"untrusted_instruments",
]
Loading
Loading