From 48274cff85e059c2617194b7b08fcd150f0d2894 Mon Sep 17 00:00:00 2001 From: di-omics Date: Tue, 28 Jul 2026 03:57:18 -0700 Subject: [PATCH 1/3] Hold a printed fixture to the same standard as an instrument The argument running through this portfolio is that a lab does not need a vendor to ship AI-native labware. That is only honest if the part you printed is held to the standard everything else here is held to, and a printed fixture is untested hardware made in-house, sitting inside the working envelope of a moving robot, sometimes near reagents. The boundary that decides most of it is physical rather than regulatory. Fused-deposition parts were measured by computed tomography at 4.05 to 6.32 percent porosity with infill fixed at 100 percent, so an FDM part cannot be validated as cleanable and does not belong in a fluid path. Post-processing that claims to close this is recorded as a claim rather than a solution, because a coating whose integrity nobody measured is an assertion about a surface. The consequences are computed rather than declared. Of the materials characterized here only a certified biocompatible photopolymer reaches culture contact, and only that plus machined stock reach sample contact: everything printable holds things and nothing printable holds liquid. A fixture with nothing declared about it is refused for every use, since passing by virtue of having no properties is the vacuous-pass failure this package exists to refuse. Four refusals carry the practical weight. A porous process in a fluid path. A material whose glass transition sits below the sterilization cycle -- PETG relaxes its frozen-in extrusion stresses at 121 C and comes out the wrong shape. A part that nothing positively locates, which needs no operator error to move and invalidates every taught position that references it once it has. And a critical dimension that was designed rather than measured, which is a fact about the model: desktop tolerance runs about half a percent with a half-millimetre floor, shrinkage is material and vendor specific, and the first layer is wider than the model by roughly the compensation a slicer applies without being asked. hardware/tilt_module.scad is the worked example, a passive fixed-angle fixture that pools residual liquid at one side of each well. No hinge and no adjustment, because an adjustable angle is an angle nobody records, and a hinge is a degree of freedom inside a moving arm's envelope. It defaults to printing the test coupon rather than the fixture so the fit is proven before hours are committed, and it states no recovery figure anywhere. The guide gives the gravimetric protocol that would produce one instead. Every material property, tolerance, and porosity figure was verified against published sources before it was encoded, and the ones that could not be confirmed are marked rather than rounded. openscad is not installed here, so the SCAD is checked for structure but has not been rendered. 71 new tests, 418 total. ruff clean. --- README.md | 44 +- autonomous_lab/__init__.py | 27 + autonomous_lab/printed.py | 1744 ++++++++++++++++++++++++++++++++++++ docs/PRINTED_FIXTURES.md | 960 ++++++++++++++++++++ hardware/README.md | 578 ++++++++++++ hardware/tilt_module.scad | 973 ++++++++++++++++++++ tests/test_printed.py | 760 ++++++++++++++++ 7 files changed, 5085 insertions(+), 1 deletion(-) create mode 100644 autonomous_lab/printed.py create mode 100644 docs/PRINTED_FIXTURES.md create mode 100644 hardware/README.md create mode 100644 hardware/tilt_module.scad create mode 100644 tests/test_printed.py diff --git a/README.md b/README.md index 3dd20e7..5824756 100644 --- a/README.md +++ b/README.md @@ -400,6 +400,48 @@ measured response data nothing here has. What it *can* say is how many plates ar flight between sensor and actuator, since every one of them is committed before the correction lands. +## Printing the fixture instead of waiting for it + +This portfolio's argument is that a lab does not need a vendor to ship AI-native labware. +That argument is only honest if a part you printed yourself is held to the same standard as +everything else here -- and a printed fixture is untested hardware, made in-house, sitting +inside the working envelope of a moving robot, sometimes near reagents. + +`printed` computes whether a specific part may be used for a specific purpose. The load- +bearing boundary is physical rather than regulatory: **fused-deposition parts were measured +by computed tomography at 4.05 to 6.32 percent porosity with infill fixed at 100 percent.** +A part that porous cannot be validated as cleanable, so it does not go in a fluid path, and +no amount of coating or annealing is accepted here as having fixed that -- post-processing +is recorded as a claim, not a solution. + +The consequences fall out rather than being asserted. Of the materials described here, only +a certified biocompatible photopolymer reaches culture contact, and only that plus machined +stock reach sample contact. Everything printable holds things; nothing printable holds +liquid. A blank fixture is refused for every use, because the failure this module exists to +prevent is a part that passes by virtue of having nothing declared about it. + +Four refusals are worth naming, because each is a real way a printed part fails on a deck: + +- **`porous_process_in_fluid_path`** -- the boundary above. +- **`softens_in_the_cycle`** -- PETG's glass transition is 69-77 C, so a 121 C autoclave + cycle relaxes its frozen-in extrusion stresses. The part comes out the wrong shape. +- **`not_positively_located`** -- an unlocated fixture is a crash waiting for the first + knock. It needs no operator error to move, and once it has moved every taught position + that references it is wrong with nothing reporting that anything changed. +- **`dimension_not_measured`** -- a designed dimension is a fact about the model. Desktop + tolerance runs about +/-0.5 percent with a +/-0.5 mm floor and scales with length, + shrinkage is material and vendor specific, and the first layer comes out wider than the + model by roughly the 0.2 mm a slicer compensates by default. + +`hardware/tilt_module.scad` is the worked example: a passive fixed-angle tilt fixture that +pools residual liquid at one side of each well so a tip can reach more of it. No hinge and +no adjustment, because an adjustable angle is an angle nobody records. It defaults to +printing the **test coupon** rather than the fixture, so the fit is proven before hours are +committed. And it states no recovery figure anywhere -- `docs/PRINTED_FIXTURES.md` gives the +gravimetric protocol that would produce one instead. An unmeasured fixture is a net addition +of a crash surface, a cleaning obligation, and an uncharacterized material to a workcell +that had none of them. + ## The RE queue is computed, not argued about ``` @@ -554,7 +596,7 @@ event receiver and silently steals the first one's callbacks. pip install -e '.[dev]' && pytest ``` -351 device-free tests. The ones that matter most try to make the layer lie: +418 device-free tests. The ones that matter most try to make the layer lie: - claim a step is automated when its command is undecoded; claim a decoded command is runnable while its siblings are not; claim a federated leg runs when no run card was ever diff --git a/autonomous_lab/__init__.py b/autonomous_lab/__init__.py index 0451de2..4794bc2 100644 --- a/autonomous_lab/__init__.py +++ b/autonomous_lab/__init__.py @@ -49,6 +49,9 @@ teaching an expert demonstrating an operation, and a machine measured against that demonstration. A demonstration is data, not authority: one performance is a value with no tolerance, and saying so is the point. + printed whether a 3D-printed fixture may be used for a given purpose. The argument + that a lab does not need a vendor to ship labware is only honest if the part + you printed is held to the same standard as everything else here. feedback whether a control loop can actually close. A sensor downstream of its own actuator is a post-mortem wearing the costume of a control loop. @@ -121,6 +124,19 @@ ) from .model import Artifact, Protocol, Role, Step, Tier, Transform, Verdict, ZeroDecodeOp from .provenance import Attestation, CustodyGap, Event, RunRecord, provenance_report +from .printed import ( + Contact, + CriticalDimension, + Fitness, + Fixture, + Material, + Process, + Sterilization, + Use, + fitness, + materials_reaching, + unqualified, +) from .qc import Basis, Criterion, Decision, Gate, Readiness, evaluate, gate_report from .teaching import ( Attainment, @@ -174,12 +190,14 @@ def loop_closure_for(protocol, workcell=None): "Closable", "Closure", "Cohort", + "Contact", "Controller", "Correction", "Cover", "CoverageReport", "CoverageRow", "Criterion", + "CriticalDimension", "CustodyGap", "Decision", "Demand", @@ -193,6 +211,8 @@ def loop_closure_for(protocol, workcell=None): "FailureMode", "FederatedSpec", "FeedbackReport", + "Fitness", + "Fixture", "Gate", "Handoff", "InstrumentConfig", @@ -211,9 +231,11 @@ def loop_closure_for(protocol, workcell=None): "LoopClosure", "MISASSIGNMENT", "MachineObservation", + "Material", "Misassignment", "NextDemonstration", "Observable", + "Process", "Protocol", "Readiness", "Role", @@ -226,6 +248,7 @@ def loop_closure_for(protocol, workcell=None): "Step", "StepResult", "StepVerdict", + "Sterilization", "Tier", "TimeBasis", "Traceability", @@ -234,6 +257,7 @@ def loop_closure_for(protocol, workcell=None): "Transform", "UndeclaredTransform", "Unlock", + "Use", "Verdict", "VisionCapability", "VisionRequirement", @@ -254,6 +278,7 @@ def loop_closure_for(protocol, workcell=None): "estimate", "evaluate", "feedback_report", + "fitness", "gate_report", "in_flight_exposure", "knowledge_summary", @@ -261,6 +286,7 @@ def loop_closure_for(protocol, workcell=None): "loop_closure", "loop_closure_for", "mandatory_gates", + "materials_reaching", "provenance_report", "rank_unlocks", "recovery_report", @@ -272,6 +298,7 @@ def loop_closure_for(protocol, workcell=None): "trusted_for", "undeclared_transforms", "unmet_demands", + "unqualified", "untaught_operations", "untrusted_instruments", ] diff --git a/autonomous_lab/printed.py b/autonomous_lab/printed.py new file mode 100644 index 0000000..052bb98 --- /dev/null +++ b/autonomous_lab/printed.py @@ -0,0 +1,1744 @@ +"""Whether a printed part may be used for a particular purpose, and every reason it may not. + +The argument this package makes is that a lab does not need a vendor to ship AI-native +labware. The instruments are already on the bench, the missing pieces are fixtures, and a +fixture can be printed in an afternoon. That argument is only honest if the printed part is +held to the standard everything else here is held to, because a printed fixture is untested +hardware, made in-house, with no lot number, sitting on a deck inside the working envelope +of a moving arm and sometimes within splash distance of a reagent. It is exactly the kind of +object this package refuses to trust by default. + +So this module computes fitness for a use. It is not a print guide. Nothing here tells +anybody what layer height to run, and the numbers it holds are thresholds for deciding, not +settings for making. + +Four ideas carry the module. + +POROSITY IS A PROPERTY OF THE PROCESS, NOT OF THE PRINT QUALITY. Fused deposition parts +carry an internal void network that survives every setting anybody has tried. It has been +measured by X-ray computed tomography at 4.05 to 6.32 percent BY VOLUME with infill fixed at +100 percent, over raster angles and extrusion widths, and no configuration approached zero. +Over 99 percent of those pores are under 0.2 mm, and they concentrate between the outer +perimeter and the infill -- which is to say the void network is connected to the outside +surface. Fluid gets in. A separate immersion study concluded in as many words that no +manufacturing settings provide adequate sealing against fluid intake, and that a part given +45 minutes of cold acetone vapor was still stained through by dye after 30 minutes of +immersion. Bacteria attach preferentially in the grooves between layers, where biofilm is +thickest. None of this is a calibration problem. A better printer prints a better-looking +porous part. So `Process.cleanable_in_principle` is a claim about the process and it is the +gate most of this module hangs from. + +The corollary matters as much as the fact: the fixes people reach for are not fixes. Vapor +smoothing reduces surface roughness by up to 90 percent and leaves the part permeable, which +changes what a person can see and not what a person can clean. Epoxy infiltration is claimed +airtight by vendors of epoxy, has no peer-reviewed cleaning validation behind it, and by +those same vendors' admission cannot reach internal channels. Parylene coating is the one +barrier with a real measurement behind it -- it restored cell viability on a cytotoxic resin +and held for days -- and it detached after five to six autoclave cycles. So +`PostProcess.closes_the_fluid_path` returns False for every member, and that uniform answer +is a measured result rather than a conservative stub. + +TG IS THE GATE FOR AN AUTOCLAVE, NOT HDT, AND THAT IS SPECIFICALLY A PRINTED-PART RULE. +Heat deflection temperature measures a standard bar deflecting 0.25 mm under an applied load. +A part in an autoclave usually carries no external load, so the intuitive reading is that +exceeding HDT is safe. For a printed part that reasoning fails, because the part contains +frozen-in extrusion stresses that relax the moment the polymer passes its glass transition: +it moves with zero load applied. The gap is not academic. One carbon-fiber nylon has a +published HDT of 170 to 194 C sitting on top of a glass transition of 70 C, because the +filler carries the load the matrix no longer can. Gate that material on HDT and a 121 C +cycle looks like a wide margin. Gate it on Tg and the part is 51 C past the line. So +`Material.survives` reads Tg and only Tg, `hdt` is carried and never consulted, and a +material with no Tg on record fails the check rather than skipping it. + +CERTIFICATION ATTACHES TO A WORKFLOW, NOT TO A BOTTLE. The photopolymer story is the +opposite shape from the fused-deposition one. Photopolymer parts are dense, so porosity is +not the objection; residual chemistry is. Typical formulations contain more than twenty +compounds and manufacturers disclose only the hazard-classified ones. Six splint materials +that all started at 1.0 to 1.47 percent photoinitiator differed by three orders of magnitude +in extractable initiator after cure, from 44.0 to 42,083.5 ng/mL, so leachate load cannot be +read off a formulation. Post-curing helps enormously -- one resin moved from 48.5 percent +cell viability at a short low-temperature cure to 89.5 percent at a long hot one -- and it +cannot help at all with the species that are not polymerizable, which is why a light +stabilizer identified at roughly 50 ug/mL in culture medium destroyed every oocyte it was +put near, after UV post-cure and after plasma treatment. Both resins in that study were +certified biocompatible at the time. That is the reason `ResinQualification` requires the +certified resin, the certified workflow, and a validated post-cure as three separate +conditions: a certification is measured on one printer at one layer height with one wash +time and one cure, and a part made outside that is not the article that was tested. + +Read the last point carefully before using this module for anything living. Reproductive and +developmental toxicity is an endpoint the biocompatibility standards do not trigger for a +benchtop fixture, so a resin can hold a full certification and have never been tested against +the thing a lab actually wants to put in it. CULTURE_CONTACT is the strictest level this +module knows and it is still not clearance for gametes or embryos; that qualification is a +mouse embryo assay run on parts from the actual build, and this module does not model it. + +UNMEASURED IS UNTRUSTED, AND A DESIGNED DIMENSION IS NOT A MEASURED ONE. This is the same +convention `intelligence.trusted_for` applies to an unbenchmarked operation and `qc.evaluate` +applies to an absent measurement, and it bites hardest on dimensions because a designed +dimension LOOKS like a measurement -- it is a number, with units, in a document. It is a +number about a model. Desktop tolerance is quoted at plus or minus 0.5 percent with a floor +of plus or minus 0.5 mm, scaling with length, so a calibration cube says nothing about a 200 +mm nest. Shrinkage is not published on any filament datasheet checked here, and the values +shipped inside slicers disagree with each other by 0.5 percent on the same polymer between +two vendors. Layer quantization forces total height to a multiple of the layer height. The +first layer is squished wider than the model by roughly the 0.2 mm of compensation the +slicer applies by default. All of that is recoverable by measuring the part, and none of it +is recoverable by reading the model. `DimensionState` keeps DESIGNED and MEASURED apart for +that reason, and `unqualified` defaults to listing everything, because a new print has +demonstrated nothing at all. + +WHAT THIS MODULE WILL NOT DO. It returns no predicted part lifetime, no autoclave cycle +count, no leachable concentration, and no probability that a part is fine. Every specific +cycle-count claim located for the high-temperature polymers traced back to commercial +aggregator pages rather than to a study or a manufacturer document, and a number of that +shape is exactly the kind that gets quoted into a specification with no author attached. +`NOT_ESTIMATED` names each refusal and what it would take to lift it, in the same form +`durability` uses -- and reuses that module's `Refusal`, because a second dataclass with the +same three fields would be a second definition of one discipline. +""" + +from __future__ import annotations + +from dataclasses import dataclass, field +from enum import Enum +from typing import Dict, List, Optional, Tuple + +from .durability import Refusal +from .qc import Basis + + +# -- processes ------------------------------------------------------------------ + + +class Process(str, Enum): + """How the part was made, which decides whether it can ever be cleaned. + + The split here is not about quality or cost. It is about whether the finished part has an + internal void network by construction, because that single physical property decides most + of what follows: whether a fluid path is permissible, whether a cleaning validation could + in principle be written, and whether a sterilization cycle that the material survives + means anything about the part being sterile. + + SLS sits in the list without a material behind it, and that is deliberate rather than an + omission. This repo holds no porosity measurement for powder-bed parts at all, so the + process cannot earn a fluid path here -- not because it is known to be porous, but because + nobody in this evidence set measured it. An uncharacterized process and a measured-porous + one are refused by different grounds, because they cost different things to clear: one + needs a different part, the other needs a measurement. + """ + + FDM = "fdm" # fused deposition; measured 4.05-6.32 percent porosity at 100 percent infill + SLA_DLP = "sla_dlp" # vat photopolymer; dense parts, and the objection is chemical instead + SLS = "sls" # powder bed fusion; no porosity measurement is held here + MACHINED = "machined" # cut from moulded or extruded stock, so the stock's datasheet governs + + @property + def cleanable_in_principle(self) -> bool: + """Whether parts are non-porous enough that a cleaning validation could be written. + + True does not mean clean. It means the part has no construction-inherent void network, + so removing a residue from it is a procedure somebody could validate rather than a + claim about a volume nothing reaches. False for FDM on measured evidence, and false for + SLS because this repo holds no evidence either way -- absence of a measurement is not + evidence of density, in exactly the way an absent QC measurement is not a passing gate. + """ + return self in (Process.SLA_DLP, Process.MACHINED) + + @property + def porosity_characterized(self) -> bool: + """Whether anybody has measured this process's void fraction in this evidence set. + + Separates the two ways `cleanable_in_principle` returns False. FDM is characterized and + fails; SLS is simply unmeasured. Collapsing them would send a lab to buy a different + printer when the honest next step is a computed-tomography scan of the part it has. + """ + return self is not Process.SLS + + +class PostProcess(str, Enum): + """A treatment applied after printing, and what it was actually shown to do. + + The members are the things labs reach for when told a printed part cannot be cleaned. None + of them is recorded here as a solution, and the module is careful to say WHY per member + rather than dismissing them as a class -- see POST_PROCESS_EVIDENCE. Two of them have real + measurements behind them and neither measurement is about sealing a porous part. + """ + + VAPOR_SMOOTHING = "vapor_smoothing" + EPOXY_INFILTRATION = "epoxy_infiltration" + PARYLENE_COATING = "parylene_coating" + ANNEALING = "annealing" + EXTENDED_POST_CURE = "extended_post_cure" + SOLVENT_EXTRACTION = "solvent_extraction" + + @property + def closes_the_fluid_path(self) -> bool: + """Whether this treatment has been SHOWN to make a porous part cleanable. + + False for every member, and that uniformity is a measured result rather than a + conservative default. An acetone-vapor-treated part was still stained through by dye + after 30 minutes of immersion in the one study that tested exactly this; epoxy + infiltration has vendor claims and no peer-reviewed cleaning validation, and cannot + reach internal channels by those vendors' own admission; a parylene barrier that did + restore cell viability detached after five to six autoclave cycles. A property that can + only return False is worth keeping as a property, because the alternative is a caller + reading the member list and assuming one of them counts. + """ + return False + + @property + def moves_dimensions(self) -> bool: + """Whether applying this invalidates a measurement taken before it. + + Solvent extraction and extended post-cure are excluded not because they are gentle but + because no dimensional change has been measured for them here. Everything else on the + list either removes material, adds a layer, or takes the part through its own glass + transition, and an annealing study measured changes of -2.50 to +2.10 percent that were + anisotropic -- height grew while width shrank -- so this is percent-level movement with + a sign that depends on the axis. + """ + return self in ( + PostProcess.VAPOR_SMOOTHING, + PostProcess.EPOXY_INFILTRATION, + PostProcess.PARYLENE_COATING, + PostProcess.ANNEALING, + ) + + +POST_PROCESS_EVIDENCE: Dict[PostProcess, str] = { + PostProcess.VAPOR_SMOOTHING: ( + "reduces surface roughness substantially and does not seal the part. In the study that " + "immersed treated parts in disinfectant, dye still penetrated a part given 45 minutes of " + "cold acetone vapor, and mechanical properties fell" + ), + PostProcess.EPOXY_INFILTRATION: ( + "claimed airtight and watertight to 65 psi by vendors of the epoxy, with no peer-reviewed " + "cleaning validation located, and conceded by those same sources to be unable to reach " + "internal channels. It also introduces a second uncharacterized contact material and can " + "seal existing bioburden underneath itself" + ), + PostProcess.PARYLENE_COATING: ( + "the one barrier with a measurement behind it: a 10 um coating restored a severely " + "cytotoxic resin to 93 and 85 percent viability on two assays and held for four days. It " + "detached after five to six autoclave cycles, so a coated part in a reusable workflow " + "needs a defined replacement interval rather than indefinite reuse" + ), + PostProcess.ANNEALING: ( + "demonstrated thermal survival and nothing else. An annealed part came through a 134 C " + "cycle with its shape and function, in a study whose authors measured no porosity, no " + "cleanability, no sterility and explicitly no crystallinity. Dimensional cost in the best " + "case was +0.47 percent on diameter and -1.43 percent on one bar width" + ), + PostProcess.EXTENDED_POST_CURE: ( + "moves photopolymer cytotoxicity a long way and cannot move all of it. Longer and hotter " + "cure took one resin from 48.5 to 89.5 percent viability, still short of control, and it " + "consumes only the polymerizable fraction -- residual initiator and light stabilizers are " + "not consumed by further irradiation at all" + ), + PostProcess.SOLVENT_EXTRACTION: ( + "the mechanism that can in principle remove non-polymerizable leachates, since post-cure " + "cannot. No protocol with verifiable parameters was located here, so a lab applying it is " + "doing something reasonable and unquantified" + ), +} + + +# -- contact levels ------------------------------------------------------------- + + +class Contact(str, Enum): + """How close to the science the part is trusted, ordered by how much each level costs. + + The ordering is the point and each level must be earned rather than assumed. NONE and + INCIDENTAL are structural claims about geometry: the part holds something, or sits near + something, and no fluid path passes through it. SAMPLE_CONTACT and CULTURE_CONTACT are + materials claims, and they are a different kind of question that no amount of dimensional + work answers. A part can hit every dimension in the microplate standards and be entirely + unfit to touch a sample, because the standards say nothing about material, leachables, + cytotoxicity, sterility or biocompatibility -- each one's scope is confined to the single + feature in its title. + + CULTURE_CONTACT is the strictest level here and is still not the strictest level that + exists. Gamete and embryo contact is qualified by a mouse embryo assay on parts from the + actual build, which this module does not model and which no photopolymer resin located here + advertises. + """ + + NONE = "none" # structural only; no fluid path, no line of sight to the sample + INCIDENTAL = "incidental" # splash or aerosol reach, with no intended contact + SAMPLE_CONTACT = "sample_contact" # the sample or its buffer touches the part + CULTURE_CONTACT = "culture_contact" # cells or media, at 37 C, for hours to days + + @property + def in_the_fluid_path(self) -> bool: + """True where a residue or a leachate reaches the sample by design rather than by accident. + + This is the line the porosity gate is drawn on. INCIDENTAL sits below it deliberately: + splash contact is a real exposure and it is bounded, transient, and not the geometry + that makes an unrecoverable void network matter. + """ + return self in (Contact.SAMPLE_CONTACT, Contact.CULTURE_CONTACT) + + @property + def rank(self) -> int: + return _CONTACT_ORDER.index(self) + + +# Strictest last, so a numeric comparison reads the same direction as the prose. +_CONTACT_ORDER: Tuple[Contact, ...] = ( + Contact.NONE, + Contact.INCIDENTAL, + Contact.SAMPLE_CONTACT, + Contact.CULTURE_CONTACT, +) + + +# -- sterilization -------------------------------------------------------------- + + +class Sterilization(str, Enum): + """How the part is decontaminated between uses. + + SINGLE_USE is a real answer and frequently the right one for a printed part: it removes + the cleaning validation question entirely rather than answering it badly. It is not a + weaker choice than autoclaving, and a module that ordered these by rigor would imply that + it is. + + The two steam cycles are kept apart because 13 C separates several materials' verdicts, + and because the cycles differ in more than temperature -- the higher one runs at roughly + 2.1 bar, which deforms unvented internal cavities on its own, independently of how + thermally capable the polymer is. + """ + + NONE = "none" # nothing is done; the part is reused as-is + SINGLE_USE = "single_use" # printed, used once, discarded + IPA_WIPE = "ipa_wipe" # 70 percent isopropanol on the surface + AUTOCLAVE_121 = "autoclave_121" # roughly 15-30 minutes at 121 C + AUTOCLAVE_134 = "autoclave_134" # roughly 3-4 minutes at 134 C, around 2.1 bar + + @property + def cycle_c(self) -> Optional[float]: + """The temperature the part is taken to, or None where no thermal load is applied.""" + if self is Sterilization.AUTOCLAVE_121: + return 121.0 + if self is Sterilization.AUTOCLAVE_134: + return 134.0 + return None + + @property + def thermal(self) -> bool: + return self.cycle_c is not None + + +# -- thermal thresholds, with the method and the provenance that produced them --- + + +@dataclass(frozen=True) +class Thermal: + """A temperature threshold as the evidence actually states it: a range and a method. + + A single number is the wrong shape for every value in this domain and publishing one is + how a guide misleads. Four grades of one polymer from ONE manufacturer span 7 C of glass + transition. Two manufacturers of nominally the same polymer differ by about 8 C. The same + polymer measured by two standard test methods on differently produced bars differs by 15 + to 20 C. Print orientation alone moved the onset of dimensional change by 17 C on one + material. So the range is stored, `conservative()` returns its low end, and the method is + carried alongside every figure because the method is part of the number. + + `basis` is `qc.Basis`, reused rather than mirrored for the same reason `durability` reuses + it: a manufacturer's datasheet figure is real evidence and it is not evidence about a part + printed on your machine at your layer height in your orientation. Datasheet values are + measured on moulded or specially printed test bars, and both of the high-temperature + datasheets read here carry explicit disclaimers to that effect. VENDOR is the honest basis + for almost everything in this module, and IN_HOUSE is what a lab would have to earn. + """ + + low_c: float + high_c: float + method: str + basis: Basis + note: str = "" + + def __post_init__(self) -> None: + if self.high_c < self.low_c: + raise ValueError( + f"thermal range has high {self.high_c} C below low {self.low_c} C" + ) + + @property + def spread_c(self) -> float: + return self.high_c - self.low_c + + def conservative(self) -> float: + """The low end, which is the only end a gate may use. + + The spread is not measurement noise. It is grade, manufacturer, test method and print + orientation, and the part in somebody's hand is one unidentified draw from it. Gating on + the middle of the range means half the possible parts are past the line. + """ + return self.low_c + + def describe(self) -> str: + if self.low_c == self.high_c: + return f"{self.low_c} C ({self.method})" + return f"{self.low_c}-{self.high_c} C ({self.method})" + + +# -- materials ------------------------------------------------------------------ + + +@dataclass(frozen=True) +class Material: + """One polymer as one process makes it, with the contact level it may reach. + + `tg` and `hdt` are both carried and only one of them is ever consulted. That asymmetry is + the module's sharpest single rule and it is enforced in `survives` rather than documented + and forgotten: a printed part relaxes its frozen-in extrusion stresses at the glass + transition with zero applied load, so heat deflection temperature -- which is measured + under load, on a moulded bar -- describes a situation the part is not in. Where a filler + carries load above Tg the two numbers can be 100 C or more apart, and reading the higher + one is how a part goes into a cycle it cannot survive. + + `ceiling` is the highest `Contact` this material may reach on the evidence held here. It is + a materials judgment and it is deliberately independent of the process porosity gate, so a + material and a process can refuse the same use for two different reasons. Both get + reported. A high-temperature engineering polymer that comfortably survives steam still + carries an FDM ceiling, because surviving a cycle says nothing about whether soil trapped + in a 4 to 6 percent void network was removed before it -- and steam does not remove + protein, it fixes it. + """ + + name: str + process: Process + ceiling: Contact + tg: Optional[Thermal] = None + hdt: Optional[Thermal] = None + melt: Optional[Thermal] = None + chemistry: str = "" + note: str = "" + + def survives(self, method: Sterilization) -> Tuple[Optional[bool], str]: + """Does the material clear this cycle on thermal grounds? None means nobody can say. + + Three-valued on purpose. True and False are answers; None is the state of a material + whose glass transition is not on record, which is common -- one polycarbonate blend's + sheet publishes heat deflection at two loads and no Tg at all, and one polypropylene + copolymer's sheet publishes neither, only a melting point. Returning False for those + would be right by accident and wrong in reasoning; returning True would be the vacuous + pass this package exists to refuse. Returning None makes the caller decide what to do + about not knowing, which is the honest position. + + Note what a True never claims. It is a statement about softening and nothing else. It is + not dimensional retention -- the only measured numbers located for any polymer through a + steam cycle still moved 0.2 to 1.4 percent in the best case, so a function depending on + a tolerance tighter than about one percent is not answered by this at all. It is not + hydrolytic stability, which is a separate failure mode that leaves shape intact while + molecular weight falls. And it is not sterility: whether steam penetrates the interior + void network of a porous part to a validated assurance level is unmeasured here. + """ + cycle = method.cycle_c + if cycle is None: + return True, f"{method.value} applies no thermal load to the part" + if self.tg is None: + carried = f"; heat deflection is on record at {self.hdt.describe()}" if self.hdt else "" + return None, ( + f"no glass transition is on record for {self.name}, so nothing here can be compared " + f"against a {cycle:.0f} C cycle{carried}. Heat deflection does not answer this " + "question for a printed part: it is measured under load on a moulded bar, and a " + "printed part moves at Tg under no load at all" + ) + tg = self.tg.conservative() + if tg <= cycle: + return False, ( + f"{self.name} has a glass transition from {self.tg.describe()}; the conservative end " + f"{tg:.0f} C is at or below the {cycle:.0f} C cycle, so the part relaxes its " + "frozen-in extrusion stresses and moves with nothing applied to it" + ) + return True, ( + f"{self.name} has a glass transition from {self.tg.describe()}, clearing a " + f"{cycle:.0f} C cycle by {tg - cycle:.0f} C at the conservative end. That is a " + "statement about softening only, not about dimensional retention, hydrolysis, or " + "whether the part came out sterile" + ) + + +# The materials this repo holds evidence about. Every figure carries the method that produced +# it, and every one is VENDOR or LITERATURE basis -- which is to say nobody here has measured +# any of them on a part from their own machine. That is the normal state and it is the reason +# Basis is on the record at all. + +MATERIALS: Dict[str, Material] = { + "pla_fdm": Material( + name="PLA (FDM)", + process=Process.FDM, + ceiling=Contact.INCIDENTAL, + tg=Thermal( + 54.0, + 61.0, + "DSC 10 C/min", + Basis.VENDOR, + "four grades from a single manufacturer span 7 C; the colorant and additive package " + "moves this number, so the spool decides and not the polymer name", + ), + hdt=Thermal(52.0, 61.0, "ISO 75 at 1.8 and 0.45 MPa", Basis.VENDOR), + chemistry=( + "ester backbone, so it is attacked by water at elevated temperature; a part can hold " + "its shape through a wet cycle and still lose molecular weight and embrittle" + ), + note=( + "visible warping was observed from approximately 70 C for a curved-beam geometry in " + "the one annealing study read here, well under the heat deflection figure. Annealed " + "parts did come through a 134 C cycle functional, which establishes thermal survival " + "and explicitly not porosity, cleanability or sterility" + ), + ), + "petg_fdm": Material( + name="PETG (FDM)", + process=Process.FDM, + ceiling=Contact.INCIDENTAL, + tg=Thermal( + 69.0, + 77.0, + "DSC 10 C/min and ISO 11357", + Basis.VENDOR, + "about 8 C between two manufacturers of nominally the same polymer", + ), + hdt=Thermal(68.0, 76.0, "ISO 75 at 1.8 and 0.455 MPa", Basis.VENDOR), + chemistry="ester linkages; hydrolysis-sensitive in wet heat", + note=( + "one preprint flags this polymer as cytotoxic while a peer-reviewed stem-cell study " + "cleared it. Cytotoxicity results are brand and lot specific and do not transfer " + "between spools of nominally identical polymer, because colorants, plasticizers and " + "processing aids are not disclosed on consumer filament" + ), + ), + "abs_fdm": Material( + name="ABS (FDM)", + process=Process.FDM, + ceiling=Contact.INCIDENTAL, + tg=Thermal( + 105.2, + 105.2, + "ASTM D7426, inflection point", + Basis.VENDOR, + "one industrial grade's figure; one desktop manufacturer's sheet for its own ABS lists " + "Tg as not available, so this single number does not describe the class", + ), + hdt=Thermal( + 84.0, + 104.4, + "ASTM D648 Method B and ISO 75", + Basis.VENDOR, + "a 20 C gap between an industrial and a desktop grade, and it is the gap that decides " + "a 121 C question", + ), + note=( + "a housing in the annealing study deformed at 121 C to the point of being unusable, " + "which corroborates the datasheet prediction experimentally. The vendor selling a " + "medical-grade ABS claims gamma and ethylene oxide sterilization for it and not steam" + ), + ), + "asa_fdm": Material( + name="ASA (FDM)", + process=Process.FDM, + ceiling=Contact.INCIDENTAL, + tg=Thermal( + 104.0, + 104.0, + "ASTM D7426", + Basis.VENDOR, + "one industrial grade; a widely used desktop grade's datasheet gives no Tg at all", + ), + hdt=Thermal(86.0, 103.0, "ISO 75 and ASTM D648 Method B", Basis.VENDOR), + note="the most defensible general-purpose choice here for a structural deck part off the fluid path", + ), + "pc_blend_fdm": Material( + name="PC blend (FDM)", + process=Process.FDM, + ceiling=Contact.INCIDENTAL, + tg=None, + hdt=Thermal(93.0, 113.0, "ISO 75 at 1.8 and 0.45 MPa", Basis.VENDOR), + chemistry="carbonate linkages; hydrolysis-sensitive in wet heat", + note=( + "the entry where a polymer name does the most damage. The blend's own datasheet " + "publishes no glass transition. Unfilled bisphenol-A resin is quoted near 147 C, and " + "this blend sits roughly 30 C below that resin at 1.8 MPa, so quoting the resin's " + "figure for this material would be about 37 C optimistic" + ), + ), + "pp_copolymer_fdm": Material( + name="PP copolymer (FDM)", + process=Process.FDM, + ceiling=Contact.INCIDENTAL, + tg=None, + hdt=None, + melt=Thermal(137.0, 137.0, "ASTM D3418", Basis.VENDOR), + note=( + "the datasheet gives a melting temperature and neither a glass transition nor a heat " + "deflection figure. A 134 C cycle runs 3 C below that melting point. The reputation of " + "moulded polypropylene labware, which is routinely autoclaved, does not carry over to " + "this filament: homopolymer melts higher, and the grade decides whether a 134 C cycle " + "is marginal or catastrophic" + ), + ), + "pa_cf_fdm": Material( + name="Carbon-fiber nylon (FDM)", + process=Process.FDM, + ceiling=Contact.INCIDENTAL, + tg=Thermal(70.0, 70.0, "DSC 10 C/min", Basis.VENDOR), + hdt=Thermal(170.0, 194.0, "ISO 75 at 1.8 and 0.45 MPa", Basis.VENDOR), + chemistry=( + "amide linkages and strongly hygroscopic. A steam cycle drives it toward saturation, " + "which swells the part and plasticizes the matrix, dropping the transition further" + ), + note=( + "the material that makes the Tg-versus-HDT rule concrete: heat deflection sits 100 to " + "124 C above the glass transition purely because carbon fiber carries the load the " + "matrix no longer can. Read the HDT column and a 121 C cycle looks like a wide margin; " + "read Tg and the part is 51 C past the line" + ), + ), + "pei_9085_fdm": Material( + name="PEI 9085 (FDM)", + process=Process.FDM, + ceiling=Contact.INCIDENTAL, + tg=Thermal(177.3, 177.3, "ASTM D7426, inflection point", Basis.VENDOR), + hdt=Thermal(170.2, 178.4, "ASTM D648 Method B, printed, XY and XZ", Basis.VENDOR), + note=( + "clears both steam cycles thermally and reaches no further up the contact ladder than " + "any other FDM material, because the void network is the same. The same datasheet's " + "thermomechanical curves show dimensional reversal beginning at about 175.9 C upright " + "against about 193.4 C flat, so orientation is worth 17 C on one material" + ), + ), + "pei_1010_fdm": Material( + name="PEI 1010 (FDM)", + process=Process.FDM, + ceiling=Contact.INCIDENTAL, + tg=Thermal( + 210.0, + 217.0, + "ASTM D7426 printed and ISO 11357 moulded", + Basis.VENDOR, + "the same polymer 7 C apart because the tests and the bars differ", + ), + hdt=Thermal(190.0, 215.0, "ISO 75 and ASTM D648 Method B", Basis.VENDOR), + note=( + "one trade name spans this material and PEI 9085 with 33 to 40 C between their glass " + "transitions, which is why a grade name is not a material specification" + ), + ), + "peek_fdm": Material( + name="PEEK (FDM)", + process=Process.FDM, + ceiling=Contact.INCIDENTAL, + tg=Thermal( + 143.0, + 150.0, + "ISO 11357-2, onset and midpoint", + Basis.VENDOR, + "measured on injection-moulding granules, not on a printed part", + ), + hdt=Thermal(152.0, 152.0, "ISO 75-2/Af, unannealed", Basis.VENDOR), + melt=Thermal(343.0, 343.0, "ISO 11357-3", Basis.VENDOR), + note=( + "the margin at 134 C is thinner than the numbers suggest: the transition onset is 9 C " + "above the cycle, and the material tolerates it because its crystalline phase carries " + "load above the transition. As-printed material from a machine without adequate chamber " + "temperature can be substantially amorphous, in which case it is not the article the " + "datasheet describes. The manufacturer's steam-sterilization statement is about moulded " + "granules" + ), + ), + "photopolymer_uncertified": Material( + name="Photopolymer, uncertified", + process=Process.SLA_DLP, + ceiling=Contact.INCIDENTAL, + tg=None, + hdt=None, + chemistry=( + "typical formulations contain more than twenty compounds and manufacturers disclose " + "only the hazard-classified ones. Documented cytotoxic leachates span crosslinkers, " + "monomers, photoinitiators and light stabilizers, and only the first three are " + "polymerizable, so post-cure cannot reach the fourth" + ), + note=( + "three consumer resins fell into the 0-25 percent viability band on two assays after " + "the full manufacturer workflow including sterilization. Parts from a vat printer were " + "measurably more toxic to fish embryos than parts from a fused-deposition printer in " + "the one study that compared them directly" + ), + ), + "photopolymer_certified": Material( + name="Photopolymer, certified biocompatible", + process=Process.SLA_DLP, + ceiling=Contact.CULTURE_CONTACT, + tg=None, + hdt=Thermal(54.0, 67.0, "ASTM D648-18 Method B at 1.8 and 0.45 MPa", Basis.VENDOR), + chemistry=( + "the certification covers a named list of endpoints -- cytotoxicity, mutagenicity, " + "irritation, sensitization, systemic toxicity, pyrogenicity -- and not reproductive or " + "developmental toxicity, which the standard's own risk logic does not trigger for a " + "benchtop fixture" + ), + note=( + "reaching this ceiling requires the resin, the workflow and the cure together; see " + "ResinQualification. Both heat deflection figures sit far below any steam cycle even " + "where the resin's chemistry is validated for one, so chemical clearance for a cycle " + "and dimensional survival of it are separate questions and this material answers only " + "the first. No glass transition is published, so this module refuses the autoclave " + "check for it rather than guessing" + ), + ), + "machined_stock": Material( + name="Machined from moulded or extruded stock", + process=Process.MACHINED, + ceiling=Contact.SAMPLE_CONTACT, + tg=None, + hdt=None, + note=( + "the defensible route for a reused sample-contacting part, and the reason the process " + "enum includes something that is not printed at all. A machined part inherits the " + "stock's datasheet, which has to be obtained per grade -- so no thermal figure is " + "asserted here and the autoclave check refuses until somebody supplies one. The " + "ceiling stops below culture contact because culture contact needs its own materials " + "qualification that this repo holds for no stock" + ), + ), +} + + +def materials_reaching(contact: Contact) -> Tuple[Material, ...]: + """Every material whose ceiling permits this contact level. + + Permits, not clears. A material at or above the level still has to get past the process + porosity gate, the resin qualification gate where one applies, and everything `fitness` + checks about the fixture itself. + """ + return tuple(m for m in MATERIALS.values() if m.ceiling.rank >= contact.rank) + + +def materials_surviving(method: Sterilization) -> Tuple[Material, ...]: + """Every material whose glass transition clears this cycle. + + A material with no Tg on record is not in the result, which is the same convention as + everywhere else here: the absence of a threshold is not a threshold that passes. + """ + return tuple(m for m in MATERIALS.values() if m.survives(method)[0] is True) + + +# -- what a photopolymer part has actually had established about it --------------- + + +@dataclass(frozen=True) +class ResinQualification: + """What is established about a specific photopolymer part, defaulting to nothing. + + Three fields rather than one because a certification is three claims and losing any of + them loses the certification. The resin carries a file of endpoint results. The + certification data were measured on one printer, at one layer height, washed for a stated + time in isopropanol of a stated purity, and cured at a stated temperature for a stated + duration -- a part made on different hardware, or washed in reused solvent, or cured in + something else, is not the article that was tested. And the cure this lab actually ran is + a separate question again from the cure the certification names, because cure energy is + the largest single lever on residual cytotoxicity and nobody outside the lab can verify + it was applied. + + None of the three, together, is clearance for gametes or embryos. Both resins that + destroyed every oocyte in the study that tested them were certified biocompatible at the + time, and the causative leachate identified there was a light stabilizer -- not a monomer + and not an initiator -- which no amount of further irradiation consumes. + """ + + certified_biocompatible: bool = False + workflow_as_certified: bool = False + post_cure_validated: bool = False + + @property + def cleared_for_culture(self) -> bool: + return ( + self.certified_biocompatible and self.workflow_as_certified and self.post_cure_validated + ) + + def missing(self) -> Tuple[str, ...]: + """Which of the three conditions is absent, named so a report can list them.""" + out: List[str] = [] + if not self.certified_biocompatible: + out.append("no certified biocompatible resin") + if not self.workflow_as_certified: + out.append("the part was not made by the workflow the certification was measured on") + if not self.post_cure_validated: + out.append("no validated post-cure") + return tuple(out) + + +# -- dimensions ----------------------------------------------------------------- + + +class DimensionState(str, Enum): + """Where a dimension's number came from. + + DESIGNED is the dangerous member and the reason this is an enum rather than a boolean. A + designed dimension is a number, in millimeters, in a document, and it reads exactly like a + measurement -- but it is a fact about a model. Desktop tolerance is quoted at plus or minus + 0.5 percent with a floor of plus or minus 0.5 mm and it scales with length, so a + well-behaved calibration cube establishes nothing about a 200 mm part. Shrinkage is not + published on the filament datasheets checked here at all, and the compensation values + shipped inside slicers disagree between two vendors by 0.5 percent on the same polymer. + + UNCHECKED is more honest than DESIGNED and less common, because somebody always has the + model. + """ + + UNCHECKED = "unchecked" # nobody has looked at the printed part + DESIGNED = "designed" # the number in the model; a fact about the file + MEASURED = "measured" # somebody put an instrument on the part that came out + + @property + def measured(self) -> bool: + return self is DimensionState.MEASURED + + +@dataclass(frozen=True) +class CriticalDimension: + """One dimension the part's function depends on, and where its number came from. + + Refuses three incoherent states at construction rather than carrying them into a report. A + MEASURED dimension with no measured value is the vacuous pass in its purest form: a state + flag that says a measurement happened, with no measurement. A non-MEASURED dimension + carrying a measured value is the same ambiguity from the other side. And a measurement with + no instrument named is an assertion, held to the same standard `qc.Criterion` holds a + threshold to when it requires a rationale and a basis. + + `measured_after_processing` exists because the order of operations decides whether a + measurement still describes the part. Annealing, smoothing, coating and steam cycles all + move dimensions by percent, anisotropically -- one study measured height growing while + width shrank in the same part -- so a caliper reading taken before any of those is a + reading of a part that no longer exists. + """ + + name: str + nominal_mm: float + state: DimensionState = DimensionState.DESIGNED + measured_mm: Optional[float] = None + measured_after_processing: bool = False + instrument: str = "" + + def __post_init__(self) -> None: + if self.state.measured: + if self.measured_mm is None: + raise ValueError( + f"dimension '{self.name}' claims MEASURED and carries no measured value; a state " + "flag is not a measurement" + ) + if not self.instrument: + raise ValueError( + f"dimension '{self.name}' claims MEASURED and names no instrument; an unattributed " + "measurement is an assertion" + ) + else: + if self.measured_mm is not None: + raise ValueError( + f"dimension '{self.name}' is {self.state.value} and carries a measured value; that " + "ambiguity is how a designed number comes to read as a measured one" + ) + if self.measured_after_processing: + raise ValueError( + f"dimension '{self.name}' is {self.state.value} and claims it was measured after " + "processing; nothing was measured at all" + ) + + def deviation_mm(self) -> Optional[float]: + """How far the part came out from the model, or None where nobody measured it. + + None rather than zero. Zero is a finding and would be the most misleading possible + default here, since the whole point of the distinction is that an unmeasured part is not + a part that came out on size. + """ + if self.measured_mm is None: + return None + return self.measured_mm - self.nominal_mm + + def deviation_percent(self) -> Optional[float]: + dev = self.deviation_mm() + if dev is None or self.nominal_mm == 0: + return None + return 100.0 * dev / self.nominal_mm + + +# -- the claimed benefit -------------------------------------------------------- + + +@dataclass(frozen=True) +class Benefit: + """What the fixture is claimed to buy, and whether anybody measured it. + + The question a printed fixture rarely gets asked. A part that saves nothing measurable is + a part on a deck for no reason, adding a crash surface, a cleaning obligation and an + uncharacterized material to a workcell that had none of them. `quantified` requires an + in-house basis specifically: a vendor's or a forum's claim about what a fixture buys is not + a measurement of what THIS fixture buys on THIS deck, which is the identical distinction + `qc.Basis` draws about a threshold. + """ + + claim: str + measured: Optional[str] = None + basis: Basis = Basis.INTUITION + + @property + def quantified(self) -> bool: + return self.measured is not None and self.basis.validated + + +# -- the fixture ---------------------------------------------------------------- + + +@dataclass(frozen=True) +class Fixture: + """A specific printed part, with every field defaulting to its untrusted value. + + Constructible with no arguments, and a fixture constructed that way is never fit for + anything. That is not a convenience -- it is the test the rest of the module is built + around. A default that reads as a pass is the failure mode this package hunts everywhere + else, and a fixture dataclass whose defaults were `positively_located=True` and + `dimensions=()` would report a part nobody has built as ready for a deck. + + `process` is a property reading the material rather than a field of its own, because two + fields that can disagree eventually will, and a fixture claiming FDM while carrying a + photopolymer would be refused by one gate and cleared by another. + """ + + name: str = "" + material: Optional[Material] = None + built_for: Contact = Contact.NONE + dimensions: Tuple[CriticalDimension, ...] = () + positively_located: bool = False + benefit: Optional[Benefit] = None + post_processing: Tuple[PostProcess, ...] = () + resin: ResinQualification = field(default_factory=ResinQualification) + demonstrated: Tuple["Qualification", ...] = () + note: str = "" + + @property + def process(self) -> Optional[Process]: + return None if self.material is None else self.material.process + + def dimensions_measured(self) -> bool: + """True only where at least one critical dimension is declared and every one was measured. + + The leading truthiness check is the whole method. `all()` over an empty tuple is True, so + the natural one-liner reports a fixture that declares no dimensions at all as fully + measured -- a check that passes when handed nothing, which is the exact shape of vacuous + pass this package refuses in `qc.evaluate` and again here. + """ + return bool(self.dimensions) and all(d.state.measured for d in self.dimensions) + + def dimension_changing_processes(self) -> Tuple[PostProcess, ...]: + return tuple(p for p in self.post_processing if p.moves_dimensions) + + def stale_dimensions(self, thermally_cycled: bool = False) -> Tuple[CriticalDimension, ...]: + """Measured dimensions whose measurement predates something that moved them. + + Returns nothing when nothing moved them, so this cannot manufacture a finding out of a + part that was simply printed and measured. + """ + if not self.dimension_changing_processes() and not thermally_cycled: + return () + return tuple( + d for d in self.dimensions if d.state.measured and not d.measured_after_processing + ) + + +# -- what has not been demonstrated ---------------------------------------------- + + +class Qualification(str, Enum): + """Something that could be demonstrated about a printed part, and usually has not been. + + The list is the standing backlog for one part. Four of the eight cannot be cleared by + anything in this module's own fields and have to be recorded by a lab that actually did the + work, because they are wet-bench results rather than facts about the build: cleanability to + a stated residue limit, a leachables characterization, sterility of the interior rather + than the surface, and a thermal measurement on a part from this machine instead of a + datasheet figure about a moulded bar. + + No published study located here has validated a fused-deposition part as cleanable to a + recognized standard. That is absence of evidence from this search rather than proof none + exists, and it is the crux of the question, so the qualification stays on the list rather + than being quietly assumed either way. + """ + + MATERIAL_DECLARED = "material_declared" + DIMENSIONS_MEASURED = "dimensions_measured" + POSITIVE_LOCATION = "positive_location" + BENEFIT = "benefit" + CLEANABILITY = "cleanability" + LEACHABLES = "leachables" + STERILITY = "sterility" + THERMAL_ON_THIS_PART = "thermal_on_this_part" + + +def unqualified(fixture: Fixture) -> Tuple[Qualification, ...]: + """Everything that has not been demonstrated about this fixture. + + Defaults to everything, in the order the enum declares. A new print is untrusted in exactly + the way an unbenchmarked operation is untrusted in `intelligence.trusted_for` -- and for a + stronger reason, since an operation at least has a script somebody wrote and read, while a + fresh print has a shape somebody hoped for. + + Four of the eight are read off the fixture's own fields, and the other four can only be + cleared by a caller recording them in `demonstrated`. That asymmetry is deliberate: this + module can see whether a dimension was measured and cannot see whether a residue limit was + met, and inventing a way to infer the second from the first is precisely the move this + package exists to refuse. + """ + held = set(fixture.demonstrated) + if fixture.material is not None: + held.add(Qualification.MATERIAL_DECLARED) + if fixture.dimensions_measured(): + held.add(Qualification.DIMENSIONS_MEASURED) + if fixture.positively_located: + held.add(Qualification.POSITIVE_LOCATION) + if fixture.benefit is not None and fixture.benefit.quantified: + held.add(Qualification.BENEFIT) + return tuple(q for q in Qualification if q not in held) + + +# -- the use --------------------------------------------------------------------- + + +@dataclass(frozen=True) +class Use: + """What the part is being asked to do, stated by the caller. + + `contact` has no default and must be named. Every other field defaults to the demanding + case, because a deck fixture sits inside an arm's sweep and gets registered against unless + somebody establishes otherwise, and a default that assumed the safe geometry would let an + under-specified use clear checks it never faced. + + `robot_registers_to_it` covers both halves of the same requirement -- a gripper closing on + the part and a tip having to reach a position defined by it -- because they fail on the + same missing evidence. Both need the printed part's real dimension rather than the model's, + and neither is helped by knowing the model's. + """ + + name: str + contact: Contact + robot_registers_to_it: bool = True + inside_moving_envelope: bool = True + sterilization: Sterilization = Sterilization.NONE + note: str = "" + + +# -- the verdict ----------------------------------------------------------------- + + +class Fitness(str, Enum): + """Whether a printed part may be used for a purpose, and if not, which refusal comes first. + + FIT is a member of the same enum as the refusals, following `feedback.Closable`, because + fitness is the absence of every refusal rather than a state of its own. There is no field + anywhere that sets it. + + `structural` draws the line the reporting order is built on. Two of these name something no + measurement, procedure or purchase clears for the part as built -- the only thing that + clears them is a different part, made a different way or out of a different polymer. + Everything else names work: measure the dimension, bolt the fixture down, certify the + resin, quantify what the thing buys. Reporting a cheap refusal ahead of a structural one + sends a lab to do work that will not matter. + """ + + MATERIAL_UNDECLARED = "material_undeclared" + POROUS_PROCESS_IN_FLUID_PATH = "porous_process_in_fluid_path" + CYTOTOXICITY_UNCLEARED = "cytotoxicity_uncleared" + SOFTENS_IN_THE_CYCLE = "softens_in_the_cycle" + UNCHARACTERIZED_PROCESS = "uncharacterized_process" + THERMAL_LIMIT_UNKNOWN = "thermal_limit_unknown" + ABOVE_MATERIAL_CEILING = "above_material_ceiling" + ABOVE_DECLARED_CONTACT = "above_declared_contact" + NOT_POSITIVELY_LOCATED = "not_positively_located" + DIMENSION_NOT_MEASURED = "dimension_not_measured" + DIMENSION_STALE_AFTER_CYCLE = "dimension_stale_after_cycle" + BENEFIT_NOT_QUANTIFIED = "benefit_not_quantified" + FIT = "fit" + + @property + def usable(self) -> bool: + return self is Fitness.FIT + + @property + def structural(self) -> bool: + return self in (Fitness.POROUS_PROCESS_IN_FLUID_PATH, Fitness.SOFTENS_IN_THE_CYCLE) + + +# Reported in this order. FIT is not in it: it is what is left when nothing else applies. +_REFUSAL_ORDER: Tuple[Fitness, ...] = ( + Fitness.MATERIAL_UNDECLARED, + Fitness.POROUS_PROCESS_IN_FLUID_PATH, + Fitness.CYTOTOXICITY_UNCLEARED, + Fitness.SOFTENS_IN_THE_CYCLE, + Fitness.UNCHARACTERIZED_PROCESS, + Fitness.THERMAL_LIMIT_UNKNOWN, + Fitness.ABOVE_MATERIAL_CEILING, + Fitness.ABOVE_DECLARED_CONTACT, + Fitness.NOT_POSITIVELY_LOCATED, + Fitness.DIMENSION_NOT_MEASURED, + Fitness.DIMENSION_STALE_AFTER_CYCLE, + Fitness.BENEFIT_NOT_QUANTIFIED, +) + + +@dataclass(frozen=True) +class Objection: + """One reason a part may not be used for a purpose, with the specific thing that is wrong.""" + + reason: Fitness + detail: str + + +@dataclass(frozen=True) +class Assessment: + """One part costed against one use. + + `objections` is every refusal the pairing earned, in reporting order, and the verdict is + the first of them. Carrying all of them rather than returning on the first is the same + choice `feedback.Closure` makes and for the same reason: a printed part is usually wrong in + more than one way, and a report showing only the headline sends somebody to fix that and + rediscover the next one after the fix. + """ + + fixture: Fixture + use: Use + objections: Tuple[Objection, ...] + + @property + def verdict(self) -> Fitness: + return self.objections[0].reason if self.objections else Fitness.FIT + + @property + def fit(self) -> bool: + return not self.objections + + @property + def needs_a_different_part(self) -> bool: + """True where at least one refusal is structural, so no amount of work on this part helps.""" + return any(o.reason.structural for o in self.objections) + + def refusals(self) -> Tuple[Fitness, ...]: + return tuple(o.reason for o in self.objections) + + def all_reasons(self) -> Tuple[str, ...]: + return tuple(o.detail for o in self.objections) + + @property + def reason(self) -> str: + if self.objections: + return self.objections[0].detail + return ( + f"'{self.fixture.name or 'this part'}' clears every refusal for '{self.use.name}' and " + "for no other use; fitness is computed per purpose and does not transfer" + ) + + +def _porosity_detail(fixture: Fixture, use: Use) -> str: + applied = fixture.post_processing + base = ( + f"{use.contact.value} puts a {Process.FDM.value} part in the fluid path. Fused deposition " + "parts were measured by computed tomography at 4.05 to 6.32 percent porosity with infill " + "fixed at 100 percent, across raster angles and extrusion widths, and the voids " + "concentrate between the outer perimeter and the infill -- so the network is connected to " + "the outside surface, fluid enters it, and nothing reaches in to remove what enters. The " + "immersion study that tested this concluded that no manufacturing settings provide " + "adequate sealing, and bacteria attach preferentially in the grooves between layers" + ) + if not applied: + return base + claims = "; ".join(f"{p.value}: {POST_PROCESS_EVIDENCE[p]}" for p in applied) + return ( + base + + ". Post-processing is applied and none of it is a demonstrated fix for this -- " + + claims + ) + + +def fitness(fixture: Fixture, use: Use) -> Assessment: + """Whether this part may be used for this purpose, with every refusal it earns. + + Every check runs. The function does not return on the first refusal, and it does not + shortcut the fixture-level checks when the material is missing, because a part with no + declared material still sits somewhere on a deck and still either is or is not bolted down. + + Fitness is computed per use and never cached on the fixture. The same part is fit as a + bracket and unfit as a reservoir, and a fixture carrying a fitness field would let the + first answer be read as the second. + """ + found: List[Objection] = [] + mat = fixture.material + + if mat is None: + found.append( + Objection( + Fitness.MATERIAL_UNDECLARED, + "no material is declared, so no thermal limit, no contact ceiling and no process " + "porosity claim can be resolved. Nothing about the material is refused here because " + "nothing about it is known, which is a worse position than a bad material honestly " + "named", + ) + ) + else: + if use.contact.in_the_fluid_path and not mat.process.cleanable_in_principle: + if mat.process.porosity_characterized: + found.append(Objection(Fitness.POROUS_PROCESS_IN_FLUID_PATH, _porosity_detail(fixture, use))) + else: + found.append( + Objection( + Fitness.UNCHARACTERIZED_PROCESS, + f"{use.contact.value} needs a process whose void fraction somebody measured, and " + f"this repo holds no porosity measurement for {mat.process.value} at all. This is " + "not a finding that the process is porous; it is the absence of a finding either " + "way, and it is cleared by a measurement rather than by a different part", + ) + ) + + if use.contact is Contact.CULTURE_CONTACT and mat.process is Process.SLA_DLP: + if not fixture.resin.cleared_for_culture: + missing = "; ".join(fixture.resin.missing()) + found.append( + Objection( + Fitness.CYTOTOXICITY_UNCLEARED, + f"culture contact with a photopolymer part requires a certified biocompatible " + f"resin, the workflow the certification was measured on, and a validated " + f"post-cure. Missing: {missing}. Incompletely cured resin is cytotoxic, and the " + "part of it that post-curing cannot reach is the non-polymerizable fraction -- " + "the leachate identified in the oocyte study was a light stabilizer at roughly " + "50 ug/mL, which further irradiation does not consume", + ) + ) + + if use.contact.rank > mat.ceiling.rank: + found.append( + Objection( + Fitness.ABOVE_MATERIAL_CEILING, + f"'{mat.name}' reaches {mat.ceiling.value} on the evidence held here and this use " + f"asks for {use.contact.value}", + ) + ) + + if use.sterilization.thermal: + ok, why = mat.survives(use.sterilization) + if ok is None: + found.append(Objection(Fitness.THERMAL_LIMIT_UNKNOWN, why)) + elif not ok: + found.append(Objection(Fitness.SOFTENS_IN_THE_CYCLE, why)) + + if use.contact.rank > fixture.built_for.rank: + found.append( + Objection( + Fitness.ABOVE_DECLARED_CONTACT, + f"the part was built for {fixture.built_for.value} and this use asks for " + f"{use.contact.value}. A fixture promoted up the ladder after the fact was designed, " + "printed and finished against a weaker requirement than the one it is now being held " + "to", + ) + ) + + if use.inside_moving_envelope and not fixture.positively_located: + found.append( + Objection( + Fitness.NOT_POSITIVELY_LOCATED, + "the part sits inside a moving arm's envelope and nothing positively locates it. An " + "unlocated fixture is a crash waiting for the first knock: it needs no operator error " + "to move, and once it has moved every taught position that references it is wrong " + "with nothing reporting that anything changed", + ) + ) + + if use.robot_registers_to_it: + if not fixture.dimensions_measured(): + declared = len(fixture.dimensions) + unmeasured = tuple(d.name for d in fixture.dimensions if not d.state.measured) + if declared == 0: + detail = ( + "a robot grips this part or a tip must reach a position it defines, and no critical " + "dimension is declared at all" + ) + else: + detail = ( + "a robot grips this part or a tip must reach a position it defines, and these " + f"dimensions were never measured on the printed part: {', '.join(unmeasured)}" + ) + found.append( + Objection( + Fitness.DIMENSION_NOT_MEASURED, + detail + + ". A designed dimension is a fact about the model. Desktop tolerance is quoted at " + "plus or minus 0.5 percent with a floor of plus or minus 0.5 mm and it scales with " + "length, shrinkage is material and vendor specific and is not published on the " + "filament datasheets checked here, and the first layer comes out wider than the " + "model by roughly the 0.2 mm of compensation a slicer applies by default", + ) + ) + else: + stale = fixture.stale_dimensions(use.sterilization.thermal) + if stale: + movers = [p.value for p in fixture.dimension_changing_processes()] + if use.sterilization.thermal: + movers.append(use.sterilization.value) + found.append( + Objection( + Fitness.DIMENSION_STALE_AFTER_CYCLE, + f"these dimensions were measured before {', '.join(movers)}: " + f"{', '.join(d.name for d in stale)}. Thermal and coating steps move printed parts " + "by percent rather than by microns, and anisotropically -- one annealing study " + "measured height growing while width shrank in the same part, with best-case " + "changes of +0.47 and -1.43 percent. A measurement taken before them describes a " + "part that no longer exists", + ) + ) + + if fixture.benefit is None or not fixture.benefit.quantified: + if fixture.benefit is None: + detail = "nothing states what this part is for" + elif fixture.benefit.measured is None: + detail = f"the claim is '{fixture.benefit.claim}' and nobody measured it" + else: + detail = ( + f"the claim is '{fixture.benefit.claim}' and the measurement behind it rests on " + f"{fixture.benefit.basis.value}, which is not a measurement of what this part buys on " + "this deck" + ) + found.append( + Objection( + Fitness.BENEFIT_NOT_QUANTIFIED, + detail + + ". An unquantified fixture is a net addition of a crash surface, a cleaning " + "obligation and an uncharacterized material to a workcell that had none of them", + ) + ) + + by_reason = {o.reason: o for o in found} + ordered = tuple(by_reason[r] for r in _REFUSAL_ORDER if r in by_reason) + return Assessment(fixture=fixture, use=use, objections=ordered) + + +# -- the microplate envelope a printed nest has to clear ------------------------- +# Included because the plate nest is the canonical printed fixture and because the number +# everybody quotes for it is the wrong one. Nothing in this section is evidence about contact +# fitness: dimensional conformity and fitness to touch a sample are independent questions, and +# the standards themselves say nothing at all about material, leachables, cytotoxicity, +# sterility or biocompatibility. + + +class Flange(str, Enum): + """Which bottom-flange variant a plate declares, of the five that are mutually exclusive. + + The standard requires a plate to state which one it meets, and a vendor page saying only + that a plate is footprint-compliant does not state it. UNDECLARED is therefore a real and + common state rather than a placeholder, and it is the state in which a printed nest cannot + be dimensioned -- a nest built around a tall flange will not reliably hold a short-flange + plate, and the difference between the extremes is over 5 mm. + """ + + SHORT = "short" # 2.41 mm +/- 0.38 + MEDIUM = "medium" # 6.10 mm +/- 0.38 + TALL = "tall" # 7.62 mm +/- 0.38 + SHORT_INTERRUPTED = "short_interrupted" # 2.41 mm +/- 0.38, one interruption per long side + DUAL = "dual" # 2.41 mm on the short sides, 7.62 mm on the long sides + UNDECLARED = "undeclared" + + @property + def declared(self) -> bool: + return self is not Flange.UNDECLARED + + +# Nominal footprint, and the two tolerances that apply to it. Both are real and only one of +# them is famous. +FOOTPRINT_LENGTH_MM = 127.76 +FOOTPRINT_WIDTH_MM = 85.48 +CORNER_ZONE_TOLERANCE_MM = 0.25 # applies only within 12.7 mm of the four outside corners +MID_SIDE_TOLERANCE_MM = 0.5 # applies anywhere else along the side +CORNER_RADIUS_MM = 3.18 +CORNER_RADIUS_TOLERANCE_MM = 1.6 # a permitted range of 1.58 to 4.78 mm, a threefold spread +TYPICAL_HEIGHT_MM = 14.35 +TYPICAL_HEIGHT_TOLERANCE_MM = 0.76 # overall plate height, Datum A to maximum protrusion + +_FLANGE_NOMINAL: Dict[Flange, Tuple[float, float]] = { + Flange.SHORT: (2.41, 2.41), + Flange.MEDIUM: (6.10, 6.10), + Flange.TALL: (7.62, 7.62), + Flange.SHORT_INTERRUPTED: (2.41, 2.41), + Flange.DUAL: (2.41, 7.62), +} +_FLANGE_TOLERANCE_MM = 0.38 + + +@dataclass(frozen=True) +class PlateEnvelope: + """The largest a conforming plate may actually be, which is what a nest has to clear. + + Every figure here is a maximum rather than a nominal, and that is the point of computing it + instead of copying the headline. A nest cut to the nominal footprint with the famous + tolerance jams a plate that bowed to the mid-side limit and still conforms. A nest cut to + the nominal corner radius binds on a plate at the upper end of a range that spans 1.58 to + 4.78 mm. Neither plate is out of specification. + + `max_height_mm` is None unless the caller states the plate is a standard-height microplate, + because height is standardized only for that format. Deep-well, PCR and reservoir plates + keep the footprint and abandon the height, and the heights they use are individual product + specifications rather than a standard, so no figure is emitted for them. + """ + + max_length_mm: float + max_width_mm: float + max_corner_radius_mm: float + flange_low_mm: Optional[float] + flange_high_mm: Optional[float] + max_height_mm: Optional[float] + note: str + + @property + def flange_known(self) -> bool: + return self.flange_low_mm is not None and self.flange_high_mm is not None + + +DRAFT_NOTE = ( + "the standards state that their dimensions and tolerances exclude draft. Moulded plates " + "carry draft angles the numbers here do not describe, and those angles are not published " + "anywhere in the standards, so a printed nest modelled on nominals with vertical walls does " + "not have the same cross-section as the plate it is meant to seat. Dimensional conformity " + "is also not evidence of fitness for contact with cells, media or samples -- the standards " + "say nothing about material, leachables, cytotoxicity, sterility or biocompatibility" +) + + +def plate_envelope( + flange: Flange = Flange.UNDECLARED, typical_height: bool = False +) -> PlateEnvelope: + """The maxima a printed nest has to clear for a plate that conforms. + + Refuses to emit flange numbers for an undeclared variant rather than picking the common one. + The variants are mutually exclusive by the standard's own text, the standard requires the + plate to declare which it meets, and guessing produces a nest that holds some plates and + drops others -- which is worse than a nest that was never built, because the failure is + intermittent. + """ + if flange.declared: + low, high = _FLANGE_NOMINAL[flange] + flange_low: Optional[float] = low - _FLANGE_TOLERANCE_MM + flange_high: Optional[float] = high + _FLANGE_TOLERANCE_MM + if flange is Flange.DUAL: + note = ( + "the two flange figures are different SIDES of one plate, not a tolerance band: the " + "short sides and the long sides differ by over 5 mm. " + DRAFT_NOTE + ) + else: + note = DRAFT_NOTE + else: + flange_low = None + flange_high = None + note = ( + "no flange variant is declared, so no flange height is emitted. Five mutually exclusive " + "variants exist, they span 2.41 to 7.62 mm nominal, and the standard requires a plate to " + "state which one it meets. A nest dimensioned around a guess holds some conforming " + "plates and drops others. " + DRAFT_NOTE + ) + return PlateEnvelope( + max_length_mm=FOOTPRINT_LENGTH_MM + MID_SIDE_TOLERANCE_MM, + max_width_mm=FOOTPRINT_WIDTH_MM + MID_SIDE_TOLERANCE_MM, + max_corner_radius_mm=CORNER_RADIUS_MM + CORNER_RADIUS_TOLERANCE_MM, + flange_low_mm=flange_low, + flange_high_mm=flange_high, + max_height_mm=(TYPICAL_HEIGHT_MM + TYPICAL_HEIGHT_TOLERANCE_MM) if typical_height else None, + note=note, + ) + + +# -- reference fixtures and uses -------------------------------------------------- +# Three parts a lab actually prints, in the state a lab actually prints them in. The first is +# the honest default: designed, printed, put on the deck. The last is what it costs to get one +# part to FIT for one use. + +FIXTURES: Dict[str, Fixture] = { + "plate_nest_as_printed": Fixture( + name="plate_nest_as_printed", + material=MATERIALS["pla_fdm"], + built_for=Contact.NONE, + dimensions=( + CriticalDimension(name="pocket_length", nominal_mm=128.5), + CriticalDimension(name="pocket_width", nominal_mm=86.2), + CriticalDimension(name="corner_relief_radius", nominal_mm=4.8), + ), + positively_located=False, + benefit=None, + note=( + "the state every printed fixture starts in and most stay in: modelled, printed, set on " + "the deck. Nothing about it is measured and nothing holds it down" + ), + ), + "reagent_trough_petg": Fixture( + name="reagent_trough_petg", + material=MATERIALS["petg_fdm"], + built_for=Contact.SAMPLE_CONTACT, + dimensions=( + CriticalDimension( + name="trough_floor_depth", + nominal_mm=12.0, + state=DimensionState.MEASURED, + measured_mm=11.82, + measured_after_processing=True, + instrument="digital caliper, three positions", + ), + ), + positively_located=True, + post_processing=(PostProcess.VAPOR_SMOOTHING,), + benefit=Benefit( + claim="holds a shared reagent volume no commercial trough matches", + measured="dead volume 1.1 mL against 4.0 mL for the commercial trough, measured over six fills", + basis=Basis.IN_HOUSE, + ), + note=( + "the part that is well built and still refused. Everything a lab controls has been done " + "here -- located, measured after processing, benefit quantified, surface smoothed -- and " + "the refusal is the process, which is why it is structural" + ), + ), + "sensor_bracket_asa": Fixture( + name="sensor_bracket_asa", + material=MATERIALS["asa_fdm"], + built_for=Contact.NONE, + dimensions=( + CriticalDimension( + name="mount_hole_pitch", + nominal_mm=40.0, + state=DimensionState.MEASURED, + measured_mm=39.86, + measured_after_processing=True, + instrument="digital caliper, five repeats", + ), + CriticalDimension( + name="face_to_deck_offset", + nominal_mm=22.0, + state=DimensionState.MEASURED, + measured_mm=21.94, + measured_after_processing=True, + instrument="height gauge against the deck datum", + ), + ), + positively_located=True, + benefit=Benefit( + claim="puts a sensor where no vendor bracket reaches", + measured="removes one manual repositioning per run, timed at 40 s over eight runs", + basis=Basis.IN_HOUSE, + ), + note="structural, off the fluid path, bolted down, and measured after it came off the bed", + ), +} + + +USES: Dict[str, Use] = { + "deck_bracket": Use( + name="deck_bracket", + contact=Contact.NONE, + robot_registers_to_it=True, + inside_moving_envelope=True, + sterilization=Sterilization.NONE, + note="holds a sensor near the deck; the arm sweeps past it and a taught position depends on it", + ), + "plate_nest": Use( + name="plate_nest", + contact=Contact.INCIDENTAL, + robot_registers_to_it=True, + inside_moving_envelope=True, + sterilization=Sterilization.IPA_WIPE, + note="a gripper places a plate into it, so the pocket dimension is the whole function", + ), + "shared_reagent_trough": Use( + name="shared_reagent_trough", + contact=Contact.SAMPLE_CONTACT, + robot_registers_to_it=True, + inside_moving_envelope=True, + sterilization=Sterilization.AUTOCLAVE_121, + note="tips enter it and buffer sits in it; the fluid path runs through the printed surface", + ), + "culture_insert": Use( + name="culture_insert", + contact=Contact.CULTURE_CONTACT, + robot_registers_to_it=True, + inside_moving_envelope=True, + sterilization=Sterilization.AUTOCLAVE_121, + note="cells and media sit against the printed surface at 37 C for days", + ), +} + + +def survey() -> Dict[str, int]: + """How much of this catalog reaches how far, for the report. + + The number worth looking at is `autoclavable_and_sample_contact`, and it is zero. No + material here both clears a steam cycle on its glass transition and reaches sample contact + on its process and materials evidence -- the polymers that survive steam are all made by the + process that cannot be validated as cleanable, and the one material that reaches sample + contact is not printed and inherits a datasheet nobody has supplied. + """ + autoclavable = set(m.name for m in materials_surviving(Sterilization.AUTOCLAVE_121)) + sample = set(m.name for m in materials_reaching(Contact.SAMPLE_CONTACT)) + return { + "materials": len(MATERIALS), + "fdm": sum(1 for m in MATERIALS.values() if m.process is Process.FDM), + "with_tg_on_record": sum(1 for m in MATERIALS.values() if m.tg is not None), + "validated_thresholds": sum( + 1 for m in MATERIALS.values() if m.tg is not None and m.tg.basis.validated + ), + "reaching_sample_contact": len(sample), + "reaching_culture_contact": len(materials_reaching(Contact.CULTURE_CONTACT)), + "autoclavable_121": len(autoclavable), + "autoclavable_and_sample_contact": len(autoclavable & sample), + } + + +# -- what this module will not estimate ------------------------------------------- + + +NOT_ESTIMATED: Tuple[Refusal, ...] = ( + Refusal( + quantity="part service life", + why=( + "a lifetime is a distribution fitted to units that failed. This module has one of each " + "part, no failure history, and a wear mechanism that differs with every geometry, " + "orientation and duty cycle. A figure computed from a part that has not broken yet is " + "arithmetic on nothing" + ), + what_it_would_take=( + "failure and censoring times across many parts of the same design, printed on the same " + "machine, in the same service -- which is a fleet operator's dataset and not one lab's" + ), + ), + Refusal( + quantity="autoclave cycles survived", + why=( + "every specific cycle-count claim located for the high-temperature polymers traced back " + "to commercial aggregator pages rather than to a study or a manufacturer document, and " + "the manufacturer documents themselves state repeated steam compatibility without a " + "number. Repeated cycling is a cumulative creep and hydrolysis problem, so a count is " + "not an extrapolation from a single cycle" + ), + what_it_would_take=( + "dimensional and mechanical measurement of this part across a cycle series on this " + "sterilizer, which is a study a lab can run and nobody in this evidence set has" + ), + ), + Refusal( + quantity="leachable release rate or concentration", + why=( + "release depends on the polymer, the additive package that is not disclosed on consumer " + "filament or in a resin's safety data, the cure state, the surface-area-to-volume ratio " + "of the specific geometry, the medium, and the time at temperature. Six materials that " + "started at essentially identical photoinitiator content differed by three orders of " + "magnitude in extractable initiator after cure" + ), + what_it_would_take=( + "an extraction study on parts from this build in the actual medium at the actual " + "temperature and duration, analyzed by mass spectrometry" + ), + ), + Refusal( + quantity="cytotoxicity of a specific printed part", + why=( + "results are material, brand and lot specific and do not transfer between spools or " + "bottles of nominally identical polymer. Two standard assays disagreed by 23 percentage " + "points on the same material in one study, so even a measured number needs more than one " + "method behind it" + ), + what_it_would_take=( + "an assay on parts from this build, by more than one method, against the cells that will " + "actually be exposed -- and for gametes or embryos, a mouse embryo assay, which no resin " + "located here advertises" + ), + ), + Refusal( + quantity="shrinkage percentage for a material", + why=( + "no filament datasheet checked here publishes one, so every figure in circulation is a " + "slicer default or a community measurement. Two vendors shipping profiles inside the same " + "slicer differ by 0.5 percent on the same polymer, and the widely repeated figure of up " + "to 11 percent for one of them is wrong by more than an order of magnitude against the " + "compensation that slicer actually applies" + ), + what_it_would_take=( + "printing and measuring a test artifact on this machine with this spool, which is the " + "measurement that turns a DESIGNED dimension into a MEASURED one and is the whole reason " + "the distinction exists" + ), + ), + Refusal( + quantity="sterility assurance level for a porous part", + why=( + "surviving the heat and being sterile are different claims. No study located here " + "cultured the interior of an autoclaved fused-deposition part or ran a biological " + "indicator inside its void network, so whether steam penetrates that network to a " + "validated level is unmeasured rather than established either way" + ), + what_it_would_take=( + "biological indicators placed in the worst-case interior geometry, which the medical " + "device guidance on additive manufacturing already defines as the greatest surface area " + "and greatest porosity configuration" + ), + ), + Refusal( + quantity="a probability that an unqualified part is fine", + why=( + "there is no population to draw it from and no prior anybody measured. A number of this " + "shape reaches a report with no author attached to it and is then quoted as though " + "somebody had computed it" + ), + what_it_would_take=( + "the qualifications `unqualified` already lists, performed rather than estimated. Every " + "one of them is a piece of work, and none of them is a number this module can supply" + ), + ), +) + + +def refusals() -> Tuple[Refusal, ...]: + """The estimates this module declines to make, with what each would require. + + Returned rather than only documented so the refusal is checkable. A module that says in its + docstring that it predicts no lifetimes and then exposes a remaining-life helper has + documented a discipline it does not have. + """ + return NOT_ESTIMATED + + +__all__ = [ + "Assessment", + "Benefit", + "Contact", + "CriticalDimension", + "DimensionState", + "FIXTURES", + "Fitness", + "Fixture", + "Flange", + "MATERIALS", + "Material", + "NOT_ESTIMATED", + "Objection", + "POST_PROCESS_EVIDENCE", + "PlateEnvelope", + "PostProcess", + "Process", + "Qualification", + "ResinQualification", + "Sterilization", + "Thermal", + "USES", + "Use", + "fitness", + "materials_reaching", + "materials_surviving", + "plate_envelope", + "refusals", + "survey", + "unqualified", +] diff --git a/docs/PRINTED_FIXTURES.md b/docs/PRINTED_FIXTURES.md new file mode 100644 index 0000000..5fadf80 --- /dev/null +++ b/docs/PRINTED_FIXTURES.md @@ -0,0 +1,960 @@ +# Printed fixtures + +How this lab uses a 3D printer, which printer to buy, and the one boundary that decides +everything downstream. + +A printer is the cheapest deck hardware a lab will ever own and the easiest to misuse. It +is worth having. It is worth having for a narrower set of parts than most people assume, +and the narrowing is not a policy choice -- it falls out of how the parts are made. + +Everything below carries where it came from. Figures are marked `[DS]` when this guide's +research opened the manufacturer datasheet directly, `[DS-2]` when the value came out of a +manufacturer document via search extraction rather than a direct read (one confidence step +lower -- re-verify against the PDF before you act on it), `[EXP]` when it was measured in a +peer-reviewed experiment, `[STD]` when it is clause text from a published standard, and +`[VENDOR]` when it is a vendor claim nobody independent has checked. Numbers with no mark +do not appear. Where a number would be useful and does not exist, this guide says what +measurement would produce it rather than filling the hole. + +--- + +## 1. What printing is for here, and what it is not + +**A printed part may hold labware. It may not be labware.** + +Fixtures, adapters, nests, risers, guards, jigs, tip-box shims, camera mounts, cable +routing, the wedge in §6 -- parts whose entire job is to put a certified consumable in a +known place and keep it there. That is the permitted set, and it is a large and genuinely +useful set. + +Anything a sample, a reagent, or a wash buffer sits in, flows through, or touches is +outside it. Not "discouraged". Outside it. + +### The reason is the void network, not a regulation + +Fused deposition parts are porous by construction, at every setting, including the one +everybody reaches for first. + +- X-ray computed tomography of FDM PLA at **100% infill**, 0.2 mm layers, 210 C nozzle, + measured porosity **4.05% to 6.32%** across raster angles and extrusion widths. No + parameter set reached zero or near-zero. The authors fixed infill at 100% specifically + "to ensure a dense structure" and got 4-6% voids anyway. `[EXP: Wang 2019]` +- The pore size distribution matters more than the percentage. Over **99% of pores were + below 0.2 mm**; the smallest resolved was 38.8 um and the largest 3.60 mm. Critically, + they concentrate in two places: between bonded layers, and **between the outer perimeter + and the infill rasters**. `[EXP: Wang 2019]` + +That second location is the whole argument. The void network is not a set of sealed +bubbles buried in the middle of the part. It is connected to the outside surface. Liquid +and organisms have a path in, and there is no line of sight, no mechanical access, and no +flow path to get them back out. + +The consequences have been measured directly: + +- Immersing FDM ABS parts in medical disinfectants across layer heights of 0.1 / 0.2 / 0.4 + mm and 1 / 2 / 3 perimeters, the authors concluded verbatim that **"no manufacturing + settings can provide enough sealing against fluid intake."** Thinner layers absorbed + less; none stopped it. Saturation was reached at 48 h with 1-2 perimeters and 72 h with + 3. `[EXP: Popescu 2021]` +- **Acetone vapor smoothing does not seal the part.** In the same study, 45 minutes of cold + acetone vapor reduced open porosity, and methylene blue dye still stained the treated + parts after 30 minutes of immersion. The treatment also degraded mechanical properties. + Smoothing changes what you can see, not what you can clean. `[EXP: Popescu 2021]` +- Bacteria preferentially colonize the layer lines. Across eight PLA variants plus + metal-filled, carbon-filled and wood-filled filaments, attachment after 2 h ran to + 6.5x10^6-1.9x10^7 CFU for *E. coli*, 3.6x10^7-6.15x10^7 for *P. aeruginosa*, and + 6.6x10^5-2.16x10^6 for *S. aureus*, with biofilms **thickest between the layers** and + bacteria filling "the valleys" of the layer structure. `[EXP: Hall 2021]` +- Roughness on an FDM part is wildly anisotropic, which is why a quoted Ra can be + meaningless. The same study measured Ra of roughly 0.64-1.15 um *along* the layers, and + **over 100 um** across them. A sub-micron Ra on an FDM part is a measurement of the easy + direction. `[EXP: Hall 2021]` + +FDA's additive-manufacturing guidance does not prohibit any of this -- and reading it +carefully is more useful than a prohibition would be. It says complex additively +manufactured geometries are "expected to increase the difficulty in removing manufacturing +material residues (cleaning) and in sterilization due to the likelihood of increased +surface area", that "sterilization process validation should account for the complex +geometry of your device under worst-case conditions", and that worst case includes the +"combination of largest surface area, greatest porosity". `[STD: FDA 2017 §VI.E]` + +The burden is validation, not prohibition. So: **this guide's research located no published +study validating an FDM-printed part as cleanable to any recognized standard** (ANSI/AAMI +ST98, ISO 17664, ISO 15883). That is absence of evidence from a targeted search rather than +proof that none exists -- but it is the crux, and it did not close. Until it does, the +defensible public claim about an FDM part is *porous by construction, not validated as +cleanable*, and this lab treats it accordingly. + +### The three things people try instead, and what happened when they were tested + +**"Autoclave it."** Surviving a cycle and being clean are different claims. Steam kills +organisms it reaches; it does not remove protein, endotoxin, nucleic acid, or chemical +residue, and it fixes protein onto surfaces. Whether steam penetrates the interior void +network of an FDM part at all is unvalidated -- no study located cultured the interior of +an autoclaved FDM part or ran biological indicators inside the voids. A part can be sterile +and still cross-contaminate the next sample. See §3 for what autoclaving does to the +dimensions, which is a separate disaster. + +**"Coat it."** Vacuum epoxy infiltration is a vendor claim (airtight and watertight to 65 +psi with one specific two-part resin) with no peer-reviewed validation of the coated part +as cleanable, and the same vendor sources concede that internal channels cannot be reached. +`[VENDOR]` A coating also introduces a second, uncharacterized contact material with its +own cure chemistry and leachables, and can trap existing bioburden underneath rather than +removing it. Barrier coatings do have one documented success: a 10 um parylene layer +restored a cytotoxic resin to 93%/85% cell viability -- **and it detached after 5-6 autoclave +cycles**. `[EXP: Kress 2020]` A coated part in a reusable workflow needs a defined +replacement interval, not indefinite reuse. + +**"Use resin instead, it's smooth."** Smoothness is not the problem, chemistry is, and +resin's chemistry is worse. Two Formlabs resins **both marketed as ISO-certified +biocompatible dental materials** were tested in a mouse oocyte maturation assay. Dental LT +Clear caused complete oocyte degeneration in every condition tested, including after UV +post-cure and after oxygen-plasma treatment -- 37.4 +/- 21.3% degenerate at 1 hour, all of +them by 16 hours. Dental SG untreated did the same; plasma treatment rescued gross survival +(75.3 +/- 10.5% reaching MII against a 74.3% polystyrene control) while leaving **57.0 +/- +37.2% with abnormal chromosome morphology against 19.4 +/- 17.3% on control plastic**. +`[EXP: Rogers 2021]` + +Two things in that result generalize. First, a material can pass a survival endpoint and be +genotoxic underneath it, so validating on viability alone misses a whole class of harm. +Second, the causative leachate was identified as **Tinuvin 292, a hindered amine light +stabilizer**, at roughly 50 ug/mL in the medium -- not a monomer and not a photoinitiator. +Post-curing cannot consume it, because it is not a polymerizable species. The standard +"just post-cure it harder" mental model does not cover the thing that actually did the +damage. + +Certification also attaches to a workflow, not to a bottle. Formlabs' BioMed Clear +certification data are tied to a named printer, a 100 um layer height, a 20-minute wash in +**99%** isopropanol, and a 60-minute cure at 60 C. `[DS]` A part made on a different +printer, washed in reused hardware-store 91% IPA, or cured in a nail lamp is not the +article that was tested. And its ISO 10993-3 claim is "not mutagenic" only -- reproductive +and developmental toxicity are endpoints the standard's own risk logic never triggers for a +benchtop lab fixture, which is exactly the gap the oocyte study fell into. + +For anything that would contact gametes or embryos, the applicable qualification is the +Mouse Embryo Assay on the actual finished parts from the actual build (1-cell format, +>=80% reaching expanded blastocyst at 96 h, per FDA guidance), not ISO 10993-5 and not a +vendor certificate. `[STD]` This guide's research found no commercially available +photopolymer resin advertising MEA testing. + +### The line, written as this repo writes lines + +Not shipped code -- nothing imports this. It is the vocabulary the rest of the guide uses, +in the shape the repo uses for boundaries that have to hold. + +```python +class Contact(str, Enum): + """How close a printed part is allowed to get to material. Ordered outward to inward. + + The load-bearing value is SAMPLE, and excluding it is a claim about the process rather + than about any one filament. An FDM part's void network connects to its outer surface + (Wang 2019), no print setting seals it (Popescu 2021), and nothing published validates + such a part as cleanable. A better printer does not move this line. A more expensive + polymer does not move it either -- PEEK survives the autoclave and is just as porous + going in. + """ + + NONE = "none" # structural; nowhere near the fluid path + ADJACENT = "adjacent" # holds certified labware; never wetted in normal operation + SPLASH = "splash" # a spill can reach it; must be wipeable and cheap to replace + SAMPLE = "sample" # sample, reagent, or buffer touches the printed surface + + @property + def permitted_for_printed_parts(self) -> bool: + return self is not Contact.SAMPLE +``` + +A printed nest holding a certified microplate is `ADJACENT` and fine. The same nest with a +well milled into it so a sample can sit directly in the print is `SAMPLE` and is not a +fixture, it is unvalidated labware wearing a fixture's name. + +One more trap worth naming before it costs somebody a plate: **dimensional conformity is +not fitness for contact.** A part can hit every dimension in ANSI/SLAS 1 through 6 and be +entirely unfit to touch a sample. None of those standards says anything about material, +leachables, extractables, cytotoxicity, sterility, or nuclease-free status -- each one's +scope is confined to the single geometric feature in its title. `[STD]` "SBS compliant" is a +statement about a footprint. + +--- + +## 2. Buying the printer + +### The recommendation + +**Buy an enclosed machine with an actively heated chamber and a hardened steel nozzle.** +Concretely: a **Bambu Lab X2D**, list **$649** (captured 2026-07-28). `[VENDOR: official +store spec table]` + +Confirmed specification, from the manufacturer's own spec table: + +| | X2D | +| --- | --- | +| build volume | 256 x 256 x 260 mm main nozzle; 235.5 x 256 x 256 mm dual | +| chamber | enclosed, **Active Chamber Heating supported, max 65 C** | +| nozzle | hardened steel, max **300 C** | +| bed | max **120 C** | +| filtration | G3 pre-filter + H12 HEPA + granulated coconut shell activated carbon | +| price | from $649.00 (captured 2026-07-28) | + +The reasoning, axis by axis. + +**Enclosed and actively heated is the axis that decides what you can print at all.** Every +material worth using for a load-bearing deck fixture -- ABS, ASA, PC -- warps at the sizes +deck fixtures actually are, and warp is driven by the maximum in-plane dimension (§5). A +microplate footprint is 127.76 x 85.48 mm `[STD: SLAS 1 §4.1.1.1]` before you add walls, +so a nest is already a large flat part on day one. A passive enclosure is a box that traps +some waste heat; an actively heated chamber is a controlled variable. Everything else on +the spec sheet is negotiable and this is not. + +**Nozzle temperature is co-limiting, not decisive on its own.** This is the mistake the +spec tables invite. A Prusa CORE One+ with the announced HT hotend reaches 400 C -- higher +than the H2D's 350 C -- while running a 55 C chamber against the H2D's 65 C, which makes it +the *worse* machine for a large warp-prone PC part. Read chamber, bed, enclosure and nozzle +together or you will buy the wrong number. + +**Hardened steel over brass, and this one is a contamination argument rather than a wear +argument.** The Prusa CORE One+ ships a brass nozzle (`High-flow Prusa Nozzle brass CHT - +0.4 mm`); brass alloys commonly contain lead. The X2D, P2S, H2D and QIDI Plus4 all ship +hardened steel. For a part that will sit on a deck next to open labware, that distinction +matters more than any temperature spec on the page. + +**Filtration is an operator-exposure control and should be read as exactly that.** ABS, ASA +and PC printing emits styrene and ultrafine particles. HEPA and activated carbon reduce +what the person in the room breathes. They do **not** make the chamber a clean environment +and they are not a sterility control -- the parts are printed in ordinary room air and +carry bioburden from manufacture, which is precisely why FDA's guidance treats the +cleanliness of build material and build environment as the control for interiors that +cannot be flushed. Site the printer somewhere samples are not open, whatever filters it has. + +### Alternatives, with the tradeoff stated + +| machine | chamber | nozzle | bed | build volume | price (captured 2026-07-28) | why you would pick it | +| --- | --- | --- | --- | --- | --- | --- | +| **Bambu X2D** | enclosed, **active 65 C** | hardened steel, 300 C | 120 C | 256 x 256 x 260 | from $649 | the default: heated chamber, steel nozzle, HEPA + carbon, lowest verified price in the class | +| **QIDI Plus4** | enclosed, **active 65 C** (PTC + circulation fan) | 370 C | 120 C | 305 x 305 x 280 | $649 sale / **$799 list** | biggest envelope and highest nozzle ceiling of the affordable set. Pick it when a fixture will not fit in 256 mm | +| **Bambu H2D** | enclosed, **active 65 C** | hardened steel, 350 C | 120 C | 325 x 320 x 325 single-nozzle | $1,549 sale / **$1,749 list** | more envelope and 350 C. Note 1320 W at 110 V -- a real circuit-loading question on a shared bench -- and a 10-30 C stated working range that rules out some siting | +| **Prusa CORE One+** | enclosed, **active 55 C** | brass CHT 0.4, 290 C stock | 120 C | 250 x 220 x 270 | see price note | pick it when repairability dominates. Coolest chamber of the enclosed set and a brass nozzle out of the box | +| **Bambu P1S / P2S** | enclosed, **passive** | 300 C | 100 / 110 C | 256 x 256 x 256 | P1S $399 sale / $699 list; P2S from $549 | cheapest way into an enclosure. Buy only if PLA and PETG are genuinely all you will ever print | +| **Prusa MK4S** | **not enclosed** (add-on) | 290 C | 120 C | 250 x 210 x 220 | see price note | do not buy for this job without the enclosure. ABS, ASA and PA are conditioned on the add-on in Prusa's own material list | +| **Intamsys Funmat HT** | enclosed, **active 90 C** | 450 C | 160 C | 260 x 260 x 260 | quote only | the only machine here in PEEK/PEKK territory. See the PEI caution below | + +Four cautions that will otherwise cost you money. + +**The Bambu X1C is end-of-life.** Manufacturing and active sales ended 2026-03-31; firmware +feature updates run to 2027-05-31, security patches to 2029-05-31, spare parts and support +to 2031-03-31. The product page redirects to the general listing. It is still the machine +most third-party guides recommend. Do not buy one new, and treat every X1C figure you find +online as unverifiable legacy -- this guide could not retrieve a single X1C specification +from a Bambu-controlled source. + +**Bambu's own site contradicts itself on the P series.** The store's buying-guide prose says +"an enclosed chamber with active heating (P series or X2D)". The P2S product page's own FAQ +says verbatim: "The P2S does not have an active chamber heating function." The P1S spec +table has no chamber-heating row at all, only a regulator fan and a carbon filter. The FAQ +and the spec tables are right and the buying guide is wrong. The marketing phrase "50 C +Chamber Ready" carries no test conditions, no tolerance and no ambient reference, and Bambu +publishes no chamber temperature for the P1S at all -- so any specific passive chamber +number, including one you might be tempted to write into a design note, would be invented. + +**"PEI plate" is not "prints PEI", and this collision will produce a real error.** On Bambu +and Prusa pages, "Textured PEI Plate" and "PEI spring steel sheet" name the polyetherimide +*build surface* the part is printed on. PEI/ULTEM as a printable engineering polymer needs +roughly 360-420 C at the nozzle plus a hot chamber. **No consumer or prosumer machine here +reaches it.** Confirmed chamber ceilings: P1S none (passive), P2S none (passive), X2D 65 C, +H2D 65 C, H2C 65 C, MK4S none as standard, CORE One+ 55 C, CORE One L 60 C, QIDI Plus4 +65 C. Only the Funmat-class machine (90 C chamber, 450 C nozzle) is in that territory -- +and Intamsys' own material list on its product page names PEEK, PEKK, PC, PPS, PPA and PA +variants and **does not name PEI or ULTEM**. Reseller copy asserts ULTEM capability. +Reseller copy is not a specification. + +**The Prusa HT hotend is announced, not in hand.** The CORE One+ page labels it "Up to +400 C (Coming Soon)" and the product page says "Estimated to ship this summer" at $184.26. +Do not build a purchase case on a 400 C capability nobody has yet received. + +### Price note, and why this guide will not print a total + +Prices captured 2026-07-28, and several are promotional -- the P1S at $399 against a $699 +list is a 43% discount, and Bambu advertises 30-day price protection, which implies active +price movement. Use list prices with the capture date. Sale prices in a document are stale +within weeks and read as a specification later. + +Prusa's figures did not resolve cleanly and are reported both ways rather than averaged. +Prusa's own comparison table gives assembled prices of **CORE One+ from $1,299**, **MK4S +from $999**, **CORE One L from $1,799**. A text render of the same pages returned CORE One+ +assembled **$1,202.78** / kit **$925**, and MK4S assembled **$925** / kit **$657.40**. +Whether the lower set is a live promotion, a VAT-excluded view, or a currency artifact could +not be determined. Do not pick one and do not split the difference. + +Also check voltage: the units above are the 100-120 VAC US variants, and Bambu explicitly +warns to buy the version matching your region. + +**The first-year requisition.** The printer is the only line on it this guide can price +honestly, and it is not the line that decides whether you get qualified parts. + +``` + printer $649 list (X2D) / $799 list (QIDI Plus4) [sourced, 2026-07-28] + filament by the kilogram, several materials PRICE LOCALLY + spare hardened nozzles consumable; buy before you need one PRICE LOCALLY + drybox + desiccant not optional -- see below PRICE LOCALLY + metrology calipers at minimum; NOT OPTIONAL PRICE LOCALLY + spare build sheet consumable PRICE LOCALLY +``` + +This guide does not publish a first-year total, and the reason is the same rule the rest of +the repo runs on: every non-printer figure it could have written is one it did not verify, +and an unmeasured number that reaches a document becomes a budget nobody measured. Price +those five lines on the day you order and put the **sum** on the requisition. What this +guide will assert is which lines cannot be dropped: + +- **Metrology is not optional and it is not an accessory.** The qualification ladder in §4 + has a MEASURED rung, and a lab with a printer and no measuring instrument physically + cannot reach it. It can only produce objects. Calipers are the floor; gauge pins or + blocks for the features that actually register are better. +- **The drybox is not optional either, and the datasheets say why.** Saturation water + absorption: PETG 0.51% max `[DS]`, PEEK 0.45% at 23 C and 0.55% at 100 C `[DS]`, ULTEM + 1010 1.25% `[DS]`. Nylon is worse -- strongly hygroscopic, and moisture both swells the + part and plasticizes the polymer, dropping its Tg. (This guide's research did not obtain a + numeric moisture-uptake figure for a PA filament, so the mechanism is stated and the + number is not.) Wet filament prints badly and prints *differently* on different days, + which quietly destroys the repeatability that §5 depends on. + +### The two axes this guide could not source + +The brief for this section named four decisive axes. Two of them -- **reliability and the +time cost of fussing**, and **repairability** -- have no verifiable public numbers, and the +recommendation above rests only on the two that do. Saying so is more useful than +laundering a forum consensus into a specification. + +What would settle them: + +- **Fuss.** Your own maintenance log. Hours of intervention per successful part over the + first fifty parts, recorded as it happens. Nobody else's log transfers, because it is + dominated by the material mix and the part geometry you personally print. +- **Repairability.** Two things a buyer can check before ordering, both of which are vendor + behavior rather than opinion: does the vendor publish spare-part availability with dates, + and does it publish part drawings and firmware. Both vendors here have a concrete data + point on record. Bambu published a dated end-of-support schedule for the X1 series with + spare parts guaranteed to 2031-03-31. Prusa offers a **$10** upgrade kit to take a CORE + One to CORE One+ spec, or the upgrade parts as files you print yourself. Weigh those as + evidence of intent; neither is a reliability measurement. + +--- + +## 3. Materials + +Read the polymer-class column last. **Grade names are not material specifications.** +"ULTEM" spans 9085 at Tg 177 C and 1010 at Tg 210-217 C -- a 33-40 C spread under one trade +name. "PC" spans unfilled resin near Tg 147 C and a printable PC Blend at HDT 93 C under +1.8 MPa. "PP" spans homopolymer and a copolymer melting at 137 C. Check the specific +product's datasheet, every time. + +### The gate is Tg, not HDT + +Heat deflection temperature measures 0.25 mm of deflection in a standard bar under an +applied load of 0.45 or 1.8 MPa. An autoclaved fixture usually carries no external load, so +the natural inference is that it can safely exceed its HDT. **That inference fails for +printed parts specifically.** An FDM part contains frozen-in extrusion stresses that relax +the moment the polymer passes its glass transition, so it shrinks and warps with zero load +applied. Use Tg as the conservative gate. This is why PLA was observed warping from +approximately 70 C in a curved-beam geometry `[EXP: Chiscop 2025]` despite HDT figures near +57 C being quoted for it. + +Two further corrections to how these numbers get read. **Orientation changes the answer** -- +one manufacturer's TMA data on printed ULTEM 9085 shows dimensional reversal beginning at +~175.9 C upright versus ~193.4 C flat, a 17 C spread from print orientation alone on one +material. `[DS]` And **fillers inflate HDT without raising the polymer's temperature limit** +-- carbon-fiber nylons show HDT sitting 100-140 C above Tg because the fiber and the +crystalline phase carry the load, not because the matrix became more thermally stable. +Chemical resistance, hydrolysis behavior and creep still track the base polymer. + +### The table + +Autoclave columns compare each polymer's Tg against the cycle. 121 C cycles typically run +15-30 min; 132-134 C cycles typically 3-4 min. Exposure time matters for creep, so a +pass/fail is only meaningful next to a cycle. + +| material | Tg | HDT (load, method) | 121 C | 134 C | chemical and process notes | contact | +| --- | --- | --- | --- | --- | --- | --- | +| **PLA** | 54-61 C `[DS]` across four grades from one maker (Basic 60, Matte 61, Translucent 54, Tough+ 61) | 52-58 C @1.8 MPa; 57-61 C @0.45 MPa, ISO 75 `[DS]` | **FAIL** | **FAIL** | Ester backbone -- hydrolyzes in wet heat. Warps from ~70 C in curved geometries `[EXP]`. Do not publish a single PLA figure; the 7 C spread across one maker's own grades is the honest measure of what colorant packages do | `NONE`, `ADJACENT` | +| **PETG** | 69-77 C (69 C one maker `[DS]`, 71.24 C another `[DS]`, 77.4 C a third `[DS-2]`) | 68 C @1.8, 71 C @0.45 `[DS]`; 65 C @1.8, 69 C @0.45 `[DS]`; 76.2 +/- 0.8 C @0.455 `[DS-2]` | **FAIL** | **FAIL** | ~8 C maker-to-maker gap on nominally the same polymer. Ester backbone, hydrolysis-sensitive. Max water absorption 0.51% `[DS]` | `NONE`, `ADJACENT`, `SPLASH` | +| **ABS** | ~105 C (105 C `[DS]`, 105.2 C `[DS-2]`); one maker lists Tg as "N/A" | 84-104 C depending entirely on grade: 104.4 C @66 psi industrial `[DS-2]` vs 84 C @1.8 MPa desktop `[DS-2]` vs 95 C @0.45 MPa `[DS]` | **FAIL** | **FAIL** | The ~20 C industrial-vs-desktop HDT gap is real and decides nothing good. An ABS housing was independently observed to deform at 121 C "to the extent that it can no longer be used" `[EXP]`. Note the maker of a biocompatible medical ABS claims gamma and EtO sterilization and **not** steam | `NONE`, `ADJACENT`, `SPLASH` | +| **ASA** | 98 C `[DS]`; ~104 C `[DS-2]`; a third datasheet gives none | 86 C @1.8, 93 C @0.45 `[DS]`; 100 C @1.8, 103 C @0.45 `[DS]`; 98.2/103 C @264/66 psi `[DS-2]` | **FAIL** | **FAIL** | The practical choice for deck fixtures that see UV or long service. Better outdoor stability than ABS, similar thermal ceiling | `NONE`, `ADJACENT`, `SPLASH` | +| **PC (printable blend)** | not given on the blend datasheet | **113 C @0.45, 93 C @1.8** `[DS]` | **MARGINAL** | **FAIL** | This is the row that misleads. A guide quoting "PC: HDT 130 C" is about 37 C optimistic for the blend somebody will actually buy | `NONE`, `ADJACENT`, `SPLASH` | +| **PC (unfilled resin)** | ~147 C `[DS-2]` | ~124-126 C @1.80, ISO 75 `[DS-2]`; Vicat B/50 ~147 C | pass on thermal grounds | pass on thermal grounds | Grade-specific figures must come from the specific grade sheet -- the source pages for these returned 404/500. Carbonate linkages hydrolyze in wet heat; the molecular-weight-loss and crazing mechanism is well known in device practice and this guide could not source it properly, so it is flagged rather than asserted | `NONE`, `ADJACENT` | +| **PP** | not on any filament datasheet reached. PP's Tg is below room temperature but no manufacturer source confirmed it -- leave it out rather than guess | none published on the filament TDS | **MARGINAL** | **DO NOT** | The single most dangerous row. One PP copolymer filament has **Tm 137 C** `[DS]` -- a 134 C cycle runs 3 C below its melting point. Homopolymer melts ~160-165 C. Do not carry the reputation of molded PP labware, which is routinely autoclaved, onto PP filament without checking that grade's Tm | `NONE`, `ADJACENT` | +| **PA / nylon (CF-filled)** | 70 C (PAHT-CF), 85 C (PPA-CF) `[DS-2]` | 170 C @1.8 / 194 C @0.45 (PAHT-CF); 196/227 C (PPA-CF) `[DS-2]` | see notes | see notes | **The HDT column is a trap here.** It sits 100-140 C above Tg because carbon fiber carries the load. Separately, PA is strongly hygroscopic and a steam cycle drives it toward saturation -- moisture swells the part *and* plasticizes the polymer, dropping unfilled nylon's Tg substantially. It will move even if it never softens | `NONE`, `ADJACENT` | +| **PEI / ULTEM 9085** | **177.3 C** `[DS]` | printed 178.2 C (XY) / 178.4 C (XZ) @66 psi; 170.2 / 172.6 @264 psi `[DS]` | pass | pass | Not printable on any machine in §2. TMA reversal at ~175.9 C upright vs ~193.4 C flat -- print orientation, not polymer, sets deformation onset | `NONE`, `ADJACENT` | +| **PEI / ULTEM 1010** | **210 C** printed `[DS-2]` / **217 C** molded `[DS]` | 200 C @0.45, 190 C @1.8 (ISO 75) `[DS]`; 215/210 C @66/264 psi (ASTM D648) `[DS-2]` | pass | pass | The ISO-versus-ASTM gap of 15-20 C on nominally identical material is why every figure here carries its method. Water absorption 1.25% at saturation `[DS]` | `NONE`, `ADJACENT` | +| **PEEK** | onset **143 C**, midpoint 150 C `[DS]` | DTUL **152 C** @1.8 MPa unannealed `[DS]`; Tm 343 C | pass | pass, thinly | Margin at 134 C is 9 C on Tg onset. It passes because the **crystalline phase** carries load above Tg -- and as-printed PEEK from a machine without adequate chamber temperature can be substantially amorphous, in which case it is not the material the datasheet describes. The maker's "suitable for steam sterilisation" statement covers injection-molded granules, not printed parts | `NONE`, `ADJACENT` | + +### What the autoclave column does not mean + +**"Survives the cycle" is a much weaker claim than "holds its dimensions."** The only +measured dimensional data this guide's research located for any polymer on that table is +for annealed PLA at 134 C for 60 min: after annealing at 120 C for 60 min encapsulated in +silicone, a hollow cylinder moved +0.47% on outer and inner diameter and +0.18% on height, +and a rectangular bar moved +0.64% height, **-1.43% width**, +0.12% length. Salt-encapsulated +annealing was markedly worse -- up to -2.50% on inner diameter and +2.10% on bar height. +`[EXP: Chiscop 2025]` + +Read that correctly. It supports *annealed PLA can survive a cycle*. It does not support +*PLA autoclaves without dimensional change* -- the best case still moved 0.2-1.4%, and the +movement is anisotropic, with Z growing while XY shrinks. **If a fixture's function depends +on a tolerance tighter than roughly 1%, no datasheet Tg or HDT will tell you whether it +survives.** That has to be measured on the part. + +The same paper explicitly did not measure crystallinity, porosity, cleanability or +sterility. It establishes thermal survival and nothing else. + +### Four more things not to write down + +- **No cycle count.** Every specific claim located about how many autoclave cycles PEEK or + PEI survives -- "1,000+ cycles", "<0.5% over 3000 cycles", "+/-0.2% for 50+ cycles" -- + traced to content-marketing or AI-generated aggregator pages, not to a study or a + manufacturer document. Repeated cycling is a cumulative creep-and-hydrolysis problem. The + honest statement is that per-cycle data exists and multi-cycle data was not located. +- **Pressurized cycles need vented geometry.** A 134 C cycle runs at roughly 2.1 bar. Sealed + or trapped-volume prints -- closed hollows, capped tubes, any unvented internal cavity -- + can deform or collapse from the pressure differential, especially during exhaust, + regardless of the polymer's thermal capability. +- **Datasheet values describe molded or specially-printed test bars, not your part.** Both + the PEEK and the ULTEM 1010 figures above describe injection-molded material and carry + explicit disclaimers that real properties depend on geometry and processing. +- **Exclude filled, metal-filled, wood-filled and antimicrobial filaments** from anything + near labware. They are worse on every axis that matters here: higher roughness, + filler-matrix interfaces, additional leachable species. Copper and silver antimicrobial + filaments are not a cleanability control -- they leach metal ions, which is a direct + problem for cell culture and for any assay sensitive to divalent cations, and bacteria + attached to metal-filled PLA at the same 10^6-10^7 CFU order as plain PLA. `[EXP: Hall + 2021]` + +One genuinely favorable finding, recorded because omitting it would be its own dishonesty: +plain FDM thermoplastics are generally **not** cytotoxic in cell contact. ABS, PETG, PLA and +Nylon 12 did not reduce human iPSC viability against control in one study, while two SLA +resins reduced it by roughly 60% and 90%. `[EXP]` That is real, and it does not license +sample contact, because cytotoxicity was never the reason for the boundary in §1 -- porosity +was. It is also not uniform: a preprint reports significant cytotoxicity for PETG and PC, +and colorants, plasticizers and processing aids differ between spools of nominally identical +polymer and are not disclosed on consumer filament. + +--- + +## 4. The qualification ladder + +The repo already refuses to let an instrument's reputation transfer to a step: a federated +step is supervised only when a run card for *that step* has been proven, and an +unbenchmarked operation is untrusted by default. A printed part gets the same treatment, +for the same reason. + +```python +class Rung(str, Enum): + """How far a printed fixture has been qualified. Ordered weakest to strongest. + + NO RUNG IMPLIES THE ONE ABOVE IT. That is the whole content of this enum and it is not + a formality -- each rung tests a different physical claim, and the claims are + independent. A model that compiles says nothing about a print. A print says nothing + about its dimensions, because nobody measured them. Dimensions say nothing about fit, + because the plate has tolerances of its own. Fit says nothing about the arm. + + The last of those is the one that hurts, so `qualified_for_arm` is a property rather + than a comment. A hand is compliant and adaptive: it feels a part binding and stops. An + arm is neither. It goes to a commanded position at a commanded force and finds out what + is there afterwards. A part that fits by hand has been qualified for a hand. + """ + + MODELED = "modeled" # the source compiles and renders; no object exists + PRINTED = "printed" # an object exists; nobody has measured it + MEASURED = "measured" # the registering features were measured, at the working temperature + FITTED = "fitted" # it seats on the deck with the real plate, robot powered down + DRY_RUN = "dry_run" # the arm ran the motion against it at reduced speed, no material + IN_USE = "in_use" # it has run with material, and the run was recorded + + @property + def qualified_for_arm(self) -> bool: + return self in (Rung.DRY_RUN, Rung.IN_USE) + + @property + def measured(self) -> bool: + """False for MODELED and PRINTED. An unmeasured part is untrusted, matching how an + unbenchmarked operation is untrusted: the absence of a measurement is not evidence + of adequacy.""" + return self is not Rung.MODELED and self is not Rung.PRINTED +``` + +### What each rung actually requires + +**MODELED.** The `.scad` compiles. Record the parameter values, because the next print will +not be from the same source unless you do. + +**PRINTED.** An object exists. This is the rung everybody stops at, and it carries no +information about the object beyond "the printer did not fail". + +**MEASURED.** The features that *register* -- the ones that decide where the plate ends up +-- were measured with a real instrument and written down against the model's nominals. Four +requirements that are easy to skip: + +- *Measure at the temperature the part will work at.* SLAS dimensions are specified at + 20 C `[STD: SLAS 1-4 §1.2]`, and ANSI/SLAS 6's test method at 25 C +/- 2 C -- so there is + not even one temperature across the standard family. Printed polymers have far higher + thermal expansion and lower creep resistance than molded PP or PS. A part in tolerance on + the bench can be out of tolerance at 37 C in an incubator. +- *Do not trust the bottom face.* The first layer is squished against the bed and comes out + wider than modeled -- PrusaSlicer ships `elefant_foot_compensation = 0.2` in several of + its own profiles, and Prusa's documentation puts values around 0.2 mm as typical for a + 0.4 mm nozzle. The bottom few layers are the least dimensionally trustworthy region of + the part, and they are also the region that seats on the deck. Put registration datums + somewhere else, or account for it explicitly. +- *A calibration cube proves nothing about a nest.* Tolerance scales with length in every + formulation anyone publishes. Desktop FDM is quoted at **+/-0.5% with a +/-0.5 mm floor**, + industrial FDM at +/-0.15% with a +/-0.2 mm floor. `[VENDOR: Protolabs Network]` An + industrial machine's own spec sheet reads "+/- .200 mm (.008 in), or +/- .002 mm/mm, + whichever is greater" -- a +/-0.2 mm floor rising to +/-0.4 mm across a 200 mm feature. + `[DS]` Express any tolerance you claim as +/-(floor) or +/-(percent x dimension), + whichever is larger, and state the reference length. +- *Exploit repeatability, which is far better than accuracy.* One study of 35 printed + cuboids reported repeatability standard deviations in the few-hundredths-of-a-millimeter + range. `[EXP, low confidence -- retrieved via search snippet; the full text was + unreachable, so treat the order of magnitude as indicative and the digits as unverified]` + The practical consequence is real either way: an FDM printer reproduces its own error + consistently, so **measure and compensate** works even when out-of-the-box accuracy is + 5-10x worse than repeatability. That loop is the only route to a part that registers + properly, and it needs the calipers from §2. + +**FITTED.** The real plate -- the one from the lot you will actually run, not a +dimensionally different one from a different vendor -- seats in the part, on the deck, in +its real position, with the robot powered down. Check it seats fully, releases without +binding, and does not rock. Then check it with the plate at working mass, full of liquid, +because an empty plate and a loaded plate sit differently. + +**DRY_RUN.** The arm executes the real motion against the part at reduced speed, with no +material anywhere. This is a distinct rung because everything the hand did for you in +FITTED is now absent: approach vector, clearance to the gripper jaw or the pipetting head, +what happens when the plate is a half-millimeter off nominal, whether the part moves when +the arm nudges it. Reduced speed exists so a collision is a scrape rather than a repair. + +**IN_USE.** It has run with material and the run was recorded. Recorded means an event with +an attestation, in the sense the repo already uses: an instrument confirmed it, or a human +witnessed it. Software having sent the command is `asserted` and is a log line about intent. + +### The refusal + +Do not report a part at a rung it has not reached, and in particular do not let FITTED read +as DRY_RUN. **A part that fits by hand has not been qualified for an arm.** The failure mode +in §5 that puts an instrument at risk -- a part shifting mid-run -- is invisible from every +rung below DRY_RUN, and a printed fixture is exactly the kind of hardware that gets promoted +straight from "I fitted it, looks great" to a production run. + +--- + +## 5. The failure modes + +Six concrete ways a printed fixture fails, each with what it does and what would catch it. + +### Warp on a large flat part, so the base rocks + +Warp is driven by the maximum in-plane dimension and by the temperature gradient during the +build -- non-uniform cooling and differential volumetric shrinkage of the extruded polymer, +plus poor bed adhesion. `[VENDOR: Protolabs Network; EXP: warpage study]` A microplate +footprint is 127.76 x 85.48 mm before you add anything `[STD]`, so a nest is a large flat +part by default. + +The magnitude is millimeters, not microns. One study printing ABS with the bed at 110 C +reported warpage "around 3.7 mm" before optimization and "around 0.8 mm after improvement". +`[EXP]` **Those numbers cannot be scaled to your part** -- the retrievable text does not +state the specimen dimensions or the measurement instrument -- so use them for the order of +magnitude and nothing else. A general "warp per 100 mm of length" rule does not exist in any +dataset this guide's research located, and anything of that form would be an estimate +dressed as a finding. + +There is also a floor you cannot design under: **a printed plate cannot be flatter than the +bed it was printed on.** Mesh bed leveling conforms the first layer *to* the bed's shape +rather than correcting it, so bed non-flatness transfers into the part's bottom face. One +measurement project reports peak-to-valley bed variation of a quarter of a millimeter -- +more than two layers at 0.1 mm. `[single anecdote, printer unspecified]` No manufacturer bed +flatness specification for the machines in §2 could be confirmed; a "0.10 mm" figure +circulating for one of them traced to forum discussion, not to a spec sheet. + +*What it does:* the base rocks, so the plate's true position depends on which corner is +loaded. *Catches it:* MEASURED, if you measure flatness rather than only length and width; +FITTED, if you check for rock with the plate at working mass. Nothing below that. +*Mitigations that are actually supported:* heated bed, rafts, radii at sharp corners +`[VENDOR]`; lower-shrinkage material (see below); and designing the part so the registering +surface is not a large unbroken flat. + +### Shrinkage moves a registration feature out of tolerance + +Filament manufacturers largely **do not publish shrinkage**. This guide's research opened +three technical datasheets -- a PETG, an ASA, and an ABS -- and none of them lists a +shrinkage figure at all. `[DS]` So every per-material shrinkage percentage circulating +online is a slicer default or a community measurement, not a manufacturer specification. + +What can be sourced is what the slicers ship. In one slicer's repository, one filament +vendor's profiles set compensation implying measured shrinkage of **PLA 0.05%, PETG 0.15%, +ABS 0.513%, ASA 0.513%, PC-CF 0.15%**. In the *same repository*, a different vendor ships +100% -- zero compensation -- for PLA, PETG, ABS and ASA, and 99.8% for PC-CF. Two vendors, +one slicer, a 0.5% disagreement on ABS. `[shipped code, read directly]` Independent +knowledge-base material puts FDM shrinkage in the **0.2-1%** range `[VENDOR]`, consistent at +the low end. + +The arithmetic is what makes this a design problem rather than a curiosity. Across the +127.76 mm long dimension of a plate footprint: + +``` + PLA 0.05% x 127.76 mm = 0.06 mm + PETG 0.15% x 127.76 mm = 0.19 mm + ABS 0.513% x 127.76 mm = 0.66 mm +``` + +That last figure is larger than the entire SLAS 1 corner-zone tolerance band of +/-0.25 mm. +An ABS nest modeled at nominal, printed without shrink compensation, is out of the plate +standard's tolerance before the plate arrives. + +Reject two figures you will find repeated everywhere: "ABS shrinks up to 11%" and "PLA 0.2% +to 3%". Neither is supported by any manufacturer datasheet or slicer default located, and +11% is off by more than an order of magnitude from the 0.513% a real shipped profile +compensates for. They appear to conflate free volumetric contraction and injection-molding +shrink with printed-part shrink. + +*What it does:* a pin, slot or shoulder ends up outside the clearance you designed, and the +plate either binds or floats. *Catches it:* MEASURED, on the registering feature +specifically -- overall dimensions can be right while a feature is not. *Mitigation:* +measure and compensate, exploiting repeatability, and re-measure after any filament, color +or lot change, because the additive package moves the number. + +### The part shifts mid-run and causes a crash + +A fixture that is friction-held or double-sided-taped will eventually move: an arm nudge, +a thermal cycle, a wipe-down with solvent. When it moves, the deck position the robot was +taught no longer describes where the plate is. + +*What it does:* severity is **mechanical** -- the instrument itself is at risk, which puts +it in a different class from every other failure here. It is also **silent by construction**: +nothing in a printed part reports its own position, and the arm has no way to know the +fixture is not where it was taught. The first evidence is the collision. + +*Catches it:* DRY_RUN catches a fixture that moves under arm contact, which is exactly why +that rung exists. Nothing else does, and in particular a camera does not, unless a validated +check with a known pose and measured sensitivity exists -- which in this workcell it does +not, for any condition. + +*Mitigation is design, not procedure:* positively locate the fixture into existing deck +features -- bolt it, key it, capture it between rails. A registration that depends on +somebody putting it back in the same place is not a registration. And re-run DRY_RUN after +any event that could have moved it. + +### Static + +Printed thermoplastics are insulators. Charge builds and then does things you did not plan: +lightweight labware sticks or jumps, tips cling, powders migrate, and on a bad day an +electrostatic discharge reaches something that minds. + +**This guide has no measured figure for it and will not invent one.** What would produce one +is a surface and volume resistivity measurement, per a named test method, **on the printed +part in its printed orientation** rather than on the pellet -- and consumer filament +datasheets do not carry that value. + +Two things are worth saying without a number. Humidity control changes the behavior +substantially, and the drybox you bought in §2 is already half of that story. And the obvious +fix -- carbon-filled "ESD-safe" filament -- is excluded here for the reasons in §3: filled +grades bring higher roughness, filler-matrix interfaces and extra leachable species to a +part that sits next to open labware. + +### A lip too shallow to retain the plate at angle + +This is the failure that ends §6's worked example if the design is careless, and it is a +standards problem more than a geometry problem. + +**The flange the lip grips is not one thing.** ANSI/SLAS 3-2004 offers **five mutually +exclusive variants**, and requires the plate to declare which one it meets: §4.1 Short = +2.41 mm +/- 0.38, §4.2 Medium = 6.10 mm +/- 0.38, §4.3 Tall = 7.62 mm +/- 0.38, §4.4 Short +with interruptions = 2.41 mm +/- 0.38, §4.5 Dual = 2.41 mm on the short sides and 7.62 mm on +the long sides. `[STD]` A lip sized against a tall 7.62 mm flange will not reliably retain a +2.41 mm short-flange plate. "SBS compliant" on a vendor page does not tell you which variant +you are getting. + +*What it does:* the plate slides, tips, or lifts off the nest at angle, which at best aborts +the step and at worst puts liquid on the deck and a plate under an arm. + +*Catches it:* not a calculation. This guide has no friction coefficients for a printed +surface against a molded plate base, so any formula for required lip height as a function of +angle would be a fabricated number wearing an equation. **The measurement that settles it** +is the direct one: the actual plate, at the actual maximum angle, at working mass with +liquid in it, on the actual printed surface, tilted past the design angle until it moves. +Record the angle at which it moves. Do that for every plate type and every flange variant +the fixture is supposed to hold, and repeat it after any material change, because surface +finish drives friction and finish varies by filament and by layer height. + +### A stack height that puts the plate out of Z range + +A riser adds its height to the plate's, and the plate's height is not the number most people +have in their head. + +- ANSI/SLAS 2-2004 specifies **14.35 mm +/- 0.25 mm** from the resting plane to the maximum + protrusion of the perimeter wells, and 14.35 mm +/- 0.76 mm overall -- **for a typical + microplate**. `[STD]` (Note the standard's own Figure 1 misprints the inch equivalent as + 0.56560 in against the correct 0.5650 in in the clause text. Use the metric value.) +- Plate types that are not typical microplates simply do not comply with SLAS 2 while still + complying with SLAS 1, 3 and 4. A 96-well 1000 uL deep-well plate is documented at length + 127.8 mm and width 85.5 mm -- SLAS 1 conformant -- with a height of **44.1 mm**, roughly + 3x the SLAS 2 figure. Another vendor's 2.2 mL deep-well plate lists 44 mm and advertises + only an "ANSI-SBS Footprint". `[DS]` + +So the accurate framing is: **height is standardized only for standard-height microplates; +other formats keep the footprint and abandon the height.** Saying flatly that height is not +standardized is wrong, and so is assuming 14.35 mm. And ~44 mm is not a standard value +either -- those are two individual products' specifications, and other vendors will differ. + +*What it does:* the pipetting head or gripper cannot reach the well bottom, or cannot clear +the plate top, or the deck position is simply outside the Z envelope. Discovered during +DRY_RUN if you are careful and during production if you are not. + +*Mitigation:* compute the stack rather than assuming it. Riser height + the specific plate's +measured height + tip length + the clearance the head needs above the plate, against the +instrument's stated Z range, for every plate format that will ever sit on that fixture. If a +deep-well plate can land there, size for the deep-well plate. + +### The plate is not the nominal plate + +Four more clauses that will bite a nest designed to nominal dimensions. + +- **There are two footprint tolerances, not one.** +/-0.25 mm applies **only** within 12.7 mm + of the four outside corners `[STD: SLAS 1 §4.1.1.1]`. Anywhere else along the side the + tolerance is **+/-0.5 mm** `[STD: §4.1.1.2]`. A nest cut to "127.76 x 85.48 +/- 0.25" as a + flat statement will jam real plates that bow mid-side and remain fully conformant. +- **Corner radius is 3.18 mm +/- 1.6 mm** -- a permitted range of 1.58 to 4.78 mm, a 3x + spread, scoped to the bottom flange corners specifically. `[STD: §4.1.2.1]` Design + clearance features against the **maximum** 4.78 mm, not the nominal, or you will bind on + plates at the top of the range. +- **Draft is excluded from the standard's numbers.** Every figure note states "Dimensions + and tolerances do not include draft." `[STD]` Molded plates carry draft angles the + standard's dimensions do not describe, and the actual angles are not published anywhere in + it. A printed nest with vertical walls cut to nominals does not have the same real + cross-section as the molded plate it is supposed to seat. +- **SLAS 2 has two alternative compliance parts** and a plate must declare which. §4.1 adds + a minimum 1 mm clearance from the resting plane to the bottom external surface of the + wells; §4.2 does not. `[STD]` If a fixture assumes 1 mm of clearance under the wells -- for + bottom-reading optics, or a heat block -- a §4.2-compliant plate is not required to give + it to you. + +And one that saves a wasted investigation: **ANSI/SLAS 6-2012 sets no limits at all.** Its +§7 states that the standard "specifies definitions and a test method only" and that it is +"not the intent of this standard to state a limit" for well bottom elevation or its +variation. `[STD]` "SLAS 6 compliant" conveys no flatness or bottom-thickness guarantee +whatsoever. For any optical path, the instrument's own specification is the requirement. + +One FDM-specific conformity note worth designing around: SLAS 1 §4.1.1.3 requires the +footprint be "continuous and uninterrupted around the base of the plate." Brim remnants, +elephant-foot bulges and support scars on a printed part's base are the same class of defect +in reverse -- they snag stage nests and gripper jaws. + +--- + +## 6. Worked example: `hardware/tilt_module.scad` + +A printed wedge that presents a plate at a fixed angle on the deck, so that residual liquid +pools toward one side of each well and more of it comes out at the aspiration step. It is +`ADJACENT`: it holds a certified plate and nothing touches it. + +**This guide states no recovery-improvement figure for the tilt module, because none has +been measured.** The module's design intent is that recovery improves. Intent is not +evidence. In the repo's own vocabulary, "the tilt module improves recovery" currently rests +on basis `intuition` -- a scientist believes it, nothing is written down -- and it stays +there until the experiment below is run and recorded. It cannot become `in_house` any other +way. + +### What the ladder says about it today + +At best it reaches **MEASURED**. The wedge angle and the lip that retains the plate can be +measured against the model. Everything above that is unproven: + +- **FITTED** needs the real plate, at working mass, at the design angle, checked for rock + and checked for retention -- and per §5 that check has to be done separately for each + flange variant it will hold. +- **DRY_RUN** needs the arm to run the real aspiration approach against a tilted plate, + at reduced speed. Tilting changes the geometry the head was taught: the well bottom is no + longer perpendicular to the approach, the clearance to the plate's high side shrinks, and + the stack height went up by the wedge. Every one of those is a fresh collision opportunity + and none of them is visible from FITTED. +- **IN_USE** needs a recorded run with material. + +The wedge also creates its own new failure mode, which the flat nest did not have: it puts +the plate at an angle where a shallow lip stops retaining. That is §5's lip failure, and it +is not hypothetical for this part -- it is the part's defining feature. + +### The acceptance test + +The claim is a **difference between two conditions**, so the design is paired and the control +is the identical protocol without the wedge. Nine requirements, each of which exists because +skipping it produces a number that looks like an answer and is not. + +**1. Endpoint must be a quantity, not an impression.** "The wells look drier" is not a +result. The natural endpoint is residual volume, measured **gravimetrically** -- weigh the +plate before and after the aspiration step on a balance whose resolution and repeatability +are adequate for the residual volume in question, and state both. If they are not adequate, +the experiment does not exist yet and buying a balance is the first step, not the wedge. + +**2. Do not route the endpoint through the plate reader.** This workcell's absorbance read +is BROKEN -- written, run on the instrument, and it times out deterministically -- so an +optical recovery endpoint is not evaluable here at all. Even repaired, A260 does not +discriminate library from primer, carrier or free nucleotide, and at low input those +dominate the signal. A gate reading that number passes an empty well confidently. Gravimetry +or an orthogonal spike-and-recover, not OD. + +**3. Measure per well, before any pooling.** The loss the wedge is supposed to reduce is a +per-well loss. A measurement taken after pooling reads 96 wells as one number, and a +single-well effect moves that number by about a ninety-sixth -- which is exactly why the +recovery layer already calls single-well loss silent. Measure while the material is still +`ADDRESSED`. + +**4. n >= 3 paired runs before any tolerance is stated.** This repo sets +`MIN_DEMONSTRATIONS = 3` for stating a spread at all, and the same arithmetic applies here: +two points have one degree of freedom, so the interval they imply *is* the two points and +nothing in them can say whether either is an outlier. One paired plate produces a value and +no tolerance. + +**5. Report the observed range, not mean plus k sigma.** The min and max of what happened is +a fact. Mean plus k standard deviations is a fact plus a k somebody chose, and k is exactly +the kind of number that reaches a report with no author attached to it. + +**6. Judge on the worst observation, not the mean.** A mean lets one excellent plate pay for +a bad one. At the bench, the bad plate is the one that costs the sample. + +**7. Randomize position, and never put wedge and control on different sides of the deck.** +Failures that appear column-wise are hardware, not biology -- biology fails randomly across +a plate and hardware fails geometrically. A wedge that occupies one deck position for the +whole experiment has confounded treatment with position, and the effect you measure may be +the position. + +**8. Hold everything else constant and say so.** Same session, same instrument, same +consumable lot, same filament lot for the wedge itself. A wedge reprinted from a different +spool is a different part until it has been measured again (§5, shrinkage). + +**9. Pre-register the threshold and record its basis.** Decide before the run what +improvement would make the wedge worth keeping, and record where that target came from. It +is `intuition` or `vendor` until this experiment runs; only afterward is there an `in_house` +number, and only for this plate type, this liquid, this volume and this aspiration height. + +### What the test cannot show + +It cannot show the wedge is safe for the arm. That is DRY_RUN, it is a separate +qualification, and a recovery result does not substitute for it in either direction -- a +wedge that improves recovery by hand and crashes the head is a worse part than no wedge. + +And a bench result obtained by a person pipetting does not transfer to the arm. The +aspiration height, the approach angle and the compliance are all different. If the wedge is +for the robot, the paired experiment has to be run by the robot, after DRY_RUN, or the +number belongs to a hand. + +--- + +## What this guide refuses to tell you + +Kept as a list rather than buried, because the gaps are the part most likely to get filled +in by somebody's search results. + +- **Whether any FDM part has ever been validated as cleanable to a recognized standard.** A + targeted search found none. That is the crux of §1 and it did not close. +- **Whether steam penetrates and sterilizes the interior voids of an FDM part.** No located + study cultured the interior of an autoclaved FDM part or ran biological indicators inside + the void network. Surviving the heat and being sterile are different claims. +- **How many autoclave cycles any high-temperature polymer survives.** Every specific count + traced to commercial aggregator pages. +- **A shrinkage percentage from any filament manufacturer.** Three datasheets were opened; + none published one. The slicer defaults in §5 are the best available and they disagree + with each other. +- **A warp-per-unit-length rule for large flat parts.** No dataset located measures flatness + across a range of part sizes. The 3.7 mm / 0.8 mm figures cannot be scaled because the + specimen dimensions were not stated. +- **A manufacturer bed-flatness specification** for any machine in §2. +- **Coefficients of thermal expansion** for ASA and PC. Conflicting values circulate for ABS + (90e-6 /C versus 120e-6 /C) and the PLA and PETG figures found came from a retailer blog. + None are printed here. +- **A static-dissipation figure** for any printable material, printed or otherwise. +- **The dimension callouts inside the SLAS engineering drawings.** Those figures are raster + images. The clause text and the figure notes were read; anything appearing only on a + drawing is not captured here. +- **Whether the SLAS standards were reaffirmed after June 2017.** The published copies remain + (R2012) revisions and the standards page still carries forward-looking text about a + reaffirmation. No ANSI record of a later action was found, so no current reaffirmation year + is asserted. There is also no ANSI/SLAS 5 in the published set (1, 2, 3, 4, 6) and no + authoritative statement explaining the gap -- do not speculate about it in print. +- **Whether the standards copies consulted are the controlling versions.** They are the + copies the standards body itself publishes, dated 2011. For anything that goes to print, + check the purchased controlled copies. +- **Any manufacturer claim of biocompatibility, cytotoxicity testing, autoclavability, or + sample-contact suitability for any printer or stock filament in §2.** There is none on any + official page read. "Supported filament" lists are printability ratings -- one vendor + literally grades them "Ideal" and "Capable" -- meaning will-it-extrude-without-clogging. + Treat the absence as absence, and never read a spec-table material list as a + fitness-for-contact statement. + +--- + +## Sources + +Standards and regulatory +- ANSI/SLAS 1-2004 (R2012) Footprint Dimensions; 2-2004 (R2012) Height Dimensions; 3-2004 + (R2012) Bottom Outside Flange Dimensions; 4-2004 (R2012) Well Positions; 6-2012 Well + Bottom Elevation. Clause text read directly from the standards body's published PDFs. +- ASME Y14.5M-1994, the drawing standard those five invoke normatively. +- ISO/ASTM 52902:2023, Additive manufacturing -- Test artefacts -- Geometric capability + assessment. Paywalled; title, scope and edition status confirmed only. This is the citable + route to a defensible in-house accuracy claim, rather than quoting a vendor number. +- ANSI/AAMI ST98:2022 cleaning validation (superseded AAMI TIR30). Confirmed acceptance + criteria: TOC <= 12 ug/cm^2, ATP <= 22 fmol/cm^2. The widely quoted 6.4 ug/cm^2 protein + criterion could not be confirmed and is not printed here. +- US FDA, *Technical Considerations for Additive Manufactured Medical Devices*, issued + 2017-12-05, §VI.E. +- US FDA guidance on the Mouse Embryo Assay for assisted reproduction devices. + +Peer-reviewed +- Wang X, Zhao L, Fuh JYH, Lee HP (2019). *Polymers* 11(7):1154. doi:10.3390/polym11071154. + Porosity at 100% infill by X-ray CT; pore size distribution and location. +- Popescu D, Baciu F, Amza CG, Cotrut CM, Marinescu R (2021). *Polymers* 13(23):4249. + doi:10.3390/polym13234249. No print setting seals against fluid intake; vapor-smoothed + parts still penetrated by dye. +- Hall DC Jr, Palmer P, Ji H-F, Ehrlich GD, Krol JE (2021). *Front Microbiol* 12:646303. + doi:10.3389/fmicb.2021.646303. Biofilm in the layer lines; anisotropic Ra. +- Chiscop F, Cazacu C-C, Cazacu D-A, Cotet CE (2025). *J Funct Biomater* 16(9):334. + doi:10.3390/jfb16090334. Annealed PLA through a 134 C / 60 min cycle, with dimensional + penalties. +- Rogers HB, Zhou LT, Kusuhara A, Zaniker E, Shafaie S, Owen BC, Duncan FE, Woodruff TK + (2021). *Chemosphere* 270:129003. ISO-certified dental resins release ovo-toxic leachates. +- Kress S, Schaller-Ammann R, Feiel J, Priedl J, Kasper C, Egger D (2020). *Materials* + 13(13):3011. doi:10.3390/ma13133011. Cytotoxicity of stereolithography photopolymers; + parylene barrier and its detachment after 5-6 autoclave cycles. +- Oskui SM et al. (2016). *Environ Sci Technol Lett*. doi:10.1021/acs.estlett.5b00249. Both + FDM and SLA parts toxic to zebrafish embryos; SLA significantly more so. + +Manufacturer datasheets read directly +- Bambu Lab PETG Basic TDS V3.0; Victrex PEEK 450G TDS (rev. March 2026); SABIC ULTEM Resin + 1010 TDS (Europe, rev. 20170620); Stratasys ULTEM 9085 MDS (2025); Prusament ASA TDS v1.1; + Prusament PC Blend TDS v1.1; PPprint P-filament TDS v1.001; Polymaker PETG TDS V2.0 + (2025-11-17); Polymaker ASA TDS; 3DXTech 3DXMAX ABS TDS Rev 3.0; Formlabs BioMed Clear TDS + (doc 2001432-TDS-ENUS-0, rev 04); Stratasys F123 series product specification. +- Values for PLA grades, ABS-M30, Stratasys ASA, Bambu PC, PA grades and FDM ULTEM 1010 came + from search extraction of manufacturer PDFs rather than direct reads and are marked + `[DS-2]` throughout. One known problem: a Bambu PC datasheet surfaced HDT 117 C @1.8 MPa + against 112 C @0.45 MPa, which is inverted relative to every other sheet read -- HDT at the + higher load should be the lower number. Those two look transposed. Do not publish that pair + until somebody opens the PDF. + +Vendor pages, captured 2026-07-28 +- Bambu Lab US store product pages for P1S, P2S, X2D, H2D, H2C and the 3D Printers category + listing; the X1-series end-of-life announcement. +- Prusa Research product pages for MK4S, CORE One+ and the HT hotend upgrade. +- QIDI Tech US store Plus4 product and technical-specification pages. Note the Plus4 + marketing page carries a comparison graphic reading "Hot bed 100 / Nozzle 360 / Chamber + 55" -- those are previous-generation figures shown for contrast, not the Plus4's. An + automated scrape of that page pulls the wrong three numbers. +- Intamsys Funmat HT product page. +- Protolabs Network (Hubs) knowledge base, dimensional accuracy of 3D printed parts. + +Shipped code, read directly +- OrcaSlicer repository: `resources/profiles/*/filament/*.json` (`filament_shrink` values) and + `src/libslic3r/PrintConfig.cpp` (the tooltip defining what those values mean). +- PrusaSlicer repository: `resources/profiles/PrusaResearch.ini` + (`elefant_foot_compensation`). +- Prusa-Firmware: `Firmware/variants/MK3S.h`, for the Z-axis step geometry behind layer-height + quantization. That granularity is (leadscrew lead)/(full steps per revolution) and must be + recomputed per machine rather than assumed. diff --git a/hardware/README.md b/hardware/README.md new file mode 100644 index 0000000..6b08c45 --- /dev/null +++ b/hardware/README.md @@ -0,0 +1,578 @@ +# Plate tilt module + +`tilt_module.scad` is a passive fixed-angle tilt fixture for a liquid handler deck. It +holds one ANSI/SLAS-footprint plate or reagent reservoir at a shallow angle so residual +liquid pools at one side of each well, where a tip can reach more of it. It is a base +plate that registers to a deck position, a solid wedge, and an upper platform with corner +retaining lips. No hinge, no adjustment, no moving parts. + +It matters where the last few microliters are the product: bead cleanups, elutions, and +reservoir dead volume. + +## What this repository will not tell you + +**No recovery figure appears anywhere in this file or in the model.** Not a percentage, +not a microliter count, not a range. Recovery depends on the well geometry, the liquid +class, the tip, the aspiration profile, and the surface energy of a molded polymer that +neither the model nor this README knows anything about. The only honest recovery number is +one you weighed on your own plates with your own liquid. The acceptance test below is how +you get it. Until it has been run, this is an untested idea with a part number, and it +should be described that way in a method section. + +The model applies the same rule to itself. Every derived quantity is a function that +echoes its result at render time, and any function whose input has not been measured +reports `NOT COMPUTED` and names the parameter rather than substituting a plausible value. +`tilt_module.scad` refuses to render the full fixture at all until `deck_measured = true`. + +## Safety + +### An unlocated fixture is a deck crash + +This part is taller and heavier than the plate the deck position was designed for, and its +center of mass is higher. If it can slide, rotate, or walk under gantry acceleration, the +head's taught coordinates point at where it used to be. A tip strike on a printed wedge +does not stop a Z axis: it snaps tips, shears the fixture off its seat, or drives the head +into the deck. + +Consequences: + +- `tilt_module.scad` will not render `part = "fixture"` until `deck_measured` is set true + and the registration dimensions are filled in. That refusal is the design, not a bug. +- `registration_style = "none"` prints a HAZARD line in the render report every time. It + is only defensible when the fixture is clamped or bolted by other means. +- **Re-teach every affected labware position after installing the fixture.** The plate no + longer sits where it sat, in any of the three axes, and it is no longer level. +- The body overhangs the SLAS footprint by roughly 4.3 mm per side in X and 4.8 mm per + side in Y at the default parameters. The render report prints the exact figures. Check + the neighboring deck positions and the head's travel path before you print, not after. +- Run the complete motion program dry, at production speed and acceleration, including any + gripper moves, before any liquid is on the deck. + +### What this material may touch + +This is a **holder**. Nothing printed from this file is qualified to contact samples, +reagents, cells, or media. Keep the certified consumable between the printed part and the +sample: the plate is the barrier. + +- Dimensional conformity to ANSI/SLAS is not evidence of fitness for contact. Those + standards specify footprint, height, flange, well positions and a well-bottom-elevation + test method, and say nothing about material, leachables, extractables, cytotoxicity, + sterility, or nuclease status. +- FDM parts are porous by construction. X-ray CT of PLA printed at 100% infill measured + 4.05 to 6.32% internal porosity across raster settings, with pores concentrated at the + shell-to-infill interface -- that is, connected to the outer surface (Wang et al., + *Polymers* 11(7):1154, 2019). Immersion testing of ABS found that no combination of + layer height, perimeter count, or infill pattern sealed a part against fluid intake, and + that acetone vapor smoothing left parts that still stained with dye after 30 minutes of + immersion (Popescu et al., *Polymers* 13(23):4249, 2021). +- Bacteria preferentially colonize the layer lines. Biofilm was thickest in the grooves + between printed layers on every polymer tested (Hall et al., *Front. Microbiol.* + 12:646303, 2021). +- Many photopolymer (SLA/DLP/MSLA) resins remain cytotoxic after full manufacturer + post-cure, including resins carrying ISO 10993 biocompatibility certification, because + the certification does not cover the exposure this fixture would represent and because + non-polymerizable additives are not consumed by further UV exposure. +- **Do not autoclave.** PLA, PETG, ABS and ASA all have a glass transition below 121 C. + Printed parts contain frozen-in extrusion stress that relaxes as soon as the polymer + passes Tg, so the part moves with no load applied at all. Wipe-down with 70% IPA is + housekeeping, not decontamination, and the part will absorb some of what you wipe it + with. + +Treat the fixture as a non-sample-contacting fixture with a defined replacement interval, +or as single-use if it is ever splashed. + +## Files + +| file | what it is | +| --- | --- | +| `tilt_module.scad` | the parametric model, the computed report, and the test coupon | +| `README.md` | this file | + +Render with OpenSCAD 2019.05 or later (`assert()` is required): + +``` +openscad -o coupon.stl -D 'part="coupon"' tilt_module.scad +openscad -o fixture.stl -D 'part="fixture"' tilt_module.scad # refuses until measured +``` + +The echoed report goes to stderr on the command line and to the console pane in the GUI. +Read it every time. It is where the numbers that decide feasibility appear. + +## Parameters + +Every dimension is a named parameter at the top of the file with its units and its reason. +The groups below are the ones you will actually change. + +### The labware being held + +| parameter | default | note | +| --- | --- | --- | +| `plate_len_mm`, `plate_wid_mm` | 127.76, 85.48 | ANSI/SLAS 1-2004 §4.1.1.1 nominals | +| `plate_footprint_tol_mm` | 0.5 | the §4.1.1.2 mid-side tolerance, not the §4.1.1.1 corner one. See below. | +| `plate_corner_r_max_mm` | 4.78 | §4.1.2.1 is 3.18 ± 1.6 mm. Use the maximum for a clearance feature. | +| `plate_flange_h_mm`, `plate_flange_tol_mm` | 6.10, 0.38 | ANSI/SLAS 3-2004 §4.2 medium. Five incompatible variants exist. | +| `plate_height_mm` | 14.35 | ANSI/SLAS 2-2004, and **only** for a standard-height microplate | +| `plate_cg_height_mm` | -1 | unmeasured sentinel; no SLAS standard gives plate mass | +| `well_cols`, `well_rows`, `well_pitch_mm` | 12, 8, 9.0 | ANSI/SLAS 4-2004 §4.1 (96-well) | + +Three of these are traps worth stating plainly: + +**There are two footprint tolerances, not one.** SLAS 1 §4.1.1.1 gives ± 0.25 mm, but only +within 12.7 mm of the four outside corners. §4.1.1.2 gives ± 0.5 mm anywhere else along +the side. A conforming plate may bow outward by 0.5 mm at mid-side. A pocket cut to the +commonly quoted ± 0.25 mm jams legal plates, in the middle, where it is not obvious why. +The model uses the loose figure. + +**The flange variant is a required input.** SLAS 3 offers five mutually exclusive variants +(short 2.41, medium 6.10, tall 7.62, short-with-interruptions 2.41, dual 2.41/7.62 mm, all +± 0.38) and requires a plate to declare which one it meets. "SBS compliant" on a vendor +page does not tell you. The retaining lip has to engage the flange and nothing above it, +so this parameter sizes the lip. On a short-flange plate the default 3.0 mm lip is too +tall and the render report says `FAIL`. + +**Height is standardized only for standard-height plates.** Deep-well plates, PCR plates +and reservoirs keep the SLAS 1 footprint and abandon the SLAS 2 height. Published deep-well +products sit near 44 mm, but that is a product spec, not a standard, and other vendors +differ. If you are tilting one, measure it and put the measurement in. + +### The angle + +| parameter | default | note | +| --- | --- | --- | +| `tilt_angle_deg` | 7 | about the plate's short axis; the column-12 end goes down | +| `cross_tilt_deg` | 0 | optional second tilt; turns an edge pool into a corner pool | + +Compound tilt is steeper than it sounds. Two 6-degree components make an 8.5-degree plane. +`max_slope_deg()` computes the true slope and the report prints it. + +### Clearance + +| parameter | default | note | +| --- | --- | --- | +| `clearance_mm` | 0.35 | per side, plate into pocket -- a hole | +| `deck_clearance_mm` | 0.30 | per side, skirt into deck nest -- a shaft | + +**Clearance is the parameter you tune after measuring a test print, not the one you trust +from the model.** The printed part is not the model: + +- Shrinkage is material and grade dependent and generally not published. Filament technical + data sheets checked directly (Polymaker PETG V2.0, Polymaker ASA, 3DXTech 3DXMAX ABS) + carry no shrinkage figure at all. Every per-material percentage circulating online is a + slicer default or a community measurement. +- Slicer defaults disagree with each other. OrcaSlicer ships `filament_shrink` of 99.95% + for one vendor's PLA, 99.85% PETG, and 99.487% for ABS and ASA, while another vendor's + profiles in the same repository ship 100% for all four. A 0.5% disagreement on ABS is + 0.64 mm across 127.76 mm -- larger than the entire SLAS footprint tolerance. +- Desktop FDM accuracy is quoted at about ± 0.5% with a ± 0.5 mm floor, and it scales with + length. A good calibration cube proves nothing about a 129 mm pocket. +- The first layers are the least trustworthy part of the print. PrusaSlicer's own profiles + ship 0.2 mm of elephant-foot compensation for a 0.4 mm nozzle, and the bottom of this + fixture is exactly where the registration feature lives. +- Repeatability is far better than accuracy: hundredths of a millimeter of run-to-run + spread against tenths of accuracy error. That asymmetry is why measure-and-compensate + works -- your printer reproduces its own error reliably enough to cancel it. + +Two clearances, not one, because a hole and a shaft err in opposite directions on the same +machine. Do not slave one to the other. + +### Lips, base, and registration + +| parameter | default | note | +| --- | --- | --- | +| `lip_style` | `"corners"` | `"corners"` or `"full"` | +| `lip_height_mm`, `lip_height_uphill_mm` | 3.0, 3.0 | above the seat plane | +| `lip_thickness_mm` | 2.4 | at the lip top; 6 walls at a 0.4 mm nozzle | +| `lip_fillet_mm` | 1.2 | extra thickness at the root, tapered outward face | +| `lip_arm_mm` | 24.0 | length of each corner-lip arm | +| `base_thickness_mm` | 4.0 | flat slab spanning the whole footprint | +| `min_lift_mm` | 2.0 | minimum wedge thickness at the low corner | +| `platform_thickness_mm` | 3.0 | slab under the seat | +| `wall_mm` | 2.4 | registration skirt wall, and the minimum wall anywhere | +| `registration_style` | `"slas_skirt"` | `"slas_skirt"`, `"post"`, or `"none"` | +| `deck_nest_depth_mm` | -1 | **MEASURE.** Unmeasured sentinel. | +| `deck_nest_rim_w_mm` | -1 | **MEASURE.** Unmeasured sentinel. | +| `seats_on` | `"nest_floor"` | which surface carries the weight; see below | +| `reg_post_*` | -1 | **MEASURE.** Only for `registration_style = "post"`. | +| `deck_measured` | `false` | the gate on rendering the fixture | + +### How the retaining lips hold the plate + +The plate on a tilted seat wants to slide downhill. **Whether friction alone would hold it +is not computed anywhere, and cannot be**: the static coefficient between a printed lip and +a molded polypropylene skirt is unpublished, changes with layer orientation and surface +finish, and drops further when either surface is wet with buffer or ethanol. Design as if +friction is zero. The lips are the only thing holding the plate. + +So the plate does not center itself in the pocket. It slides until it touches the two +downhill corner lips, and **those lips are the datum**. Every reach number in the model is +referenced to that seated position. + +The lip height has a window and both ends are real: + +- **Lower bound**, `lip_min_engage_mm` = 1.2 mm. The lip must stand proud enough to engage + the flange face rather than the plate's bottom edge break. SLAS figure Note 4 in all four + dimensional standards states that dimensions and tolerances do not include draft. Molded + plates carry draft angles the standard excludes and does not publish, so the flange face + is not vertical and the exact contact height cannot be computed from the standard. A + shorter lip is contacting whatever the draft leaves it, which is not a designed contact. +- **Upper bound**, minimum flange height minus `lip_flange_margin_mm`. Above the flange the + lip stops pushing on the flange and starts pushing on the plate body -- into gripper jaw + travel, into a skirt chamfer that will cam the plate upward, or into the well region. + Minimum flange height is nominal minus 0.38 mm: 5.72 mm for a medium flange, 2.03 mm for + a short one. + +`lip_window_ok()` checks the window and the report prints `PASS` or `FAIL` with the reason. + +Corner lips rather than continuous rails, for three reasons. SLAS 3 §4.4 permits a single +interruption on center of each long side; a continuous long-side rail can land on that +interruption, but a corner lip cannot, because the interruption edges are at least 47.8 mm +from the nearest part edge. The mid-side gaps are also where a gripper puts its jaws. And +the downhill mid-side gap is the drain path for anything spilled. + +Tipping is not the failure mode. `tip_over_angle_deg()` computes the angle at which the +plate's center of mass passes over the downhill lip contact line, and it reports +`NOT COMPUTED` until you supply `plate_cg_height_mm` -- but for any plausible CG height the +answer is far above any usable tilt. The lip's job is purely to arrest sliding. That means +lip **thickness and root strength** matter, not lip height. The lip load itself is also +`NOT COMPUTED`: no SLAS standard specifies plate mass, so weigh your loaded plate if you +want a number. + +### Deck registration and the two seating conditions + +`registration_style = "slas_skirt"` gives the fixture's own base an ANSI/SLAS-1 footprint +skirt so it drops into the deck nest the way a plate does. This reuses location the deck +already provides and needs one measurement (the nest depth) instead of a hole pattern. The +skirt is continuous and uninterrupted around the base, as SLAS 1 §4.1.1.3 requires of a +plate footprint, and for the same reason -- a gap snags a nest. + +The fixture body is wider than the skirt, so its shoulder sits above the nest rim. Only one +of those two surfaces can carry the weight, and a fixture that is unsure which one will +rock: + +- `seats_on = "nest_floor"` runs the skirt `seat_relief_mm` deeper than the nest, so the + skirt bottom lands on the nest floor and the shoulder floats clear of the rim. +- `seats_on = "nest_rim"` runs the skirt shallower, so the shoulder lands on the rim and + the skirt bottom floats. This needs `deck_nest_rim_w_mm` to be wide enough to carry the + shoulder. + +Pick one. Do not size both to zero clearance and hope. + +`registration_style = "post"` uses **two** posts on the long axis, not four. Four posts in +four holes over-constrain the part: with desktop FDM positional error across a 100 mm span, +at least one post binds and the fixture seats on three, rocking. Two posts fix position and +rotation, which is all that is needed. + +## The two numbers that decide feasibility + +Both are echoed on every render. At the default parameters (96-well, 7 degrees, no cross +tilt, `registration_style` engagement excluded because it depends on your measured nest): + +### `tip_reach_delta_mm()` -- 12.07 mm + +How much further down the head must travel at the far corner than at the near one. + +Two well bottoms separated by `dx` along the plate's long axis and `dy` along its short +axis are separated vertically, once tilted, by + +``` +dz = dx * sin(tilt) + dy * sin(cross) * cos(tilt) +``` + +`sin`, not `tan`: the wells are separated by `dx` measured **along the plate**, which is +now the hypotenuse rather than the deck projection. Using `tan` overstates the drop, and at +shallow angles the two differ by under a percent -- exactly the kind of error that survives +review. + +The span is between well **centers**, from SLAS 4: `(cols - 1) * pitch` by +`(rows - 1) * pitch`, which is 99.0 x 63.0 mm for a 96-well plate, not the 127.76 x 85.48 +footprint. Using the footprint overstates the drop by about 30%, which sounds conservative +until it talks somebody out of a workable angle. + +If the head is near the end of its Z stroke at the near corner -- which is where a +deep-well plate on a tall fixture puts it -- the far corner is where it runs out. Confirm +the commanded Z at the far corner is inside the head's envelope before running liquid. + +### `stack_top_mm()` -- 39.38 mm above the deck datum, versus 14.35 mm flat + +25.03 mm of extra Z consumed, decomposed by `stack_budget()`: + +| contribution | mm | reducible? | +| --- | --- | --- | +| fixed: base slab + min lift + platform slab | 8.98 | yes, `low_profile` takes it to 5.59 | +| angle-driven: pocket length x sin(tilt) | 15.72 | **no** | +| the plate itself | 14.24 | no | + +`low_profile = true` thins the base slab, the lift, and the platform slab, and saves +3.40 mm out of 25.03. **The angle-driven term is untouchable.** You cannot thin your way +out of the angle. If the stack does not fit the head envelope or the gripper approach, the +angle comes down or the fixture does not get used. + +Add your measured `deck_nest_depth_mm` on top of all of this if the skirt seats on the nest +floor. + +## Print settings that matter + +Material-dependent settings are marked. Everything else is geometry. + +| setting | value | why | +| --- | --- | --- | +| layer height | 0.15 to 0.2 mm | finer layers reduce fluid ingress at the surface, but do not seal the part | +| perimeters / walls | 5 or more | the lip is a cantilever loaded across the layers; walls carry that load, infill does not | +| top / bottom layers | 5 or more | the seat is a functional surface | +| infill | 15 to 25% gyroid | **do not model internal voids.** An enclosed cavity in a fixture that gets wiped cannot be dried. Let the slicer make the sparse structure so it drains and dries through the walls. | +| elephant-foot compensation | as your profile normally uses | the registration skirt is on the first layers; if you disable it, the skirt is oversize | +| brim | only if adhesion fails, and **remove it completely** | SLAS 1 §4.1.1.3 requires a continuous uninterrupted footprint. Brim and elephant-foot remnants on the skirt snag deck nests and gripper jaws. | +| supports | see below | | +| nozzle | steel, not brass | **material-adjacent.** Brass alloys commonly contain lead. Even for a non-contacting fixture this is the cheaper choice to make correctly. | + +Material choice is yours, with these constraints: + +- **Nothing here is autoclavable.** PLA (Tg roughly 54 to 61 C across four grades from one + manufacturer), PETG (Tg roughly 69 to 77 C, with an 8 C spread between manufacturers on + nominally the same polymer), ABS (HDT 84 to 104 C between desktop and industrial grades) + and ASA (HDT 86 to 103 C) are all below a 121 C cycle. For a printed part Tg is the + correct gate, not HDT: an autoclaved fixture carries no external load, but it does carry + frozen-in extrusion stress that relaxes above Tg and moves the part with zero load + applied. +- PLA warps at bench-relevant temperatures. One study observed visible warping of an + as-printed PLA curved beam from approximately 70 C. A fixture that lives near a heated + block or in a warm room is a PLA problem. +- Shrinkage differs by material **and by grade and vendor within a material** (see + Clearance above). If you change filament, reprint the coupon. The clearance you + calibrated does not transfer. + +## Orientation and supports + +**Print base down, seat up, wedge as printed.** The part is modeled sitting on the deck; it +prints in that same orientation with no rotation. + +Why: + +- The registration skirt and the base slab are the datum surfaces of the whole fixture, and + they end up flat on the bed. That is the most dimensionally repeatable orientation + available, and the one where the skirt walls are vertical rather than stepped. +- The seat plane becomes a shallow top surface at the tilt angle. At 7 degrees it needs no + support, and it comes out as an ironed-quality top rather than a support-scarred + underside. The seat is the surface the plate registers against, so it should not be the + supported one. +- The lips end up as vertical walls printed as walls, so their thickness comes from + perimeters and not from infill. +- The wedge's outer faces lean **inward** going up (the body footprint is deliberately + larger than the platform's projection by `body_margin_mm`), so there are no overhangs + anywhere on the exterior. + +Supports: **none required** at the default parameters. Check the preview if you raise +`tilt_angle_deg` past roughly 45 degrees, which is not a sensible tilt for this application +anyway, or if you widen `lip_arm_mm` so far that the corner lips bridge. + +Where the layer lines end up relative to load, and what to do about it: + +The layer lines run horizontally, parallel to the deck. The plate pushes the downhill lip +sideways, so the lip is a transverse cantilever and the tensile stress at its root acts +**across** the layer interfaces. That is the weakest direction in an FDM part, and it is +where this fixture will fail if it fails. + +Rotating the part to fix that would put layer lines the right way for the lip and the wrong +way for everything else: the registration skirt would print as stepped walls, the seat would +need support, and the base would lose its flatness reference. The skirt and the seat matter +more, because a lip that flexes lets a plate creep, while a skirt that is out of round lets +the whole fixture sit wrong. So the orientation stays, and the lip is compensated +geometrically instead: + +- `lip_thickness_mm` = 2.4 mm at the tip, which is 6 perimeters at a 0.4 mm nozzle. +- `lip_fillet_mm` = 1.2 mm of extra thickness at the root, with the outer face tapered so + the section modulus is highest exactly where the bending moment is. +- The inner face stays vertical. That is the plate contact and it should be a face, not a + slope. +- Push the coupon's lip hard with a thumb, sideways, before you print the fixture. If it + creaks or whitens at the root, add perimeters or thickness. This is the one place the + coupon tests strength rather than fit. + +## The test coupon: print this first + +``` +openscad -o coupon.stl -D 'part="coupon"' tilt_module.scad +``` + +The coupon is a small pad carrying three things: + +1. **The downhill corner of the real fixture**, cut out of the same `upper_solid()` module + the fixture uses -- the corner lip, its root fillet, its lead-in chamfer, the pocket + corner radius, and the seat at the design angle. It is not a re-modeled approximation. A + coupon built from its own geometry only validates the coupon. +2. **A corner of the registration skirt** at its real wall thickness and corner radius. +3. **A witness block**, 20.0 mm nominal in X and Y, 10.3 mm in Z. + +What to do with it: + +- Measure the witness block on all three axes. X and Y give your printer's actual in-plane + scale error, which is what `clearance_mm` and `deck_clearance_mm` exist to cancel. Z is + different in kind: it is layer quantization plus first-layer squish, which is a constant + offset on total height rather than a percentage, so do not convert it to a percent and + apply it to lengths. The Z nominal is 10.3 mm on purpose -- it is not a multiple of any + common layer height, and total height is forced to `first_layer + n * layer`, so a + nominal that lands between two layers is the only one that shows you the quantization at + all. A 20.0 mm gauge at a 0.2 mm layer hides it completely. +- Offer a real plate corner to the lip. It should drop in past the lead-in without force + and sit flat, and it should not be able to lift over the lip. +- Push the lip sideways, hard. +- Drop the skirt corner into the corresponding corner of your deck nest. + +**What the coupon cannot tell you.** It validates features, not the whole-footprint fit, +because it is not 127.76 mm long. Take your measured scale error from the witness block, +apply it to the full length, and only then commit to printing the fixture. If you change +filament, spool, or printer, the coupon is invalidated and you print it again. + +## Fitting the fixture to your deck + +1. Print and check the coupon. Set `clearance_mm` and `deck_clearance_mm` from what you + measured. +2. Measure the deck position: nest depth to the surface a plate's own footprint rests on, + and rim width if you intend to seat on the rim. Set `deck_nest_depth_mm`, + `deck_nest_rim_w_mm`, and `seats_on`. +3. Set `plate_flange_h_mm` from your plate's declared SLAS 3 variant and confirm the + report says the lip window is `PASS`. +4. Set `plate_height_mm` to your plate's real height if it is not a standard-height + microplate. +5. Read `tip_reach_delta_mm()` and `stack_top_mm()` from the report and check both against + the head's Z envelope and the gripper's approach. Do this before printing, not after. +6. Set `deck_measured = true` and render. +7. Install it, then re-teach every affected labware position. +8. Run the full motion program dry at production speed and acceleration. +9. Only then run the acceptance test. + +## Acceptance test + +**An unmeasured fixture is an unproven one.** The claim being tested is that the fixture +reduces residual volume. That claim is gravimetric, not visual, and it has to be measured +against the same protocol without the fixture. + +### What to measure + +Residual volume: the mass of liquid remaining in the vessel after a defined recovery +aspiration, converted to volume using the density of the actual liquid at the measured +temperature. + +Use a balance with 0.1 mg readability or better. At the density of water, 0.1 mg is +approximately 0.1 uL, which is the resolution the question deserves. A 1 mg balance cannot +answer it. + +### Procedure, per plate + +1. Weigh the empty, dry plate. `m_empty`. +2. Dispense a defined starting volume into a defined set of wells with the liquid handler, + using the same tips, liquid class, and aspiration profile you will use in production. +3. Weigh. `m_filled`. The delivered volume is `(m_filled - m_empty) / density`, and this + also tells you whether the dispense itself was in spec. +4. Run the recovery aspiration program. +5. Weigh. `m_after`. Residual mass is `m_after - m_empty`. + +### The control most people skip + +Run an **evaporation control plate** in parallel: filled identically, weighed on the same +schedule, never aspirated. At microliter scale over a several-minute program, evaporative +loss is comparable in magnitude to the effect being measured. Subtract it. Without this +control, a slow protocol looks like a good recovery. + +Record ambient temperature and humidity with each run. Use the density of your actual +liquid at that temperature, not water at 20 C, unless the liquid is water. + +### Replicates, and what the replicate is + +**The experimental unit is the plate or the run, not the well.** The fixture is applied per +plate, so wells within one plate are pseudo-replicates for the fixture effect: they share +the plate, the seating, the tip box, and the moment in time. Averaging 96 wells and +reporting n = 96 overstates the precision by roughly an order of magnitude. + +Minimum defensible design: + +- **At least 3 independent runs per arm, on at least 2 different days**, so between-day + variation is inside the estimate rather than hidden by it. +- **Paired and interleaved**: within each run, alternate fixture and no-fixture plates + rather than doing all of one arm then all of the other. Tip lots, room temperature, and + operator technique all drift. +- Compute the mean residual across wells **within** a plate, then treat those plate means as + the replicates. +- Report per-well residuals too, as a distribution. A fixture that lowers the mean while + widening the spread has traded one problem for another, and the mean alone hides it. +- Randomize or counterbalance plate positions and which physical plates go in which arm. + +### Pre-register the threshold + +Decide, **before running**, the smallest reduction in residual volume that would change +what you do. That is a scientific judgment about your assay, not a statistical one, and +writing it down afterward is how a null result becomes a positive one. + +Adopt the fixture only if all of the following hold: + +1. The paired mean reduction in residual volume exceeds your pre-stated minimum useful + difference, and the 95% confidence interval on the paired difference excludes it. +2. The coefficient of variation of recovered volume is not worse with the fixture than + without. +3. Zero collisions, zero mispicks, and zero plate movements across the full test, including + every gripper move. +4. The plate is in the same position at the end of the run as at the start. Mark the plate + and the fixture, and check. +5. The commanded Z at the far corner stayed inside the head's envelope, with margin. + +Any of 3, 4, or 5 failing is disqualifying regardless of how good the recovery number is. A +fixture that recovers more liquid and occasionally crashes is worse than no fixture. + +### Reporting + +Report the measured reduction with its confidence interval, the replicate count as **runs** +(not wells), the balance readability, the liquid and its density basis, the evaporation +correction, the plate type and its declared SLAS 3 flange variant, the tilt angle, and the +print material and printer. A recovery figure without the liquid, the plate, and the tip is +not transferable to anyone else's bench, including your own next month. + +If the result does not clear your threshold, say so and leave the fixture out. That is a +result. + +## Verification status of this file + +The model was checked for syntactic balance and for undefined function and module +references by static analysis. **It has not been rendered**: `openscad` is not installed on +the machine where this was written, so no F5 preview, no F6 render, no STL, and no +manifold check has been performed. Render it yourself before printing, read the echoed +report, and treat any geometry surprise as a bug in this file rather than in your setup. + +## Sources + +Dimensional standards, clause text read directly: + +- ANSI/SLAS 1-2004 (R2012) Microplates -- Footprint Dimensions +- ANSI/SLAS 2-2004 (R2012) Microplates -- Height Dimensions +- ANSI/SLAS 3-2004 (R2012) Microplates -- Bottom Outside Flange Dimensions +- ANSI/SLAS 4-2004 (R2012) Microplates -- Well Positions +- ANSI/SLAS 6-2012 Microplates -- Well Bottom Elevation (defines a test method and sets no + limits: §7 states explicitly that it is not the intent of the standard to state a limit) + +All five are published by SLAS at `slas.org`. Dimensions in SLAS 1 through 4 apply at 20 C; +SLAS 6's test method specifies 25 ± 2 C. Do not quote a single temperature for the family. +Before publishing anything that depends on these clauses, check the purchased ANSI copies. + +Porosity, cleanability, and material behavior: + +- Wang X, Zhao L, Fuh JYH, Lee HP (2019). *Polymers* 11(7):1154. doi:10.3390/polym11071154 +- Popescu D, Baciu F, Amza CG, Cotrut CM, Marinescu R (2021). *Polymers* 13(23):4249. + doi:10.3390/polym13234249 +- Hall DC Jr, Palmer P, Ji H-F, Ehrlich GD, Krol JE (2021). *Front. Microbiol.* 12:646303. + doi:10.3389/fmicb.2021.646303 +- Chiscop F, Cazacu C-C, Cazacu D-A, Cotet CE (2025). *J. Funct. Biomater.* 16(9):334. + doi:10.3390/jfb16090334 +- US FDA (2017). *Technical Considerations for Additive Manufactured Medical Devices*, + Section VI.E, on validating cleaning and sterilization against the worst-case + configuration. + +Dimensional accuracy and shrinkage: + +- Protolabs Network (Hubs) knowledge base, dimensional accuracy of 3D printed parts +- OrcaSlicer, `resources/profiles/*/filament/*.json` (`filament_shrink`) and + `src/libslic3r/PrintConfig.cpp` (its definition) +- PrusaSlicer, `resources/profiles/PrusaResearch.ini` (`elefant_foot_compensation`) +- Polymaker PETG V2.0, Polymaker ASA, and 3DXTech 3DXMAX ABS technical data sheets, checked + directly for a shrinkage figure and found to carry none + +Style precedent: `hardware/tube_nest.scad` in the `di-omics/bay-hack` repository -- every +dimension a named parameter at the top, and a fit coupon before anything large. diff --git a/hardware/tilt_module.scad b/hardware/tilt_module.scad new file mode 100644 index 0000000..91e3492 --- /dev/null +++ b/hardware/tilt_module.scad @@ -0,0 +1,973 @@ +// ============================================================================= +// tilt_module.scad -- a passive fixed-angle tilt fixture for a liquid handler +// deck. It holds one ANSI/SLAS-footprint plate or reservoir at a shallow angle +// so residual liquid pools at one side of each well and a tip can reach more of +// it. There is no hinge, no actuator, and no adjustment: the angle is a printed +// constant, which is the point. An adjustable angle is an angle nobody records. +// +// WHY THIS EXISTS +// +// On a bead cleanup, an elution, or a reservoir with 400 uL left in it, the last +// few microliters are the product. A flat plate on a flat deck leaves them +// spread across the whole well bottom in a film the tip cannot chase. Tilting +// the vessel concentrates the same volume into a smaller footprint at one edge. +// That is the entire mechanism. It is geometry, not chemistry. +// +// WHAT THIS FILE DOES NOT DO +// +// It does not tell you how much liquid you will recover. No number in this file +// is a recovery figure, and none should be added, because recovery depends on +// the well geometry, the liquid class, the tip, the aspiration profile, and the +// surface energy of a molded polymer this model knows nothing about. The only +// honest recovery number is one you weighed. hardware/README.md gives the +// acceptance test that produces it. Until that test has been run, this fixture +// is an untested idea with a part number. +// +// WHAT IT COMPUTES +// +// Everything derived here is a function, and every function echoes its result at +// render time (see report() at the bottom). Two of those numbers decide whether +// the fixture is usable at all before any liquid is involved: +// +// tip_reach_delta_mm() how much deeper the head must travel at the far corner +// stack_top_mm() how much taller the loaded deck position becomes +// +// Where an input has not been measured, the corresponding function returns +// nothing and the report says NOT COMPUTED. It does not substitute a plausible +// value. An unmeasured fixture is an unproven one, and a model that quietly +// invents the missing measurement makes that impossible to notice. +// +// SAFETY -- DECK CRASH +// +// A fixture that is not positively located on the deck is a crash. It is heavier +// and taller than the plate the deck position was built for, its center of mass +// is higher, and if it can slide, rotate, or walk under gantry acceleration then +// the head's taught coordinates point at where it used to be. A tip strike on a +// printed wedge does not stop a Z axis; it snaps tips, shears the fixture, or +// drives the head into the deck. This file therefore refuses to render the full +// fixture until deck_measured is set true (see the assert in tilt_fixture()). +// That refusal is deliberate. Print the coupon, measure your deck position, then +// come back. Also re-teach every affected labware position after installing the +// fixture -- the plate no longer sits where it sat, in any of three axes. +// +// SAFETY -- WHAT THIS MATERIAL MAY TOUCH +// +// This is a HOLDER. Nothing printed here is qualified to contact samples, +// reagents, cells, or media, and dimensional conformity to ANSI/SLAS is not +// evidence of fitness for contact -- those standards say nothing about material, +// leachables, cytotoxicity, sterility, or nuclease status. FDM parts carry 4-6% +// internal porosity even at 100% infill (Wang et al., Polymers 11(7):1154, +// 2019, X-ray CT, 4.05-6.32% across raster settings) and the voids connect to +// the outer surface, so they wick liquid in and cannot be reliably cleaned; +// vapor smoothing reduces roughness without sealing the part (Popescu et al., +// Polymers 13(23):4249, 2021). Many photopolymer resins remain cytotoxic after +// full post-cure. Keep the certified consumable between this part and the +// sample: the plate is the barrier. Wipe-down of the fixture is housekeeping, +// not decontamination. Do not autoclave -- PLA, PETG, ABS and ASA all have a +// glass transition below 121 C and will relax printed-in stress and move. +// +// PRECEDENT +// +// Parameter-block style follows hardware/tube_nest.scad in the bay-hack repo: +// every dimension named at the top, nothing magic buried in the geometry, and a +// fit coupon printed before anything large. +// ============================================================================= + +$fa = 2; +$fs = 0.4; + +// Which solid to emit. Print "coupon" first. Always. +part = "coupon"; // "coupon" | "fixture" | "both" + + +// ============================================================================= +// 1. THE LABWARE BEING HELD +// +// Defaults are the ANSI/SLAS nominals. They are nominals: a conforming plate is +// allowed to differ from every one of them, and the tolerance that governs a +// pocket is not the one people quote. +// ============================================================================= + +// ANSI/SLAS 1-2004 (R2012) 4.1.1.1 nominal footprint. mm. +plate_len_mm = 127.76; +plate_wid_mm = 85.48; + +// THE TOLERANCE THAT ACTUALLY SIZES THE POCKET. mm, applies to the overall +// dimension (not per side). +// +// SLAS 1 states TWO footprint tolerances and the commonly quoted one is the +// wrong one for this job. 4.1.1.1 gives +/- 0.25 mm, but only within 12.7 mm of +// the four outside corners. 4.1.1.2 gives +/- 0.5 mm anywhere else along the +// side. A conforming plate may bow outward by 0.5 mm at mid-side. A pocket cut +// to 127.76 +/- 0.25 jams a legal plate, and it jams it in the middle where you +// cannot see why. Use the loose figure. +plate_footprint_tol_mm = 0.5; + +// SLAS 1 4.1.2.1 corner radius is 3.18 +/- 1.6 mm, a permitted range of 1.58 to +// 4.78 mm -- a 3x spread. A clearance feature cut to the 3.18 nominal binds on +// every plate at the top of the range. Use the maximum for anything the corner +// has to fit into. mm. +plate_corner_r_max_mm = 4.78; + +// ANSI/SLAS 3-2004 bottom outside flange. FIVE mutually exclusive variants +// exist and a plate is required to declare which one it meets: +// 4.1 short 2.41 4.2 medium 6.10 4.3 tall 7.62 +// 4.4 short with interruptions 2.41 4.5 dual 2.41 short sides / 7.62 long +// all +/- 0.38 mm from Datum A. "SBS compliant" on a vendor page does not tell +// you which. The retaining lip has to engage this flange and nothing above it, +// so this is a load-bearing input, not documentation. Set it from the plate's +// own datasheet. mm. +plate_flange_h_mm = 6.10; +plate_flange_tol_mm = 0.38; + +// Plate height from Datum A. ANSI/SLAS 2-2004 4.1.1.2 gives 14.35 +/- 0.76 mm +// overall -- but ONLY for a standard-height microplate. Deep-well plates, +// PCR plates and reservoirs keep the SLAS 1 footprint and abandon the SLAS 2 +// height entirely; published deep-well products sit near 44 mm, which is a +// product spec and not a standard. If you are tilting a deep-well plate or a +// reservoir, put ITS measured height here, because this number and the tilt +// angle together set how much Z the fixture consumes. mm. +plate_height_mm = 14.35; + +// Height of the loaded plate's center of mass above its own bottom face. mm. +// NOT SPECIFIED BY ANY SLAS STANDARD -- no standard gives plate mass or mass +// distribution, and a full deep-well plate is a different object from an empty +// one. Leave negative and the tip-over check reports NOT COMPUTED rather than +// inventing a number. Measure it if you care about the answer. +plate_cg_height_mm = -1; + +// Well grid, ANSI/SLAS 4-2004 4.1 (96), 4.2 (384), 4.3 (1536). Used only to +// compute the span between the first and last well CENTERS, which is the span +// the tips actually have to cover. Offsets are from the outside edge and are +// listed for completeness; the reach math needs the pitch and the counts. +// 96: offsets 14.38 / 11.24, pitch 9.0 +// 384: offsets 12.13 / 8.99, pitch 4.5 +// 1536: offsets 11.005 / 7.865, pitch 2.25 +well_cols = 12; +well_rows = 8; +well_pitch_mm = 9.0; +well_a1_x_mm = 14.38; +well_a1_y_mm = 11.24; + + +// ============================================================================= +// 2. THE ANGLE +// ============================================================================= + +// Tilt about the plate's SHORT axis: the +x end of the plate goes down, so +// liquid runs toward the column-12 end. deg. +// +// This is the single most expensive parameter in the file. Every millimeter of +// extra deck height and every millimeter of extra Z travel scales with it, and +// no amount of thinning the base buys any of that back (see stack_budget()). +// Shallow is not timid: at 7 deg the far well bottom is already 12 mm below the +// near one across a 96-well plate. +tilt_angle_deg = 7; + +// Optional second tilt about the plate's LONG axis: the +y edge goes down. +// Nonzero turns the pooling point from an edge into a corner, at the cost of +// making the true slope steeper than either component (see max_slope_deg()). +// Default 0 -- single-axis tilt is easier to print, easier to seat, and easier +// to reason about, and most well geometries pool acceptably on one axis. deg. +cross_tilt_deg = 0; + + +// ============================================================================= +// 3. CLEARANCE -- the parameter you tune, and the one nobody tunes enough +// +// The printed part is not the model. It comes out of the machine a different +// size, and the difference is not a constant you can look up: +// +// - Shrinkage is material and grade dependent, and it is not published. +// Filament technical data sheets generally do not carry a shrinkage figure +// at all (checked directly on Polymaker PETG V2.0, Polymaker ASA, and +// 3DXTech 3DXMAX ABS -- none list one). Every per-material percentage +// circulating online is a slicer default or a community measurement. +// - Slicer defaults disagree with each other. OrcaSlicer ships filament_shrink +// of 99.95% for one vendor's PLA, 99.85% PETG, 99.487% ABS and ASA, while +// another vendor's profiles in the same repository ship 100% for all four. +// That is a 0.5% disagreement on ABS between two profiles in one slicer. +// 0.5% of 127.76 mm is 0.64 mm, which is larger than the entire SLAS +// footprint tolerance. +// - Desktop FDM accuracy is quoted at about +/- 0.5% with a +/- 0.5 mm floor, +// and tolerance scales with length, so a good calibration cube proves +// nothing about a 128 mm pocket. +// - The first layers are the least trustworthy region of the part. Squish +// widens them; PrusaSlicer's own profiles ship 0.2 mm of elephant-foot +// compensation for a 0.4 mm nozzle. The bottom of this fixture is exactly +// where the registration feature lives. +// - Repeatability is far better than accuracy -- a few hundredths of a +// millimeter of run-to-run spread against tenths of accuracy error. That +// asymmetry is what makes measure-and-compensate work: your printer +// reproduces its own error reliably enough to cancel it. +// +// So: do not trust the nominal. Print the coupon, measure it with calipers +// against coupon_gauge_mm, and set the two clearances below from what you +// measured. Then reprint the coupon and confirm. That loop is the design +// process; the numbers below are only its starting point. +// ============================================================================= + +// Per side, between the plate and the pocket walls. This is a HOLE the plate +// drops into, and printers usually make holes undersize. mm. +clearance_mm = 0.35; + +// Per side, between the fixture's registration skirt and the deck nest. This is +// a SHAFT going into someone else's hole, and printers usually make shafts +// oversize. Opposite sign of error, so it gets its own parameter. Do not slave +// it to clearance_mm even though the first-guess magnitude is similar. mm. +deck_clearance_mm = 0.30; + + +// ============================================================================= +// 4. PLATFORM AND RETAINING LIPS +// +// LIP GEOMETRY -- what holds the plate, and what it is holding against +// +// The plate on the tilted seat wants to slide downhill. Whether friction alone +// would hold it is NOT COMPUTED anywhere in this file, and cannot be: the static +// coefficient between a printed lip and a molded polypropylene skirt is not +// published, changes with layer orientation and surface finish, and drops +// further when either surface is wet with buffer or ethanol. Design as if +// friction is zero. The lips are the only thing holding the plate. +// +// The plate therefore does not center itself in the pocket. It slides until it +// touches the two DOWNHILL corner lips, and those lips are the datum. Every +// reach number in this file is referenced to that seated position, not to a +// centered one. +// +// The lip height has a window, and both ends of it are real: +// +// lower bound the lip must stand proud enough to engage the flange face +// rather than the plate's bottom edge break. SLAS figure Note 4 +// in all four dimensional standards states "Dimensions and +// tolerances do not include draft" -- molded plates carry draft +// angles the standard excludes and does not publish, so the +// flange face is not vertical and the exact contact height +// cannot be computed from the standard. A lip shorter than about +// 1.2 mm is contacting whatever the draft leaves it, which is +// not a designed contact. +// +// upper bound the lip must stay BELOW the minimum flange height for the +// declared variant, or it stops pushing on the flange and starts +// pushing on the plate body -- into gripper jaw travel, into a +// skirt chamfer that will cam the plate upward, or into the well +// region. Minimum flange height is nominal minus 0.38 mm. +// +// lip_window_ok() checks that window against the declared variant and the +// report prints PASS or FAIL. On a SLAS 3 4.1 short-flange plate (2.41 nominal, +// 2.03 minimum) the default 3.0 mm lip FAILS, correctly -- short-flange plates +// need a shorter lip and there is not much window left. +// +// Corner lips rather than continuous rails, for three reasons: +// 1. SLAS 3 4.4 permits a single interruption on center of each long side. +// A continuous long-side rail can land on that interruption; a corner lip +// cannot, because the interruption edges are at least 47.8 mm from the +// nearest part edge. +// 2. The mid-side gaps are where a gripper puts its jaws. +// 3. The downhill mid-side gap is the drain path for anything spilled. +// +// The lip is loaded as a transverse cantilever, and its root is a layer +// interface -- the weakest direction in an FDM part. That is why lip_thickness +// is generous and why the outer face tapers: the section modulus is highest +// exactly where the bending moment is. See README for the orientation argument. +// ============================================================================= + +// Corner lips are the default and the argued-for choice above. "full" exists +// because some reservoirs have no flange worth catching at a corner, and it +// keeps its own downhill drain slot. +lip_style = "corners"; // "corners" | "full" + +lip_height_mm = 3.0; // downhill and side lips, above the seat plane, mm +lip_height_uphill_mm = 3.0; // uphill lips -- reduce for gripper clearance, mm +lip_thickness_mm = 2.4; // at the lip top; 6 walls at a 0.4 mm nozzle, mm +lip_fillet_mm = 1.2; // extra thickness at the lip root, tapered, mm +lip_arm_mm = 24.0; // length of each corner-lip arm along the side, mm +lip_lead_in_mm = 1.0; // chamfer at the lip top inner edge, guides entry, mm +lip_min_engage_mm = 1.2; // lower bound of the lip height window, mm +lip_flange_margin_mm = 0.5; // how far the lip stays below the minimum flange, mm + +// Platform slab under the seat. Thin enough not to waste Z, thick enough that +// the seat is not a drum. mm. +platform_thickness_mm = 3.0; + +// The seat is relieved in the middle so the plate lands on a perimeter band +// rather than on whatever high spot the print left. A printed plane cannot be +// flatter than the bed it was printed on, mesh leveling conforms the first layer +// TO the bed rather than correcting it, and a 128 mm seat that bulges 0.2 mm in +// the middle makes the plate rock. A relieved center converts an unknown +// flatness into a known three-sided contact. mm. +seat_band_mm = 12.0; +seat_relief_depth_mm = 0.6; + +// Drain channel from the relieved center out through the downhill edge. A spill +// that pools under a plate in a deck nest is both a contamination problem and a +// seating problem -- the plate floats on the film and stops sitting where it was +// taught. Liquid must have somewhere to go that is not "under the plate". mm. +drain_slot_w_mm = 8.0; + + +// ============================================================================= +// 5. BASE AND DECK REGISTRATION +// +// UNMEASURED IS UNTRUSTED. deck_measured stays false until you have put calipers +// on your own deck position, and tilt_fixture() will not render until you flip +// it. See the SAFETY -- DECK CRASH note in the header for why this is a hard +// stop and not a warning. +// ============================================================================= + +deck_measured = false; + +// "slas_skirt" -- the fixture's own base carries an ANSI/SLAS-1 footprint skirt +// and drops into the deck nest exactly the way a plate does. +// Preferred: it reuses location the deck already provides, and +// needs one measurement (the nest depth) rather than a hole +// pattern. +// "post" -- locating posts drop into measured holes in the deck. +// "none" -- NOTHING LOCATES THE FIXTURE. Only defensible when the fixture +// is clamped or bolted by other means. The report says so out +// loud every render. +registration_style = "slas_skirt"; + +// MEASURE. Depth from the deck nest rim down to the surface a plate's own +// footprint rests on. Negative means unmeasured. mm. +deck_nest_depth_mm = -1; + +// MEASURE. Width of the nest rim, if the fixture is to rest on it. Negative +// means unmeasured. mm. +deck_nest_rim_w_mm = -1; + +// Which surface carries the fixture's weight. These are mutually exclusive and +// a fixture that is unsure which one it is will rock, because both cannot be in +// contact at once on a printed part. +// "nest_floor" -- the skirt runs deeper than the nest and the body shoulder +// floats clear of the rim by seat_relief_mm. +// "nest_rim" -- the skirt runs shallower and the body shoulder lands on the +// rim, with the skirt bottom floating. +seats_on = "nest_floor"; +seat_relief_mm = 0.4; // the deliberate gap on whichever surface is NOT carrying, mm + +// Registration skirt wall. Also the minimum wall anywhere in the part. mm. +wall_mm = 2.4; + +// Post registration, if registration_style is "post". All MEASURED off your +// deck. Negative means unmeasured and the model will refuse. mm. +// +// TWO posts, not four. Four posts in four holes over-constrain the part: with +// desktop FDM positional error on a 100 mm span, at least one post binds and the +// fixture seats on three of them, rocking. Two posts locate position and +// rotation, which is all that is needed. If the deck offers four holes, use two +// and leave the others empty; if you must use more, one hole has to become a +// slot, and this model does not cut slots for you. +reg_post_d_mm = -1; +reg_post_pitch_x_mm = -1; +reg_post_h_mm = -1; + +// How deep the registration feature engages the deck. Derived for the skirt +// case, explicit for posts. Reported so it can be sanity-checked against the +// deck: a 1 mm engagement is decoration, not registration. mm. +reg_engage_min_mm = 2.0; + +// Flat slab spanning the whole footprint, under the wedge. Takes the +// registration loads and gives the wedge something continuous to sit on. mm. +base_thickness_mm = 4.0; + +// Minimum wedge thickness at the LOW corner, above the base slab. Without it the +// wedge tapers to a knife edge that will not print and will not survive being +// picked up. mm. +min_lift_mm = 2.0; + +// Extra footprint on the body beyond the projected platform, so the wedge walls +// slope slightly inward going up and never overhang. mm. +body_margin_mm = 0.6; + + +// ============================================================================= +// 6. LOW PROFILE +// +// Why total stack height matters, and what it is actually made of. +// +// Raising the plate spends three budgets at once: +// - the head's Z envelope. The tip has to reach the bottom of the LOWEST well +// while the head still has travel left, and the head has to clear the +// HIGHEST point of the plate on every move in and out. +// - the gripper's approach. A plate mover taught to grab a plate at deck +// height will now find it 25 mm higher and tilted. The jaws must still close +// on the flange, and the approach path must clear the lips. +// - the neighbors. This body is wider than the plate it holds. Check the +// adjacent deck positions before you print, not after. +// +// The load-bearing fact is in stack_budget(): the fixed part of the stack (base, +// lift, platform slab) is a few millimeters, and the angle-driven part is the +// plate length times sin(tilt). At 7 degrees across a 128 mm pocket that is +// about 15.7 mm, and low_profile does not touch it. You cannot thin your way +// out of the angle. If the stack does not fit, the angle comes down or the +// fixture does not get used. +// ============================================================================= + +low_profile = false; + +low_profile_base_thickness_mm = 2.4; +low_profile_min_lift_mm = 1.2; +low_profile_platform_thickness_mm = 2.0; + + +// ============================================================================= +// 7. TEST COUPON +// +// A 90 x 55 mm part that carries the registration skirt corner, one downhill +// corner lip cut from the real geometry, and a caliper witness block. It prints +// in minutes instead of hours and it is the only thing standing between a wrong +// clearance and a wasted afternoon. +// +// The coupon validates FEATURES, not the whole-footprint fit. It tells you what +// your printer did to a wall thickness, a corner radius, a vertical face and a +// 20 mm nominal. It cannot tell you whether a 127.76 mm skirt fits your nest, +// because it is not 127.76 mm long. Measure the gauge, compute your machine's +// actual scale error, apply it to the full length, and only then print the +// fixture. +// ============================================================================= + +coupon_gauge_mm = 20.0; // nominal witness block in X and Y. Measure both. mm +coupon_gauge_z_mm = 10.3; // deliberately not a multiple of a common layer height, mm +coupon_pad_t_mm = 2.0; // common pad so the pieces stay together, mm +coupon_sample_mm = 34.0; // how much of each real corner to keep, mm + + +// ============================================================================= +// 8. DERIVED GEOMETRY -- functions, never stored constants +// ============================================================================= + +function r2(v) = round(v * 100) / 100; + +// Low-profile substitutions. +function base_t() = low_profile ? low_profile_base_thickness_mm : base_thickness_mm; +function lift() = low_profile ? low_profile_min_lift_mm : min_lift_mm; +function plat_t() = low_profile ? low_profile_platform_thickness_mm : platform_thickness_mm; + +// The pocket. The footprint tolerance is on the OVERALL dimension so it is added +// once; the clearance is per side so it is added twice. Conflating the two is +// the most common way to get this wrong by half a millimeter. +function pocket_len() = plate_len_mm + plate_footprint_tol_mm + 2 * clearance_mm; +function pocket_wid() = plate_wid_mm + plate_footprint_tol_mm + 2 * clearance_mm; +function pocket_r() = plate_corner_r_max_mm + clearance_mm; + +// The platform carrying the pocket and the lips, measured IN the seat plane. +function plat_len() = pocket_len() + 2 * (lip_thickness_mm + lip_fillet_mm); +function plat_wid() = pocket_wid() + 2 * (lip_thickness_mm + lip_fillet_mm); +function plat_r() = pocket_r() + lip_thickness_mm + lip_fillet_mm; + +// Unit normal of the seat plane, z component. Everything vertical scales by it. +function seat_nz() = cos(cross_tilt_deg) * cos(tilt_angle_deg); + +// TRUE slope of the seat plane. With both tilts nonzero this is steeper than +// either component, which is the trap in compound tilt: two comfortable-sounding +// 6 degree tilts make an 8.5 degree plane. +function max_slope_deg() = acos(seat_nz()); + +// --------------------------------------------------------------------------- +// tip_reach_delta_mm() +// +// THE number that decides whether the tips can still bottom out. +// +// Two well bottoms separated by dx along the plate's long axis and dy along its +// short axis are separated vertically, once the plate is tilted, by +// +// dz = dx * sin(tilt) + dy * sin(cross) * cos(tilt) +// +// sin, not tan: the wells are separated by dx measured ALONG the plate, which is +// now the hypotenuse rather than the deck projection. Using tan overstates the +// drop, and at shallow angles the two differ by less than a percent, which is +// exactly the kind of error that survives review. +// +// The plate seats against the DOWNHILL corner lips, so the near corner is the A1 +// end and the far corner is the last well of the last column. The head has to +// travel tip_reach_delta_mm() further down at the far corner than at the near +// one, on top of whatever the fixture already added. If the head is near the end +// of its Z stroke at the near corner -- which is where a deep-well plate on a +// tall fixture puts it -- the far corner is where it runs out. +// +// Span is between well CENTERS, from SLAS 4: (cols - 1) * pitch by +// (rows - 1) * pitch. For a 96-well plate that is 99.0 x 63.0 mm, not the +// 127.76 x 85.48 footprint. Using the footprint overstates the drop by about +// 30%, which sounds conservative until it talks somebody out of a workable +// angle. +// --------------------------------------------------------------------------- +function well_span_x_mm() = (well_cols - 1) * well_pitch_mm; +function well_span_y_mm() = (well_rows - 1) * well_pitch_mm; + +function drop_along_seat_mm(dx, dy) = + dx * sin(tilt_angle_deg) + dy * sin(cross_tilt_deg) * cos(tilt_angle_deg); + +function tip_reach_delta_mm() = + drop_along_seat_mm(well_span_x_mm(), well_span_y_mm()); + +// Same computation across the plate footprint rather than the well grid. This is +// the one that matters for the plate's own corners and for the lips. +function footprint_drop_mm() = drop_along_seat_mm(pocket_len(), pocket_wid()); + +// Height of the seat plane under a given well center, in the deck frame. The +// plate seats against the DOWNHILL lips, so a well's height is set by how far it +// sits from the downhill edges -- which is where the SLAS 4 first-well offsets +// stop being trivia and start being the thing that positions the tip. Column and +// row are 1-based; column 1 is the A1 end, which is uphill. +function well_dist_from_downhill_x_mm(col) = + plate_len_mm - (well_a1_x_mm + (col - 1) * well_pitch_mm); +function well_dist_from_downhill_y_mm(row) = + plate_wid_mm - (well_a1_y_mm + (row - 1) * well_pitch_mm); + +function well_seat_z_mm(col, row) = + plate_low_seat_z() + + drop_along_seat_mm(well_dist_from_downhill_x_mm(col), + well_dist_from_downhill_y_mm(row)); + +// --------------------------------------------------------------------------- +// Vertical stack, referenced to the plane a plate's own footprint would rest on +// in this deck position -- so a plate sitting flat has its bottom at z = 0 and +// every number below is directly comparable to "no fixture". +// --------------------------------------------------------------------------- + +// How far the registration feature engages the deck, and where the body starts. +function skirt_h_mm() = + registration_style != "slas_skirt" ? 0 + : deck_nest_depth_mm < 0 ? 0 + : seats_on == "nest_floor" ? deck_nest_depth_mm + seat_relief_mm + : deck_nest_depth_mm - seat_relief_mm; + +function body_bottom_z() = registration_style == "slas_skirt" ? skirt_h_mm() : 0; +function body_top_z() = body_bottom_z() + base_t(); + +// Seat plane at the low (downhill) corner of the PLATFORM. +function seat_low_z() = body_top_z() + lift() + plat_t() * seat_nz(); + +// Seat plane under the plate's own low and high corners. The plate is inset from +// the platform edge by one lip thickness plus its fillet, so its low corner sits +// a hair above the platform's. +function plate_low_seat_z() = + seat_low_z() + drop_along_seat_mm(lip_thickness_mm + lip_fillet_mm, + lip_thickness_mm + lip_fillet_mm); + +function plate_high_seat_z() = plate_low_seat_z() + footprint_drop_mm(); + +// Highest point of the loaded plate above the deck datum. Valid only if +// plate_height_mm is this plate's real height -- see the SLAS 2 note in section 1. +function stack_top_mm() = plate_high_seat_z() + plate_height_mm * seat_nz(); + +// What the stack is made of. Returns [fixed, angle_driven, plate]. +function stack_budget() = [ + seat_low_z(), // skirt + base + lift + platform slab + footprint_drop_mm(), // set by the angle alone + plate_height_mm * seat_nz() // the plate itself +]; + +// How much taller this deck position became versus a plate sitting flat. +function stack_penalty_mm() = stack_top_mm() - plate_height_mm; + +// Body footprint. Sized from the projection of the tilted platform plus a +// margin, so the wedge walls lean inward going up and never overhang. +function body_len() = plat_len() * cos(tilt_angle_deg) + 2 * body_margin_mm; +function body_wid() = plat_wid() * cos(cross_tilt_deg) + 2 * body_margin_mm; +function body_r() = plat_r(); + +// How far the body overhangs the SLAS skirt on each side. This is what collides +// with the next deck position. +function overhang_x_mm() = (body_len() - plate_len_mm) / 2; +function overhang_y_mm() = (body_wid() - plate_wid_mm) / 2; + +// Registration skirt outline: an SLAS footprint minus clearance, so it drops +// into a nest built for a plate. +function skirt_len() = plate_len_mm - 2 * deck_clearance_mm; +function skirt_wid() = plate_wid_mm - 2 * deck_clearance_mm; +function skirt_r() = 3.18; // SLAS 1 nominal; a shaft wants the SMALL radius + +// The lip height window, from the declared flange variant. +function flange_min_mm() = plate_flange_h_mm - plate_flange_tol_mm; +function lip_max_mm() = flange_min_mm() - lip_flange_margin_mm; +function lip_window_ok() = + lip_height_mm >= lip_min_engage_mm && lip_height_mm <= lip_max_mm(); + +// Tip-over about the downhill lip contact line. Refuses without a measured CG. +function tip_over_angle_deg() = + plate_cg_height_mm <= 0 ? undef + : atan((plate_len_mm / 2) / plate_cg_height_mm); + +// Pivot height for the platform rotation, chosen so the platform's bottom-face +// low corner lands exactly lift() above the base slab. +function pivot_z() = + body_top_z() + lift() + + (plat_len() / 2) * sin(tilt_angle_deg) + + (plat_wid() / 2) * sin(cross_tilt_deg) * cos(tilt_angle_deg) + + plat_t() * seat_nz(); + + +// ============================================================================= +// 9. REPORT +// +// Printed on every render, including the coupon. Anything that depends on an +// unmeasured input says NOT COMPUTED and names the parameter. +// ============================================================================= + +module report() { + echo(""); + echo("=== tilt_module.scad ==="); + echo(str("part : ", part, + low_profile ? " (low profile)" : "")); + echo(""); + echo("--- angle ---"); + echo(str("tilt / cross : ", tilt_angle_deg, " / ", cross_tilt_deg, " deg")); + echo(str("true seat plane slope : ", r2(max_slope_deg()), " deg")); + echo(""); + echo("--- reach ---"); + echo(str("well grid : ", well_cols, " x ", well_rows, + " at ", well_pitch_mm, " mm pitch")); + echo(str("well center span : ", r2(well_span_x_mm()), " x ", + r2(well_span_y_mm()), " mm")); + echo(str("TIP REACH DELTA : ", r2(tip_reach_delta_mm()), + " mm extra Z at the far corner")); + echo(str("seat under first well : ", r2(well_seat_z_mm(1, 1)), + " mm above the deck datum (uphill, shallowest)")); + echo(str("seat under last well : ", r2(well_seat_z_mm(well_cols, well_rows)), + " mm above the deck datum (downhill, deepest)")); + echo(str("plate corner drop : ", r2(footprint_drop_mm()), " mm")); + echo(""); + echo("--- vertical stack, versus a plate sitting flat ---"); + echo(str(" fixed (base+lift+slab): ", r2(stack_budget()[0]), " mm")); + echo(str(" angle (irreducible) : ", r2(stack_budget()[1]), " mm")); + echo(str(" plate : ", r2(stack_budget()[2]), " mm")); + echo(str("top of plate above datum : ", r2(stack_top_mm()), " mm")); + echo(str("EXTRA Z CONSUMED : ", r2(stack_penalty_mm()), + " mm -- head envelope and gripper approach must both still clear")); + echo(str("plate height basis : ", plate_height_mm == 14.35 + ? "SLAS 2 standard-height nominal -- WRONG for deep-well/PCR/reservoir" + : "operator supplied")); + echo(""); + echo("--- footprint ---"); + echo(str("pocket : ", r2(pocket_len()), " x ", r2(pocket_wid()), + " mm (tol once, clearance twice)")); + echo(str("body : ", r2(body_len()), " x ", r2(body_wid()), " mm")); + echo(str("overhang beyond SLAS : ", r2(overhang_x_mm()), " mm per side in X, ", + r2(overhang_y_mm()), " mm per side in Y -- CHECK NEIGHBORING DECK POSITIONS")); + echo(""); + echo("--- lips ---"); + echo(str("flange variant declared : ", plate_flange_h_mm, " +/- ", plate_flange_tol_mm, + " mm (minimum ", r2(flange_min_mm()), ")")); + echo(str("lip height window : ", lip_min_engage_mm, " to ", r2(lip_max_mm()), " mm")); + echo(str("lip height ", lip_height_mm, " mm : ", + lip_window_ok() ? "PASS" + : lip_height_mm < lip_min_engage_mm + ? "FAIL -- too short to engage the flange face" + : "FAIL -- taller than the minimum flange, will push on the plate body")); + echo(str("static hold without lips: NOT COMPUTED -- friction coefficient between a ", + "printed lip and a molded skirt is unpublished and changes when wet. ", + "Design as if it is zero.")); + echo(str("lip load : NOT COMPUTED -- no SLAS standard specifies ", + "plate mass. Weigh your loaded plate.")); + echo(str("tip-over angle : ", + plate_cg_height_mm <= 0 + ? "NOT COMPUTED -- set plate_cg_height_mm from a measurement" + : str(r2(tip_over_angle_deg()), " deg, versus ", r2(max_slope_deg()), + " deg of tilt"))); + echo(""); + echo("--- deck registration ---"); + echo(str("style : ", registration_style)); + echo(str("deck measured : ", deck_measured ? "yes" : "NO")); + echo(str("nest depth : ", deck_nest_depth_mm < 0 + ? "NOT MEASURED" : str(deck_nest_depth_mm, " mm"))); + echo(str("skirt engagement : ", skirt_h_mm() <= 0 + ? "NOT COMPUTED" + : str(r2(skirt_h_mm()), " mm", + skirt_h_mm() < reg_engage_min_mm + ? " -- BELOW reg_engage_min_mm, this is decoration not registration" + : ""))); + echo(str("carried by : ", seats_on)); + echo(str("shoulder on rim : ", + seats_on != "nest_rim" ? "n/a -- shoulder floats by seat_relief_mm" + : deck_nest_rim_w_mm < 0 + ? "NOT COMPUTED -- set deck_nest_rim_w_mm from a measurement" + : str(r2(min(overhang_x_mm(), overhang_y_mm())), " mm of shoulder over a ", + deck_nest_rim_w_mm, " mm rim: ", + min(overhang_x_mm(), overhang_y_mm()) <= deck_nest_rim_w_mm + ? "PASS" : "FAIL -- shoulder overhangs the rim and will rock"))); + if (registration_style == "none") + echo("HAZARD : registration_style is none. Nothing locates this fixture. A fixture that can move under gantry acceleration is a deck crash and a re-teach."); + echo(""); + echo("--- what is not modeled ---"); + echo("recovery improvement : NOT MODELED. Weigh it. See hardware/README.md."); + echo("sample contact : NOT QUALIFIED. This is a holder; the plate is the barrier."); + echo(""); +} + + +// ============================================================================= +// 10. GEOMETRY +// ============================================================================= + +// Rounded rectangle, centered. Radius is clamped so a small part cannot produce +// a degenerate profile. +module rrect(l, w, r) { + rr = min(r, l / 2 - 0.01, w / 2 - 0.01); + offset(r = rr) square([l - 2 * rr, w - 2 * rr], center = true); +} + +// Places children from the seat-plane frame (origin at the platform center, seat +// plane at z = 0, +x downhill) into the deck frame. +module tilt_place() { + translate([0, 0, pivot_z()]) + rotate([-cross_tilt_deg, tilt_angle_deg, 0]) + children(); +} + +// The platform slab. Seat surface is its top face, at z = 0 in the seat frame. +module platform_slab() { + translate([0, 0, -plat_t()]) + linear_extrude(height = plat_t()) + rrect(plat_len(), plat_wid(), plat_r()); +} + +// Full lip ring at height h, outer face tapering inward going up so the root is +// thicker than the tip. The inner face stays vertical: that is the plate contact +// and it should be a face, not a slope. +module lip_ring(h) { + difference() { + linear_extrude(height = h, + scale = [(plat_len() - 2 * lip_fillet_mm) / plat_len(), + (plat_wid() - 2 * lip_fillet_mm) / plat_wid()]) + rrect(plat_len(), plat_wid(), plat_r()); + translate([0, 0, -1]) + linear_extrude(height = h + 2) + rrect(pocket_len(), pocket_wid(), pocket_r()); + } +} + +// Lead-in chamfer at the top inner edge of a lip of height h. Cut per lip rather +// than once globally, because the uphill lips may be shorter and a single cut at +// one height would either miss them or take a second bite out of the tall ones. +module lip_lead_in_cut(h) { + translate([0, 0, h - lip_lead_in_mm]) + linear_extrude(height = lip_lead_in_mm + 0.1, + scale = [(pocket_len() + 2 * lip_lead_in_mm) / pocket_len(), + (pocket_wid() + 2 * lip_lead_in_mm) / pocket_wid()]) + rrect(pocket_len(), pocket_wid(), pocket_r()); +} + +// One corner lip, in the (+x, +y) quadrant, cut out of the full ring so its +// cross-section and corner radius are exactly the ring's. +module lip_corner_pp(h) { + ox = plat_len() / 2; + oy = plat_wid() / 2; + band = lip_thickness_mm + lip_fillet_mm; + intersection() { + difference() { + lip_ring(h); + lip_lead_in_cut(h); + } + union() { + translate([ox - lip_arm_mm, oy - band - 1, -1]) + cube([lip_arm_mm + 1, band + 2, h + 2]); + translate([ox - band - 1, oy - lip_arm_mm, -1]) + cube([band + 2, lip_arm_mm + 1, h + 2]); + } + } +} + +// Four corner lips. Downhill corners (+x) get lip_height_mm; uphill corners get +// lip_height_uphill_mm, which can be dropped for gripper clearance because the +// uphill lip only does work during transport and placement. +module lips_all() { + if (lip_style_is_corners()) { + lip_corner_pp(lip_height_mm); // +x +y downhill + mirror([0, 1, 0]) lip_corner_pp(lip_height_mm); // +x -y downhill + mirror([1, 0, 0]) lip_corner_pp(lip_height_uphill_mm); // -x +y uphill + mirror([1, 0, 0]) mirror([0, 1, 0]) + lip_corner_pp(lip_height_uphill_mm); // -x -y uphill + } else { + difference() { + lip_ring(lip_height_mm); + lip_lead_in_cut(lip_height_mm); + // Keep the downhill mid-side drain open even on a continuous ring. + translate([plat_len() / 2 - 20, -drain_slot_w_mm / 2, -1]) + cube([40, drain_slot_w_mm, lip_height_mm + 2]); + } + } +} + +function lip_style_is_corners() = lip_style == "corners"; + +// Relieved seat center plus the downhill drain channel. +module seat_relief_cut() { + translate([0, 0, -seat_relief_depth_mm]) + linear_extrude(height = seat_relief_depth_mm + 1) + union() { + rrect(pocket_len() - 2 * seat_band_mm, + pocket_wid() - 2 * seat_band_mm, + max(pocket_r() - seat_band_mm, 1)); + // channel out through the downhill edge + translate([plat_len() / 4, 0]) + square([plat_len() / 2 + 2, drain_slot_w_mm], center = true); + } +} + +// Flat slab spanning the whole footprint. +module body_slab() { + translate([0, 0, body_bottom_z()]) + linear_extrude(height = base_t()) + rrect(body_len(), body_wid(), body_r()); +} + +// The wedge: a hull from the body slab top to the tilted platform's bottom face. +module wedge() { + hull() { + tilt_place() + translate([0, 0, -plat_t()]) + linear_extrude(height = 0.01) + rrect(plat_len(), plat_wid(), plat_r()); + translate([0, 0, body_top_z() - 0.01]) + linear_extrude(height = 0.01) + rrect(body_len(), body_wid(), body_r()); + } +} + +// SLAS-footprint registration skirt: a continuous hollow wall under the body. +// SLAS 1 4.1.1.3 requires the footprint be continuous and uninterrupted around +// the base. So does this one, for the same reason -- a gap snags a nest. +module skirt() { + h = skirt_h_mm(); + if (h > 0) + linear_extrude(height = h) + difference() { + rrect(skirt_len(), skirt_wid(), skirt_r()); + rrect(skirt_len() - 2 * wall_mm, skirt_wid() - 2 * wall_mm, + max(skirt_r() - wall_mm, 0.5)); + } +} + +// Locating posts, when the deck offers holes rather than a nest. Two, on the +// long axis. See the reg_post_* comment for why not four. +module posts() { + if (registration_style == "post" && reg_post_d_mm > 0) + for (sx = [-1, 1]) + translate([sx * reg_post_pitch_x_mm / 2, 0, -reg_post_h_mm]) + cylinder(h = reg_post_h_mm + 0.01, + d = reg_post_d_mm - 2 * deck_clearance_mm); +} + +module registration_feature() { + if (registration_style == "slas_skirt") skirt(); + else if (registration_style == "post") posts(); +} + +// Everything above the deck, without the registration feature. Factored out so +// the coupon can cut a corner from the identical geometry rather than a +// re-derivation of it -- a coupon that models the feature separately validates +// the coupon, not the fixture. +module upper_solid() { + difference() { + union() { + body_slab(); + wedge(); + tilt_place() union() { + platform_slab(); + lips_all(); + } + } + tilt_place() seat_relief_cut(); + } +} + +module tilt_fixture() { + assert(deck_measured, + "REFUSING TO RENDER: deck_measured is false. An unlocated fixture is a deck crash. Put calipers on the deck position, set deck_nest_depth_mm (or the reg_post_* values), then set deck_measured = true. Render part=\"coupon\" first -- it has no such requirement."); + assert(registration_style != "slas_skirt" || deck_nest_depth_mm > 0, + "REFUSING TO RENDER: registration_style is \"slas_skirt\" but deck_nest_depth_mm is unmeasured."); + assert(registration_style != "post" || + (reg_post_d_mm > 0 && reg_post_pitch_x_mm > 0 && reg_post_h_mm > 0), + "REFUSING TO RENDER: registration_style is \"post\" but one or more reg_post_* values are unmeasured."); + + difference() { + union() { + registration_feature(); + upper_solid(); + } + // Trim anything that strayed below the registration datum. + translate([-400, -400, -400 + min(0, -reg_post_h_mm)]) + cube([800, 800, 400]); + } +} + + +// ============================================================================= +// 11. TEST COUPON +// ============================================================================= + +// The downhill (+x, +y) corner of the real fixture, trimmed to a printable pad. +// Cut from upper_solid() rather than re-modeled, so what you measure on the +// coupon is what the fixture will do. A coupon built from its own geometry +// validates the coupon. +module coupon_seat_corner() { + s = coupon_sample_mm; + z0 = body_bottom_z(); + ztop = seat_low_z() + lip_height_mm + 2; + translate([-(body_len() / 2 - s), -(body_wid() / 2 - s), -z0]) + intersection() { + upper_solid(); + translate([body_len() / 2 - s, body_wid() / 2 - s, z0 - 0.01]) + cube([s + 2, s + 2, ztop - z0 + 1]); + } +} + +// A corner of the registration skirt at its real wall thickness and radius. +module coupon_skirt_corner() { + s = coupon_sample_mm; + h = max(skirt_h_mm(), 4.0); // if the nest is unmeasured, print a 4 mm sample + translate([-(skirt_len() / 2 - s), -(skirt_wid() / 2 - s), 0]) + intersection() { + linear_extrude(height = h) + difference() { + rrect(skirt_len(), skirt_wid(), skirt_r()); + rrect(skirt_len() - 2 * wall_mm, skirt_wid() - 2 * wall_mm, + max(skirt_r() - wall_mm, 0.5)); + } + translate([skirt_len() / 2 - s, skirt_wid() / 2 - s, -1]) + cube([s + 2, s + 2, h + 2]); + } +} + +// Witness block. Measure all three axes with calipers. X and Y give the printer's +// in-plane scale error, which is what clearance_mm and deck_clearance_mm are +// compensating for. Z gives layer quantization plus first-layer squish, which is +// a constant offset on total height rather than a percentage, so do not convert +// the Z error to a percent and apply it to a length. +// +// coupon_gauge_z_mm is deliberately NOT a multiple of any common layer height. +// A 20.0 mm tall gauge at a 0.2 mm layer is an exact multiple and hides the +// quantization entirely; total height is forced to first_layer + n * layer, so a +// height that lands between two layers is the only one that shows you the error. +module coupon_gauge() { + cube([coupon_gauge_mm, coupon_gauge_mm, coupon_gauge_z_mm]); +} + +module test_coupon() { + s = coupon_sample_mm + 2; + pad_l = 2 * s + coupon_gauge_mm + 22; + pad_w = s + 8; + difference() { + union() { + cube([pad_l, pad_w, coupon_pad_t_mm]); + translate([4, 4, coupon_pad_t_mm]) coupon_seat_corner(); + translate([s + 10, 4, coupon_pad_t_mm]) coupon_skirt_corner(); + translate([2 * s + 16, 4, coupon_pad_t_mm]) coupon_gauge(); + } + // Nothing below the pad. + translate([-10, -10, -20]) cube([pad_l + 20, pad_w + 20, 20]); + } +} + + +// ============================================================================= +// 12. TOP LEVEL +// ============================================================================= + +report(); + +if (part == "fixture" || part == "both") tilt_fixture(); +if (part == "coupon" || part == "both") + translate([part == "both" ? body_len() / 2 + 30 : 0, 0, 0]) test_coupon(); diff --git a/tests/test_printed.py b/tests/test_printed.py new file mode 100644 index 0000000..2d35ceb --- /dev/null +++ b/tests/test_printed.py @@ -0,0 +1,760 @@ +"""Device-free tests for the printed-fixture fitness layer. + +The tests that matter here try to get a printed part reported as fit when it is not, and the +richest source of that is the vacuous pass: a check that passes when handed nothing. A +fixture with no fields set. A fitness computation over a part with no declared dimensions, +where `all()` over an empty tuple is True and a part nobody measured reads as fully measured. +A material with no glass transition on record going into an autoclave check that has nothing +to compare and returns a cheerful default. A designed dimension, which is a number with units +in a document and is a fact about a model rather than about the part that came out. + +The other half of the file guards the refusals. Five of them are load-bearing -- a porous +process in a fluid path, an uncleared photopolymer at culture contact, a cycle above the +glass transition, an unmeasured dimension a robot registers against, and an unlocated part +inside a moving envelope -- and each has to keep refusing when the obvious workaround is +applied to it. Post-processing must not clear porosity. Heat deflection must not clear an +autoclave. A benefit claim resting on somebody's opinion must not clear a measurement. +""" + +from __future__ import annotations + +import pytest + +from autonomous_lab import printed +from autonomous_lab.printed import ( + FIXTURES, + MATERIALS, + NOT_ESTIMATED, + POST_PROCESS_EVIDENCE, + USES, + Assessment, + Basis, + Benefit, + Contact, + CriticalDimension, + DimensionState, + Fitness, + Fixture, + Flange, + Material, + Objection, + PostProcess, + Process, + Qualification, + ResinQualification, + Sterilization, + Thermal, + Use, + fitness, + materials_reaching, + materials_surviving, + plate_envelope, + refusals, + survey, + unqualified, +) + + +def _measured(name="d", nominal=10.0, value=9.9, after=True): + return CriticalDimension( + name=name, + nominal_mm=nominal, + state=DimensionState.MEASURED, + measured_mm=value, + measured_after_processing=after, + instrument="digital caliper", + ) + + +def _benefit(): + return Benefit(claim="saves a hand motion", measured="40 s over eight runs", basis=Basis.IN_HOUSE) + + +def _good_fixture(material=None, **kw): + """A fixture with everything a lab controls already done, so a test can vary one thing.""" + defaults = dict( + name="controlled", + material=material if material is not None else MATERIALS["asa_fdm"], + built_for=Contact.CULTURE_CONTACT, + dimensions=(_measured(),), + positively_located=True, + benefit=_benefit(), + ) + defaults.update(kw) + return Fixture(**defaults) + + +def _benign_use(**kw): + """The least demanding use this module can express.""" + defaults = dict( + name="shelf", + contact=Contact.NONE, + robot_registers_to_it=False, + inside_moving_envelope=False, + sterilization=Sterilization.NONE, + ) + defaults.update(kw) + return Use(**defaults) + + +# -- vacuous passes -------------------------------------------------------------- + + +def test_a_fixture_with_no_fields_set_is_never_fit(): + """The single most important test in the file. + + A dataclass of booleans and empty tuples is the easiest thing in this module to get wrong, + because every default that reads as True or as 'nothing to check' turns a part nobody has + built into a part cleared for a deck. Even against the least demanding use expressible + here, a bare fixture has to come back unfit. + """ + result = fitness(Fixture(), _benign_use()) + assert not result.fit + assert result.verdict is not Fitness.FIT + assert Fitness.MATERIAL_UNDECLARED in result.refusals() + assert Fitness.BENEFIT_NOT_QUANTIFIED in result.refusals() + + +def test_a_bare_fixture_is_unqualified_in_every_respect(): + """`unqualified` defaults to everything, so a new print is untrusted the way a new op is.""" + standing = unqualified(Fixture()) + assert set(standing) == set(Qualification) + + +def test_a_fixture_that_declares_no_dimensions_does_not_read_as_all_measured(): + """`all()` over an empty tuple is True; a part nobody measured must not inherit that.""" + assert not Fixture().dimensions_measured() + assert Qualification.DIMENSIONS_MEASURED in unqualified(Fixture()) + + use = _benign_use(name="gripped", robot_registers_to_it=True) + result = fitness(Fixture(material=MATERIALS["asa_fdm"]), use) + assert Fitness.DIMENSION_NOT_MEASURED in result.refusals() + assert "no critical dimension is declared" in result.all_reasons()[ + result.refusals().index(Fitness.DIMENSION_NOT_MEASURED) + ] + + +def test_a_designed_dimension_is_not_a_measured_one(): + designed = CriticalDimension(name="pocket", nominal_mm=128.5) + assert designed.state is DimensionState.DESIGNED + assert not designed.state.measured + assert designed.deviation_mm() is None + assert designed.deviation_percent() is None + + fixture = _good_fixture(dimensions=(designed,)) + assert not fixture.dimensions_measured() + result = fitness(fixture, _benign_use(robot_registers_to_it=True)) + assert Fitness.DIMENSION_NOT_MEASURED in result.refusals() + + +def test_an_unchecked_dimension_is_not_a_measured_one_either(): + unchecked = CriticalDimension( + name="pocket", nominal_mm=128.5, state=DimensionState.UNCHECKED + ) + assert not unchecked.state.measured + assert not _good_fixture(dimensions=(unchecked,)).dimensions_measured() + + +def test_a_dimension_cannot_claim_measured_without_a_number(): + with pytest.raises(ValueError, match="carries no measured value"): + CriticalDimension(name="pocket", nominal_mm=10.0, state=DimensionState.MEASURED) + + +def test_a_dimension_cannot_claim_measured_without_naming_an_instrument(): + with pytest.raises(ValueError, match="names no instrument"): + CriticalDimension( + name="pocket", nominal_mm=10.0, state=DimensionState.MEASURED, measured_mm=9.9 + ) + + +def test_a_designed_dimension_cannot_carry_a_measured_number(): + """The ambiguity from the other side: a value present under a state that did not earn it.""" + with pytest.raises(ValueError, match="carries a measured value"): + CriticalDimension(name="pocket", nominal_mm=10.0, measured_mm=9.9) + + +def test_an_unmeasured_dimension_cannot_claim_it_was_measured_after_processing(): + with pytest.raises(ValueError, match="nothing was measured at all"): + CriticalDimension(name="pocket", nominal_mm=10.0, measured_after_processing=True) + + +def test_a_material_with_unknown_tg_does_not_pass_an_autoclave_check(): + """The other headline vacuous pass: nothing to compare, so nothing refuses. + + Two real materials sit in this state -- a polycarbonate blend whose sheet publishes heat + deflection at two loads and no glass transition, and a polypropylene copolymer whose sheet + publishes only a melting point. Both would sail through a check that skipped what it could + not evaluate. + """ + for key in ("pc_blend_fdm", "pp_copolymer_fdm", "photopolymer_certified", "machined_stock"): + mat = MATERIALS[key] + assert mat.tg is None + ok, why = mat.survives(Sterilization.AUTOCLAVE_121) + assert ok is None, f"{key} must not return a boolean verdict on a threshold it does not have" + assert "no glass transition is on record" in why + + result = fitness( + _good_fixture(material=mat), _benign_use(sterilization=Sterilization.AUTOCLAVE_121) + ) + assert Fitness.THERMAL_LIMIT_UNKNOWN in result.refusals() + assert not result.fit + + +def test_a_material_with_unknown_tg_is_not_listed_as_surviving(): + surviving = {m.name for m in materials_surviving(Sterilization.AUTOCLAVE_121)} + assert MATERIALS["pc_blend_fdm"].name not in surviving + assert MATERIALS["pp_copolymer_fdm"].name not in surviving + + +def test_a_fixture_cannot_declare_itself_fit(): + """There is no field that sets fitness; it is only ever computed, per use. + + The same rule the ledger applies to automation and `qc` applies to gate readiness. If this + ever fails, somebody has added an override and the report can be made to lie. + """ + fixture = FIXTURES["plate_nest_as_printed"] + assert not hasattr(fixture, "fit") + assert not hasattr(fixture, "fitness") + assert not hasattr(fixture, "verdict") + + +def test_a_benefit_nobody_measured_is_not_quantified(): + assert not Benefit(claim="feels faster").quantified + claimed = Benefit(claim="faster", measured="a forum post says 30 percent", basis=Basis.LITERATURE) + assert not claimed.quantified, "somebody else's measurement is not a measurement of this part" + assert Benefit(claim="faster", measured="40 s over eight runs", basis=Basis.IN_HOUSE).quantified + + +def test_an_unquantified_benefit_refuses_even_a_perfect_part(): + result = fitness(_good_fixture(benefit=None), _benign_use()) + assert Fitness.BENEFIT_NOT_QUANTIFIED in result.refusals() + assert not result.fit + + +def test_no_threshold_in_the_catalog_is_validated_here(): + """Every figure is a datasheet figure, and a datasheet is not evidence about your part.""" + assert survey()["validated_thresholds"] == 0 + for mat in MATERIALS.values(): + for thermal in (mat.tg, mat.hdt, mat.melt): + if thermal is not None: + assert not thermal.basis.validated + assert thermal.method, "a temperature with no test method behind it is not a threshold" + + +# -- refusal 1: a porous process in the fluid path -------------------------------- + + +def test_fdm_at_sample_contact_is_refused(): + result = fitness(_good_fixture(), Use(name="trough", contact=Contact.SAMPLE_CONTACT)) + assert Fitness.POROUS_PROCESS_IN_FLUID_PATH in result.refusals() + assert result.needs_a_different_part + + +def test_fdm_at_culture_contact_is_refused(): + result = fitness(_good_fixture(), Use(name="insert", contact=Contact.CULTURE_CONTACT)) + assert Fitness.POROUS_PROCESS_IN_FLUID_PATH in result.refusals() + + +def test_fdm_below_the_fluid_path_is_not_refused_for_porosity(): + """The gate is drawn at intended contact, not at proximity, so a bracket is not refused.""" + for contact in (Contact.NONE, Contact.INCIDENTAL): + result = fitness(_good_fixture(), _benign_use(contact=contact)) + assert Fitness.POROUS_PROCESS_IN_FLUID_PATH not in result.refusals() + + +def test_no_post_processing_clears_the_porosity_refusal(): + """The workaround every lab reaches for, refused per treatment with what it actually showed.""" + for treatment in PostProcess: + assert not treatment.closes_the_fluid_path + fixture = _good_fixture(post_processing=(treatment,)) + result = fitness(fixture, Use(name="trough", contact=Contact.SAMPLE_CONTACT)) + assert Fitness.POROUS_PROCESS_IN_FLUID_PATH in result.refusals() + detail = result.all_reasons()[result.refusals().index(Fitness.POROUS_PROCESS_IN_FLUID_PATH)] + assert treatment.value in detail, "the refusal must name the claimed fix it is rejecting" + + +def test_every_post_process_carries_what_it_was_actually_shown_to_do(): + for treatment in PostProcess: + assert treatment in POST_PROCESS_EVIDENCE + assert POST_PROCESS_EVIDENCE[treatment].strip() + + +def test_the_porosity_refusal_is_structural(): + """No measurement, procedure or purchase clears it for the part as built.""" + assert Fitness.POROUS_PROCESS_IN_FLUID_PATH.structural + result = fitness(_good_fixture(), Use(name="trough", contact=Contact.SAMPLE_CONTACT)) + assert result.needs_a_different_part + + +def test_an_uncharacterized_process_is_refused_on_a_different_ground(): + """Measured-porous and never-measured cost different things to clear and must not collapse.""" + assert not Process.SLS.cleanable_in_principle + assert not Process.SLS.porosity_characterized + assert Process.FDM.porosity_characterized + + powder = Material( + name="a powder-bed part", process=Process.SLS, ceiling=Contact.CULTURE_CONTACT + ) + result = fitness(_good_fixture(material=powder), Use(name="trough", contact=Contact.SAMPLE_CONTACT)) + assert Fitness.UNCHARACTERIZED_PROCESS in result.refusals() + assert Fitness.POROUS_PROCESS_IN_FLUID_PATH not in result.refusals() + assert not Fitness.UNCHARACTERIZED_PROCESS.structural + + +def test_no_sls_material_is_declared_here(): + """The enum member exists and no material claims it, because nothing here measured one.""" + assert not any(m.process is Process.SLS for m in MATERIALS.values()) + + +# -- refusal 2: photopolymer at culture contact ----------------------------------- + + +def test_photopolymer_at_culture_contact_without_a_certified_resin_is_refused(): + fixture = _good_fixture( + material=MATERIALS["photopolymer_certified"], + resin=ResinQualification( + certified_biocompatible=False, workflow_as_certified=True, post_cure_validated=True + ), + ) + result = fitness(fixture, Use(name="insert", contact=Contact.CULTURE_CONTACT)) + assert Fitness.CYTOTOXICITY_UNCLEARED in result.refusals() + detail = result.all_reasons()[result.refusals().index(Fitness.CYTOTOXICITY_UNCLEARED)] + assert "no certified biocompatible resin" in detail + + +def test_photopolymer_at_culture_contact_without_a_validated_post_cure_is_refused(): + fixture = _good_fixture( + material=MATERIALS["photopolymer_certified"], + resin=ResinQualification( + certified_biocompatible=True, workflow_as_certified=True, post_cure_validated=False + ), + ) + result = fitness(fixture, Use(name="insert", contact=Contact.CULTURE_CONTACT)) + assert Fitness.CYTOTOXICITY_UNCLEARED in result.refusals() + assert "no validated post-cure" in result.reason + + +def test_a_certified_resin_made_outside_the_certified_workflow_is_refused(): + """The certification is measured on one printer, one wash and one cure. That is the article.""" + fixture = _good_fixture( + material=MATERIALS["photopolymer_certified"], + resin=ResinQualification( + certified_biocompatible=True, workflow_as_certified=False, post_cure_validated=True + ), + ) + result = fitness(fixture, Use(name="insert", contact=Contact.CULTURE_CONTACT)) + assert Fitness.CYTOTOXICITY_UNCLEARED in result.refusals() + assert "not made by the workflow" in result.reason + + +def test_a_default_resin_qualification_clears_nothing(): + assert not ResinQualification().cleared_for_culture + assert len(ResinQualification().missing()) == 3 + + +def test_a_fully_cleared_photopolymer_passes_the_cytotoxicity_ground(): + """The module is not a machine that always says no; the path exists and it is narrow.""" + fixture = _good_fixture( + material=MATERIALS["photopolymer_certified"], + resin=ResinQualification( + certified_biocompatible=True, workflow_as_certified=True, post_cure_validated=True + ), + ) + result = fitness(fixture, Use(name="insert", contact=Contact.CULTURE_CONTACT)) + assert Fitness.CYTOTOXICITY_UNCLEARED not in result.refusals() + assert Fitness.POROUS_PROCESS_IN_FLUID_PATH not in result.refusals() + assert result.fit + + +def test_an_uncertified_photopolymer_is_refused_by_its_ceiling_as_well(): + fixture = _good_fixture( + material=MATERIALS["photopolymer_uncertified"], + resin=ResinQualification(True, True, True), + ) + result = fitness(fixture, Use(name="insert", contact=Contact.CULTURE_CONTACT)) + assert Fitness.ABOVE_MATERIAL_CEILING in result.refusals() + + +# -- refusal 3: an autoclave above the glass transition --------------------------- + + +def test_autoclaving_a_material_whose_tg_is_below_the_cycle_is_refused(): + for key in ("pla_fdm", "petg_fdm", "abs_fdm", "asa_fdm"): + mat = MATERIALS[key] + for cycle in (Sterilization.AUTOCLAVE_121, Sterilization.AUTOCLAVE_134): + ok, _why = mat.survives(cycle) + assert ok is False, f"{key} must not clear {cycle.value}" + result = fitness(_good_fixture(material=mat), _benign_use(sterilization=cycle)) + assert Fitness.SOFTENS_IN_THE_CYCLE in result.refusals() + + +def test_the_autoclave_check_reads_tg_and_not_hdt(): + """The material that makes the rule concrete: HDT clears 121 C by 49 C and Tg fails it by 51. + + A carbon-fiber nylon carries a published heat deflection temperature of 170 to 194 C on top + of a glass transition of 70 C, because the filler carries the load the matrix no longer can. + A check that read the HDT column would wave this part into a cycle it cannot survive. + """ + mat = MATERIALS["pa_cf_fdm"] + assert mat.hdt is not None and mat.hdt.conservative() > 121.0 + assert mat.tg is not None and mat.tg.conservative() < 121.0 + ok, why = mat.survives(Sterilization.AUTOCLAVE_121) + assert ok is False + assert "glass transition" in why + + +def test_the_high_temperature_polymers_clear_the_cycle_thermally(): + for key in ("pei_9085_fdm", "pei_1010_fdm", "peek_fdm"): + assert MATERIALS[key].survives(Sterilization.AUTOCLAVE_134)[0] is True + + +def test_a_non_thermal_method_does_not_invoke_the_thermal_gate(): + for method in (Sterilization.NONE, Sterilization.SINGLE_USE, Sterilization.IPA_WIPE): + assert not method.thermal + assert method.cycle_c is None + result = fitness(_good_fixture(material=MATERIALS["pla_fdm"]), _benign_use(sterilization=method)) + assert Fitness.SOFTENS_IN_THE_CYCLE not in result.refusals() + assert Fitness.THERMAL_LIMIT_UNKNOWN not in result.refusals() + + +def test_a_thermal_range_uses_its_low_end_for_the_gate(): + """Half the possible parts are past the line if the gate reads the middle of the spread.""" + band = Thermal(54.0, 61.0, "DSC 10 C/min", Basis.VENDOR) + assert band.conservative() == 54.0 + assert band.spread_c == 7.0 + + +def test_an_inverted_thermal_range_is_refused(): + with pytest.raises(ValueError, match="below low"): + Thermal(120.0, 90.0, "DSC", Basis.VENDOR) + + +# -- refusal 4: an unmeasured dimension where a robot registers ------------------- + + +def test_an_unmeasured_dimension_is_refused_where_a_robot_registers_to_it(): + fixture = _good_fixture(dimensions=(CriticalDimension(name="pocket", nominal_mm=128.5),)) + result = fitness(fixture, _benign_use(robot_registers_to_it=True)) + assert Fitness.DIMENSION_NOT_MEASURED in result.refusals() + assert "pocket" in result.reason + + +def test_a_partially_measured_set_of_dimensions_does_not_pass_on_the_half_that_was_measured(): + fixture = _good_fixture( + dimensions=(_measured(name="pitch"), CriticalDimension(name="offset", nominal_mm=22.0)) + ) + assert not fixture.dimensions_measured() + result = fitness(fixture, _benign_use(robot_registers_to_it=True)) + detail = result.all_reasons()[result.refusals().index(Fitness.DIMENSION_NOT_MEASURED)] + assert "offset" in detail + assert "pitch" not in detail + + +def test_a_measured_dimension_clears_the_gate(): + result = fitness(_good_fixture(), _benign_use(robot_registers_to_it=True)) + assert Fitness.DIMENSION_NOT_MEASURED not in result.refusals() + assert result.fit + + +def test_a_dimension_measured_before_a_thermal_cycle_is_stale(): + """Thermal and coating steps move printed parts by percent, and the sign depends on the axis.""" + fixture = _good_fixture( + material=MATERIALS["pei_1010_fdm"], dimensions=(_measured(name="pitch", after=False),) + ) + result = fitness( + fixture, _benign_use(robot_registers_to_it=True, sterilization=Sterilization.AUTOCLAVE_134) + ) + assert Fitness.DIMENSION_STALE_AFTER_CYCLE in result.refusals() + + +def test_a_dimension_measured_before_a_dimension_changing_post_process_is_stale(): + fixture = _good_fixture( + dimensions=(_measured(name="pitch", after=False),), post_processing=(PostProcess.ANNEALING,) + ) + result = fitness(fixture, _benign_use(robot_registers_to_it=True)) + assert Fitness.DIMENSION_STALE_AFTER_CYCLE in result.refusals() + + +def test_nothing_that_moves_dimensions_means_nothing_is_stale(): + """Staleness must be computed from something that actually happened, not asserted.""" + fixture = _good_fixture( + dimensions=(_measured(name="pitch", after=False),), + post_processing=(PostProcess.SOLVENT_EXTRACTION,), + ) + assert fixture.stale_dimensions(thermally_cycled=False) == () + result = fitness(fixture, _benign_use(robot_registers_to_it=True)) + assert Fitness.DIMENSION_STALE_AFTER_CYCLE not in result.refusals() + + +def test_a_dimension_gate_is_not_invoked_where_nothing_registers_to_the_part(): + fixture = _good_fixture(dimensions=()) + result = fitness(fixture, _benign_use(robot_registers_to_it=False)) + assert Fitness.DIMENSION_NOT_MEASURED not in result.refusals() + + +# -- refusal 5: an unlocated part inside a moving envelope ------------------------ + + +def test_an_unlocated_fixture_inside_a_moving_envelope_is_refused(): + fixture = _good_fixture(positively_located=False) + result = fitness(fixture, _benign_use(inside_moving_envelope=True)) + assert Fitness.NOT_POSITIVELY_LOCATED in result.refusals() + assert "crash" in result.reason + + +def test_an_unlocated_fixture_outside_any_envelope_is_not_refused_for_location(): + fixture = _good_fixture(positively_located=False) + result = fitness(fixture, _benign_use(inside_moving_envelope=False)) + assert Fitness.NOT_POSITIVELY_LOCATED not in result.refusals() + + +def test_a_use_defaults_to_the_demanding_geometry(): + """An under-specified use must not clear checks it never faced.""" + use = Use(name="unspecified", contact=Contact.NONE) + assert use.inside_moving_envelope + assert use.robot_registers_to_it + + +# -- the contact ladder ---------------------------------------------------------- + + +def test_contact_levels_are_ordered_and_the_fluid_path_line_is_where_it_says(): + order = [Contact.NONE, Contact.INCIDENTAL, Contact.SAMPLE_CONTACT, Contact.CULTURE_CONTACT] + assert [c.rank for c in order] == [0, 1, 2, 3] + assert not Contact.NONE.in_the_fluid_path + assert not Contact.INCIDENTAL.in_the_fluid_path + assert Contact.SAMPLE_CONTACT.in_the_fluid_path + assert Contact.CULTURE_CONTACT.in_the_fluid_path + + +def test_a_part_used_above_what_it_was_built_for_is_refused(): + fixture = _good_fixture(built_for=Contact.INCIDENTAL, material=MATERIALS["machined_stock"]) + result = fitness(fixture, Use(name="trough", contact=Contact.SAMPLE_CONTACT)) + assert Fitness.ABOVE_DECLARED_CONTACT in result.refusals() + + +def test_a_part_used_at_or_below_what_it_was_built_for_is_not_refused_for_that(): + fixture = _good_fixture(built_for=Contact.SAMPLE_CONTACT, material=MATERIALS["machined_stock"]) + result = fitness(fixture, Use(name="trough", contact=Contact.SAMPLE_CONTACT)) + assert Fitness.ABOVE_DECLARED_CONTACT not in result.refusals() + + +def test_the_material_ceiling_and_the_declared_contact_are_separate_refusals(): + """One is a fact about the polymer and the process; the other is about what was designed.""" + fixture = _good_fixture(built_for=Contact.NONE, material=MATERIALS["pla_fdm"]) + result = fitness(fixture, Use(name="insert", contact=Contact.CULTURE_CONTACT)) + assert Fitness.ABOVE_MATERIAL_CEILING in result.refusals() + assert Fitness.ABOVE_DECLARED_CONTACT in result.refusals() + + +# -- how refusals are reported ---------------------------------------------------- + + +def test_every_refusal_is_carried_not_just_the_first(): + """A printed part is usually wrong in several ways, and fixing the headline reveals the next.""" + result = fitness(Fixture(), Use(name="insert", contact=Contact.CULTURE_CONTACT)) + assert len(result.objections) >= 4 + assert result.verdict is result.objections[0].reason + assert len(result.all_reasons()) == len(result.objections) + + +def test_structural_refusals_are_reported_ahead_of_the_ones_money_can_move(): + fixture = _good_fixture(built_for=Contact.NONE, positively_located=False, benefit=None) + result = fitness( + fixture, + Use(name="trough", contact=Contact.SAMPLE_CONTACT, sterilization=Sterilization.AUTOCLAVE_121), + ) + reasons = result.refusals() + assert Fitness.POROUS_PROCESS_IN_FLUID_PATH in reasons + assert Fitness.NOT_POSITIVELY_LOCATED in reasons + assert reasons.index(Fitness.POROUS_PROCESS_IN_FLUID_PATH) < reasons.index( + Fitness.NOT_POSITIVELY_LOCATED + ) + assert result.verdict.structural + + +def test_only_the_two_refusals_a_different_part_fixes_are_structural(): + structural = {f for f in Fitness if f.structural} + assert structural == {Fitness.POROUS_PROCESS_IN_FLUID_PATH, Fitness.SOFTENS_IN_THE_CYCLE} + assert not Fitness.FIT.structural + + +def test_fit_is_the_absence_of_refusals_and_nothing_sets_it(): + empty = Assessment(fixture=Fixture(), use=_benign_use(), objections=()) + assert empty.verdict is Fitness.FIT + assert empty.fit + assert Fitness.FIT.usable + assert not empty.needs_a_different_part + + one = Assessment( + fixture=Fixture(), use=_benign_use(), objections=(Objection(Fitness.BENEFIT_NOT_QUANTIFIED, "x"),) + ) + assert not one.fit + assert not one.verdict.usable + + +def test_fitness_is_computed_per_use_and_does_not_transfer(): + """The same part is fit as a bracket and unfit as a reservoir.""" + bracket = FIXTURES["sensor_bracket_asa"] + assert fitness(bracket, USES["deck_bracket"]).fit + assert not fitness(bracket, USES["shared_reagent_trough"]).fit + + +# -- the catalog's findings -------------------------------------------------------- + + +def test_no_fdm_material_reaches_sample_contact(): + """Thermal capability does not buy a fluid path, and the two are easy to conflate.""" + for mat in MATERIALS.values(): + if mat.process is Process.FDM: + assert mat.ceiling.rank < Contact.SAMPLE_CONTACT.rank, mat.name + + +def test_the_polymers_that_survive_steam_still_do_not_reach_sample_contact(): + for mat in materials_surviving(Sterilization.AUTOCLAVE_134): + assert mat.process is Process.FDM + assert not mat.ceiling.in_the_fluid_path + + +def test_nothing_here_is_both_autoclavable_and_sample_contact_capable(): + """The finding this catalog exists to produce, pinned so a regression is loud.""" + assert survey()["autoclavable_and_sample_contact"] == 0 + autoclavable = {m.name for m in materials_surviving(Sterilization.AUTOCLAVE_121)} + sample = {m.name for m in materials_reaching(Contact.SAMPLE_CONTACT)} + assert autoclavable and sample + assert not (autoclavable & sample) + + +def test_a_new_print_is_untrusted(): + result = fitness(FIXTURES["plate_nest_as_printed"], USES["plate_nest"]) + assert not result.fit + assert set(result.refusals()) == { + Fitness.ABOVE_DECLARED_CONTACT, + Fitness.NOT_POSITIVELY_LOCATED, + Fitness.DIMENSION_NOT_MEASURED, + Fitness.BENEFIT_NOT_QUANTIFIED, + } + assert not result.needs_a_different_part, "every one of those is work, not a different part" + + +def test_a_well_built_part_can_still_be_refused_structurally(): + """Everything a lab controls, done, and the process is still the refusal.""" + fixture = FIXTURES["reagent_trough_petg"] + assert fixture.positively_located + assert fixture.dimensions_measured() + assert fixture.benefit is not None and fixture.benefit.quantified + result = fitness(fixture, USES["shared_reagent_trough"]) + assert result.verdict is Fitness.POROUS_PROCESS_IN_FLUID_PATH + assert result.needs_a_different_part + + +def test_a_structural_bracket_reaches_fit(): + result = fitness(FIXTURES["sensor_bracket_asa"], USES["deck_bracket"]) + assert result.fit + assert result.verdict is Fitness.FIT + assert "for no other use" in result.reason + + +def test_a_fit_part_is_still_unqualified_in_the_wet_bench_respects(): + """FIT is fitness for a use, not a clean bill. Four qualifications this module cannot see.""" + standing = unqualified(FIXTURES["sensor_bracket_asa"]) + assert set(standing) == { + Qualification.CLEANABILITY, + Qualification.LEACHABLES, + Qualification.STERILITY, + Qualification.THERMAL_ON_THIS_PART, + } + + +# -- the microplate envelope ------------------------------------------------------- + + +def test_a_nest_must_clear_the_loose_mid_side_tolerance_not_the_famous_one(): + """The commonly quoted tolerance holds only within 12.7 mm of the four outside corners.""" + env = plate_envelope(Flange.MEDIUM) + assert env.max_length_mm == pytest.approx(128.26) + assert env.max_width_mm == pytest.approx(85.98) + assert env.max_length_mm > 127.76 + 0.25 + + +def test_the_corner_relief_uses_the_maximum_radius_not_the_nominal(): + """The permitted range spans 1.58 to 4.78 mm; a nest cut to 3.18 binds on the upper end.""" + assert plate_envelope(Flange.SHORT).max_corner_radius_mm == pytest.approx(4.78) + + +def test_an_undeclared_flange_variant_refuses_to_emit_a_number(): + env = plate_envelope() + assert not env.flange_known + assert env.flange_low_mm is None and env.flange_high_mm is None + assert "no flange variant is declared" in env.note + assert not Flange.UNDECLARED.declared + + +def test_a_declared_flange_variant_emits_its_band(): + env = plate_envelope(Flange.TALL) + assert env.flange_known + assert env.flange_low_mm == pytest.approx(7.24) + assert env.flange_high_mm == pytest.approx(8.00) + + +def test_the_dual_variant_reports_two_sides_rather_than_a_tolerance_band(): + env = plate_envelope(Flange.DUAL) + assert env.flange_low_mm == pytest.approx(2.03) + assert env.flange_high_mm == pytest.approx(8.00) + assert "different SIDES" in env.note + + +def test_no_height_is_emitted_for_a_plate_that_is_not_a_standard_height_microplate(): + """Deep-well, PCR and reservoir plates keep the footprint and abandon the height.""" + assert plate_envelope(Flange.MEDIUM).max_height_mm is None + assert plate_envelope(Flange.MEDIUM, typical_height=True).max_height_mm == pytest.approx(15.11) + + +def test_the_envelope_says_dimensional_conformity_is_not_contact_fitness(): + for flange in (Flange.UNDECLARED, Flange.SHORT, Flange.DUAL): + assert "not evidence of fitness for contact" in plate_envelope(flange).note + + +# -- the refusal to estimate -------------------------------------------------------- + + +def test_the_module_returns_no_lifetime_leachable_rate_or_probability(): + """Guards the discipline the docstring claims, in the only way that stays true under edits.""" + forbidden = ( + "predict", + "estimate", + "lifetime", + "remaining_life", + "expected", + "probability", + "likely", + "cycles_survived", + "leachable_rate", + ) + # The refusal ledger is allowed to name the quantities it declines to return; nothing else is. + for name in printed.__all__: + if name in ("NOT_ESTIMATED", "refusals"): + continue + lowered = name.lower() + for token in forbidden: + assert token not in lowered, f"'{name}' looks like a predicted quantity" + + +def test_every_refusal_names_what_it_would_take_to_lift_it(): + assert refusals() is NOT_ESTIMATED + assert len(NOT_ESTIMATED) >= 5 + quantities = {r.quantity for r in NOT_ESTIMATED} + assert "autoclave cycles survived" in quantities + assert "leachable release rate or concentration" in quantities + assert "part service life" in quantities + for r in NOT_ESTIMATED: + assert r.why.strip() + assert r.what_it_would_take.strip() + + +def test_no_cycle_count_is_published_anywhere_in_the_catalog(): + """Every specific count located for these polymers traced to commercial aggregator pages.""" + for mat in MATERIALS.values(): + text = " ".join(filter(None, (mat.note, mat.chemistry))).lower() + for claim in ("cycles at", "autoclave cycles", "1,000", "2,000", "3000"): + assert claim not in text, f"{mat.name} carries an unsourced cycle count" From 23434cd6c17fc74f925076b4cd8e60e83e26ef3b Mon Sep 17 00:00:00 2001 From: di-omics Date: Tue, 28 Jul 2026 03:59:36 -0700 Subject: [PATCH 2/3] Keep the printed-fixture docs ASCII The repo runs an ascii job that refuses non-ASCII bytes in tracked text, and the two new markdown files carried section signs and plus-minus signs. Replaced with '+/-' and plain cross-references. Worth keeping: a tolerance written as a plus-minus glyph reads fine in a browser and badly in a terminal, a diff, and half the tools that will ever open a hardware note. --- docs/PRINTED_FIXTURES.md | 102 +++++++++++++++++++-------------------- hardware/README.md | 36 +++++++------- 2 files changed, 69 insertions(+), 69 deletions(-) diff --git a/docs/PRINTED_FIXTURES.md b/docs/PRINTED_FIXTURES.md index 5fadf80..7b60345 100644 --- a/docs/PRINTED_FIXTURES.md +++ b/docs/PRINTED_FIXTURES.md @@ -23,7 +23,7 @@ measurement would produce it rather than filling the hole. **A printed part may hold labware. It may not be labware.** Fixtures, adapters, nests, risers, guards, jigs, tip-box shims, camera mounts, cable -routing, the wedge in §6 -- parts whose entire job is to put a certified consumable in a +routing, the wedge in 6 -- parts whose entire job is to put a certified consumable in a known place and keep it there. That is the permitted set, and it is a large and genuinely useful set. @@ -55,15 +55,15 @@ The consequences have been measured directly: mm and 1 / 2 / 3 perimeters, the authors concluded verbatim that **"no manufacturing settings can provide enough sealing against fluid intake."** Thinner layers absorbed less; none stopped it. Saturation was reached at 48 h with 1-2 perimeters and 72 h with - 3. `[EXP: Popescu 2021]` + 3. `[EXP: Popescu 2021]` - **Acetone vapor smoothing does not seal the part.** In the same study, 45 minutes of cold acetone vapor reduced open porosity, and methylene blue dye still stained the treated parts after 30 minutes of immersion. The treatment also degraded mechanical properties. Smoothing changes what you can see, not what you can clean. `[EXP: Popescu 2021]` - Bacteria preferentially colonize the layer lines. Across eight PLA variants plus metal-filled, carbon-filled and wood-filled filaments, attachment after 2 h ran to - 6.5x10^6-1.9x10^7 CFU for *E. coli*, 3.6x10^7-6.15x10^7 for *P. aeruginosa*, and - 6.6x10^5-2.16x10^6 for *S. aureus*, with biofilms **thickest between the layers** and + 6.5x10^6-1.9x10^7 CFU for *E. coli*, 3.6x10^7-6.15x10^7 for *P. aeruginosa*, and + 6.6x10^5-2.16x10^6 for *S. aureus*, with biofilms **thickest between the layers** and bacteria filling "the valleys" of the layer structure. `[EXP: Hall 2021]` - Roughness on an FDM part is wildly anisotropic, which is why a quoted Ra can be meaningless. The same study measured Ra of roughly 0.64-1.15 um *along* the layers, and @@ -76,7 +76,7 @@ manufactured geometries are "expected to increase the difficulty in removing man material residues (cleaning) and in sterilization due to the likelihood of increased surface area", that "sterilization process validation should account for the complex geometry of your device under worst-case conditions", and that worst case includes the -"combination of largest surface area, greatest porosity". `[STD: FDA 2017 §VI.E]` +"combination of largest surface area, greatest porosity". `[STD: FDA 2017 VI.E]` The burden is validation, not prohibition. So: **this guide's research located no published study validating an FDM-printed part as cleanable to any recognized standard** (ANSI/AAMI @@ -92,7 +92,7 @@ organisms it reaches; it does not remove protein, endotoxin, nucleic acid, or ch residue, and it fixes protein onto surfaces. Whether steam penetrates the interior void network of an FDM part at all is unvalidated -- no study located cultured the interior of an autoclaved FDM part or ran biological indicators inside the voids. A part can be sterile -and still cross-contaminate the next sample. See §3 for what autoclaving does to the +and still cross-contaminate the next sample. See 3 for what autoclaving does to the dimensions, which is a separate disaster. **"Coat it."** Vacuum epoxy infiltration is a vendor claim (airtight and watertight to 65 @@ -200,8 +200,8 @@ The reasoning, axis by axis. **Enclosed and actively heated is the axis that decides what you can print at all.** Every material worth using for a load-bearing deck fixture -- ABS, ASA, PC -- warps at the sizes -deck fixtures actually are, and warp is driven by the maximum in-plane dimension (§5). A -microplate footprint is 127.76 x 85.48 mm `[STD: SLAS 1 §4.1.1.1]` before you add walls, +deck fixtures actually are, and warp is driven by the maximum in-plane dimension (5). A +microplate footprint is 127.76 x 85.48 mm `[STD: SLAS 1 4.1.1.1]` before you add walls, so a nest is already a large flat part on day one. A passive enclosure is a box that traps some waste heat; an actively heated chamber is a controlled variable. Everything else on the spec sheet is negotiable and this is not. @@ -306,17 +306,17 @@ and an unmeasured number that reaches a document becomes a budget nobody measure those five lines on the day you order and put the **sum** on the requisition. What this guide will assert is which lines cannot be dropped: -- **Metrology is not optional and it is not an accessory.** The qualification ladder in §4 +- **Metrology is not optional and it is not an accessory.** The qualification ladder in 4 has a MEASURED rung, and a lab with a printer and no measuring instrument physically cannot reach it. It can only produce objects. Calipers are the floor; gauge pins or blocks for the features that actually register are better. - **The drybox is not optional either, and the datasheets say why.** Saturation water absorption: PETG 0.51% max `[DS]`, PEEK 0.45% at 23 C and 0.55% at 100 C `[DS]`, ULTEM - 1010 1.25% `[DS]`. Nylon is worse -- strongly hygroscopic, and moisture both swells the + 1010 1.25% `[DS]`. Nylon is worse -- strongly hygroscopic, and moisture both swells the part and plasticizes the polymer, dropping its Tg. (This guide's research did not obtain a numeric moisture-uptake figure for a PA filament, so the mechanism is stated and the number is not.) Wet filament prints badly and prints *differently* on different days, - which quietly destroys the repeatability that §5 depends on. + which quietly destroys the repeatability that 5 depends on. ### The two axes this guide could not source @@ -383,7 +383,7 @@ pass/fail is only meaningful next to a cycle. | **PC (unfilled resin)** | ~147 C `[DS-2]` | ~124-126 C @1.80, ISO 75 `[DS-2]`; Vicat B/50 ~147 C | pass on thermal grounds | pass on thermal grounds | Grade-specific figures must come from the specific grade sheet -- the source pages for these returned 404/500. Carbonate linkages hydrolyze in wet heat; the molecular-weight-loss and crazing mechanism is well known in device practice and this guide could not source it properly, so it is flagged rather than asserted | `NONE`, `ADJACENT` | | **PP** | not on any filament datasheet reached. PP's Tg is below room temperature but no manufacturer source confirmed it -- leave it out rather than guess | none published on the filament TDS | **MARGINAL** | **DO NOT** | The single most dangerous row. One PP copolymer filament has **Tm 137 C** `[DS]` -- a 134 C cycle runs 3 C below its melting point. Homopolymer melts ~160-165 C. Do not carry the reputation of molded PP labware, which is routinely autoclaved, onto PP filament without checking that grade's Tm | `NONE`, `ADJACENT` | | **PA / nylon (CF-filled)** | 70 C (PAHT-CF), 85 C (PPA-CF) `[DS-2]` | 170 C @1.8 / 194 C @0.45 (PAHT-CF); 196/227 C (PPA-CF) `[DS-2]` | see notes | see notes | **The HDT column is a trap here.** It sits 100-140 C above Tg because carbon fiber carries the load. Separately, PA is strongly hygroscopic and a steam cycle drives it toward saturation -- moisture swells the part *and* plasticizes the polymer, dropping unfilled nylon's Tg substantially. It will move even if it never softens | `NONE`, `ADJACENT` | -| **PEI / ULTEM 9085** | **177.3 C** `[DS]` | printed 178.2 C (XY) / 178.4 C (XZ) @66 psi; 170.2 / 172.6 @264 psi `[DS]` | pass | pass | Not printable on any machine in §2. TMA reversal at ~175.9 C upright vs ~193.4 C flat -- print orientation, not polymer, sets deformation onset | `NONE`, `ADJACENT` | +| **PEI / ULTEM 9085** | **177.3 C** `[DS]` | printed 178.2 C (XY) / 178.4 C (XZ) @66 psi; 170.2 / 172.6 @264 psi `[DS]` | pass | pass | Not printable on any machine in 2. TMA reversal at ~175.9 C upright vs ~193.4 C flat -- print orientation, not polymer, sets deformation onset | `NONE`, `ADJACENT` | | **PEI / ULTEM 1010** | **210 C** printed `[DS-2]` / **217 C** molded `[DS]` | 200 C @0.45, 190 C @1.8 (ISO 75) `[DS]`; 215/210 C @66/264 psi (ASTM D648) `[DS-2]` | pass | pass | The ISO-versus-ASTM gap of 15-20 C on nominally identical material is why every figure here carries its method. Water absorption 1.25% at saturation `[DS]` | `NONE`, `ADJACENT` | | **PEEK** | onset **143 C**, midpoint 150 C `[DS]` | DTUL **152 C** @1.8 MPa unannealed `[DS]`; Tm 343 C | pass | pass, thinly | Margin at 134 C is 9 C on Tg onset. It passes because the **crystalline phase** carries load above Tg -- and as-printed PEEK from a machine without adequate chamber temperature can be substantially amorphous, in which case it is not the material the datasheet describes. The maker's "suitable for steam sterilisation" statement covers injection-molded granules, not printed parts | `NONE`, `ADJACENT` | @@ -426,13 +426,13 @@ sterility. It establishes thermal survival and nothing else. filaments are not a cleanability control -- they leach metal ions, which is a direct problem for cell culture and for any assay sensitive to divalent cations, and bacteria attached to metal-filled PLA at the same 10^6-10^7 CFU order as plain PLA. `[EXP: Hall - 2021]` + 2021]` One genuinely favorable finding, recorded because omitting it would be its own dishonesty: plain FDM thermoplastics are generally **not** cytotoxic in cell contact. ABS, PETG, PLA and Nylon 12 did not reduce human iPSC viability against control in one study, while two SLA resins reduced it by roughly 60% and 90%. `[EXP]` That is real, and it does not license -sample contact, because cytotoxicity was never the reason for the boundary in §1 -- porosity +sample contact, because cytotoxicity was never the reason for the boundary in 1 -- porosity was. It is also not uniform: a preprint reports significant cytotoxicity for PETG and PC, and colorants, plasticizers and processing aids differ between spools of nominally identical polymer and are not disclosed on consumer filament. @@ -494,14 +494,14 @@ information about the object beyond "the printer did not fail". requirements that are easy to skip: - *Measure at the temperature the part will work at.* SLAS dimensions are specified at - 20 C `[STD: SLAS 1-4 §1.2]`, and ANSI/SLAS 6's test method at 25 C +/- 2 C -- so there is + 20 C `[STD: SLAS 1-4 1.2]`, and ANSI/SLAS 6's test method at 25 C +/- 2 C -- so there is not even one temperature across the standard family. Printed polymers have far higher thermal expansion and lower creep resistance than molded PP or PS. A part in tolerance on the bench can be out of tolerance at 37 C in an incubator. - *Do not trust the bottom face.* The first layer is squished against the bed and comes out wider than modeled -- PrusaSlicer ships `elefant_foot_compensation = 0.2` in several of its own profiles, and Prusa's documentation puts values around 0.2 mm as typical for a - 0.4 mm nozzle. The bottom few layers are the least dimensionally trustworthy region of + 0.4 mm nozzle. The bottom few layers are the least dimensionally trustworthy region of the part, and they are also the region that seats on the deck. Put registration datums somewhere else, or account for it explicitly. - *A calibration cube proves nothing about a nest.* Tolerance scales with length in every @@ -517,8 +517,8 @@ requirements that are easy to skip: unreachable, so treat the order of magnitude as indicative and the digits as unverified]` The practical consequence is real either way: an FDM printer reproduces its own error consistently, so **measure and compensate** works even when out-of-the-box accuracy is - 5-10x worse than repeatability. That loop is the only route to a part that registers - properly, and it needs the calipers from §2. + 5-10x worse than repeatability. That loop is the only route to a part that registers + properly, and it needs the calipers from 2. **FITTED.** The real plate -- the one from the lot you will actually run, not a dimensionally different one from a different vendor -- seats in the part, on the deck, in @@ -540,7 +540,7 @@ witnessed it. Software having sent the command is `asserted` and is a log line a Do not report a part at a rung it has not reached, and in particular do not let FITTED read as DRY_RUN. **A part that fits by hand has not been qualified for an arm.** The failure mode -in §5 that puts an instrument at risk -- a part shifting mid-run -- is invisible from every +in 5 that puts an instrument at risk -- a part shifting mid-run -- is invisible from every rung below DRY_RUN, and a printed fixture is exactly the kind of hardware that gets promoted straight from "I fitted it, looks great" to a production run. @@ -571,7 +571,7 @@ bed it was printed on.** Mesh bed leveling conforms the first layer *to* the bed rather than correcting it, so bed non-flatness transfers into the part's bottom face. One measurement project reports peak-to-valley bed variation of a quarter of a millimeter -- more than two layers at 0.1 mm. `[single anecdote, printer unspecified]` No manufacturer bed -flatness specification for the machines in §2 could be confirmed; a "0.10 mm" figure +flatness specification for the machines in 2 could be confirmed; a "0.10 mm" figure circulating for one of them traced to forum discussion, not to a spec sheet. *What it does:* the base rocks, so the plate's true position depends on which corner is @@ -600,9 +600,9 @@ The arithmetic is what makes this a design problem rather than a curiosity. Acro 127.76 mm long dimension of a plate footprint: ``` - PLA 0.05% x 127.76 mm = 0.06 mm - PETG 0.15% x 127.76 mm = 0.19 mm - ABS 0.513% x 127.76 mm = 0.66 mm + PLA 0.05% x 127.76 mm = 0.06 mm + PETG 0.15% x 127.76 mm = 0.19 mm + ABS 0.513% x 127.76 mm = 0.66 mm ``` That last figure is larger than the entire SLAS 1 corner-zone tolerance band of +/-0.25 mm. @@ -654,20 +654,20 @@ part in its printed orientation** rather than on the pellet -- and consumer fila datasheets do not carry that value. Two things are worth saying without a number. Humidity control changes the behavior -substantially, and the drybox you bought in §2 is already half of that story. And the obvious -fix -- carbon-filled "ESD-safe" filament -- is excluded here for the reasons in §3: filled +substantially, and the drybox you bought in 2 is already half of that story. And the obvious +fix -- carbon-filled "ESD-safe" filament -- is excluded here for the reasons in 3: filled grades bring higher roughness, filler-matrix interfaces and extra leachable species to a part that sits next to open labware. ### A lip too shallow to retain the plate at angle -This is the failure that ends §6's worked example if the design is careless, and it is a +This is the failure that ends 6's worked example if the design is careless, and it is a standards problem more than a geometry problem. **The flange the lip grips is not one thing.** ANSI/SLAS 3-2004 offers **five mutually -exclusive variants**, and requires the plate to declare which one it meets: §4.1 Short = -2.41 mm +/- 0.38, §4.2 Medium = 6.10 mm +/- 0.38, §4.3 Tall = 7.62 mm +/- 0.38, §4.4 Short -with interruptions = 2.41 mm +/- 0.38, §4.5 Dual = 2.41 mm on the short sides and 7.62 mm on +exclusive variants**, and requires the plate to declare which one it meets: 4.1 Short = +2.41 mm +/- 0.38, 4.2 Medium = 6.10 mm +/- 0.38, 4.3 Tall = 7.62 mm +/- 0.38, 4.4 Short +with interruptions = 2.41 mm +/- 0.38, 4.5 Dual = 2.41 mm on the short sides and 7.62 mm on the long sides. `[STD]` A lip sized against a tall 7.62 mm flange will not reliably retain a 2.41 mm short-flange plate. "SBS compliant" on a vendor page does not tell you which variant you are getting. @@ -692,11 +692,11 @@ have in their head. - ANSI/SLAS 2-2004 specifies **14.35 mm +/- 0.25 mm** from the resting plane to the maximum protrusion of the perimeter wells, and 14.35 mm +/- 0.76 mm overall -- **for a typical microplate**. `[STD]` (Note the standard's own Figure 1 misprints the inch equivalent as - 0.56560 in against the correct 0.5650 in in the clause text. Use the metric value.) + 0.56560 in against the correct 0.5650 in in the clause text. Use the metric value.) - Plate types that are not typical microplates simply do not comply with SLAS 2 while still complying with SLAS 1, 3 and 4. A 96-well 1000 uL deep-well plate is documented at length - 127.8 mm and width 85.5 mm -- SLAS 1 conformant -- with a height of **44.1 mm**, roughly - 3x the SLAS 2 figure. Another vendor's 2.2 mL deep-well plate lists 44 mm and advertises + 127.8 mm and width 85.5 mm -- SLAS 1 conformant -- with a height of **44.1 mm**, roughly + 3x the SLAS 2 figure. Another vendor's 2.2 mL deep-well plate lists 44 mm and advertises only an "ANSI-SBS Footprint". `[DS]` So the accurate framing is: **height is standardized only for standard-height microplates; @@ -718,11 +718,11 @@ deep-well plate can land there, size for the deep-well plate. Four more clauses that will bite a nest designed to nominal dimensions. - **There are two footprint tolerances, not one.** +/-0.25 mm applies **only** within 12.7 mm - of the four outside corners `[STD: SLAS 1 §4.1.1.1]`. Anywhere else along the side the - tolerance is **+/-0.5 mm** `[STD: §4.1.1.2]`. A nest cut to "127.76 x 85.48 +/- 0.25" as a + of the four outside corners `[STD: SLAS 1 4.1.1.1]`. Anywhere else along the side the + tolerance is **+/-0.5 mm** `[STD: 4.1.1.2]`. A nest cut to "127.76 x 85.48 +/- 0.25" as a flat statement will jam real plates that bow mid-side and remain fully conformant. - **Corner radius is 3.18 mm +/- 1.6 mm** -- a permitted range of 1.58 to 4.78 mm, a 3x - spread, scoped to the bottom flange corners specifically. `[STD: §4.1.2.1]` Design + spread, scoped to the bottom flange corners specifically. `[STD: 4.1.2.1]` Design clearance features against the **maximum** 4.78 mm, not the nominal, or you will bind on plates at the top of the range. - **Draft is excluded from the standard's numbers.** Every figure note states "Dimensions @@ -730,19 +730,19 @@ Four more clauses that will bite a nest designed to nominal dimensions. standard's dimensions do not describe, and the actual angles are not published anywhere in it. A printed nest with vertical walls cut to nominals does not have the same real cross-section as the molded plate it is supposed to seat. -- **SLAS 2 has two alternative compliance parts** and a plate must declare which. §4.1 adds +- **SLAS 2 has two alternative compliance parts** and a plate must declare which. 4.1 adds a minimum 1 mm clearance from the resting plane to the bottom external surface of the - wells; §4.2 does not. `[STD]` If a fixture assumes 1 mm of clearance under the wells -- for - bottom-reading optics, or a heat block -- a §4.2-compliant plate is not required to give + wells; 4.2 does not. `[STD]` If a fixture assumes 1 mm of clearance under the wells -- for + bottom-reading optics, or a heat block -- a 4.2-compliant plate is not required to give it to you. And one that saves a wasted investigation: **ANSI/SLAS 6-2012 sets no limits at all.** Its -§7 states that the standard "specifies definitions and a test method only" and that it is +7 states that the standard "specifies definitions and a test method only" and that it is "not the intent of this standard to state a limit" for well bottom elevation or its variation. `[STD]` "SLAS 6 compliant" conveys no flatness or bottom-thickness guarantee whatsoever. For any optical path, the instrument's own specification is the requirement. -One FDM-specific conformity note worth designing around: SLAS 1 §4.1.1.3 requires the +One FDM-specific conformity note worth designing around: SLAS 1 4.1.1.3 requires the footprint be "continuous and uninterrupted around the base of the plate." Brim remnants, elephant-foot bulges and support scars on a printed part's base are the same class of defect in reverse -- they snag stage nests and gripper jaws. @@ -768,7 +768,7 @@ At best it reaches **MEASURED**. The wedge angle and the lip that retains the pl measured against the model. Everything above that is unproven: - **FITTED** needs the real plate, at working mass, at the design angle, checked for rock - and checked for retention -- and per §5 that check has to be done separately for each + and checked for retention -- and per 5 that check has to be done separately for each flange variant it will hold. - **DRY_RUN** needs the arm to run the real aspiration approach against a tilted plate, at reduced speed. Tilting changes the geometry the head was taught: the well bottom is no @@ -778,7 +778,7 @@ measured against the model. Everything above that is unproven: - **IN_USE** needs a recorded run with material. The wedge also creates its own new failure mode, which the flat nest did not have: it puts -the plate at an angle where a shallow lip stops retaining. That is §5's lip failure, and it +the plate at an angle where a shallow lip stops retaining. That is 5's lip failure, and it is not hypothetical for this part -- it is the part's defining feature. ### The acceptance test @@ -827,7 +827,7 @@ the position. **8. Hold everything else constant and say so.** Same session, same instrument, same consumable lot, same filament lot for the wedge itself. A wedge reprinted from a different -spool is a different part until it has been measured again (§5, shrinkage). +spool is a different part until it has been measured again (5, shrinkage). **9. Pre-register the threshold and record its basis.** Decide before the run what improvement would make the wedge worth keeping, and record where that target came from. It @@ -853,19 +853,19 @@ Kept as a list rather than buried, because the gaps are the part most likely to in by somebody's search results. - **Whether any FDM part has ever been validated as cleanable to a recognized standard.** A - targeted search found none. That is the crux of §1 and it did not close. + targeted search found none. That is the crux of 1 and it did not close. - **Whether steam penetrates and sterilizes the interior voids of an FDM part.** No located study cultured the interior of an autoclaved FDM part or ran biological indicators inside the void network. Surviving the heat and being sterile are different claims. - **How many autoclave cycles any high-temperature polymer survives.** Every specific count traced to commercial aggregator pages. - **A shrinkage percentage from any filament manufacturer.** Three datasheets were opened; - none published one. The slicer defaults in §5 are the best available and they disagree + none published one. The slicer defaults in 5 are the best available and they disagree with each other. - **A warp-per-unit-length rule for large flat parts.** No dataset located measures flatness across a range of part sizes. The 3.7 mm / 0.8 mm figures cannot be scaled because the specimen dimensions were not stated. -- **A manufacturer bed-flatness specification** for any machine in §2. +- **A manufacturer bed-flatness specification** for any machine in 2. - **Coefficients of thermal expansion** for ASA and PC. Conflicting values circulate for ABS (90e-6 /C versus 120e-6 /C) and the PLA and PETG figures found came from a retailer blog. None are printed here. @@ -882,7 +882,7 @@ in by somebody's search results. copies the standards body itself publishes, dated 2011. For anything that goes to print, check the purchased controlled copies. - **Any manufacturer claim of biocompatibility, cytotoxicity testing, autoclavability, or - sample-contact suitability for any printer or stock filament in §2.** There is none on any + sample-contact suitability for any printer or stock filament in 2.** There is none on any official page read. "Supported filament" lists are printability ratings -- one vendor literally grades them "Ideal" and "Capable" -- meaning will-it-extrude-without-clogging. Treat the absence as absence, and never read a spec-table material list as a @@ -904,7 +904,7 @@ Standards and regulatory criteria: TOC <= 12 ug/cm^2, ATP <= 22 fmol/cm^2. The widely quoted 6.4 ug/cm^2 protein criterion could not be confirmed and is not printed here. - US FDA, *Technical Considerations for Additive Manufactured Medical Devices*, issued - 2017-12-05, §VI.E. + 2017-12-05, VI.E. - US FDA guidance on the Mouse Embryo Assay for assisted reproduction devices. Peer-reviewed @@ -921,14 +921,14 @@ Peer-reviewed - Rogers HB, Zhou LT, Kusuhara A, Zaniker E, Shafaie S, Owen BC, Duncan FE, Woodruff TK (2021). *Chemosphere* 270:129003. ISO-certified dental resins release ovo-toxic leachates. - Kress S, Schaller-Ammann R, Feiel J, Priedl J, Kasper C, Egger D (2020). *Materials* - 13(13):3011. doi:10.3390/ma13133011. Cytotoxicity of stereolithography photopolymers; + 13(13):3011. doi:10.3390/ma13133011. Cytotoxicity of stereolithography photopolymers; parylene barrier and its detachment after 5-6 autoclave cycles. - Oskui SM et al. (2016). *Environ Sci Technol Lett*. doi:10.1021/acs.estlett.5b00249. Both FDM and SLA parts toxic to zebrafish embryos; SLA significantly more so. Manufacturer datasheets read directly - Bambu Lab PETG Basic TDS V3.0; Victrex PEEK 450G TDS (rev. March 2026); SABIC ULTEM Resin - 1010 TDS (Europe, rev. 20170620); Stratasys ULTEM 9085 MDS (2025); Prusament ASA TDS v1.1; + 1010 TDS (Europe, rev. 20170620); Stratasys ULTEM 9085 MDS (2025); Prusament ASA TDS v1.1; Prusament PC Blend TDS v1.1; PPprint P-filament TDS v1.001; Polymaker PETG TDS V2.0 (2025-11-17); Polymaker ASA TDS; 3DXTech 3DXMAX ABS TDS Rev 3.0; Formlabs BioMed Clear TDS (doc 2001432-TDS-ENUS-0, rev 04); Stratasys F123 series product specification. @@ -945,7 +945,7 @@ Vendor pages, captured 2026-07-28 - Prusa Research product pages for MK4S, CORE One+ and the HT hotend upgrade. - QIDI Tech US store Plus4 product and technical-specification pages. Note the Plus4 marketing page carries a comparison graphic reading "Hot bed 100 / Nozzle 360 / Chamber - 55" -- those are previous-generation figures shown for contrast, not the Plus4's. An + 55" -- those are previous-generation figures shown for contrast, not the Plus4's. An automated scrape of that page pulls the wrong three numbers. - Intamsys Funmat HT product page. - Protolabs Network (Hubs) knowledge base, dimensional accuracy of 3D printed parts. diff --git a/hardware/README.md b/hardware/README.md index 6b08c45..2ab0034 100644 --- a/hardware/README.md +++ b/hardware/README.md @@ -59,7 +59,7 @@ sample: the plate is the barrier. test method, and say nothing about material, leachables, extractables, cytotoxicity, sterility, or nuclease status. - FDM parts are porous by construction. X-ray CT of PLA printed at 100% infill measured - 4.05 to 6.32% internal porosity across raster settings, with pores concentrated at the + 4.05 to 6.32% internal porosity across raster settings, with pores concentrated at the shell-to-infill interface -- that is, connected to the outer surface (Wang et al., *Polymers* 11(7):1154, 2019). Immersion testing of ABS found that no combination of layer height, perimeter count, or infill pattern sealed a part against fluid intake, and @@ -67,7 +67,7 @@ sample: the plate is the barrier. immersion (Popescu et al., *Polymers* 13(23):4249, 2021). - Bacteria preferentially colonize the layer lines. Biofilm was thickest in the grooves between printed layers on every polymer tested (Hall et al., *Front. Microbiol.* - 12:646303, 2021). + 12:646303, 2021). - Many photopolymer (SLA/DLP/MSLA) resins remain cytotoxic after full manufacturer post-cure, including resins carrying ISO 10993 biocompatibility certification, because the certification does not cover the exposure this fixture would represent and because @@ -107,25 +107,25 @@ The groups below are the ones you will actually change. | parameter | default | note | | --- | --- | --- | -| `plate_len_mm`, `plate_wid_mm` | 127.76, 85.48 | ANSI/SLAS 1-2004 §4.1.1.1 nominals | -| `plate_footprint_tol_mm` | 0.5 | the §4.1.1.2 mid-side tolerance, not the §4.1.1.1 corner one. See below. | -| `plate_corner_r_max_mm` | 4.78 | §4.1.2.1 is 3.18 ± 1.6 mm. Use the maximum for a clearance feature. | -| `plate_flange_h_mm`, `plate_flange_tol_mm` | 6.10, 0.38 | ANSI/SLAS 3-2004 §4.2 medium. Five incompatible variants exist. | +| `plate_len_mm`, `plate_wid_mm` | 127.76, 85.48 | ANSI/SLAS 1-2004 4.1.1.1 nominals | +| `plate_footprint_tol_mm` | 0.5 | the 4.1.1.2 mid-side tolerance, not the 4.1.1.1 corner one. See below. | +| `plate_corner_r_max_mm` | 4.78 | 4.1.2.1 is 3.18 +/- 1.6 mm. Use the maximum for a clearance feature. | +| `plate_flange_h_mm`, `plate_flange_tol_mm` | 6.10, 0.38 | ANSI/SLAS 3-2004 4.2 medium. Five incompatible variants exist. | | `plate_height_mm` | 14.35 | ANSI/SLAS 2-2004, and **only** for a standard-height microplate | | `plate_cg_height_mm` | -1 | unmeasured sentinel; no SLAS standard gives plate mass | -| `well_cols`, `well_rows`, `well_pitch_mm` | 12, 8, 9.0 | ANSI/SLAS 4-2004 §4.1 (96-well) | +| `well_cols`, `well_rows`, `well_pitch_mm` | 12, 8, 9.0 | ANSI/SLAS 4-2004 4.1 (96-well) | Three of these are traps worth stating plainly: -**There are two footprint tolerances, not one.** SLAS 1 §4.1.1.1 gives ± 0.25 mm, but only -within 12.7 mm of the four outside corners. §4.1.1.2 gives ± 0.5 mm anywhere else along +**There are two footprint tolerances, not one.** SLAS 1 4.1.1.1 gives +/- 0.25 mm, but only +within 12.7 mm of the four outside corners. 4.1.1.2 gives +/- 0.5 mm anywhere else along the side. A conforming plate may bow outward by 0.5 mm at mid-side. A pocket cut to the -commonly quoted ± 0.25 mm jams legal plates, in the middle, where it is not obvious why. +commonly quoted +/- 0.25 mm jams legal plates, in the middle, where it is not obvious why. The model uses the loose figure. **The flange variant is a required input.** SLAS 3 offers five mutually exclusive variants (short 2.41, medium 6.10, tall 7.62, short-with-interruptions 2.41, dual 2.41/7.62 mm, all -± 0.38) and requires a plate to declare which one it meets. "SBS compliant" on a vendor ++/- 0.38) and requires a plate to declare which one it meets. "SBS compliant" on a vendor page does not tell you. The retaining lip has to engage the flange and nothing above it, so this parameter sizes the lip. On a short-flange plate the default 3.0 mm lip is too tall and the render report says `FAIL`. @@ -162,8 +162,8 @@ from the model.** The printed part is not the model: - Slicer defaults disagree with each other. OrcaSlicer ships `filament_shrink` of 99.95% for one vendor's PLA, 99.85% PETG, and 99.487% for ABS and ASA, while another vendor's profiles in the same repository ship 100% for all four. A 0.5% disagreement on ABS is - 0.64 mm across 127.76 mm -- larger than the entire SLAS footprint tolerance. -- Desktop FDM accuracy is quoted at about ± 0.5% with a ± 0.5 mm floor, and it scales with + 0.64 mm across 127.76 mm -- larger than the entire SLAS footprint tolerance. +- Desktop FDM accuracy is quoted at about +/- 0.5% with a +/- 0.5 mm floor, and it scales with length. A good calibration cube proves nothing about a 129 mm pocket. - The first layers are the least trustworthy part of the print. PrusaSlicer's own profiles ship 0.2 mm of elephant-foot compensation for a 0.4 mm nozzle, and the bottom of this @@ -223,7 +223,7 @@ The lip height has a window and both ends are real: `lip_window_ok()` checks the window and the report prints `PASS` or `FAIL` with the reason. -Corner lips rather than continuous rails, for three reasons. SLAS 3 §4.4 permits a single +Corner lips rather than continuous rails, for three reasons. SLAS 3 4.4 permits a single interruption on center of each long side; a continuous long-side rail can land on that interruption, but a corner lip cannot, because the interruption edges are at least 47.8 mm from the nearest part edge. The mid-side gaps are also where a gripper puts its jaws. And @@ -242,7 +242,7 @@ want a number. `registration_style = "slas_skirt"` gives the fixture's own base an ANSI/SLAS-1 footprint skirt so it drops into the deck nest the way a plate does. This reuses location the deck already provides and needs one measurement (the nest depth) instead of a hole pattern. The -skirt is continuous and uninterrupted around the base, as SLAS 1 §4.1.1.3 requires of a +skirt is continuous and uninterrupted around the base, as SLAS 1 4.1.1.3 requires of a plate footprint, and for the same reason -- a gap snags a nest. The fixture body is wider than the skirt, so its shoulder sits above the nest rim. Only one @@ -321,7 +321,7 @@ Material-dependent settings are marked. Everything else is geometry. | top / bottom layers | 5 or more | the seat is a functional surface | | infill | 15 to 25% gyroid | **do not model internal voids.** An enclosed cavity in a fixture that gets wiped cannot be dried. Let the slicer make the sparse structure so it drains and dries through the walls. | | elephant-foot compensation | as your profile normally uses | the registration skirt is on the first layers; if you disable it, the skirt is oversize | -| brim | only if adhesion fails, and **remove it completely** | SLAS 1 §4.1.1.3 requires a continuous uninterrupted footprint. Brim and elephant-foot remnants on the skirt snag deck nests and gripper jaws. | +| brim | only if adhesion fails, and **remove it completely** | SLAS 1 4.1.1.3 requires a continuous uninterrupted footprint. Brim and elephant-foot remnants on the skirt snag deck nests and gripper jaws. | | supports | see below | | | nozzle | steel, not brass | **material-adjacent.** Brass alloys commonly contain lead. Even for a non-contacting fixture this is the cheaper choice to make correctly. | @@ -546,10 +546,10 @@ Dimensional standards, clause text read directly: - ANSI/SLAS 3-2004 (R2012) Microplates -- Bottom Outside Flange Dimensions - ANSI/SLAS 4-2004 (R2012) Microplates -- Well Positions - ANSI/SLAS 6-2012 Microplates -- Well Bottom Elevation (defines a test method and sets no - limits: §7 states explicitly that it is not the intent of the standard to state a limit) + limits: 7 states explicitly that it is not the intent of the standard to state a limit) All five are published by SLAS at `slas.org`. Dimensions in SLAS 1 through 4 apply at 20 C; -SLAS 6's test method specifies 25 ± 2 C. Do not quote a single temperature for the family. +SLAS 6's test method specifies 25 +/- 2 C. Do not quote a single temperature for the family. Before publishing anything that depends on these clauses, check the purchased ANSI copies. Porosity, cleanability, and material behavior: From 0832622c5d24f7091de700ad89d704a829c7c2d1 Mon Sep 17 00:00:00 2001 From: di-omics Date: Tue, 28 Jul 2026 18:35:43 -0700 Subject: [PATCH 3/3] Join the camera, the arm, and the fixture into one question The portfolio had a vision layer, an arm, and a printed-fixture layer with nothing between them. The join is the flow this work is actually for: capture a demonstration, train an arm from it, and use a printed fixture to make the task tractable. This computes whether a capture can train a policy at all, and whether the policy may then be handed material. A video is not a demonstration. Teleoperation and kinesthetic capture record the arm's own joint states, so the action stream is measured; a monocular human video has none and needs pose estimation followed by a retargeting model onto different kinematics, which is two estimators in series rather than a conversion. No verified figure for human-to-robot retargeting error was located at all, and the published magnitudes for the hand-pose stage alone reach 185.67 mm, so an unmeasured retargeting is not a small unknown. The fixture is inside what the policy learned, which is the connection to the printed layer and the reason a reprint is not a neutral act. A policy carries the fixture revision it learned on, and a capture that recorded no fixture does not count as agreeing with one that did: the part was physically there whether or not anybody wrote it down. A properly run evaluation is not a good result, and this is the distinction the first draft got wrong. Trust.MEASURED means the evaluation was conducted correctly, and a policy that succeeded in zero of twenty held-out trials satisfied it. The enum docstring already said MEASURED does not mean good; the boolean said otherwise to anyone who branched on it rather than reading the prose. evidence_complete and trusted are now separate, and trusted requires a declared acceptance rate that the interval's lower bound clears. There is no default rate, because what is tolerable belongs to the task rather than to this module. A run needs a bound that lives outside the policy, since a learned policy carries no guarantee about an input it has not seen. Interlock demands workspace, speed and force bounds, something named as enforcing them, and a measured miss rate obtained by driving the violation rather than by observing that nothing went wrong. Three further audit findings are closed with regression tests: an unrecorded fixture coupling silently agreeing with a recorded one, a retargeting error accepting zero and negative distances, and success figures quoted as bare percentages off twelve and twenty-four trials, which are counts this module's own success_rate refuses. Those now appear as counts with their intervals, and the simulation result that had been welded onto a real-robot sentence is marked as simulation and reported across all nine models rather than cherry-picked to the worst. 55 new tests, 474 total. ruff clean against the pinned version, ASCII clean. --- README.md | 78 +- autonomous_lab/__init__.py | 33 + autonomous_lab/cli.py | 2 +- autonomous_lab/imitation.py | 1896 +++++++++++++++++++++++++++++++++++ docs/CAPTURE_TO_POLICY.md | 1389 +++++++++++++++++++++++++ tests/test_feedback.py | 4 +- tests/test_imitation.py | 756 ++++++++++++++ 7 files changed, 4141 insertions(+), 17 deletions(-) create mode 100644 autonomous_lab/imitation.py create mode 100644 docs/CAPTURE_TO_POLICY.md create mode 100644 tests/test_imitation.py diff --git a/README.md b/README.md index 5824756..c4cc2d9 100644 --- a/README.md +++ b/README.md @@ -47,13 +47,13 @@ autonomous-lab feedback single_cell_genomics # can a control loop actually clos ## What it reports today Costing the single-cell genomics reference protocol (Namocell sort -> STAR whole-genome sequencing -> -ODTC PCR1 -> STAR library -> AVITI sequencing -> run-folder readout), with a plr-tested -checkout wired in via `--plr-tested`: +ODTC PCR1 -> STAR library -> AVITI sequencing -> run-folder readout), with a run-card +checkout wired in via `--run-cards`: | | steps | | | --- | --- | --- | | automated | 3 of 18 | run headless today: two link preflights and the AVITI run-folder read | -| supervised | 2 of 18 | a validated run card exists in plr-tested, gated on a confirm token and an operator | +| supervised | 2 of 18 | a validated run card exists in the checkout, gated on a confirm token and an operator | | blocked | 8 of 18 | the command is undecoded; the coverage gate refuses the run | | manual | 4 of 18 | seating a cartridge, loading a flow cell, and two STAR steps nobody has written a validated script for | | broken | 1 of 18 | the run card exists, was run on the instrument, and failed | @@ -442,6 +442,54 @@ gravimetric protocol that would produce one instead. An unmeasured fixture is a of a crash surface, a cleaning obligation, and an uncharacterized material to a workcell that had none of them. +## From a video to an arm that can be trusted + +The camera layer, the arm, and the printed fixture had nothing joining them. `imitation` +is the join: capture a demonstration, train an arm from it, and use a printed fixture to +make the task tractable. It computes whether a given capture can train a policy at all, and +whether the resulting policy may be handed material. + +Three things carry it, and each is where the naive version fails. + +**A video is not a demonstration.** A policy needs actions, not pixels. Teleoperation and +kinesthetic capture record the arm's own joint states, so the action stream is measured. A +monocular human video has no joint states: it needs pose estimation and then a retargeting +model onto a gripper with different kinematics, which is two estimators in series rather +than a format conversion. No verified millimetre figure for human-to-robot retargeting error +was found at all. The two published magnitudes belong to the hand-pose stage alone, at +185.67 mm mean per-joint error for one monocular estimator, so an unmeasured retargeting is +not a small unknown. + +**The fixture is inside what the policy learned.** A printed nest removes degrees of freedom +the policy would otherwise learn from data, which is why fixtures and learning belong in one +story. It cuts both ways: reprint that nest with different shrinkage and the demonstrations +collected against the old one may be stale. So a policy carries the fixture revision it +learned on, and a capture that recorded no fixture at all does not count as agreeing with +one that did. The part was physically there whether or not anybody wrote it down. + +**A properly run evaluation is not a good result.** This is the distinction the module got +wrong first and now enforces hardest. `Trust.MEASURED` means the evaluation was conducted +correctly: held out, externally scored, above the trial floor. It says nothing about whether +the policy works, and a policy that succeeded in **zero of twenty** held-out trials is +MEASURED. So `evidence_complete` and `trusted` are separate, and `trusted` additionally +requires a **declared acceptance rate that the interval's lower bound clears**. There is no +default rate, because what is tolerable is a property of the task: a failure rate a retry +fixes is not a failure rate for a transfer that consumes the last of a sample. Judging on +the lower bound rather than the point estimate is what makes 18 of 20 fail an 80 percent +bar that 90 of 100 clears, at the same point estimate. + +A run needs a fourth thing none of those provide: a bound that lives **outside** the policy. +A learned policy carries no guarantee about an input it has not seen, so the limit on what +it can reach, how fast and how hard cannot come from the policy or from a model checking the +policy. `Interlock` demands workspace, speed and force bounds, something named as enforcing +them, and -- following `vision`'s rule for detectors -- a **measured miss rate**, obtained by +deliberately driving the violation rather than by observing that nothing went wrong. + +`docs/CAPTURE_TO_POLICY.md` is the practical side: which capture modality to choose and why +that single decision determines how much of the rest is solved work, what a usable capture +contains, where the printed fixture earns its place, and the ladder from simulation to +material where no rung implies the one above it. + ## The RE queue is computed, not argued about ``` @@ -532,12 +580,13 @@ narrow and carry their own caveats. ## Three things it refuses to do -1. **Let an instrument's reputation transfer to a step.** plr-tested has a validated - whole-genome sequencing preparation addition and a validated PCR enrichment choreography; it has no validated bead cleanup - and no validated library pooling. So those cost out as manual even though they name a - validated instrument. A federated step is supervised only when a run card for *that - step* has been proven. The whole-genome sequencing leg that does count is dry-validated, and the ledger - says so in the same breath: its wet form has never run. +1. **Let an instrument's reputation transfer to a step.** The proven run cards are a + whole-genome sequencing preparation addition and a PCR enrichment choreography; there + is no validated bead cleanup and no validated library pooling. So those cost out as + manual even though they name a validated instrument. A federated step is supervised + only when a run card for *that step* has been proven. The whole-genome sequencing leg + that does count is dry-validated, and the ledger says so in the same breath: its wet + form has never run. 2. **Model only part of what would refuse a run.** `GuardedReplayer.setup()` has three preconditions, not one: coverage, an endpoint, and a transport a connection class can open. `DEFAULT_TRANSPORT` is UNKNOWN for three of these instruments by design, so a @@ -585,10 +634,11 @@ and there is no flag that moves an instrument. Anything that does goes through p controllers, behind their own `armed` and `allow_actuation` switches, with a human present. -Note also plr-tested's hard constraint, which any scheduler built on this must respect: -one driver process per instrument. Two STAR clients raise `USBError [Errno 16] Resource -busy`, and on the ODTC the collision is quieter, because a second process re-registers the -event receiver and silently steals the first one's callbacks. +Note also the hard constraint the instruments impose, which any scheduler built on this +must respect: one driver process per instrument. Two STAR clients raise +`USBError [Errno 16] Resource busy`, and on the ODTC the collision is quieter, because a +second process re-registers the event receiver and silently steals the first one's +callbacks. ## Tests @@ -596,7 +646,7 @@ event receiver and silently steals the first one's callbacks. pip install -e '.[dev]' && pytest ``` -418 device-free tests. The ones that matter most try to make the layer lie: +477 device-free tests. The ones that matter most try to make the layer lie: - claim a step is automated when its command is undecoded; claim a decoded command is runnable while its siblings are not; claim a federated leg runs when no run card was ever diff --git a/autonomous_lab/__init__.py b/autonomous_lab/__init__.py index 4794bc2..29054f9 100644 --- a/autonomous_lab/__init__.py +++ b/autonomous_lab/__init__.py @@ -49,6 +49,9 @@ teaching an expert demonstrating an operation, and a machine measured against that demonstration. A demonstration is data, not authority: one performance is a value with no tolerance, and saying so is the point. + imitation from a captured demonstration to a policy that may be trusted with + material. A video is not a demonstration, the fixture is inside what the + policy learned, and a properly run evaluation is not a good result. printed whether a 3D-printed fixture may be used for a given purpose. The argument that a lab does not need a vendor to ship labware is only honest if the part you printed is held to the same standard as everything else here. @@ -97,6 +100,22 @@ feedback_report, in_flight_exposure, ) +from .imitation import ( + Capture, + Evaluation, + FixtureCoupling, + Interlock, + Modality, + Policy, + Retargeting, + Trust, + demonstrations_needed, + fixture_coupled, + may_run, + success_rate, + trusted_with_material, + usable_for_training, +) from .intelligence import ( Benchmark, BenchmarkStatus, @@ -186,6 +205,7 @@ def loop_closure_for(protocol, workcell=None): "BenchmarkStatus", "BoundaryReport", "Campaign", + "Capture", "Ceiling", "Closable", "Closure", @@ -205,6 +225,7 @@ def loop_closure_for(protocol, workcell=None): "Detection", "Duration", "Entitlement", + "Evaluation", "Event", "Executor", "FEDERATED", @@ -213,11 +234,13 @@ def loop_closure_for(protocol, workcell=None): "FeedbackReport", "Fitness", "Fixture", + "FixtureCoupling", "Gate", "Handoff", "InstrumentConfig", "InstrumentHealth", "InstrumentSpec", + "Interlock", "Interval", "IntervalKind", "Judgment", @@ -233,11 +256,14 @@ def loop_closure_for(protocol, workcell=None): "MachineObservation", "Material", "Misassignment", + "Modality", "NextDemonstration", "Observable", + "Policy", "Process", "Protocol", "Readiness", + "Retargeting", "Role", "RunRecord", "RunReport", @@ -255,6 +281,7 @@ def loop_closure_for(protocol, workcell=None): "TransferReport", "TransferRow", "Transform", + "Trust", "UndeclaredTransform", "Unlock", "Use", @@ -273,12 +300,14 @@ def loop_closure_for(protocol, workcell=None): "crosses_boundary", "declared", "demonstration_queue", + "demonstrations_needed", "entitlement_summary", "envelope_for", "estimate", "evaluate", "feedback_report", "fitness", + "fixture_coupled", "gate_report", "in_flight_exposure", "knowledge_summary", @@ -287,18 +316,22 @@ def loop_closure_for(protocol, workcell=None): "loop_closure_for", "mandatory_gates", "materials_reaching", + "may_run", "provenance_report", "rank_unlocks", "recovery_report", "registry", "sota_lift", "spec", + "success_rate", "taught", "transfer_report", "trusted_for", + "trusted_with_material", "undeclared_transforms", "unmet_demands", "unqualified", "untaught_operations", "untrusted_instruments", + "usable_for_training", ] diff --git a/autonomous_lab/cli.py b/autonomous_lab/cli.py index 4a0d2e7..4290fb4 100644 --- a/autonomous_lab/cli.py +++ b/autonomous_lab/cli.py @@ -77,7 +77,7 @@ def _stock(args) -> int: print(" controller none in plr-re") if s.note: print(f" note {s.note}") - print("\nfederated instruments (driven from di-omics/plr-tested):") + print("\nfederated instruments (di-omics/plr-tested):") for key in sorted(FEDERATED): f = FEDERATED[key] wired = "wired" if (key in wc.federated and wc.plr_tested_root) else "not wired" diff --git a/autonomous_lab/imitation.py b/autonomous_lab/imitation.py new file mode 100644 index 0000000..748f4c0 --- /dev/null +++ b/autonomous_lab/imitation.py @@ -0,0 +1,1896 @@ +"""Whether a captured demonstration can train an arm, and whether the policy may touch material. + +Three layers in this package each hold a piece of this question and none of them states it. +`vision` says what a camera can see and what a visual check needs before it is a check +rather than an aspiration. `printed` says whether a fixture may be used for a purpose, and +refuses a part whose dimensions were never measured. `teaching` says what an expert +demonstrated and refuses to state a tolerance under three demonstrations. What nobody +composes is the sentence that sells the whole idea to a lab: record the expert once on +video, train the arm from the recording, print a fixture to make the task easy enough to +learn. This module computes whether that sentence holds for a particular capture, and the +answer turns on three distinctions the sentence elides. + +A VIDEO IS NOT A DEMONSTRATION. A policy needs ACTIONS, and pixels are not actions. A +teleoperated or hand-guided capture records the robot's own joint states at controller rate, +so the action stream is MEASURED -- it is what the machine did, in the machine's own action +space, with its gripper width and its proprioception alongside. A third-person human video +contains no joint states at any resolution. Getting actions out of it takes hand pose +estimation and then retargeting across an embodiment gap, and BOTH of those are steps that +can be wrong, rather than a format conversion that cannot. The measured cost of pretending +otherwise: injecting calibrated noise into triangulated hand labels moved mean success from +48.3 percent to 30.0 percent at half a sigma and to 20.0 percent at one sigma, and swapping +multi-view triangulation for a real monocular estimator produced 185.67 mm mean per-joint +error and dropped mean success from 41.5 to 24.7 percent. A second group reports 3D hand +pose errors from video "as large as 20-30 cm" and responds by discarding the human's actions +entirely. There is also a case where no estimator helps: recovering an action from an +observed state transition needs the dynamics to be injective, and a redundant arm posture or +a force pressed against a rigid constraint changes no visible state at all, so the action is +unrecoverable in principle rather than in practice. The honest head-to-head number, from the +one work that ran it, is an exchange rate and not an equivalence: on the same sweeping task, +50 teleoperated episodes reached 52 percent and 50 human-video episodes reached 44 percent, +while 100 teleoperated reached 88 percent and it took 300 human-video episodes to reach 84. +So `Modality.measured_action_stream` names which captures carry actions rather than +estimates of actions, and `equivalent_to_teleop` returns False for every retargeted capture +no matter how good the pipeline is, because the pipeline is the thing that was never +measured here. + +THE FIXTURE IS PART OF THE POLICY'S WORLD, AND THAT CUTS BOTH WAYS. This is the connection +to `printed` and it is the non-obvious one. A nest that constrains an object's pose removes +degrees of freedom the policy would otherwise have to learn from data, and there is real +published evidence for the direction of that effect -- see `FixtureCoupling`, which is also +careful about how much of the usual claim the evidence actually supports. The part nobody +says out loud is the other direction: the demonstrations were collected against a SPECIFIC +fixture, so the fixture is inside the policy's observation and inside its action +distribution. Reprint that fixture with a tuned clearance, a new spool, a different machine +or a different shrinkage and the policy is now running against a world it never saw, with +every demonstration behind it collected against the old one. Nothing in the run reports +this. The arm moves to the same taught coordinates, the plate is a millimetre high in the +pocket, and the policy is confidently doing the wrong thing. So `FixtureCoupling` carries +the revision the captures were taken against and `fixture_coupled` computes whether a +proposed run is still coupled to it -- and it will not accept a matching revision STRING as +proof, because a revision names a model file, `printed` refuses to estimate shrinkage for a +material at all, and two prints of one file are not established to be one geometry until +somebody puts an instrument on both. + +A POLICY IS UNTRUSTED UNTIL MEASURED ON HELD-OUT TRIALS, AGAINST AN EXTERNAL CRITERION. Two +different numbers get reported as success rates and neither one is one. The first is a rate +computed on the initial conditions the demonstrations were collected from, which is +structurally optimistic: behavior cloning is trained on the expert's state distribution and +its own errors carry it off that distribution, with suboptimality compounding as the square +of the horizon rather than linearly, and the effect only appears when the policy is rolled +out from starts nobody demonstrated. The measured version of that gap: 11 of 12 on training conditions against 11 of 24 at a new camera position (95 percent intervals 61.5 to 99.8 and 25.6 to 67.2), which +is the only one of the reported shifts distinguishable from baseline at that trial count. A new +table texture. The second is any criterion that reads the policy's own state. A policy +reporting high confidence on a transfer that failed is precisely the silent destructive +failure `coverage` is built around -- a number arrives, the check passes, and the material is +already gone -- and offline loss is barely better as a proxy, at a measured Pearson r of +0.308 against real-world ranking where a visual-matching simulator reached 0.924. So `Trust` +has exactly one trusted state and it requires held-out trials AND a criterion that does not +read the policy. Underneath it, `success_rate` refuses a point estimate below a declared +trial floor and reports an exact interval rather than a bare percentage where it reports +anything, because at 20 trials one trial is 5 percentage points and a 90 percent observation +carries a 95 percent interval of 68.3 to 98.8 -- which does not distinguish a policy that +works from one that fails one attempt in three. + +WHAT THIS MODULE WILL NOT DO. It returns no demonstration count for a new task, no exchange +rate between a human video and a teleoperated episode, no success rate from a handful of +trials, and no probability that the next attempt works. `demonstrations_needed` is the +sharpest case: the published real-robot per-task counts span 20 to 8,658, a spread of more +than two orders of magnitude, and the only controlled study of the question found that +environment and object DIVERSITY dominates raw count -- so a single number would be a +convention dressed as a measurement. `NOT_ESTIMATED` names each refusal and what would lift +it, in the form `durability` uses, and reuses that module's `Refusal` rather than defining a +second dataclass with the same three fields. + +AUTHORITY, IN PROSE. A learned policy PROPOSES a motion; a deterministic interlock DISPOSES. +Workspace, speed and force belong to something that does not read the policy's output, has +no confidence to be wrong about, and cannot be argued out of a limit by a fluent one. That +boundary is not modelled here and is deliberately not imported: see `INTERLOCK_NOTE` for +what such a layer has to be, and for why the vendor features usually pointed at -- a reduced +mode, a Cartesian safety boundary, current-and-model collision detection on dimensionless +sensitivity levels -- are documented conveniences rather than rated safety functions. +""" + +from __future__ import annotations + +import math +from dataclasses import dataclass, field +from enum import Enum +from typing import Dict, List, Optional, Tuple + +from .durability import Refusal +from .printed import Fixture +from .qc import Basis +from .vision import VisionRequirement + + +# -- how a capture carries actions ------------------------------------------------- + + +class Modality(str, Enum): + """How a demonstration was captured, ordered by how directly it carries actions. + + The ordering is the argument. At the top the action stream is a recording of what a + machine did; at the bottom it is the output of an estimator run over pixels, and every + member in between trades some of the first thing for some of the convenience of the + second. Nothing here says the bottom of the list does not work -- a video-only pipeline + reached 92 percent on pick-and-place of a book and 64 percent on tying a rope, which is + a working policy by any standard. It says the bottom of the list is not the top of it, + and that a module which stores both in one field will eventually compare them as though + they were. + + MONOCULAR_HUMAN_VIDEO is the member the pitch usually means by "just record it on your + phone", and it is the one with a measured penalty rather than a suspected one: the study + that swapped multi-view triangulation for a real monocular hand estimator measured 185.67 + mm mean per-joint position error and mean success falling from 41.5 to 24.7 percent. + """ + + TELEOPERATION = "teleoperation" # leader-follower or jog; the robot executed it + KINESTHETIC = "kinesthetic" # hand-guided on the arm itself; its own encoders recorded it + HANDHELD_GRIPPER = "handheld_gripper" # real gripper hardware in the human's hand + INSTRUMENTED_HUMAN_VIDEO = "instrumented_human_video" # RGBD or multi-view; pose estimated + MONOCULAR_HUMAN_VIDEO = "monocular_human_video" # one third-person camera, no depth + + @property + def measured_action_stream(self) -> bool: + """Whether the action was RECORDED by an instrument rather than estimated from pixels. + + True for the three hardware-in-the-loop captures and false for both video ones. The line + is not "human versus robot" and not "cheap versus expensive": it is whether an estimator + with an unmeasured error sits between the capture and the action label. A handheld + gripper clears it because the jaw width is read off the device and the end-effector pose + comes from on-device geometry with a stated trajectory error of 6.1 mm and 3.5 deg, + which is a measurement of the capture rather than a hope about it. + """ + return self in ( + Modality.TELEOPERATION, + Modality.KINESTHETIC, + Modality.HANDHELD_GRIPPER, + ) + + @property + def on_the_robots_own_kinematics(self) -> bool: + """Whether the recorded action is in the action space of the robot that will replay it. + + The stricter line, and the one `equivalent_to_teleop` keys off. A handheld gripper + matches the OBSERVATION space by construction and still hands the controller a pose + recorded by a device with different mass, different reach and no joint limits -- both + failures in the one twenty-trial cross-embodiment transfer counted here were joint limit + violations. It also cannot measure force: grasp force is inferred from finger deformation, + and the one clean ablation of the signal a video physically cannot contain -- relative + inter-gripper proprioception -- took bimanual cloth folding from 70 to 30 percent. + """ + return self in (Modality.TELEOPERATION, Modality.KINESTHETIC) + + +MODALITY_EVIDENCE: Dict[Modality, str] = { + Modality.TELEOPERATION: ( + "actions are recorded in the robot's own action space at controller rate, with " + "proprioception, gripper width and joint configuration alongside. This is what the " + "frontier models are still built on: one generalist policy is pretrained on roughly " + "10,000 hours of teleoperated data across 68 tasks and 7 robot configurations. What the " + "usual phrase 'zero embodiment gap' overstates is that teleoperated data is still " + "OFF-POLICY with respect to the state distribution a trained policy visits, which is the " + "covariate-shift problem behavior cloning has had since it was named" + ), + Modality.KINESTHETIC: ( + "hand-guiding records the same joint states as teleoperation, on the arm that will " + "replay them, and adds the constraint that a human hand was inside the envelope while it " + "happened. No published comparison of kinesthetic against teleoperated data was located " + "here, so the two are treated as one class on the grounds that both read the robot's own " + "encoders -- which is an argument from mechanism and not a measurement" + ), + Modality.HANDHELD_GRIPPER: ( + "the real gripper in the human's hand, so observation and action spaces match by " + "construction. Measured capture precision is 6.1 mm and 3.5 deg mean absolute trajectory " + "error, 10.1 mm and 0.8 deg inter-gripper relative pose. Task rates from 250 to 305 " + "episodes: cup arrangement 20 of 20, dynamic tossing 105 of 120, bimanual cloth folding " + "14 of 20, dish washing 14 of 20. It does not measure force -- grasp force is inferred " + "from finger deformation -- and bare-hand collection, the thing this replaces, runs at " + "only 48 to 64 percent of the handheld device's own collection speed, so the cost " + "argument for going further down this list is weaker than it sounds" + ), + Modality.INSTRUMENTED_HUMAN_VIDEO: ( + "RGBD or multi-view capture with hand pose estimation and retargeting. This is the " + "configuration behind every human-video result that actually executes a task, and it is " + "not action-free -- it is action-INFERRED, with explicit end-effector pose and gripper " + "width derived from depth and a hand mesh. Head to head against teleoperation on one " + "sweeping task: 50 teleoperated 52 percent, 50 human-video 44 percent, 100 teleoperated " + "88 percent, 300 human-video 84 percent. Stated scope limits from the same work: pinch " + "grasps only, quasi-static tasks only, and it works only where the robot can follow the " + "same strategy the human did" + ), + Modality.MONOCULAR_HUMAN_VIDEO: ( + "one third-person camera and no depth, which is what 'internet-scale video' actually " + "means. It is the measured degraded case rather than the suspected one: substituting a " + "real monocular hand estimator for multi-view triangulation gave 185.67 mm mean " + "per-joint position error and took mean success from 41.5 to 24.7 percent. No published " + "policy trained with zero action information anywhere in the pipeline was located here " + "that performs a real manipulation task at useful rates -- an absence of evidence from " + "this search rather than a proof of impossibility, and worth stating as the first, " + "because it is the claim the pitch rests on" + ), +} + + +@dataclass(frozen=True) +class Retargeting: + """The step that turns a human's hand into a robot's action, and what measured it. + + Carried as its own object because the alternative is a boolean on `Capture` reading + `retargeted=True`, and a boolean invites the reading that retargeting is a conversion + somebody performed rather than a model somebody fitted. It is the second of two estimators + in series -- hand pose, then the map onto a gripper that has different kinematics -- and + each has its own error. + + `error_measured_mm` is the field that is almost never filled in. No verified millimetre + figure for human-to-robot retargeting error was located here at all: the work that defines + the right metrics reports them only as unlabelled bar charts. The two magnitudes that ARE + published belong to the hand-pose stage, at 185.67 mm mean per-joint error for a monocular + estimator and 3D errors "as large as 20-30 cm" from another group, so an unmeasured + retargeting is not a small unknown. Kinematic validity is also not the end of it: hands + retargeted to six dexterous robot hands produced human-like motions that were "not + feasible for completing the task" when replayed in simulation. + + `basis` is `qc.Basis` and `validated` demands IN_HOUSE for the same reason a threshold + does. A retargeting error measured on somebody else's arm, hand and camera is evidence + about their pipeline. + + A measured error must be a positive distance. Zero and negative are refused at + construction rather than reported as an unusually good pipeline, because both are the + signature of a value that was defaulted or sign-flipped rather than measured, and this + field's whole job is to be the one number nobody filled in. + """ + + method: str = "" + depth_or_multiview: bool = False + error_measured_mm: Optional[float] = None + basis: Basis = Basis.INTUITION + note: str = "" + + def __post_init__(self) -> None: + if self.error_measured_mm is not None and self.error_measured_mm <= 0: + raise ValueError( + f"a retargeting error of {self.error_measured_mm} mm is not a measurement. An error " + "is a positive distance; zero and negative are the signature of a defaulted or " + "sign-flipped value, and this field exists precisely because it is the one nobody " + "fills in. Leave it None if it was not measured" + ) + + @property + def declared(self) -> bool: + return bool(self.method) + + @property + def validated(self) -> bool: + return self.declared and self.error_measured_mm is not None and self.basis.validated + + def missing(self) -> Tuple[str, ...]: + """Exactly what stands between this pipeline and `validated`, and nothing else. + + The monocular caveat is deliberately not in this list. Being monocular is a property of + the capture rather than a gap in the validation, and a measured error on a monocular rig + is still a measurement -- it is reported where it belongs, next to the refusal that uses + it, rather than here where it would read as a requirement nobody can meet after the fact. + """ + out: List[str] = [] + if not self.declared: + out.append("no retargeting method is named") + if self.error_measured_mm is None: + out.append("no retargeting error was measured on this capture") + elif not self.basis.validated: + out.append( + f"the retargeting error rests on {self.basis.value}, which is evidence about somebody " + "else's pipeline" + ) + return tuple(out) + + +# -- the capture ------------------------------------------------------------------- + + +@dataclass(frozen=True) +class Capture: + """One recording somebody hopes to train from, with every field at its untrusted value. + + Constructible with no arguments, and a capture constructed that way trains nothing. That + is the test this dataclass exists to pass, and it is the same discipline `printed.Fixture` + applies: a default that reads as a pass turns a recording nobody checked into a training + set. + + `episodes` counts recordings, not usable ones, and nothing here converts a count into a + sufficiency claim -- see `demonstrations_needed`, which refuses to. + + `coupling` is on the capture rather than on the policy because the fixture is a property + of the session that was recorded, and a policy trained on two sessions against two fixture + revisions is coupled to neither. `Policy.training_coupling` enforces exactly that, in the + way `teaching.Envelope` refuses to pool two sets of conditions into one range. + """ + + name: str = "" + modality: Optional[Modality] = None + episodes: int = 0 + retargeting: Retargeting = field(default_factory=Retargeting) + timestamps_synchronized: bool = False + camera_calibrated: bool = False + gripper_state_recorded: bool = False + coupling: Optional["FixtureCoupling"] = None + note: str = "" + + def __post_init__(self) -> None: + if self.episodes < 0: + raise ValueError(f"capture '{self.name or 'unnamed'}' declares {self.episodes} episodes") + measured = self.modality is not None and self.modality.measured_action_stream + if measured and self.retargeting.declared: + raise ValueError( + f"capture '{self.name or 'unnamed'}' is {self.modality.value}, which records the " + "action directly, and also declares a retargeting method. One capture cannot both " + "measure its actions and estimate them, and a record that claims both will be read " + "as whichever is more convenient" + ) + + @property + def carries_actions(self) -> bool: + """Whether an action stream exists at all, measured or inferred. + + An inferred one counts here and does not count in `equivalent_to_teleop`. Those are + different questions: whether there is anything to train on, and whether what there is + can stand in for the thing it is being compared to. + """ + if self.modality is None: + return False + return self.modality.measured_action_stream or self.retargeting.declared + + +# -- what stands between a capture and a training set -------------------------------- + + +class Trainable(str, Enum): + """Whether a capture can train a policy, and if not, what specifically is missing. + + USABLE is a member of the same enum as the refusals, following `printed.Fitness`, because + usability is the absence of every refusal rather than a state of its own. No field sets it. + + `needs_recapture` is the line worth acting on. Some of these are cleared by doing work on + data that already exists -- write down the modality, calibrate the camera that has not + moved, measure the retargeting error. The rest are cleared only by recording the session + again, because the missing signal was never written to disk and no processing recovers it. + A report that mixes the two sends a lab to reprocess a dataset that has to be recollected. + """ + + MODALITY_UNDECLARED = "modality_undeclared" + NO_EPISODES = "no_episodes" + NO_ACTION_STREAM = "no_action_stream" + RETARGETING_UNVALIDATED = "retargeting_unvalidated" + TIMESTAMPS_UNSYNCHRONIZED = "timestamps_unsynchronized" + CAMERA_UNCALIBRATED = "camera_uncalibrated" + GRIPPER_STATE_MISSING = "gripper_state_missing" + USABLE = "usable" + + @property + def usable(self) -> bool: + return self is Trainable.USABLE + + @property + def needs_recapture(self) -> bool: + return self in ( + Trainable.NO_EPISODES, + Trainable.NO_ACTION_STREAM, + Trainable.TIMESTAMPS_UNSYNCHRONIZED, + Trainable.GRIPPER_STATE_MISSING, + ) + + +# Reported in this order. USABLE is not in it: it is what is left when nothing else applies. +_MISSING_ORDER: Tuple[Trainable, ...] = ( + Trainable.MODALITY_UNDECLARED, + Trainable.NO_EPISODES, + Trainable.NO_ACTION_STREAM, + Trainable.RETARGETING_UNVALIDATED, + Trainable.TIMESTAMPS_UNSYNCHRONIZED, + Trainable.CAMERA_UNCALIBRATED, + Trainable.GRIPPER_STATE_MISSING, +) + + +@dataclass(frozen=True) +class Missing: + """One thing a capture lacks, named specifically enough to go and get.""" + + reason: Trainable + detail: str + + +@dataclass(frozen=True) +class Usability: + """One capture, costed against what training actually requires. + + Carries every refusal rather than returning on the first, for the reason + `printed.Assessment` does: a capture is usually short of more than one thing, and a report + showing only the headline sends somebody to fix that and rediscover the next one after a + week of reprocessing. + """ + + capture: Capture + missing: Tuple[Missing, ...] + + @property + def usable(self) -> bool: + return not self.missing + + @property + def verdict(self) -> Trainable: + return self.missing[0].reason if self.missing else Trainable.USABLE + + def reasons(self) -> Tuple[Trainable, ...]: + return tuple(m.reason for m in self.missing) + + def details(self) -> Tuple[str, ...]: + return tuple(m.detail for m in self.missing) + + @property + def needs_recapture(self) -> bool: + return any(m.reason.needs_recapture for m in self.missing) + + @property + def reason(self) -> str: + if self.missing: + return self.missing[0].detail + return ( + f"'{self.capture.name or 'this capture'}' carries an action stream, synchronized " + "timestamps, a calibrated camera and a gripper state. That is a training set and not a " + "trained policy, and it says nothing about what the policy will do" + ) + + +def usable_for_training(capture: Capture) -> Usability: + """Everything standing between this recording and a training set. + + Every check runs. Nothing shortcuts on a missing modality, because a capture with no + declared modality still either has synchronized timestamps or does not, and a report that + hid the rest behind the first refusal would understate the work by four items. + """ + found: List[Missing] = [] + + if capture.modality is None: + found.append( + Missing( + Trainable.MODALITY_UNDECLARED, + "no modality is declared, so whether the action stream was recorded or estimated " + "cannot be resolved, and that is the distinction the whole question turns on. An " + "undeclared capture is not a neutral one: it will be read as whichever modality the " + "reader was hoping for", + ) + ) + + if capture.episodes <= 0: + found.append( + Missing( + Trainable.NO_EPISODES, + "the capture holds no episodes. A session that was configured and never recorded " + "looks identical in a manifest to one that was recorded, which is why the count is " + "checked rather than assumed from the record existing", + ) + ) + + if not capture.carries_actions: + found.append( + Missing( + Trainable.NO_ACTION_STREAM, + "there is no action stream: the modality does not record one and no retargeting is " + "declared to infer one. A policy is a map from observation to ACTION, and pixels are " + "the observation half. Recovering the missing half from video needs the dynamics to " + "be injective, and a redundant arm posture or a force pressed against a rigid " + "constraint changes no visible state at all, so for those the action is unrecoverable " + "in principle and not merely unrecovered here", + ) + ) + elif capture.modality is not None and not capture.modality.measured_action_stream: + if not capture.retargeting.validated: + detail = ( + "the action stream is inferred by a retargeting pipeline that nobody measured here: " + + "; ".join(capture.retargeting.missing()) + + ". Retargeting is an additional fitted step, not a format conversion. Calibrated " + "noise on hand labels moved mean success from 48.3 to 30.0 percent at half a sigma " + "and to 20.0 at one sigma, and supplying ground-truth actions for as little as 2.5 " + "percent of a dataset improved downstream performance 4.2-fold -- which is a " + "measurement of what action labels buy that video cannot supply" + ) + if not capture.retargeting.depth_or_multiview: + detail += ( + ". The capture is also monocular, which is the configuration that measured 185.67 mm " + "mean per-joint error and mean success falling from 41.5 to 24.7 percent" + ) + found.append(Missing(Trainable.RETARGETING_UNVALIDATED, detail)) + + if not capture.timestamps_synchronized: + found.append( + Missing( + Trainable.TIMESTAMPS_UNSYNCHRONIZED, + "camera frames and the action stream are not synchronized to a common clock, so every " + "training pair is an observation matched with an action from an unknown time offset. " + "The magnitude is not subtle: in a hand-eye study the same calibration procedure gave " + "1.57 mm of residual error in static mode and 18.89 mm in motion mode, and the authors " + "attributed the roughly fortyfold degradation to imprecise timestamp synchronization " + "rather than to the mathematics. A constant offset teaches the policy to act late; a " + "varying one teaches it noise", + ) + ) + + if not capture.camera_calibrated: + found.append( + Missing( + Trainable.CAMERA_UNCALIBRATED, + "the camera pose is not calibrated or recorded -- the requirement `vision` already " + f"names as {VisionRequirement.POSE.value}. Two separate things need it. Retargeting " + "human video into the robot's frame is impossible without metric camera geometry. And " + "a pixel-space policy, which is what the standard architectures are, never reads the " + "extrinsics at all: it infers the viewpoint from the scene and silently encodes it, so " + "what the calibration buys is the ability to REPRODUCE the training viewpoint at " + "rollout. Camera pose is the measured-hardest shift there is -- 11 of 12 on training conditions against 11 of 24 at a new camera position (95 percent intervals 61.5 to 99.8 and 25.6 to 67.2). " + "Separately and in SIMULATION, orientation " + "perturbations of 2 to 10 deg, which is a sloppy remount, collapsed one model from " + "perturbations of 2 to 10 deg dropped nine models by between 19 and 91 points, a " + "spread wider than the effect for several of them", + ) + ) + + if not capture.gripper_state_recorded: + found.append( + Missing( + Trainable.GRIPPER_STATE_MISSING, + "no gripper state is recorded, so the capture cannot say when the hand closed. Every " + "pick is a reach the policy can imitate and a grasp it cannot, and the failure this " + "produces is specific and known: the reported failure modes on the worst task in the " + "reference set were the gripper closing too early and imprecise insertion. Estimating " + "the width from a thumb-to-fingertip distance is the standard heuristic and is " + "convention rather than a validated mapping -- the work that uses it restricts itself " + "to pinch grasps precisely because of it", + ) + ) + + by_reason = {m.reason: m for m in found} + ordered = tuple(by_reason[r] for r in _MISSING_ORDER if r in by_reason) + return Usability(capture=capture, missing=ordered) + + +def equivalent_to_teleop(capture: Capture) -> Tuple[bool, str]: + """Does this capture stand in for a teleoperated one, episode for episode? + + Answers no for everything that is not recorded on the robot's own kinematics, including a + fully validated retargeting pipeline and including a handheld gripper. The refusal is not + a judgment about quality. It is that the one published head-to-head exchange measured + roughly three human-video episodes to approach -- not beat -- what one teleoperated + episode bought, on a quasi-static sweeping task with pinch grasps, which is the friendliest + possible setting for the comparison. Treating the two as interchangeable turns that + exchange rate into a silent multiplier on every count downstream. + """ + if capture.modality is None: + return False, ( + "no modality is declared, so this capture is not established to be equivalent to " + "anything. An undeclared capture defaults to untrusted here for the same reason an " + "unbenchmarked operation does in `intelligence.trusted_for`" + ) + if capture.modality.on_the_robots_own_kinematics: + return True, ( + f"{capture.modality.value} records the robot's own joint states at controller rate, so " + "the action stream is what the machine did. This is equivalence of FORMAT and not of " + "quality: teleoperated data is still off-policy with respect to the distribution the " + "trained policy will visit" + ) + if capture.modality is Modality.HANDHELD_GRIPPER: + return False, ( + "a handheld gripper matches the observation and action spaces by construction and is " + "still not a teleoperated episode: the pose is recorded by a device with different mass " + "and reach and no joint limits, force is inferred from finger deformation rather than " + "measured, and the one clean ablation of relative inter-gripper proprioception -- a " + "signal this capture has and a video does not -- took bimanual cloth folding from 70 to " + "30 percent success" + ) + detail = ( + f"{capture.modality.value} carries no measured action stream. The published exchange rate " + "on the one head-to-head comparison located here is roughly 3 human-video episodes to " + "approach what 1 teleoperated episode buys -- 50 teleoperated 52 percent against 50 " + "human-video 44, and 100 teleoperated 88 against 300 human-video 84 -- inside stated " + "limits of pinch grasps and quasi-static tasks" + ) + if capture.retargeting.validated: + return False, ( + detail + + ". This capture's retargeting IS measured, which makes it usable for training and " + "does not make it teleoperation: a measured estimator is still an estimator, and the " + "exchange rate was measured against pipelines at least this good" + ) + return False, ( + detail + + ". This capture's retargeting is also unmeasured, so the exchange rate above is an " + "optimistic bound on it rather than a description of it" + ) + + +# -- the fixture the demonstrations were collected against ---------------------------- + + +class Coupling(str, Enum): + """Whether a proposed run faces the fixture the policy learned against. + + `holds` is true for exactly one member, and UNVERIFIED is the member that makes the enum + worth having. Two parts printed from one revision of one file are not established to be + one geometry: `printed` refuses to estimate shrinkage for a material at all, on the + grounds that no filament datasheet checked there publishes one and two vendors' slicer + profiles disagree by 0.5 percent on the same polymer. A revision string is a fact about a + file. The policy was trained against a part. + """ + + COUPLED = "coupled" # same fixture, same revision, and both parts were measured + UNVERIFIED = "unverified" # revision strings match and at least one part was never measured + REVISED = "revised" # same fixture, different revision + DIFFERENT_FIXTURE = "different_fixture" + MIXED = "mixed" # the captures were collected against more than one fixture revision + UNIDENTIFIED = "unidentified" # a side names no fixture, or names one with no revision + + @property + def holds(self) -> bool: + return self is Coupling.COUPLED + + +@dataclass(frozen=True) +class FixtureCoupling: + """The fixture a set of demonstrations was collected against, and the revision of it. + + WHAT THE EVIDENCE SUPPORTS, since the variance-reduction claim is usually stated far more + strongly than it has been measured. The direction is real and measured twice. A simulated + grasping study fitted required data against the volume the object's position was randomized + over and got an exponent of 0.35 on volume, with measured points running 728 trajectories + at a fixed point up to 24,005 over a 41 by 30 by 28 cm space. A generation study held the + demonstration budget fixed at 1,000 and varied only the reset distribution: with the + receptacle rigidly fixed -- a fixture in all but name -- square-peg assembly reached 90.7 + percent, and with both objects free it fell to 49.3; three-piece assembly went 82.0 to 13.3. + + WHAT THE SAME EVIDENCE COSTS THE CLAIM. That 0.35 exponent on VOLUME is about 1.05 in + linear extent, so cutting placement scatter by a factor of two in each axis -- eight times + less volume -- buys roughly two times less data, not eight. In the same fixed-budget table, + block stacking went from 100.0 percent in a 16 cm region to 99.3 in a 40 cm one: on an easy + low-precision task, constraining pose bought seven tenths of a point. And the strongest + version of the pitch is contradicted outright -- agents trained on the 10 human source + demonstrations and evaluated INSIDE the narrow fixtured region they were collected in + scored 11.3 and 19.3 and 1.3 percent in simulation, and 0 percent on the real robot. A + fixture shifts the data curve. It does not collapse the requirement to a handful of + demonstrations. + + WHAT IS ENGINEERING PRACTICE AND NOT EVIDENCE. The 3-2-1 locating principle, which is the + textbook justification for every jig ever cut, has no publication behind it that quantifies + its effect on a learned policy's demonstration count. Neither does any figure of the form + "a nest cuts demonstrations by N times" for a laboratory workcell. Every quantified source + located here substitutes something for the physical fixture -- a randomization range in + simulation, a reset distribution at a fixed budget, or software pose canonicalization, + where anchoring the frame to the object reached with 10 demonstrations what an image-based + policy needed 305 to match. No study manipulates a PHYSICAL fixture with an unfixtured + control and counts human demonstrations. That study does not exist in this evidence set. + + `fixture` is the `printed.Fixture` itself where a caller has one, so the question of + whether the part was ever measured resolves in the module that owns it rather than being + re-asserted here with a boolean that can drift. + """ + + fixture_name: str = "" + revision: str = "" + fixture: Optional[Fixture] = None + constrains: Tuple[str, ...] = () # which degrees of freedom the part removes, in plain words + note: str = "" + + @property + def identified(self) -> bool: + """A fixture with no revision is not identified. A revision is what changes underneath.""" + return bool(self.fixture_name) and bool(self.revision) + + @property + def anchored(self) -> bool: + """Whether the revision names a part somebody measured, rather than only a model file. + + Resolved through `printed.Fixture.dimensions_measured`, which already refuses the vacuous + pass where a part declaring no critical dimensions reads as fully measured. + """ + return self.fixture is not None and self.fixture.dimensions_measured() + + def key(self) -> Tuple[str, str]: + return (self.fixture_name, self.revision) + + def describe(self) -> str: + if not self.identified: + return "an unidentified fixture" + return f"'{self.fixture_name}' rev {self.revision}" + + +@dataclass(frozen=True) +class CouplingCheck: + """One policy's training fixture against the fixture a proposed run would face.""" + + training: Optional[FixtureCoupling] + proposed: Optional[FixtureCoupling] + state: Coupling + reason: str + + @property + def holds(self) -> bool: + return self.state.holds + + +def fixture_coupled(policy: "Policy", proposed: Optional[FixtureCoupling]) -> CouplingCheck: + """Is a proposed run still coupled to the fixture the policy learned on? + + Refuses in five distinguishable ways rather than returning a boolean, because the five cost + different work: recollect, reprint to the old revision, measure both parts, name the + fixture, or split the training set. A boolean would send every one of them to the same + place. + """ + training = policy.training_coupling() + + if training is None: + if len(policy.distinct_couplings()) > 1: + names = "; ".join(c.describe() for c in policy.distinct_couplings()) + return CouplingCheck( + None, + proposed, + Coupling.MIXED, + f"this policy's captures were collected against {len(policy.distinct_couplings())} " + f"different fixture revisions ({names}), so it is coupled to none of them. Pooling " + "them is the same error `teaching.Envelope` refuses when it will not put two sets of " + "conditions in one range: the union is wider than either, and wide is not the safe " + "direction here -- it hides which part the policy actually learned against", + ) + return CouplingCheck( + None, + proposed, + Coupling.UNIDENTIFIED, + "no capture behind this policy names the fixture it was collected against. The fixture " + "is inside the policy's observations whether or not anybody wrote it down, so an " + "unrecorded one cannot be reproduced and cannot be shown to have changed", + ) + + if proposed is None or not proposed.identified: + return CouplingCheck( + training, + proposed, + Coupling.UNIDENTIFIED, + f"the policy was trained against {training.describe()} and the proposed run names no " + "identified fixture. An unnamed part is not the same part; it is an unknown one", + ) + + if not training.identified: + return CouplingCheck( + training, + proposed, + Coupling.UNIDENTIFIED, + f"the captures name '{training.fixture_name or 'no fixture'}' with no revision, so there " + "is nothing for the proposed run to match. A fixture without a revision cannot be shown " + "to be unchanged, which is the only thing this check can establish", + ) + + if training.fixture_name != proposed.fixture_name: + return CouplingCheck( + training, + proposed, + Coupling.DIFFERENT_FIXTURE, + f"the policy learned against {training.describe()} and the run proposes " + f"{proposed.describe()}. This is not a degraded coupling, it is a different world", + ) + + if training.revision != proposed.revision: + return CouplingCheck( + training, + proposed, + Coupling.REVISED, + f"the policy learned against {training.describe()} and the run proposes revision " + f"{proposed.revision}. Every demonstration behind this policy was collected against the " + "old part, and a reprint changes clearance, shrinkage and first-layer width together -- " + "`printed` will not estimate shrinkage for a material at all, and quotes desktop " + "tolerance at plus or minus 0.5 percent with a floor of plus or minus 0.5 mm, scaling " + "with length. Nothing in the run reports this: the arm moves to the same taught " + "coordinates and the pose it was trained to expect is somewhere else", + ) + + if not (training.anchored and proposed.anchored): + unmeasured = [] + if not training.anchored: + unmeasured.append("the part the captures were taken against") + if not proposed.anchored: + unmeasured.append("the part the run would use") + return CouplingCheck( + training, + proposed, + Coupling.UNVERIFIED, + f"both sides name {training.describe()} and no measurement anchors that to a part: " + f"{' and '.join(unmeasured)} carries no measured critical dimension. A revision is a " + "fact about a model file. Two prints of one file are not established to be one geometry, " + "which is why `printed.DimensionState` keeps DESIGNED and MEASURED apart -- and this is " + "cleared by a caliper rather than by a different part", + ) + + return CouplingCheck( + training, + proposed, + Coupling.COUPLED, + f"the run faces {proposed.describe()}, the revision the demonstrations were collected " + "against, and both parts carry measured critical dimensions. This says the world the " + "policy learned is the world it will meet. It says nothing about whether it learned it", + ) + + +# -- measuring a policy --------------------------------------------------------------- + + +class SuccessCriterion(str, Enum): + """What counted as a success, which is the part of a success rate nobody writes down. + + `external` is the load-bearing line and it is drawn at whether the criterion reads the + policy's own state. TRAINING_LOSS is on the wrong side of it and belongs in the enum + because it is the most respectable-looking way to land there: offline validation error + against real-world policy ranking measured a Pearson r of 0.308, where a visual-matching + simulator on the same comparison reached 0.924. + + OPERATOR_SCORED clears the line and is the weakest thing that does. It is a human, and a + human is outside the policy, which is the whole requirement. It is not free of error: a + quality audit of 27 percent of roughly 2,700 rollouts in the most careful published + campaign found a 2.31 percent discrepancy on the success label itself. + """ + + POLICY_CONFIDENCE = "policy_confidence" # the policy grading its own rollout + TRAINING_LOSS = "training_loss" # offline error standing in for real success + OPERATOR_SCORED = "operator_scored" # a person watching the rollout and calling it + EXTERNAL_MEASUREMENT = "external_measurement" # an instrument or assay outside the policy + + @property + def external(self) -> bool: + return self in (SuccessCriterion.OPERATOR_SCORED, SuccessCriterion.EXTERNAL_MEASUREMENT) + + +# How much a single trial is allowed to move the reported number, in percentage points. This +# is a convention and is stated rather than buried: no published result establishes any +# minimum trial count, one careful group states outright that the required N depends on the +# true performance gap and so cannot be fixed in advance, and the canonical +# unreported-variance paper explicitly declines to name a number. What the convention buys is +# that the arithmetic stops being absurd -- at 5 trials one trial is 20 points, which is the +# published row where "40 percent" is 2 successes out of 5. Raising this would be defensible. +# Having no floor is not. +MAX_TRIAL_WEIGHT_PP = 5.0 + +# The floor that follows from it, computed rather than picked: one trial is 100/n points. +MIN_HELDOUT_TRIALS = int(math.ceil(100.0 / MAX_TRIAL_WEIGHT_PP)) + +# Field practice, for calibration rather than as a target: the modal real-robot evaluation in +# the papers surveyed here runs 10 to 20 trials per condition, and the reference works run 25, +# 20 and 20. This floor sits at the top of that range and is still wide enough to be +# uninformative about a 5 point difference -- separating 85 percent from 90 percent at +# conventional power needs 686 trials per policy, which is not a number any of them ran. + +# The coverage every interval in this module is reported at. +CONFIDENCE = 0.95 + + +def _log_choose(n: int, k: int) -> float: + return math.lgamma(n + 1.0) - math.lgamma(k + 1.0) - math.lgamma(n - k + 1.0) + + +def _binom_at_most(k: int, n: int, p: float) -> float: + """P(X <= k) for X ~ Binomial(n, p), summed in log space so large n does not overflow.""" + if p <= 0.0: + return 1.0 + if p >= 1.0: + return 1.0 if k >= n else 0.0 + lp = math.log(p) + lq = math.log1p(-p) + total = 0.0 + for i in range(0, k + 1): + total += math.exp(_log_choose(n, i) + i * lp + (n - i) * lq) + return min(1.0, total) + + +def _binom_at_least(k: int, n: int, p: float) -> float: + """P(X >= k) for X ~ Binomial(n, p).""" + if p <= 0.0: + return 1.0 if k <= 0 else 0.0 + if p >= 1.0: + return 1.0 + lp = math.log(p) + lq = math.log1p(-p) + total = 0.0 + for i in range(k, n + 1): + total += math.exp(_log_choose(n, i) + i * lp + (n - i) * lq) + return min(1.0, total) + + +def _bisect(fn, target: float, decreasing: bool) -> float: + lo, hi = 0.0, 1.0 + for _ in range(80): + mid = 0.5 * (lo + hi) + value = fn(mid) + if (value > target) == decreasing: + lo = mid + else: + hi = mid + return 0.5 * (lo + hi) + + +def success_interval( + successes: int, trials: int, confidence: float = CONFIDENCE +) -> Tuple[float, float]: + """The exact Clopper-Pearson interval on an observed rate, as fractions. + + Exact rather than normal-approximate because every trial count in this domain is small + enough for the approximation to be the wrong tool, and conservative rather than + average-coverage because the direction of the error matters here: this module exists to + stop a policy being trusted with material, so an interval that is slightly too wide fails + in the safe direction and one that is slightly too narrow does not. That choice is a + convention -- a statistical one this evidence set does not adjudicate -- and it is stated + rather than implied. + """ + if trials <= 0: + raise ValueError("an interval needs at least one trial; there is nothing to bound") + if successes < 0 or successes > trials: + raise ValueError(f"{successes} successes in {trials} trials is not an observation") + alpha = (1.0 - confidence) / 2.0 + low = 0.0 if successes == 0 else _bisect( + lambda p: _binom_at_least(successes, trials, p), alpha, decreasing=False + ) + high = 1.0 if successes == trials else _bisect( + lambda p: _binom_at_most(successes, trials, p), alpha, decreasing=True + ) + return low, high + + +@dataclass(frozen=True) +class Evaluation: + """Rollouts of a trained policy, and what counted as a success. + + Every default is the untrusted one. `held_out` defaults False because the natural thing to + do with a trained policy is run it from the setups that were demonstrated, and that number + is the one that gets quoted; `criterion` defaults to the policy grading itself because that + is the reading a script produces with no extra work. + + `blinded` is recorded and is NOT required for trust. Blind, interleaved evaluation is + advocated everywhere and no published robotics study located here measures the size of the + bias non-blind scoring introduces, so requiring it would be inventing a threshold out of + hygiene. Reporting it is free; demanding it on this evidence would be the same overreach + this module refuses elsewhere. + """ + + trials: int = 0 + successes: int = 0 + held_out: bool = False + criterion: SuccessCriterion = SuccessCriterion.POLICY_CONFIDENCE + blinded: bool = False + note: str = "" + + def __post_init__(self) -> None: + if self.trials < 0: + raise ValueError(f"an evaluation cannot have {self.trials} trials") + if self.successes < 0: + raise ValueError(f"an evaluation cannot have {self.successes} successes") + if self.successes > self.trials: + raise ValueError( + f"{self.successes} successes in {self.trials} trials is not an observation; a rate " + "above 1 is the arithmetic telling you the two numbers came from different places" + ) + + +@dataclass(frozen=True) +class RateReport: + """A success rate with its interval, or the refusal to state one. + + `point` is None wherever the rate is refused, never 0.0 and never the raw fraction with a + warning attached, for the reason `printed.CriticalDimension.deviation_mm` returns None + rather than zero: a number in the field will be read, and any string beside it will not. + """ + + evaluation: Evaluation + stated: bool + point: Optional[float] + low: Optional[float] + high: Optional[float] + reason: str + + @property + def width_pp(self) -> Optional[float]: + if self.low is None or self.high is None: + return None + return 100.0 * (self.high - self.low) + + def describe(self) -> str: + if not self.stated or self.point is None: + return f"no rate: {self.reason}" + return ( + f"{self.evaluation.successes} of {self.evaluation.trials} " + f"({100.0 * self.point:.1f} percent, 95 percent interval " + f"{100.0 * (self.low or 0.0):.1f} to {100.0 * (self.high or 0.0):.1f})" + ) + + +def success_rate(evaluation: Evaluation) -> RateReport: + """A rate, or a refusal, with the raw counts either way. + + Four things are refused before any arithmetic happens, and the order is from cheapest to + most expensive to fix. No trials at all. A rate computed on the conditions the + demonstrations came from, which is not a success rate: behavior cloning is trained on the + expert's state distribution and its own errors carry it off that distribution, with + suboptimality compounding as the square of the horizon rather than linearly, and the gap + only shows up from starts nobody demonstrated -- measured at 11 of 12 on training conditions against 11 of 24 at a new camera position (95 percent intervals 61.5 to 99.8 and 25.6 to 67.2). + A criterion that reads the policy's own + state. And too few trials to say anything, which is a refusal of the POINT ESTIMATE + specifically: the counts and the interval are the honest report, and the percentage is the + thing that gets quoted without them. + """ + ev = evaluation + if ev.trials <= 0: + return RateReport( + ev, + False, + None, + None, + None, + "no rollout happened. A policy nobody ran has no rate, and the absence is not a zero or " + "a pending value -- it is the untrusted default this package applies to every unmeasured " + "thing", + ) + + if not ev.held_out: + return RateReport( + ev, + False, + None, + None, + None, + f"{ev.successes} of {ev.trials} were scored on the conditions the demonstrations were " + "collected from, which is not a success rate. The policy's own errors are what carry it " + "off the demonstrated state distribution, and that effect is invisible from a start the " + "expert also started from. Measured elsewhere at 11 of 12 on training conditions against 11 of 24 at a new camera position (95 percent intervals 61.5 to 99.8 and 25.6 to 67.2). The reported " + "table-texture and lighting drops are not distinguishable from baseline at that trial " + "count and are not quoted here", + ) + + if not ev.criterion.external: + return RateReport( + ev, + False, + None, + None, + None, + f"success was scored by {ev.criterion.value}, which reads the policy's own state. A " + "policy reporting high confidence on a transfer that failed is the silent destructive " + "failure `coverage` is built around: a number arrives, the check passes, and the " + "material is already gone. Offline loss is not the safer version of this -- validation " + "error against real-world policy ranking measured a Pearson r of 0.308", + ) + + point = ev.successes / ev.trials + low, high = success_interval(ev.successes, ev.trials) + + if ev.trials < MIN_HELDOUT_TRIALS: + weight = 100.0 / ev.trials + return RateReport( + ev, + False, + None, + None, + None, + f"{ev.successes} of {ev.trials} held-out trials: one trial is {weight:.1f} percentage " + f"points, above the declared ceiling of {MAX_TRIAL_WEIGHT_PP:.1f}, and the exact 95 " + f"percent interval spans {100.0 * low:.1f} to {100.0 * high:.1f} percent. A point " + "estimate off this many trials is a number that will be quoted without its interval. The " + "counts are reported instead, which is what lets a reader reconstruct the interval " + "themselves. No published result establishes a minimum trial count -- the floor here is " + "a stated convention, not a measured threshold", + ) + + return RateReport( + ev, + True, + point, + low, + high, + f"{ev.successes} of {ev.trials} held-out trials scored by {ev.criterion.value}" + + ("" if ev.blinded else ", unblinded") + + f". The interval is {100.0 * (high - low):.1f} percentage points wide; narrowing it is " + "roughly fifteen times the rollouts for a fivefold narrowing, so this width is a budget " + "decision that was already made when the trial count was chosen", + ) + + +class Trust(str, Enum): + """Whether a trained policy may be handed material, and if not, what is missing. + + Ordered by how far the evidence gets. Only MEASURED is trusted, and it requires BOTH a + held-out trial count at or above the declared floor AND a criterion that does not read the + policy's own state. Neither condition substitutes for the other, which is why they are two + conditions: a thousand trials scored by the policy's own confidence measure the policy's + confidence, and one external measurement on one rollout is an anecdote with an instrument + attached. + + MEASURED does not mean good. It means a number exists with an interval around it and an + outside criterion behind it, and the number can perfectly well be 20 percent -- which is + what the hardest task in the reference set actually scored, on twice the demonstrations of + any other. + """ + + UNEVALUATED = "unevaluated" # no rollouts at all + TRAINING_SET_ONLY = "training_set_only" # rolled out from the conditions it was trained on + SELF_SCORED = "self_scored" # the criterion reads the policy's own state + UNDERPOWERED = "underpowered" # held out and externally scored, too few trials to state a rate + MEASURED = "measured" # held out, external criterion, at or above the declared floor + + @property + def evidence_complete(self) -> bool: + """The evaluation was done properly. NOT a statement that the policy works. + + Named for what it is, because the previous name was `trusted` and returned True for a + policy that succeeded in zero of twenty held-out trials. Every word in this enum's + docstring said MEASURED does not mean good, and the boolean said the opposite to anyone + who read it instead of the prose. A property that contradicts its own docstring is worse + than an absent one: it is an authorization-shaped answer to a question nobody asked. + """ + return self is Trust.MEASURED + + +@dataclass(frozen=True) +class TrustAssessment: + """One policy's standing, with the rate report that produced it.""" + + policy: "Policy" + verdict: Trust + rate: RateReport + reason: str + # The rate this task was declared to tolerate, and whether the interval's lower bound + # cleared it. Both are None/False unless a caller supplied a threshold, because whether a + # rate is good enough is a property of the task rather than of the evaluation. + acceptable_rate: Optional[float] = None + clears: bool = False + + @property + def evidence_complete(self) -> bool: + """The evaluation was run properly. Says nothing about whether the policy works.""" + return self.verdict.evidence_complete + + @property + def trusted(self) -> bool: + """Properly evaluated AND measured against a declared acceptance rate that it cleared. + + Both halves are required. The first without the second returns False, which is the whole + fix: this used to be `verdict.trusted`, and a policy that succeeded in zero of twenty + held-out trials came back True because the evaluation had been conducted correctly. + """ + return self.evidence_complete and self.clears + + +def trusted_with_material( + policy: "Policy", acceptable_rate: Optional[float] = None +) -> TrustAssessment: + """Has this policy earned an unattended run on real material? + + A missing evaluation is not a pass, in exactly the way `intelligence.trusted_for` refuses an + operation nobody benchmarked and `qc.evaluate` refuses a gate whose measurement never + arrived. The default is untrusted and the burden is on the measurement. + + `acceptable_rate` is required for a trusted answer and has no default, because what rate is + acceptable is a property of the TASK and not of this module. A tolerable failure rate for a + tip pickup that a retry fixes is not a tolerable failure rate for a transfer that consumes + the last of a sample. Defaulting it would be this package answering, on the caller's behalf + and invisibly, the only question in the function that requires knowing what the robot is + holding. + + It is compared against the LOWER bound of the interval rather than the point estimate. A + point estimate is what the policy did on those trials; the lower bound is what the trials + support saying about the next one, and the gap between them is the entire reason the trial + count is in the report. + """ + ev = policy.evaluation + if ev is None or ev.trials <= 0: + report = success_rate(ev if ev is not None else Evaluation()) + return TrustAssessment( + policy, + Trust.UNEVALUATED, + report, + f"'{policy.name or 'this policy'}' has never been rolled out. Training converged is not " + "a result; it is a statement about a loss curve, and offline loss against real-world " + "ranking measured a Pearson r of 0.308", + ) + + report = success_rate(ev) + if not ev.held_out: + return TrustAssessment(policy, Trust.TRAINING_SET_ONLY, report, report.reason) + if not ev.criterion.external: + return TrustAssessment(policy, Trust.SELF_SCORED, report, report.reason) + if ev.trials < MIN_HELDOUT_TRIALS: + return TrustAssessment(policy, Trust.UNDERPOWERED, report, report.reason) + if acceptable_rate is None: + return TrustAssessment( + policy, + Trust.MEASURED, + report, + f"{report.describe()}. The evaluation is complete and nothing here says whether that " + "is good enough, because no acceptable rate was declared. State the rate this task " + "tolerates and ask again", + acceptable_rate=None, + clears=False, + ) + if report.low is None or report.low < acceptable_rate: + shown = "no interval" if report.low is None else f"{100.0 * report.low:.1f} percent" + return TrustAssessment( + policy, + Trust.MEASURED, + report, + f"{report.describe()}. The interval's lower bound is {shown}, below the " + f"{100.0 * acceptable_rate:.1f} percent this task was declared to tolerate. The trials " + "do not support running it unattended on material", + acceptable_rate=acceptable_rate, + clears=False, + ) + return TrustAssessment( + policy, + Trust.MEASURED, + report, + f"{report.describe()}. This is a measured rate on held-out trials against an outside " + "criterion, which is the most this module can say. It is a statement about this task, this " + "fixture revision and this bench, and it does not transfer to another of any of them", + acceptable_rate=acceptable_rate, + clears=True, + ) + + +# -- the policy ----------------------------------------------------------------------- + + +@dataclass(frozen=True) +class Policy: + """A trained policy, its captures, and its evaluation if it has one. + + `method` is carried because the demonstration requirement is method-dependent and not only + task-dependent: at an identical 50 demonstrations on one task, one architecture reached 95 + percent where another reached 65, and the paper's own explanation was that 50 is not enough + for the more expressive model. A record that stores a count without the method has thrown + away half of what the count meant. + + Everything derived is a method rather than a field, so nothing here can disagree with the + captures it was derived from. + """ + + name: str = "" + method: str = "" + captures: Tuple[Capture, ...] = () + evaluation: Optional[Evaluation] = None + note: str = "" + + def episodes(self) -> int: + return sum(c.episodes for c in self.captures) + + def usability(self) -> Tuple[Usability, ...]: + return tuple(usable_for_training(c) for c in self.captures) + + def unusable(self) -> Tuple[Usability, ...]: + return tuple(u for u in self.usability() if not u.usable) + + def trainable(self) -> bool: + """True only where at least one capture exists and every one of them is usable. + + The leading truthiness check is the method. `all()` over an empty tuple is True, so the + one-liner reports a policy with no captures at all as trained on flawless data. + """ + return bool(self.captures) and all(u.usable for u in self.usability()) + + def distinct_couplings(self) -> Tuple[FixtureCoupling, ...]: + seen: List[FixtureCoupling] = [] + keys = set() + for c in self.captures: + if c.coupling is None: + continue + if c.coupling.key() in keys: + continue + keys.add(c.coupling.key()) + seen.append(c.coupling) + return tuple(seen) + + def training_coupling(self) -> Optional[FixtureCoupling]: + """The one fixture revision every capture was collected against, or None. + + None for zero, None for two, and None when ANY capture recorded no fixture at all. + + That last case is the one that was wrong. `distinct_couplings` skips a capture whose + coupling is None, so a policy trained on one fixtured session and two unrecorded ones + reported a single clean coupling and a run against that fixture was permitted. Two thirds + of the training data had been collected against a world nobody wrote down, and the report + said COUPLED. An unrecorded fixture is not an absent one: the part was physically there + and is inside the policy's observations whether or not it was named, so it cannot be + reproduced and cannot be shown to have changed. + """ + if any(c.coupling is None for c in self.captures): + return None + couplings = self.distinct_couplings() + return couplings[0] if len(couplings) == 1 else None + + +@dataclass(frozen=True) +class Interlock: + """A bound enforced outside the policy, by something that cannot learn. + + This is the leg the module was missing, and its absence is why a policy with no + demonstrated capability could return an authorization-shaped permitted=True. A learned + policy carries no guarantee about what it will do on an input it has not seen, so the + bound on what it can physically reach, how fast, and how hard cannot come from the policy + and cannot come from a model checking the policy. It comes from a coded limit, a + controller setting, or a mechanical stop. + + `miss_rate_measured` is the load-bearing field and follows `vision`'s rule for detectors: + a bound whose failure rate nobody measured is not a safety device, it is a belief about + one. Measuring it means deliberately driving the violation and counting what the interlock + actually stops, not observing that nothing bad happened during normal operation. + """ + + workspace_bounded: bool = False + speed_bounded: bool = False + force_bounded: bool = False + enforced_by: str = "" # what enforces it, outside the policy + miss_rate_measured: bool = False + + @property + def complete(self) -> bool: + return ( + self.workspace_bounded + and self.speed_bounded + and self.force_bounded + and bool(self.enforced_by.strip()) + and self.miss_rate_measured + ) + + def missing(self) -> Tuple[str, ...]: + out: List[str] = [] + if not self.workspace_bounded: + out.append("no workspace bound") + if not self.speed_bounded: + out.append("no speed bound") + if not self.force_bounded: + out.append("no force bound") + if not self.enforced_by.strip(): + out.append("nothing named as enforcing the bounds outside the policy") + if not self.miss_rate_measured: + out.append("the interlock's own miss rate has never been measured against a driven violation") + return tuple(out) + + +@dataclass(frozen=True) +class RunStanding: + """One policy against one proposed run: may it move, and what stops it. + + The composition is the point, in the way `coverage` composes vision and gates. Each of the + three questions is answerable and none of them is the question a lab asks. Trained on + usable data, measured on held-out trials, and still facing the fixture it learned on -- + all three, or the run is not permitted. + """ + + policy: "Policy" + proposed: Optional[FixtureCoupling] + trust: TrustAssessment + coupling: CouplingCheck + captures: Tuple[Usability, ...] + interlock: Interlock = field(default_factory=Interlock) + + @property + def permitted(self) -> bool: + return ( + bool(self.captures) + and all(u.usable for u in self.captures) + and self.trust.trusted + and self.coupling.holds + and self.interlock.complete + ) + + def blockers(self) -> Tuple[str, ...]: + out: List[str] = [] + if not self.captures: + out.append("this policy declares no captures at all, so nothing trained it") + for u in self.captures: + if not u.usable: + out.append( + f"capture '{u.capture.name or 'unnamed'}' is not usable for training: {u.reason}" + ) + if not self.trust.trusted: + out.append(f"trust is {self.trust.verdict.value}: {self.trust.reason}") + if not self.coupling.holds: + out.append(f"fixture coupling is {self.coupling.state.value}: {self.coupling.reason}") + if not self.interlock.complete: + out.append( + "no complete interlock stands outside this policy: " + + "; ".join(self.interlock.missing()) + ) + return tuple(out) + + +def may_run( + policy: Policy, + proposed: Optional[FixtureCoupling] = None, + acceptable_rate: Optional[float] = None, + interlock: Optional[Interlock] = None, +) -> RunStanding: + """Compose the three questions into the one a lab actually asks. + + Nothing is re-derived here. Every arm is a report computed by the function that owns it, so + this cannot disagree with any of them -- the same construction `intelligence.loop_closure` + uses, and for the same reason. + """ + return RunStanding( + policy=policy, + proposed=proposed, + trust=trusted_with_material(policy, acceptable_rate), + coupling=fixture_coupled(policy, proposed), + captures=policy.usability(), + interlock=interlock if interlock is not None else Interlock(), + ) + + +# -- how many demonstrations ------------------------------------------------------------ + + +@dataclass(frozen=True) +class PublishedCount: + """One published per-task demonstration count, with the method, task and trial count named. + + `trials` is beside `success` deliberately. Every success rate in this table came off 5 to 25 + rollouts, so the third digit of any of them is decoration: at 25 evaluations one trial is 4 + percentage points, at 20 it is 5, and at 5 it is 20. None of the source papers reports a + confidence interval on a real-robot success rate. + """ + + system: str + task: str + demonstrations: Optional[int] + trials: Optional[int] + success: Optional[float] + source: str + note: str = "" + + +# The published real-robot per-task counts, with their tasks and methods named. The point of +# the table is its spread rather than any row in it. +PUBLISHED_COUNTS: Tuple[PublishedCount, ...] = ( + PublishedCount( + "ACT on a bimanual low-cost platform", "Slide Ziploc", 50, 25, 0.88, "arXiv:2304.13705" + ), + PublishedCount( + "ACT on a bimanual low-cost platform", "Slot Battery", 50, 25, 0.96, "arXiv:2304.13705" + ), + PublishedCount( + "ACT on a bimanual low-cost platform", "Open Cup", 50, 25, 0.84, "arXiv:2304.13705" + ), + PublishedCount( + "ACT on a bimanual low-cost platform", + "Thread Velcro", + 100, + 25, + 0.20, + "arXiv:2304.13705", + note=( + "the counter-evidence inside the reference paper itself: twice the demonstrations of " + "every other task and the worst result of any. The authors attribute it to perception " + "and precision rather than data volume, with success roughly halving at each stage from " + "92 percent at the first to 20 final" + ), + ), + PublishedCount( + "ACT on a bimanual low-cost platform", "Prep Tape", 50, 25, 0.64, "arXiv:2304.13705" + ), + PublishedCount( + "ACT on a bimanual low-cost platform", "Put On Shoe", 50, 25, 0.92, "arXiv:2304.13705" + ), + PublishedCount( + "Diffusion Policy, real robot", + "Push-T", + 136, + None, + 0.95, + "arXiv:2303.04137", + note="proficient-human demonstrations; baselines on the same task scored 0 and 20 percent", + ), + PublishedCount("Diffusion Policy, real robot", "Mug Flip", 250, 20, 0.90, "arXiv:2303.04137"), + PublishedCount( + "Diffusion Policy, real robot", + "Sauce Pour", + None, + None, + 0.79, + "arXiv:2303.04137", + note=( + "the count is unresolved and reported as unresolved: the table lists 90 " + "proficient-human demonstrations and the appendix says 50 were collected with 90 percent " + "used for training. The paper contradicts itself in both versions" + ), + ), + PublishedCount( + "Diffusion Policy, bimanual", "Egg Beater", 210, 20, 0.55, "arXiv:2303.04137" + ), + PublishedCount( + "Diffusion Policy, bimanual", "Mat Unrolling", 162, 20, 0.75, "arXiv:2303.04137" + ), + PublishedCount( + "Diffusion Policy, bimanual", "Shirt Folding", 284, 20, 0.75, "arXiv:2303.04137" + ), + PublishedCount( + "Mobile ALOHA with co-training", + "Wipe Wine", + 50, + 20, + 0.95, + "arXiv:2401.02117", + note=( + "the same 50 demonstrations WITHOUT co-training on a large auxiliary dataset scored 50 " + "percent. The widely quoted 50-demo figure holds only with the auxiliary data, and 35 " + "demonstrations with co-training beat 50 without, 70 percent against 50" + ), + ), + PublishedCount( + "Mobile ALOHA with co-training", "High Five", 20, 20, 0.85, "arXiv:2401.02117" + ), + PublishedCount( + "Mobile ALOHA with co-training", + "Cook Shrimp", + 20, + 5, + 0.40, + "arXiv:2401.02117", + note="40 percent of 5 trials is 2 successes, which is essentially uninformative", + ), + PublishedCount( + "3D Diffusion Policy, real robot", + "4-task average", + 40, + None, + 0.85, + "arXiv:2403.03954", + note=( + "the frequently repeated 10-demonstration figure from this work is a SIMULATION result " + "across 72 sim tasks. On real hardware it used 40 per task" + ), + ), + PublishedCount( + "ALOHA Unleashed", + "Shirt (easy)", + 8658, + None, + 0.75, + "arXiv:2410.13126", + note=( + "over 26,000 real demonstrations across 5 tasks, collected on 10 robots over eight " + "months, on the same hardware lineage as the 50-demonstration rows above" + ), + ), + PublishedCount("ALOHA Unleashed", "Lace (messy)", 5133, None, 0.40, "arXiv:2410.13126"), + PublishedCount( + "ALOHA Unleashed", + "Gear insert", + 4005, + None, + 0.95, + "arXiv:2410.13126", + note=( + "95 percent for one gear, 75 for two, 40 for three. Demonstration count buys task " + "difficulty and robustness rather than a march to 100 percent" + ), + ), + PublishedCount( + "Data-scaling study, real robot", + "Fold Towels / Unplug Charger", + 1600, + None, + 0.90, + "arXiv:2410.18647", + note=( + "roughly 50 demonstrations per environment-object pair across 32 pairs. The measured " + "finding is that diversity dominates: with environments and objects held fixed there was " + "no clear relationship between demonstration count and generalization" + ), + ), + PublishedCount( + "A frontier vision-language-action model", + "per task", + None, + 10, + None, + "arXiv:2410.24164", + note=( + "reports HOURS rather than demonstration counts -- 5 hours for the simplest task and 100 " + "or more for the most complex -- and scores a normalized score over 10 episodes rather " + "than binary success. The unit of account changed, so these do not compare with the rows " + "above at all" + ), + ), +) + + +# What the count actually depends on, each with the measurement behind it. Ordered by how +# large the measured effect was, not by how often the factor is mentioned. +DETERMINANTS: Tuple[str, ...] = ( + "environment and object DIVERSITY, which the only controlled real-robot study of the " + "question found dominates raw count: performance plateaued at 400, 800 and 1,600 total " + "demonstrations for 8, 16 and 32 environment-object pairs, and with those held fixed there " + "was no clear relationship between count and generalization at all", + "whether auxiliary or pre-training data is used. Co-training moved one task from roughly 0 " + "to 95 percent at the same demonstration count, and 35 demonstrations with it beat 50 " + "without", + "task precision and horizon. The task that received twice the demonstrations of any other in " + "the reference set produced the worst result of any, at 20 percent, and the failure was " + "attributed to perception and precision rather than to data volume", + "the method. At an identical 50 demonstrations on one task, one architecture reached 95 " + "percent and another reached 65, with the paper's own explanation being that 50 is not " + "enough for the more expressive model", + "how much pose variance the setup leaves in, which is where a fixture enters. At a fixed " + "1,000-demonstration budget, rigidly fixing the receptacle moved square-peg assembly from " + "49.3 to 90.7 percent -- and moved an easy stacking task by seven tenths of one point", +) + + +@dataclass(frozen=True) +class DemonstrationEstimate: + """The answer to "how many demonstrations", which is not a number. + + `estimate` is always None and the field exists so a caller cannot read the absence as an + oversight. The refusal is carried as data for the same reason `printed.NOT_ESTIMATED` is: + a module that declines a number in its docstring and exposes a helper returning one has + documented a discipline it does not have. + """ + + task: str + estimate: Optional[int] + published: Tuple[PublishedCount, ...] + determinants: Tuple[str, ...] + refusal: Refusal + + def span(self) -> Tuple[int, int]: + counts = [p.demonstrations for p in self.published if p.demonstrations is not None] + return (min(counts), max(counts)) + + def describe(self) -> str: + low, high = self.span() + return ( + f"no count for '{self.task}'. The published real-robot per-task counts located here span " + f"{low} to {high} demonstrations, a factor of {high / low:.0f}, and both ends are real " + f"results on working hardware. {self.refusal.why}" + ) + + +def demonstrations_needed(task: str) -> DemonstrationEstimate: + """How many demonstrations a new task needs, which nobody can state in advance. + + Returns the published range with its tasks and methods named, the determinants, and an + explicit refusal -- never a number. The refusal is not caution. It is the shape of the + evidence: the counts that exist run from 20 demonstrations at 85 percent to 8,658 at 75 + percent, the one controlled study found diversity rather than count carries the gains, the + frontier work changed the unit to hours and stopped reporting counts at all, and the single + work in the set with an internal demonstration-count comparison found the task with twice + the data scoring worst. A number produced from that would be a convention with a citation + stapled to it. + + What a lab can do instead is the thing the evidence supports: collect, measure on held-out + trials, and watch the interval. That is `success_rate`, and it is the reason it exists in + this module rather than in a notebook. + """ + return DemonstrationEstimate( + task=task, + estimate=None, + published=PUBLISHED_COUNTS, + determinants=DETERMINANTS, + refusal=Refusal( + quantity=f"demonstrations required for '{task}'", + why=( + "the count for a NEW task is not knowable in advance from any published result. The " + "real-robot per-task counts span more than two orders of magnitude on comparable " + "hardware, the widely repeated 50-demonstration figure is a data-collection budget " + "from one paper that ran no ablation over demonstration count, and the only controlled " + "study measured environment and object diversity -- not count -- as the thing that " + "carries generalization. What determines it is measurable only on the task itself" + ), + what_it_would_take=( + "collecting for this task on this bench and measuring success on held-out trials as " + "the count grows, which is the demonstration-count ablation that exactly two of the " + "works surveyed here ran on real hardware. The measurement is cheap in equipment and " + "expensive in rollouts, and there is no way to buy it with a citation" + ), + ), + ) + + +def published_span() -> Tuple[int, int]: + """The measured spread of published per-task counts, computed from the table. + + Computed rather than asserted so the sentence in the docstring cannot drift from the data + under it. + """ + counts = [p.demonstrations for p in PUBLISHED_COUNTS if p.demonstrations is not None] + return (min(counts), max(counts)) + + +# -- authority, in prose ------------------------------------------------------------------ + + +INTERLOCK_NOTE = ( + "A learned policy PROPOSES a motion and a deterministic interlock DISPOSES. Workspace, " + "speed and force are bounded by something that does not read the policy's output, has no " + "confidence to be wrong about, and cannot be talked out of a limit by a fluent one -- which " + "is the same separation this package draws everywhere between a thing that suggests and a " + "thing that permits. The concept is referenced here and deliberately not imported, so " + "nothing in this module depends on a layer that may not have landed. Two honest cautions " + "about what such an interlock is usually built from. First, the arm-side features that get " + "pointed at -- a reduced mode limiting Cartesian and joint speed and joint range, a " + "Cartesian safety boundary that stops motion when the tool centre point leaves it, and " + "collision detection on dimensionless sensitivity levels -- are vendor-documented " + "conveniences: the common implementation compares modelled joint current against measured " + "current, so a wrong payload mass, a wrong mounting direction or a stale friction parameter " + "degrades it, and one manual carries no power-and-force-limiting validation and states in " + "as many words that the emergency stop should not be used as a risk reduction measure. " + "Second, an interlock that reads the policy's own confidence is not an interlock. It is the " + "policy again, wearing a second name" +) + + +# -- what this module will not estimate ---------------------------------------------------- + + +NOT_ESTIMATED: Tuple[Refusal, ...] = ( + Refusal( + quantity="demonstrations required for a new task", + why=( + "the published real-robot per-task counts span 20 to 8,658 demonstrations on comparable " + "hardware, the canonical 50-demonstration figure is a collection budget from a paper " + "with no demonstration-count ablation in it, and the only controlled study of the " + "question found environment and object diversity rather than count carrying the gains" + ), + what_it_would_take=( + "the ablation on this task and this bench: collect, evaluate on held-out trials, and " + "watch the interval rather than the point estimate" + ), + ), + Refusal( + quantity="the exchange rate between a human video and a teleoperated episode", + why=( + "one head-to-head comparison exists and it measured roughly 3 to 1 on a quasi-static " + "sweeping task with pinch grasps, inside a pipeline using depth and a hand mesh. Its own " + "authors bound it to that setting: no dexterity, no contact-rich work, and only where " + "the robot can follow the same strategy the human did. A ratio quoted outside those " + "bounds is a number with the caveats removed" + ), + what_it_would_take=( + "the same head-to-head on this task and this arm, both arms of it collected and both " + "evaluated on the same held-out conditions" + ), + ), + Refusal( + quantity="a success rate from fewer than the declared trial floor", + why=( + "the point estimate is what gets quoted and the interval is what carries the " + "information. At 10 trials a 90 percent observation carries a 95 percent interval from " + "55.5 to 99.7 percent, which cannot distinguish a policy that works from one that fails " + "one attempt in three" + ), + what_it_would_take=( + "more rollouts, and the arithmetic is unforgiving: narrowing a 10-point interval to 2 " + "points takes roughly fifteen times the trials, and separating 85 percent from 90 at " + "conventional power takes 686 per policy" + ), + ), + Refusal( + quantity="the probability that a policy succeeds on the next attempt", + why=( + "a rate measured on held-out trials describes the distribution those trials were drawn " + "from. The next attempt is on this bench, at this hour, against this fixture revision, " + "with whatever the last run left on the deck, and none of those were sampled. This is " + "the same refusal `printed` makes about the probability that an unqualified part is fine" + ), + what_it_would_take=( + "trials drawn from the conditions the run will actually face, which for a lab means the " + "evaluation is a rehearsal of the run rather than a benchmark of the policy" + ), + ), + Refusal( + quantity="retargeting error in millimetres for a human-to-robot mapping", + why=( + "no verified figure was located here. The work that defines the right metrics -- " + "fingertip position error in the global frame, relative to the wrist, relative to the " + "thumb, and fingertip orientation error -- reports them only as unlabelled bar charts. " + "The magnitudes that ARE published belong to the hand-pose stage upstream of " + "retargeting, at 185.67 mm mean per-joint error for a monocular estimator and 3D errors " + "as large as 20 to 30 cm" + ), + what_it_would_take=( + "measuring the retargeted end-effector pose against a tracked ground truth on this " + "capture rig, which is `Retargeting.error_measured_mm` and is almost never filled in" + ), + ), + Refusal( + quantity="how many demonstrations a printed fixture saves", + why=( + "no study located here manipulates a PHYSICAL fixture with an unfixtured control and " + "counts human demonstrations. Every quantified source substitutes something for the " + "fixture: a randomization range in simulation, a reset distribution at a fixed budget, " + "or software pose canonicalization. The multipliers in circulation -- 30-fold, 38-fold " + "-- come from those substitutes and would be an extrapolation onto a lab bench" + ), + what_it_would_take=( + "the experiment nobody has run: the same task, the same arm, demonstrations collected " + "with and without the nest, both evaluated on held-out trials to the same success target" + ), + ), +) + + +def refusals() -> Tuple[Refusal, ...]: + """The estimates this module declines to make, with what each would require. + + Returned rather than only documented so the refusal is checkable, following `printed` and + `durability`. + """ + return NOT_ESTIMATED + + +def summary() -> Dict[str, int]: + """Counts for the report, computed from this module's own tables. + + `published_task_rows_with_a_count` is smaller than `published_task_rows` on purpose: two of + the rows report hours instead of demonstrations and one reports a count its own paper + contradicts, and a summary that quietly dropped them would report a cleaner literature than + the one that exists. + """ + with_count = [p for p in PUBLISHED_COUNTS if p.demonstrations is not None] + low, high = published_span() + return { + "modalities": len(Modality), + "modalities_with_a_measured_action_stream": sum( + 1 for m in Modality if m.measured_action_stream + ), + "modalities_on_the_robots_own_kinematics": sum( + 1 for m in Modality if m.on_the_robots_own_kinematics + ), + "published_task_rows": len(PUBLISHED_COUNTS), + "published_task_rows_with_a_count": len(with_count), + "published_lowest_count": low, + "published_highest_count": high, + "min_heldout_trials": MIN_HELDOUT_TRIALS, + "refusals": len(NOT_ESTIMATED), + } + + +__all__ = [ + "CONFIDENCE", + "Capture", + "Coupling", + "CouplingCheck", + "DETERMINANTS", + "DemonstrationEstimate", + "Evaluation", + "FixtureCoupling", + "Interlock", + "INTERLOCK_NOTE", + "MAX_TRIAL_WEIGHT_PP", + "MIN_HELDOUT_TRIALS", + "MODALITY_EVIDENCE", + "Missing", + "Modality", + "NOT_ESTIMATED", + "PUBLISHED_COUNTS", + "Policy", + "PublishedCount", + "RateReport", + "Retargeting", + "RunStanding", + "SuccessCriterion", + "Trainable", + "Trust", + "TrustAssessment", + "Usability", + "demonstrations_needed", + "equivalent_to_teleop", + "fixture_coupled", + "may_run", + "published_span", + "refusals", + "success_interval", + "success_rate", + "summary", + "trusted_with_material", + "usable_for_training", +] diff --git a/docs/CAPTURE_TO_POLICY.md b/docs/CAPTURE_TO_POLICY.md new file mode 100644 index 0000000..23b250c --- /dev/null +++ b/docs/CAPTURE_TO_POLICY.md @@ -0,0 +1,1389 @@ +# Capture to policy + +How to get from a camera pointed at a bench to an arm that performs a lab manipulation +task, and the one decision at the front that determines whether the rest is engineering or +research. + +This guide assumes a bench scientist with a task worth automating, an arm on the bench or +on a quote, a camera, and the printer from `docs/PRINTED_FIXTURES.md`. It is opinionated, +because the failure mode here is not a wrong parameter -- it is spending four months +collecting the wrong kind of data. + +Everything below carries where it came from. `[EXP]` means measured in a published +experiment on real hardware; `[EXP-1]` means the same, from a single fetch that was not +independently re-verified -- one confidence step lower; `[SIM]` means measured in +simulation only, which is a different claim and gets confused with the first one constantly; +`[DS]` means read directly out of a manufacturer document; `[VENDOR]` means a vendor or +third-party claim nobody independent has checked; `[ARITH]` means this guide computed it +from stated inputs and you can redo it; `[PRACTICE]` means engineering convention with no +measurement behind it. Numbers with no mark do not appear. Where a number would be useful +and does not exist, this guide says what measurement would produce it rather than filling +the hole. + +Two refusals are in force throughout, and they are the reason several obvious sentences are +missing. **This guide states no success rate for any task you might attempt**, because none +has been published for a plate on a deck and extrapolating one from a paper about folding +shirts would be a fabricated number wearing a citation. And **it presents no demonstration +count as a requirement**, because the only controlled study of that question found the +count is not the variable that matters. + +--- + +## 1. The capture decision + +This is the first decision and the largest one. Everything in 2 through 8 is downstream of +it, and it is the one that gets made by accident -- usually by whoever already owns a +camera. + +**Recommendation: capture with actions. Teleoperate the arm, or hand-guide it, and record +what the controller commanded. Do not plan to recover actions from third-person video of a +person.** + +### What the two options actually are + +A policy is a function from observation to action. Training one by imitation needs pairs. +The question is where the action column comes from. + +*Capture with actions.* A person drives the arm -- a leader arm, a gamepad, a handheld +gripper, or a hand on the link with the brakes released -- and the system logs the +commanded action in the robot's own action space at the controller's rate, alongside the +images. The action column is recorded, not inferred. Every ambiguity about what the person +"meant" was resolved by the arm actually moving. + +*Third-person video.* A camera watches a person do the task. The recording contains pixels. +It contains no actions at all, in the literal sense: there is no field in it that says what +joint command or end-effector delta produced the next frame. Getting one requires +estimating hand pose, mapping a human hand onto a robot gripper, and assuming the resulting +trajectory is something the arm can execute. + +**That second step is not a conversion. It is an additional inference with its own error +and no validation on your task.** The distinction matters because "convert the video to +actions" is how the pipeline gets described, and a conversion sounds like a unit change. + +### Why retargeting is a research step, not a preprocessing step + +Three independent lines of evidence, of decreasing abstraction. + +**It is not always possible, in principle.** The gap between learning from demonstrations +(state-action pairs) and learning from observation (states only) is exactly the disagreement +between the imitator's inverse dynamics model and the expert's, upper-bounded by a negative +causal entropy term `[EXP: NeurIPS 2019, arXiv:1910.04417]`. And an inverse dynamics model +is not well-defined at all unless the transition is injective -- unless no two actions can +produce the same next state from the same state `[arXiv:2102.10769]`. Where the dynamics are +non-injective the action is unrecoverable from video as a matter of mathematics, not of +model quality. + +Read that against a lab bench, because the non-injective cases are the ones you care about: + +- A redundant arm. A 7-DOF arm reaches one end-effector pose through a continuum of joint + configurations. The video shows the pose; it does not determine the configuration. +- Force applied against a rigid constraint. Pressing a plate down into a nest at 5 N and at + 25 N produce the same next image. The plate is seated in both. +- Anything where the visible state does not change. Holding a seal, maintaining a grip, + waiting for a magnet. The action is nonzero and the frames are identical. + +Seating labware, retaining a plate, and holding a lid are the operations a lab fixture +exists to make repeatable, and they are precisely the ones video cannot label. + +**The measured exchange rate.** Phantom (CoRL 2025, arXiv:2503.00779) trains policies from +human RGBD video with zero robot demonstrations, by extracting end-effector pose from hand +mesh estimation plus segmentation plus ICP against the depth cloud, then inpainting a +rendered robot over the human arm. On a sweeping task it reports a direct head-to-head +`[EXP]`: + +| capture | episodes | reported success | +| --- | --- | --- | +| teleoperated | 50 | 52% | +| human video | 50 | 44% | +| teleoperated | 100 | 88% | +| human video | 300 | 84% | + +Roughly 3x the human episodes to *approach*, not beat, the teleop number. This guide could +not confirm the trial count behind those four rows specifically; the paper's other per-task +rates are over 25 rollouts each, where one trial moves a rate by 4 percentage points +`[ARITH]`. Treat the ordering as the finding and the digits as coarse. + +Note also what the numbers do to the cost argument, which is the usual reason people choose +video. Handheld-gripper capture in UMI (RSS 2024, arXiv:2402.10329) ran *faster* than +bare-hand human collection, which came in at 48% and 64% of the gripper's rate on two tasks +`[EXP]`. Cheap-to-collect is not the same as cheap-per-unit-of-policy. + +**Hand pose quality is the binding constraint, and it is where "just use internet video" +fails.** A 2026 cotraining study (arXiv:2606.06627) injected calibrated Gaussian noise into +its triangulated hand labels and watched mean success fall from 48.3 +/- 19.6% to 30.0 +/- +13.1% at 0.5 sigma to 20.0 +/- 6.5% at 1.0 sigma. Substituting a real monocular hand +estimator for multi-view triangulation produced 185.67 mm mean per-joint position error and +dropped mean success from 41.5% to 24.7% `[EXP]`. Independently, Human2Sim2Robot (CoRL 2025, +arXiv:2504.12609) reports hand pose errors from human video "as large as 20-30 cm" `[EXP]` +and responds by discarding the human's actions entirely -- using only the object's 6D pose +trajectory as a reward and a single pre-manipulation hand pose to seed reinforcement +learning. + +The multi-view rig, the depth camera, or the glasses are doing the work in every one of +these results. Monocular third-person video is the degraded case, and it is the case a lab +would actually have. + +**Even a perfect estimate does not give you an executable trajectory.** Human2Sim2Robot +states that even with perfect pose estimates, direct retargeting "often results in +suboptimal robot trajectories due to morphological differences." DexMachina (arXiv:2505.24853) +puts it physically: across six dexterous hands, "kinematic retargeting can produce +human-like hand motions, but when we play the retargeting results in simulation, they are +not feasible for completing the task" `[SIM]`. A whole method (SPIDER, arXiv:2511.09484) +exists to repair exactly this gap. + +**The one signal video cannot contain, isolated and measured.** UMI puts the actual robot +gripper in the human's hand so observation and action spaces match by construction. +Ablating relative inter-gripper proprioception dropped bimanual cloth folding from 70% to +30% success `[EXP]`. Proprioception is not in any video of anything. + +### The honest counterweight + +Human video is not useless, and the recommendation is not "never touch it." What the +measurements support is that it helps as *pretraining or cotraining*, and that its help +shrinks as your own action-labeled data grows. + +- The 2026 cotraining study measured the gain against robot-only training as +29.7 points + at 3 training environments, +21.1 at 5, and +12.5 at 10 -- where the last one is roughly + one standard deviation of its own error bars `[EXP]`. The widely quoted "+29.7%" is the + low-robot-data regime only. +- LAPA (ICLR 2025, arXiv:2410.11758) learns latent actions from video and pretrains on + them, then *finetunes on action-labeled robot data*. Its own stated limitation is that it + "underperforms compared to action pretraining when it comes to fine-grained motion + generation tasks like grasping" `[EXP]` -- the residual deficit lands exactly on fine + motor control, which is what a tip in a well is. +- Latent actions are not reliably identifiable from video alone: with action-correlated + distractors present, supplying ground-truth actions for as few as 2.5% of the dataset + improved downstream performance by 4.2x on average `[SIM: ICML 2025, arXiv:2502.00379]`. + +And the summary that should settle the decision: **every published result this guide could +verify, in which a video-trained policy actually executes a real task, injects action +information somewhere.** Phantom derives explicit end-effector pose and gripper width from +RGBD plus a hand mesh. LAPA finetunes on action-labeled robot data. UMI and its successors +put real gripper or exoskeleton hardware on the human hand. Human2Sim2Robot discards the +human's actions and runs reinforcement learning in simulation. The cotraining study never +runs without robot demonstrations at all. A targeted search found no published case of a +useful manipulation policy trained with zero action information anywhere in the pipeline. +That is an absence of evidence rather than a proof of impossibility, and it is the state of +the field you would be betting against. + +### What the recommendation buys, stated precisely + +Choosing action-labeled capture does not make the task work. It moves the unsolved part of +the project from "recover actions from video" -- which is an open research problem with a +185 mm error bar on its input -- to "collect enough of the right data and evaluate it +honestly," which is 5 and 6 and is ordinary work. + +Two caveats so the recommendation is not oversold. + +*"Teleoperation has zero embodiment gap" is marketing.* That phrasing appears on +data-vendor pages, not in a peer-reviewed result `[VENDOR]`. It is true only in the narrow +sense that the recorded actions were executed on that robot by that controller. Teleop data +is still off-policy with respect to the states a learned policy visits on its own, which is +the classic behavior-cloning covariate shift: suboptimality compounds as O(eps * T^2) in +horizon for behavior cloning against O(eps * T) for an interactive reduction +`[AISTATS 2011, DAgger -- widely restated; this guide did not read the original theorem +statement]`. That is why 6 insists on autonomous rollouts from held-out starts. + +*Hand-guiding puts a human in every frame.* If you capture kinesthetically, the images +contain an arm and two hands that will not be there at inference. Phantom's pipeline +inpaints a rendered robot over the human arm specifically to remove them `[EXP]`, which +tells you how the field treats a human limb in the observation. Whether this costs you +anything on a given setup is unmeasured here `[PRACTICE]`; the cheap mitigations are to +place the cameras so the operator's arm is out of frame, or to use a leader arm or gamepad +so the operator is not in the workspace at all. + +--- + +## 2. What a usable capture contains + +A "recording" that is missing any one of the following is not a demonstration, and the time +to discover that is before the first one, not after two hundred. + +### The streams, and what each is for + +| stream | what it is | why it is not optional | +| --- | --- | --- | +| action | what the controller was commanded, in the robot's own action space | the label. Without it there is no supervised pair (1) | +| proprioception | achieved joint positions, gripper width or force | commanded and achieved differ; the policy needs to know where the arm actually is. Ablating it cost 40 points on one bimanual task `[EXP: UMI]` | +| wrist camera | view from the end effector | moves with the gripper, so its relationship to the grasp is fixed by construction | +| scene camera | fixed view of the workspace | provides the context the wrist view loses at contact. Its pose is the fragile thing -- see below | +| timestamps | one clock, per frame and per action sample | the failure that silently destroys everything else | +| episode boundaries | start, end, and a success flag scored externally | an unsegmented log is not a dataset, and self-scored success is not success (6) | +| force, if it matters | six-axis wrench at the flange | the only stream that distinguishes "seated" from "crushed", and it is absent by default | + +Note what is *not* on that list: a calibration matrix consumed by the policy. Standard +behavior-cloning visuomotor policies do not read one. That has consequences, below. + +### Frame rate + +The published real-robot datasets that people train on were collected at controller rates, +not video rates. ACT records at 50 Hz control, with episodes of 8-14 seconds giving 400-700 +timesteps per demonstration `[EXP: RSS 2023, arXiv:2304.13705]`. The pi-0 pretraining +corpus spans control rates up to 50 Hz `[EXP: arXiv:2410.24164]`. + +The rule that follows is not a number, it is a constraint: **the action stream's rate is set +by your controller, and the camera stream has to be timestamped against it rather than +assumed to line up.** Record the rate you used. The timestep count is the unit every +downstream quantity is expressed in, and a dataset that does not state its rate cannot be +compared to anything. + +### Calibration, before the first demonstration + +Two separate things get called calibration and only one of them is a matrix. + +**Hand-eye calibration.** The classical formulation is AX=XB, where A is end-effector +motion, B is the corresponding camera motion, and X is the unknown rigid transform between +camera and hand (or camera and base). Three things about it are worth knowing before you +collect poses: + +- **Observability is a hard requirement, not a quality knob.** A unique solution needs at + least two motions whose rotation axes are non-parallel. If all rotation axes are parallel, + the translation of X along the common axis is unobservable. If most are near-parallel the + problem is ill-conditioned -- it still returns a unique-looking answer, just an inaccurate + one `[EXP: IEEE T-RA 1989; restated arXiv:2308.06045]`. +- **Which poses you choose matters more than which solver you use.** Holding the algorithm + fixed and changing only the pose-selection strategy moved translation error from 4.74 mm + to 2.95 mm, a 37.7% reduction -- a larger swing than the spread between most AX=XB solvers + in the same papers `[EXP: arXiv:2303.06766]`. There is also no single best solver: + simultaneous methods resist rotation noise better, separated methods resist translation + noise better, and the ranking flips with the noise profile `[EXP: PLOS ONE 2022]`. +- **The reported accuracy is a self-consistency residual, not error against truth.** + Ground-truth hand-eye transform is not available with real data, stated outright by the + authors who tested five solvers on 101 poses `[EXP: PLOS ONE 2022]`. A competently + executed calibration on an industrial arm with a good camera lands roughly in the 0.5-3 mm + and 0.1-1 deg band across the studies located `[EXP]`. + +Two consequences that get missed. First, **a translation residual quoted without its +rotation residual and its working distance is close to meaningless.** Small-angle geometry: +a rotational error of d-theta contributes about (working distance) x d-theta to +end-effector position error. Taking one vendor's documented worked example of 0.83 deg = +0.0145 rad at the 600-1000 mm working distance that same documentation describes gives 8.7 +to 14.5 mm from the rotation term alone, against its own 0.67 mm translation residual +`[ARITH -- this guide's calculation, not a published end-to-end measurement]`. Check this +on your own numbers; it is the kind of arithmetic that changes a decision. + +Second, **hand-eye accuracy is bounded by the arm's own absolute accuracy**, because the A +matrices come from forward kinematics and the standard formulations do not model kinematic +error -- it accumulates into X. For magnitude: laser-tracker kinematic calibration of an +industrial arm moved mean and max absolute position error from 2.628 mm / 6.282 mm to +0.208 mm / 0.482 mm `[EXP]`. An uncalibrated arm contributes millimeters before the vision +system contributes anything. See 3, because this is the number people substitute the +datasheet repeatability figure for. + +**Camera pose repeatability.** This is the other thing, and for a pixel-space policy it is +the one that matters. + +### What happens to a trained policy when the camera is moved + +Here is the reframe, and it is the single most useful sentence in this section: + +> **Standard behavior-cloning visuomotor policies do not consume hand-eye extrinsics at +> all.** ACT, Diffusion Policy, OpenVLA and SmolVLA map pixels to actions end-to-end. A 2026 +> paper presenting extrinsics conditioning as a novel contribution over exactly these +> baselines confirms the baselines lack it `[EXP: arXiv:2510.02268, ICRA 2026]`. + +So the millimeter accuracy of your AX=XB solve is not what determines whether the policy +works. **Camera pose repeatability relative to the training distribution is.** And +recalibrating a moved camera does not restore a pixel-space policy, because the policy never +read the calibration. (No study located tests recalibrate-versus-not as an intervention; the +inference is from the fact that these architectures take no extrinsics input. Flagged in the +refusals.) + +What moving the camera costs, measured: + +- On a real robot, against its own training-condition baseline of 91.7%, the drops were: + camera position **-45.9 points**, table texture -38.9, distractor objects -11.1, lighting + -8.4, background -2.8. Camera pose was the hardest factor tested `[EXP: ICRA 2024, + arXiv:2307.03659]`. Apply 6 before believing all of that: those rest on 12-36 trials each. + Exact 95% Clopper-Pearson intervals from their own fractions put the training baseline + (11/12) at [61.5%, 99.8%] and the new-camera condition (11/24) at [25.6%, 67.2%] -- the + camera drop survives because the intervals do not overlap, while the lighting and + distractor drops are not distinguishable from baseline at that trial count `[ARITH]`. +- In simulation, applying camera orientation perturbations of 2-10 deg -- **within the range + of a sloppy remount** -- alongside distance and spherical-position changes, nine + vision-language-action models dropped by between 19 and 91 points; several collapsed to + near zero `[SIM: LIBERO-Plus, arXiv:2510.13626]`. The variance across models is enormous, + so "VLAs are viewpoint-fragile" is only true on average. +- **Architecture decides sensitivity.** A 14-factor perturbation benchmark reports that 2D + image-based models are affected by camera pose while 3D models operating on a voxel or + point-cloud representation are robust to it, because they do not learn directly on the + captured view `[EXP: RSS 2024, arXiv:2402.08191]`. If your policy consumes a + reconstruction built from extrinsics, calibration accuracy matters and viewpoint does not. + If it consumes raw pixels, the reverse. +- **The mechanism.** Policies without extrinsics infer camera pose from visual cues in + static backgrounds, and that shortcut collapses when workspace geometry or camera placement + shifts. Conditioning on extrinsics recovers much of the loss -- gains from +0.4 to +34.8 + points across six simulated tasks under randomized camera poses `[SIM: arXiv:2510.02268]`. + +**Practical rules that follow.** + +1. Bolt the camera. Positively locate it into deck features the way `docs/PRINTED_FIXTURES.md` + 5 says to locate a fixture: not tape, not friction, not a clamp somebody re-tightens. + A mount that depends on putting it back in the same place is not a mount. +2. Measure and record the extrinsics anyway, even though the policy will not read them. It + is the only way to later answer "did the camera move?", and it is what lets you switch to + an extrinsics-conditioned or 3D architecture without recollecting. +3. Treat "the camera moved or was remounted" as an event that sends you back down the ladder + in 8, not as a recalibration ticket. +4. Do not build a recalibration schedule out of drift fear. The only directly measured drift + figure located is 1.49 pixels maximum thermal image drift across 5 C to 45 C, with 90 + minutes to reach equilibrium at each step `[EXP: Sensors 2022]` -- small, and not by + itself a justification. Every recalibration-frequency figure in circulation (weekly, + monthly, quarterly, seasonally) is vendor guidance with no measured basis `[VENDOR]`. + Recalibrating after a collision or a physical change is a sensible precaution and is also + unquantified `[PRACTICE]`. + +One more that belongs here because it is a timestamp problem wearing a calibration costume. +The same flange-based study that reached 1.57 / -1.07 / -1.12 mm residuals in static mode +reached 18.89 / -44.90 / 6.22 mm when the camera captured while the arm was moving -- roughly +a 40x degradation -- and the authors attribute it to imprecise timestamp synchronization +between camera capture and robot pose, not to the AX=XB math `[EXP: Frontiers in Robotics +and AI 7:65, 2020]`. **Synchronization is not a detail of the capture. It is the thing most +likely to be wrong.** + +--- + +## 3. The arm + +### Confirmed specification + +Read directly from the official xArm User Manual V2.0.0 (254 pp), Appendix 2, pp. 217-221 +`[DS]`. + +| | xArm 5 | xArm 6 | xArm 7 | +| --- | --- | --- | --- | +| DOF | 5 | 6 | 7 | +| payload | 3 kg | 5 kg | 3.5 kg | +| arm weight | 11.2 kg | 12.2 kg | 13.7 kg | + +Shared across all three, from the manual's common-specifications table (p. 217) `[DS]`: +reach **700 mm**; repeatability **+/-0.1 mm**; max end-effector speed **1 m/s**; max joint +speed **180 deg/s**; Cartesian range X and Y +/-700 mm, Z -400 to 951.5 mm, roll/yaw/pitch ++/-180 deg; ISO Class 5 cleanroom; mounting "any"; tool flange DIN ISO 9409-1-A50/63 (M5*6); +input 24 V DC, 16.5 A; power min 8.4 W, typical 200 W, max 500 W. + +Joint ranges differ per model (pp. 219-221) `[DS]`. All three share J2 at -118 to 120 deg. +xArm 5 and xArm 6 both carry J3 at -225 to 11 deg; xArm 7's corresponding joint is J4 at -11 +to 225 deg, which is the mirrored range and not the same limit. Read the per-model table for +the arm you buy rather than carrying one model's numbers across. + +**Two flags on that table.** Reach and repeatability come from the *shared* table; the +per-model tables for xArm 5 and xArm 7 list payload, DOF and weight and do not restate +repeatability. And payload as a flat number is a simplification -- the manual's own +"Maximum Payload" section states that maximum allowed payload depends on the center of +gravity offset from the flange, and the derating curve is published only as an image +`[DS]`. A gripper plus a wrist camera plus a full deep-well plate on a bracket is exactly +the case where the flat rating stops holding. + +### What makes it usable for capture + +The Python SDK is first-party and genuinely open: package `xarm-python-sdk`, repository +`xArm-Developer/xArm-Python-SDK`, version 1.18.4 released 2026-05-21, declared support +Python 3.5 through 3.13, licensed BSD 3-Clause `[DS]`. One `XArmAPI` surface covers xArm +5/6/7 plus the 850 and Lite 6, and exposes motion (`set_position`, `set_position_aa`, +`set_servo_angle`), kinematics (`get_forward_kinematics`, `get_inverse_kinematics`), +state/mode, grippers, controller and tool GPIO including analog, Modbus RS485 and TCP +passthrough, and safety configuration `[DS]`. + +That is what makes action-labeled capture cheap on this arm: you can read state and command +poses from the same process that writes your dataset. The manual also documents +**hand-teaching with its own teach sensitivity, levels 1-5** `[DS]`, which is the +kinesthetic capture path from 1 with no extra hardware at all. + +A first-party six-axis force/torque sensor exists as a flange accessory, compatible with +xArm 5/6/7 and 850. From its official manual, Section 5 `[DS]`: load capacity Fx, Fy = +**150 N**, Fz = **200 N**, Tx/Ty/Tz = **4 Nm**; resolution 100 mN (Fx, Fy), 150 mN (Fz), +5 mNm (torque); hysteresis 2.5% FS (Fx, Fy) and 1% FS (Fz, torque); crosstalk 3% FS; +overload 150%, and for Fz 150% positive / 300% negative; weight 445 g. It is exposed +through the SDK (`ft_sensor_enable`, `ft_sensor_set_zero`, `get_ft_sensor_data`, +`ft_ext_force`, `ft_raw_force`, plus admittance control, force control, and FT-based +collision detection with rebound) `[DS]`. + +**Correct one thing before you spec it.** The figures circulating as this sensor's *range* +-- +/-225 N, +300/-600 N, +/-6 Nm -- are the **overload** values, i.e. 150%/300% of the load +capacities. One reseller listing of "400 N / 20 Nm" matches no figure in the official manual +at all `[VENDOR]`. Use 150 N / 200 N / 4 Nm as the working spec. + +### Repeatability versus accuracy, and why +/-0.1 mm is not the number + +**Repeatability** is how closely the arm returns to the same commanded pose across +repetitions. **Accuracy** is how close the pose it reaches is to the pose the coordinate +actually names. They are different quantities, they differ by roughly an order of magnitude +on industrial arms, and the datasheet publishes only the first one. + +The +/-0.1 mm figure is a vendor-declared number with no stated measurement standard. The +string "9283" -- ISO 9283, the robot performance test standard -- appears **zero times** in +the 254-page manual, and no test load, speed, temperature or pose set is given `[DS]`. It is +not a published measured result, and real-world figures will depend on payload, speed, +thermal state and pose. + +**The number that decides whether a tip lands in a well is a stack, and repeatability is +one small term in it.** Going from the arm outward: + +1. The arm's absolute accuracy in the frame you taught in. Unpublished by this vendor and by + most others. Magnitude anchor: laser-tracker kinematic calibration of an industrial arm + moved mean/max absolute position error from 2.628 / 6.282 mm to 0.208 / 0.482 mm `[EXP]` + -- millimeters before calibration, an order of magnitude above the datasheet + repeatability. ("Absolute accuracy is about 20x repeatability" is a widely repeated rule + of thumb; the phenomenon is real and the specific multiplier is not measured `[PRACTICE]`.) +2. The hand-eye residual, if vision is in the loop: 0.5-3 mm and 0.1-1 deg for a competent + calibration, with the rotation term amplified by working distance (2). +3. Where the fixture actually is, which is a printed-part question and lives in + `docs/PRINTED_FIXTURES.md` 5. +4. The plate's own tolerance. ANSI/SLAS 1 allows +/-0.25 mm only within 12.7 mm of the four + outside corners and **+/-0.5 mm** elsewhere along the side, and corner radius is 3.18 mm + +/-1.6 mm -- a 1.58 to 4.78 mm permitted range `[STD, via the sibling guide]`. +5. Thermal state. Printed polymers move far more than molded ones, and SLAS dimensions are + specified at 20 C while its own test method for one part of the family is at 25 C +/- 2 C. + +Terms 3, 4 and 5 are frequently larger than terms 1 and 2. **A guide that answers "will the +tip land in the well?" with the arm's repeatability figure has answered a different +question.** + +There is no published figure for the end-to-end stack on a lab deck, and this guide will +not compose one out of the terms above -- they do not add independently and their signs are +not known. **The measurement that produces it** is direct: with the fixture located and the +plate seated, drive the arm to the taught pose and measure the tip's actual offset from the +well center, at the working temperature, over enough repeats to state a range rather than a +value (`autonomous_lab/teaching.py` sets that floor at `MIN_DEMONSTRATIONS = 3` and refuses +a tolerance below it). Repeat after any event in 8's demotion list. + +### Choosing the DOF + +This is the only arm choice this guide will make an argument about, and it is kinematics +rather than a benchmark. + +- **5 DOF cannot reach an arbitrary 6-DOF pose.** A full pose is three positions and three + orientations. With five joints the reachable set is constrained, and the constraint bites + the moment a fixture presents labware at an angle it did not anticipate -- which is + exactly what the tilt module in `hardware/tilt_module.scad` does. +- **6 DOF is the general case** and is the smallest arm that can put the flange anywhere in + its workspace at any orientation. It also carries the largest payload of the three (5 kg), + and it is the only model for which this manual documents a CE certification (below). +- **7 DOF adds a redundant joint.** That buys obstacle avoidance and joint-limit avoidance + around a fixed end-effector pose. It also means one end-effector pose corresponds to a + continuum of joint configurations -- which is the non-injective case from 1, and it makes + video retargeting strictly worse while making teleoperation mapping ambiguous. Redundancy + is a feature you should want deliberately, not by default. + +### What could not be confirmed about this arm + +Kept here rather than in the refusals list because it is buying advice. + +- **No independent metrology.** Every accuracy, repeatability and payload figure above + traces to UFACTORY's own documentation or storefront. No third-party study was located. +- **Price is inconsistent across the vendor's own US web properties.** Individual product + pages returned $6,000 (xArm 5), $9,500 (xArm 6), $11,000 (xArm 7); an aggregate fetch of + the same site returned $5,799 and $10,499 for two of them. The force/torque sensor lists + at $3,500 on a page that omits its force ranges entirely. These could not be reconciled. + **Get a quote.** Which control box, cables and region are included at each figure is not + determinable from the public pages. +- **F/T sensor accuracy** (as distinct from resolution and hysteresis) and its data rate are + not in the official specification table. The commonly cited 200 Hz comes from reseller + pages `[VENDOR]`. +- **"Collaborative robot" is reseller framing.** The word does not appear in the manual's + specification or safety sections. See 7. +- **Firmware and controller software are not open.** Only the Python SDK, C++ SDK and ROS + packages are confirmed BSD-3-Clause. Assume the motion and kinematics core is closed. +- **24/7 duty-cycle suitability** is a product-page claim `[VENDOR]`, while the manual itself + warns to reduce temperature for continuous high-speed operation in a 0-50 C ambient range + `[DS]`. + +--- + +## 4. Where the printed fixture earns its place + +The argument for a printed fixture in a learning pipeline is one sentence: **it removes +degrees of freedom from what has to be learned.** A plate that can be anywhere in a +40 x 40 cm region is a pose the policy must infer from pixels every episode. A plate in a +nest is a pose the policy can assume. + +That argument is correct in direction. The honest accounting of how correct, and of how +much, is below -- and the second half of this section is the trap, which is worse than most +people expect. + +### What is published evidence + +**Nothing published uses a physical fixture as the manipulated variable.** This guide's +research located no study that fixtures a part, runs an unfixtured control, and measures +demonstrations required to reach a target success rate. Every quantified source substitutes +something for the fixture: a simulated randomization range, a simulated reset distribution, +or software pose canonicalization. Everything below is therefore evidence *about pose +variance*, which is what a fixture removes -- not evidence about fixtures. + +**The scaling law.** ManiBox (arXiv:2411.01850) measured trajectories needed to reach 80% +grasp success as a function of the volume over which the object's position was randomized, +in simulation `[SIM]`: + +| randomization volume | trajectories to 80% | +| --- | --- | +| 1 cm3 (effectively a fixed point) | 728 | +| 125 cm3 (5 x 5 x 5 cm) | 1,951 | +| 1,000 cm3 (10 x 10 x 10 cm) | 8,098 | +| 8,000 cm3 (20 x 20 x 20 cm) | 14,638 | +| 34,400 cm3 (41 x 30 x 28 cm, about full reach) | 24,005 | + +Fitted: data = 640.32 x volume^0.35. Fixed point to full workspace is 33x in the raw +measured points, and the authors compute 34,400^0.35 = 38x from the fit. + +**Now do the arithmetic that changes the design decision** `[ARITH, on their fit]`. An +exponent of 0.35 on *volume* is about L^1.05 in *linear extent*, since (L^3)^0.35 = L^1.05. +So required data is roughly **linear in the linear span** of the placement region, not in +its volume: + +``` + halving tolerance on each axis = 8x less volume = 8^0.35 = about 2.1x less data + +/-20 mm scatter to +/-2 mm = 1000x less volume = 1000^0.35 = about 11x less data +``` + +A fixture that cuts placement scatter by an order of magnitude in each axis predicts roughly +10x less data, **not 1000x**. That is a large and worthwhile win and it is not the win people +describe when they say a fixture "makes it trivial." + +Caveats that must travel with those numbers: the fit is good at 1,000 / 8,000 / 34,400 cm3 +(predicted 7,186 / 14,853 / 24,332 against measured 8,098 / 14,638 / 24,005) and **poor at +125 cm3** (predicted 3,471 against measured 1,951). It is simulation, one grasping task +family, a bounding-box state representation, simulated trajectories rather than human +demonstrations, and **3-DoF translation only** -- orientation variance is not in the volume +term at all. The authors call it "preliminary verification." + +**Fixed budget, varying pose spread.** MimicGen (CoRL 2023, arXiv:2310.17596) holds the +demonstration budget at 1,000, holds the architecture fixed, runs 3 seeds, and changes only +the object reset distribution. Its D0 condition has the receptacle **rigidly fixed** -- a +fixture in all but name -- while D1 and D2 unfix it and D2 adds free rotation `[SIM]`: + +| task | D0 (receptacle fixed) | drop at D1 | drop at D2 | +| --- | --- | --- | --- | +| Square | 90.7 +/- 1.9% | -17.4 pts | -41.4 pts | +| Threading | 98.0 +/- 1.6% | -37.3 pts | -60.0 pts | +| Three Piece Assembly | 82.0 +/- 1.6% | -19.3 pts | -68.7 pts | +| Coffee | 100.0 +/- 0.0% | -9.3 pts | -22.7 pts | +| Nut-and-Bolt (Factory) | 92.7% | -11.4 pts | -20.0 pts | +| Gear Assembly (Factory) | 98.7% | -24.7 pts | -42.0 pts | +| Frame Assembly (Factory) | 82.0% | -13.3 pts | -45.3 pts | +| **Stack** | **100.0 +/- 0.0%** | **-0.7 pts** | -- | + +**Read the last row.** Widening the placement region from 16 x 16 cm to 40 x 40 cm -- 6.25x +the area -- cost 0.7 points on Stack. Constraining pose bought essentially nothing on an +easy, low-precision task. **The benefit is not task-independent.** It is large on +contact-rich, high-precision and long-horizon tasks and near zero on coarse ones. Which of +those your task is decides whether the fixture is a data-efficiency measure or just a +tidiness measure. + +The same paper also carries the negative result that kills the usual fantasy. Its 10 real +human source demonstrations, **evaluated in the same narrow fixtured D0 region they were +collected in**, scored 11.3% (Square), 19.3% (Threading), 1.3% (Three Piece Assembly) and +26.0% (Stack) in simulation; on the real robot the 10-demo source agents scored **0% on +Stack and 0% on Coffee** (94% pod grasp, 0% insertion) `[EXP]`. Reaching the D0 numbers in +the table took 1,000 demonstrations. **A fixture shifts the data curve. It does not collapse +the requirement to a handful of demonstrations.** + +Pose spread also costs you at collection time, not only at training time: MimicGen's own +rate of attempts yielding a usable demonstration fell from 73.7% to 48.9% to 31.8% (Square) +and 51.0% to 39.2% to 21.6% (Threading) across D0/D1/D2, which is 2.3-2.4x more attempts to +bank the same 1,000 usable demonstrations `[SIM]`. + +**Software isolates pose as the cause.** Two results remove pose variance in the +representation rather than the world and get the same shape of win, which is what makes the +causal story credible. An oriented-affordance-frame method needed 305 image-based +demonstrations to match what its frame-canonicalized policy reached from 10 -- the authors +call it 30x -- and ablating just the *orientation* component of the frame nearly halved +success at 10 demonstrations `[EXP: CoRL 2025, arXiv:2410.12124]`. An SO(2)-equivariant +diffusion policy averaged its result at 100 demonstrations above all baselines trained with +1,000, across 12 tasks `[SIM: IJRR 2026, DOI 10.1177/02783649261424445]`. + +### What is engineering practice + +- The **3-2-1 locating principle** -- three locators for the primary datum, two for the + secondary, one for the tertiary, constraining six degrees of freedom with minimum contacts + -- is the textbook justification for jigs and fixtures. No publication quantifies its + effect on demonstration count for a learned policy `[PRACTICE]`. +- **"Fixture it and hard-code the waypoints, skip learning entirely."** This is the dominant + approach in laboratory automation and it is often the right answer. Published lab-robotics + work uses fiducials on statically mounted labware and teach-pendant reference points and + reports no demonstration-count comparison against an unfixtured baseline `[PRACTICE]`. If + your task is a fixed pick-and-place between two known positions, a policy is the wrong + tool and this guide's honest advice is to teach waypoints. +- **Any multiplier of the form "a nest cuts demonstrations by N x" for a lab workcell.** No + such published figure exists. The defensible numbers -- 33-38x fixed-point-to-full-workspace, + 30x for frame canonicalization, 2.3-2.4x collection overhead -- come from tabletop grasping + and assembly in other domains, mostly in simulation, and applying them to labware handling + is extrapolation `[PRACTICE]`. + +### The trap: the fixture is part of what the policy learned + +A hard-coded waypoint program depends on the fixture's *geometry*. A learned policy depends +on the fixture's **appearance and position in every training frame**. Its color, its layer +lines, its shadow, where its edge falls in the image, and where the plate ends up because of +it are all training signal, and none of it is recorded anywhere unless you record it. + +**So a reprint can invalidate the demonstrations.** Concretely, from +`docs/PRINTED_FIXTURES.md` 5: + +- Filament manufacturers largely **do not publish shrinkage**; three datasheets were opened + and none listed a figure `[DS]`. Every per-material percentage in circulation is a slicer + default or a community measurement. +- Inside one slicer's shipped profiles, one vendor's compensation implies measured shrinkage + of PLA 0.05%, PETG 0.15%, ABS 0.513%, ASA 0.513%, PC-CF 0.15%, while a different vendor in + the *same repository* ships zero compensation for PLA, PETG, ABS and ASA. Two vendors, one + slicer, a 0.5% disagreement on ABS `[shipped code, read directly]`. +- Across the 127.76 mm long dimension of a plate footprint, 0.513% is **0.66 mm** -- larger + than the entire ANSI/SLAS 1 corner-zone tolerance band of +/-0.25 mm `[ARITH + STD]`. + +A change of that size moves where the plate sits. For a waypoint program you re-teach the +point. For a policy, the demonstrations were collected against a fixture that no longer +exists, and nothing in the dataset says so. + +This is the same distinction the code already draws. `autonomous_lab/printed.py` keeps +`DimensionState.DESIGNED` and `DimensionState.MEASURED` apart precisely because a designed +dimension is a number in millimeters in a document that reads exactly like a measurement and +is a fact about a model. **A fixture whose registering features are DESIGNED has not +established that the plate is where the demonstrations put it.** The sibling guide's rule +applies verbatim here: *a part reprinted from a different spool is a different part until it +has been measured again.* + +**The tilt module is the worked case.** `hardware/tilt_module.scad` and +`docs/PRINTED_FIXTURES.md` 6 describe a wedge that presents a plate at a fixed angle. Today +it reaches `MEASURED` at best, and it is instructive for a policy pipeline because it changes +several things the policy would have learned at once: + +- the well bottom is no longer perpendicular to the approach; +- the clearance to the plate's high side shrinks; +- the stack height rises by the wedge, and the sibling guide computes 39.38 mm above the + deck datum against 14.35 mm flat; +- it introduces a retention failure the flat nest did not have -- a lip too shallow to hold + the plate at angle -- for which there is no friction coefficient and therefore no formula, + only a direct measurement of the angle at which the loaded plate moves, per flange variant. + +Every one of those is inside the policy's observation and action distribution. A tilt module +reprinted at a different shrinkage is not a cosmetic change to the deck; it is a change to +the dataset, applied retroactively, silently. + +### What to do about it + +1. **Version the fixture with the dataset.** Filament, spool and lot, slicer profile, + compensation values, print date, and the *measured* registering dimensions with the + instrument named. A dataset whose fixture provenance is unrecorded cannot be diagnosed + when success drops after a reprint -- you will not be able to tell a reprint from a camera + nudge from a bad training run. +2. **Treat a reprint as a demotion event** (8). It sends the part back to `MEASURED` in the + sibling guide's ladder and sends the policy back to a fixture-with-no-material rung. + Re-run the dry run: a printed fixture that moves under arm contact is **silent by + construction** -- nothing in a printed part reports its own position, and the first + evidence is the collision. +3. **Locate it positively.** Bolt, key, or capture it between deck rails. A registration that + depends on somebody putting it back in the same place is not a registration, and for a + policy it is also a slow corruption of the training distribution. +4. **Or deliberately train across fixture variation**, which converts the problem into the + diversity lever in 5. That costs demonstrations across variants and **there is no + published number for how many.** The measurement that would produce one: print the fixture + n times across the lots you will actually use, measure the registering features on each, + collect demonstrations against each, and evaluate on a held-out reprint per 6. + +### A note on the apparent contradiction with 5 + +Section 5 reports that the only controlled study of demonstration counts found *diversity* +dominates raw count. This section says to *remove* variance. Both are right, and the +resolution is that they are about different variance: + +- Pose variance within a fixed task is variance the policy must **solve**. Removing it + reduces what has to be learned for the task you have (ManiBox, MimicGen). +- Environment and object diversity is variance that makes a policy **transfer** to + situations you did not collect. Adding it buys generalization (Lin et al.). + +Pinning the plate does not stop you varying lighting, consumable lot, plate vendor, liquid, +operator, time of day, or -- per the point above -- the fixture reprint itself. Fixture the +thing whose pose you want assumed; diversify everything you want survived. + +--- + +## 5. How many demonstrations + +There is no answer to this question, and the useful thing this section can do is show you +exactly how absent the answer is, then tell you what actually determines it. + +### The published per-task counts, with methods and tasks named + +**ACT / ALOHA** (RSS 2023, arXiv:2304.13705), 6 real bimanual tasks on a roughly USD 20k +platform, verbatim: "We record 50 demonstrations for each task, except for Thread Velcro +which has 100." Episodes 8-14 s at 50 Hz control = 400-700 timesteps. "The total amount for +demonstrations is thus around 10-20 minutes of data for each task, and 30-60 minutes in +wall-clock time because of resets and teleoperator mistakes." `[EXP]` + +Per-task final success on real hardware, 1 seed x **25 evaluations** per task: Slide Ziploc +88%, Slot Battery 96%, Open Cup 84%, Thread Velcro 20%, Prep Tape 64%, Put On Shoe 92% +`[EXP]`. At 25 trials one success is 4 percentage points `[ARITH]`; these are coarse +estimates. + +Three things about that paper matter more than the 50: + +- **Its own abstract is selective against its own table.** "80-90% success with only 10 + minutes worth of demonstrations" describes four of six tasks (84-96%). Prep Tape is 64% + and Thread Velcro is 20%. +- **Thread Velcro got twice the demonstrations and produced the worst result.** The paper + attributes this to perception and precision, not data volume -- success roughly halved at + each stage, "from 92% success at the first stage to 20% final success," with failures from + the gripper closing too early and imprecise insertion. **Demonstration count was not the + binding constraint on the hardest task in the paper.** +- **ACT contains no ablation over number of demonstrations at all.** Its Figure 8(a) x-axis + values (1, 10, 100, 200, 400) are action-chunk size k, not demonstration count. This is a + very easy misread and it is the origin of a lot of confident advice. + +**Diffusion Policy** (RSS 2023 and the extended IJRR version, arXiv:2303.04137v5), real +robot, Table 3: Push-T **136** proficient-human demonstrations; Mug Flip **250**; 6DoF Pour +**90**; Periodic Spread **90**. The journal version's bimanual tasks: Egg Beater **210** +demonstrations for 55% success over 20 trials; Mat Unrolling **162** for 75% over 20; Shirt +Folding **284** for 75% over 20 `[EXP]`. All real-robot evaluations use 20 trials, so one +trial is 5 percentage points `[ARITH]`. + +Two traps. The paper **contradicts itself** on the sauce tasks -- Table 3 lists 90 +proficient-human demonstrations while Appendix C.2.1 states "50 demonstrations are +collected, and 90% are used for training for each task." Unresolved; both are reported here. +And its **simulation** counts (200 proficient-human + 300 multi-human for several Robomimic +tasks, 200 for simulated Push-T, 1,000 scripted for BlockPush) are routinely quoted as if +they were the real-robot counts. + +**Mobile ALOHA** (CoRL 2024, arXiv:2401.02117) uses "50 in-domain demonstrations, or 20 in +the case of High Five" (Cook Shrimp also 20) `[EXP]`. **The caveat is load-bearing and it is +usually dropped: the 50 only holds with co-training on a large static-ALOHA dataset.** +Table 4 gives Wipe Wine at 95% with co-training against 50% without; Call Elevator goes from +roughly 0% to 95%. Without the auxiliary dataset, 50 is not sufficient for several of these +tasks. + +Its Figure 4 is the cleanest real-hardware demonstration-count ablation in this set (Wipe +Wine, ACT, 25/35/50 demonstrations, 20 trials each): "With co-training, the policy trained +with 35 in-domain demonstrations can outperform the no co-training policy trained with 50 +in-domain demonstrations, by 20% (70% vs. 50%)." `[EXP]` + +The same paper contains the cleanest evidence that **the requirement is method-dependent, +not only task-dependent**: at an identical 50 demonstrations on Wipe Wine, ACT reaches 95% +and Diffusion Policy reaches 65%. + +**DP3 / 3D Diffusion Policy** (RSS 2024, arXiv:2403.03954). The frequently repeated "10 +demonstrations" is a **simulation** result across 72 simulated tasks `[SIM]`. On real +hardware the paper used **40 demonstrations per task** across 4 tasks, averaging 85.0 +/- +11.2% `[EXP]`. + +**ALOHA Unleashed** (CoRL 2024, arXiv:2410.13126) is the decisive counterexample to the +whole "50 demos" framing: same hardware lineage, per-task counts three orders of magnitude +larger. Over **26,000** real demonstrations across 5 tasks, collected on 10 robots over +eight months -- ShirtEasy/Messy 8,658; LaceEasy/Messy 5,133; FingerReplace 5,247; GearInsert +4,005; RandomKitchen 3,198 `[EXP-1]`. And even at thousands per task the results are **not +saturated**: 75% / 70% on shirts, 40% on the harder lace variant, 40% on the third gear +insert, and a kitchen task degrading from 95% with one object to 65% with two and 25% with +three. **Demonstration count buys task difficulty and robustness, not a march to 100%.** +(This one rests on a single HTML fetch and was not re-extracted from the PDF -- re-verify +before quoting.) + +**Frontier VLAs changed the unit of account**, which breaks direct comparison. pi-0 +(arXiv:2410.24164) reports **hours**: "the simplest of the tasks necessitating only 5 hours +and the most complex tasks using 100 or more hours of data," and abandons binary success for +a normalized score averaged over 10 episodes per task. pi-0.5 (arXiv:2504.16054) reports +"about 400 hours of data of mobile manipulators performing household tasks in about 100 +different home environments," with no per-task breakdown at all `[EXP]`. + +Consequently: **any per-task demonstration count attributed to pi-0 or pi-0.5 comes from +outside the papers.** The figures in circulation ("100-200 to fine-tune pi-0 to a new task"; +"200-500 for pi-0.5 single-task adaptation, 1,000-5,000 across 5-20 tasks for a new robot") +surfaced only on a commercial third-party model-documentation site `[VENDOR, unverified +provenance]`. + +### The one work that actually measures the question + +Data Scaling Laws in Imitation Learning (ICLR 2025, arXiv:2410.18647), over 40,000 +demonstrations and more than 15,000 real-world rollouts. Headline finding, verbatim: "the +diversity of environments and objects is far more important than the absolute number of +demonstrations; once the number of demonstrations per environment or object reaches a certain +threshold, additional demonstrations have minimal effect." `[EXP]` + +Concretely: roughly 50 demonstrations per environment-object pair, across 32 distinct pairs, +totalling about 1,600, reached around 90% success in novel environments with unseen objects +-- measured on 2 tasks only. Reported plateaus at 400 / 800 / 1,600 total demonstrations for +8 / 16 / 32 pairs. Generalization follows a roughly power-law relationship with the number of +environments and objects, **not with raw demonstration count**; with environments and objects +held fixed there was no clear power law between count and generalization at all (correlation +coefficients -0.62 and -0.79 for the two tasks). + +### The synthesis, and the refusal + +The published real-robot range spans **20 demonstrations** (Mobile ALOHA High Five, with +co-training) to **8,658** (ALOHA Unleashed shirt) -- about a 400x spread, across tasks that +are all more similar to each other than any of them is to a lab bench. + +**The number for a new task is not knowable in advance.** What the measurements say +determines it: + +1. **Task precision and horizon.** The strongest single signal in the set: Thread Velcro at + 2x the demonstrations and the worst result, with the paper's own diagnosis being precision + and perception. +2. **Whether auxiliary or pre-training data is used.** The same 50 demonstrations mean + different things with and without co-training, by up to 95 points on one task. +3. **Environment and object diversity**, which is the only factor a controlled study found + dominant. +4. **The method.** ACT and Diffusion Policy differ by 30 points at identical count on the + same task. + +Notably absent from that list: a threshold. + +**What is repeated as evidence and is not** `[PRACTICE]`: + +- *"50 per task is the standard starting budget."* This generalizes one collection choice in + one paper into a sufficiency threshold. ACT collected 50 for all six tasks and ran no + ablation; its own Thread Velcro at 100 shows the number does not carry that meaning. +- *"Diffusion Policy needs roughly 250+."* Directionally consistent with its own real-robot + counts (136-284), but as a rule it traces to a **hypothesis sentence** in Mobile ALOHA + about prior practice, not to a controlled measurement. +- *"DP3 only needs 10."* Simulation-only. The real-robot experiments used 40. +- *"Start at 50, scale to a few hundred for long-horizon or contact-rich, thousands for + deformables."* A sensible heuristic that happens to match the spread, and no paper measures + those as thresholds. It is pattern-matching across papers with different robots, tasks, + success criteria and trial counts. +- *"Success improves smoothly with count, so collect until it plateaus."* The only two + real-hardware curves measured (Lin et al.; Mobile ALOHA Fig. 4) find closer to the + opposite: per-pair count saturates quickly and diversity carries the gains. + +### What to do instead + +Treat the count as something you **measure on your task**, not something you look up. + +- Collect in blocks. Freeze a checkpoint. Evaluate under the protocol in 6. Add a block. + Re-evaluate. +- **Recognize that this is a sequential comparison and that the naive version is invalid.** + Extending a fixed-batch test with a few more trials "constitute[s] p-hacking that + invalidates statistical assurances" `[EXP: RSS 2025, arXiv:2503.10966]`. Either pre-commit + the batch size or use a procedure designed for sequential stopping. +- Budget wall-clock honestly. ACT's own accounting is 10-20 minutes of *data* per task + against 30-60 minutes of *wall clock*, "because of resets and teleoperator mistakes" + `[EXP]` -- roughly a 2-3x overhead. That ratio is about collection, not about outcomes, and + it is the most transferable number in this section. +- Spend the marginal hour on diversity before spending it on count, because that is the only + place a controlled study found the gains. + +--- + +## 6. Evaluation + +A success rate is a point estimate of a Bernoulli parameter. Almost everything that goes +wrong in this section comes from forgetting that sentence. + +### Held-out trials, rolled out autonomously + +Two separate requirements that get collapsed. + +**Held out.** Evaluate from initial conditions the policy was not trained on -- object poses, +lighting, lot, operator, plate instance, time of day. There is no standard taxonomy for what +counts as held out; each lab picks its own factors and its own seen/unseen boundary +`[PRACTICE]`. Pick yours before you evaluate and write it down, because the boundary is what +the number means. + +**Autonomous.** The policy must drive the whole episode from its own states. Behavior cloning +trained i.i.d. on expert states incurs cost that compounds quadratically in horizon -- +O(eps * T^2) against O(eps * T) for an interactive reduction -- because the policy's own +errors move it off the training distribution, and that effect only appears under autonomous +rollout `[DAgger, AISTATS 2011; standard and widely restated, primary theorem statement not +read here]`. **Success measured from training initial conditions is structurally +optimistic**, and the structure is known. + +Related, and directly measured: the distribution shifts you will actually hit are not equal. +Against a 91.7% training-condition baseline on a real robot, camera position cost 45.9 +points, table texture 38.9, distractors 11.1, lighting 8.4, background 2.8 `[EXP: +arXiv:2307.03659]`. In simulation, increasing training environment configurations from 5 to +100 shrank the maximum generalization gap from 0.40 to under 0.10 `[SIM, same paper]` -- +which is 5's diversity lever, measured on the evaluation side. + +### An external success criterion, not the policy's confidence + +**Offline loss barely predicts real-world performance.** Across roughly 1,500 paired +simulation-and-real evaluation episodes on 2 embodiments, a behavior-cloning "validation +MSE" baseline scored Pearson r = **0.308** as a predictor of real-world policy ranking, +against r = 0.924 for a visual-matching simulator `[EXP: CoRL 2024, arXiv:2405.05941]`. A +training curve is not an evaluation and a validation loss is not a success rate. + +So the criterion has to be a physical outcome, read by something that is not the policy. In +a lab that is the assay, which is the top rung in 8. Three requirements on it: + +- **It must be defined before the run.** Written down, with the pass condition explicit. +- **It should be scored by someone who does not know which policy ran.** The most rigorous + published campaign used fully blind evaluation with randomized policy ordering and initial + conditions matched by image overlay `[EXP: arXiv:2507.05331]`. Note honestly that **no + published robotics study measures how much bias non-blind evaluation introduces**, so this + is hygiene borrowed from clinical trials rather than a quantified correction `[PRACTICE]`. +- **Expect label noise even so.** A quality audit of 27% of about 2,700 rollouts in that same + campaign found a **2.31%** discrepancy in success-rate labelling and **6.25%** on + rubric questions `[EXP]`. Careful human scoring carries a couple of points of noise. + +### The trial count needed for a meaningful interval + +This is arithmetic, so it is available even though the robotics literature declines to name a +threshold. Exact Clopper-Pearson 95% intervals `[ARITH]`: + +| trials | interval at observed 90% | width | interval at observed 50% | +| --- | --- | --- | --- | +| 10 | [55.5%, 99.7%] | 44.2 pts | [18.7%, 81.3%] | +| 20 | [68.3%, 98.8%] | 30.5 pts | [27.2%, 72.8%] | +| 50 | [78.2%, 96.7%] | 18.5 pts | [35.5%, 64.5%] | +| 100 | [82.4%, 95.1%] | 12.7 pts | [39.8%, 60.2%] | +| 200 | [85.0%, 93.8%] | 8.8 pts | -- | +| 1000 | [88.0%, 91.8%] | 3.8 pts | -- | + +The consequence in one line: **a 10-trial evaluation cannot distinguish a 60% policy from a +95% one.** A vendor engineering blog independently reports 80.5-95.9% at 70 rollouts and +88.0-91.8% at 1,030, and notes that going from a 10-point interval to a 2-point interval +takes about 15x more rollouts; the 1,000-row above reproduces its figure exactly `[VENDOR, +arithmetic independently verified]`. + +**Comparing two policies is much more expensive than measuring one.** Two-proportion normal +approximation, two-sided alpha = 0.05, 80% power, trials **per policy** `[ARITH]`: + +| comparison | trials per policy | +| --- | --- | +| 50% vs 80% | 39 | +| 50% vs 70% | 93 | +| 80% vs 90% | 199 | +| 70% vs 80% | 294 | +| 85% vs 90% | 686 | + +Inverted, against a 50% baseline: 20 trials per arm detects only about a 40-point difference; +50 detects about 27; 100 detects about 19. **That is the quantitative reason a 20-trial A/B +cannot support a claimed 5-point improvement**, which is the size of improvement most +comparisons claim. + +For calibration against what the field does: the modal per-condition trial count in a +13-paper convenience sample of recent real-robot VLA papers is **10-20**, and none of the 13 +reported confidence intervals or paired tests `[EXP, single-author preprint, convenience +sample -- treat as illustrative, not representative]`. Recent benchmarks use 10 rollouts per +task; the scale in them comes from task count, not trials per cell. The exception is the +large-behavior-model campaign at roughly 1,800 controlled real trials with 50 rollouts per +task per policy per condition `[EXP]`. + +**This guide names no minimum trial count**, and the reason is that no published result +establishes one. The required n depends on the unknown true performance gap, which is not +knowable before the experiment; one widely cited methods paper explicitly declines to name a +number `[EXP]`. The table above is offered instead of a threshold: pick the interval width +you need to make your decision, read across, and that is your n. + +### Why a rate without its interval or its trial count is not a result + +Because it is unfalsifiable and uninterpretable at the same time. Read the table above +sideways: an observed "90%" is compatible with [55.5%, 99.7%] and with [88.0%, 91.8%]. Same +point estimate, and the two support opposite decisions. The trial count is not metadata about +the number; it is half of the number. + +The critique from the field, near-verbatim: evaluation "often solely focuses on success rate +with little description of the experimental conditions, number of evaluations, success +criteria, performance, failure modes, and typically without any statistical analysis" +`[EXP: arXiv:2409.09491]`. + +**Report k out of n, not a percentage.** The raw counts let a reader reconstruct the interval; +a percentage destroys the information irreversibly. Report the success criterion, the +held-out factors, who scored it and whether they were blind, and the failure modes. + +Three more disciplines with real evidence behind them: + +- **Freeze the checkpoint before evaluation begins.** Do not select the best-performing + checkpoint by evaluation success rate. Top-N trial selection is documented as a real problem + in reinforcement learning `[EXP: AAAI 2018, arXiv:1709.06560]`; the inflation it causes on + robot hardware is unmeasured `[PRACTICE]`. +- **Seeds alone can make an algorithm look better than itself.** Ten trials of one + configuration, varying only the random seed, split into two groups of five, produced + learning curves that "do not fall within the same distribution at all" `[EXP, same paper]`. +- **Do not extend a finished batch.** Use a pre-committed batch with an exact test, or a + sequential procedure built for it -- published sequential methods report up to 32% fewer + trials than batch baselines, and a 2026 follow-up using anytime-valid inference reports up + to 70% versus batch and up to 50% versus binary-outcome sequential methods, while also + finding that **policies separate faster on fine-grained task-progress scores than on binary + success** `[EXP: arXiv:2503.10966; arXiv:2603.13616]`. In a lab, a graded score is often + free: residual volume, recovery fraction, and fraction of wells addressed are all continuous. + +### The lab-specific part + +The assay is the external criterion, and it has its own n and its own basis. Two rules from +the repo apply unchanged: + +- `autonomous_lab/teaching.py` refuses to state a tolerance below `MIN_DEMONSTRATIONS = 3`, + and judges parity on the **worst** observation rather than the mean, because a mean lets one + excellent run pay for a bad one and at the bench the bad plate is the one that costs the + sample. A policy that ran once, well, is `Attainment.INDISTINGUISHABLE_FROM_UNMEASURED` -- + the numbers exist and cannot yet be told from luck. +- The acceptance-test design in `hardware/README.md` applies to a policy exactly as it applies + to a wedge: paired against a control, endpoint a quantity rather than an impression, + measured per well before pooling, positions randomized so treatment is not confounded with + deck position, threshold pre-registered with its basis recorded, and the observed range + reported rather than mean plus k sigma. + +And note what your assay is *not* allowed to be here: this workcell's absorbance read is +BROKEN, and even repaired, A260 does not discriminate library from primer, carrier or free +nucleotide `[repo: qc.py]`. A gate reading that number passes an empty well confidently. + +--- + +## 7. The safety boundary + +**A learned policy has no guarantees.** It is a function fit to data. Nothing in training +bounds what it can output, nothing in it detects that it is out of distribution, and its +confidence is not a safety signal (6: validation loss correlates with real performance at +r = 0.308). Under a distribution shift as small as a 2-10 deg camera remount, published +models have gone to near-zero success `[SIM]` -- and "near-zero success" describes the +outcome, not the trajectory the arm took getting there. + +**So the policy proposes. Deterministic interlocks dispose.** Everything below must be +enforced outside the policy, by something the policy cannot write to. A limit that the +policy's own process sets is a limit the policy can remove. + +### What must be bounded outside the policy + +**Workspace.** A Cartesian limit enforced by the controller, not a region the policy learned +to stay inside. The xArm firmware documents a **Safety Boundary** that stops motion when the +TCP exceeds a configured Cartesian boundary `[DS]`. Configure it to the smallest box that +contains the task, and set it before the first autonomous rollout, not after the first +surprise. + +**Speed.** The arm's documented maxima are 1 m/s at the end effector and 180 deg/s per joint +`[DS]` -- those are what an unbounded policy can command. **Reduced Mode** limits max +Cartesian linear speed, max joint speed and joint range `[DS]`. Rung 2 and rung 3 of the +ladder in 8 run in it. + +**Force and contact.** Collision detection exists with sensitivity levels 0-5, where **0 +disables it** `[DS]`. Understand what it is before relying on it: UFACTORY's own support +documentation states that "by comparing the theoretical current and actual current of each +joint, the system determines whether a collision has occurred," with the expected current +predicted from a dynamic model using joint position, speed, acceleration, load weight, center +of mass, mounting direction and friction parameters. The vendor documents its own failure +modes: wrong payload mass or center of gravity, wrong mounting direction, and friction +parameter mismatch after a controller replacement all cause false triggers or missed +detection, and **dynamic payloads must be reported to the controller during pick-and-place** +`[DS]`. + +That last clause is the lab case exactly. A plate that is empty on the way in and full on the +way out is a payload change mid-task, and a collision detector running on the wrong mass is a +detector with an unknown threshold. + +An independent limit is better: the first-party F/T sensor (150 N Fx/Fy, 200 N Fz, 4 Nm, +7) exposes FT-based collision detection with configurable thresholds and rebound through the +SDK `[DS]`. Use it as a bound the policy cannot see, not as an input the policy learns to +game. + +**Stop paths.** An emergency stop button on the control box, plus a protective stop input. +Documented behavioural difference: on E-stop the program stops and reset is manual; on +protective stop the program suspends and reset can be automatic or manual `[DS]`. Timing: +pressing the control box E-stop sends a software deceleration command, clears cached +commands, and "the power supply for the robotic arm will be removed within 300ms" `[DS]`. +Safety I/O exists in redundant pairs which "must be kept in two separate branches. A single +I/O failure should not result in the loss of safety features." `[DS]` A three-position +enabling switch is documented as an accessory `[DS]`. + +Print the vendor's own caveat above the bench: **"The emergency stop should not be used as a +risk reduction measure."** `[DS]` It is what you press when the bounding already failed. + +### What is not established, and must not be assumed + +- **ISO/TS 15066** -- the technical specification governing collaborative power-and-force + limiting -- is **not referenced anywhere in the manual**: zero hits for "15066", zero for + "ISO/TS". No power-and-force-limiting validation for any xArm model could be confirmed + `[DS]`. +- **EN ISO 10218-1:2011 is listed under "Applied Standards" for CE purposes only.** Zero hits + for "PL d", "Category 3" or "safety-rated" in the manual. **Do not assume the E-stop, + protective stop, reduced mode or safety boundary are rated safety functions** `[DS]`. +- **The CE statement names xArm 6 only** (models XI1300-XI1305, certified and tested by SGS). + No equivalent certification statement for xArm 5 or xArm 7 was located in that manual + `[DS]`. +- **Collision detection sensitivity 0-5 are dimensionless levels.** There is no published + force in newtons, no pressure, and no stopping-distance or stopping-time data anywhere in + the documentation. Treating it as an operator-safety function is engineering practice, not a + validated safety result `[PRACTICE]`. +- **"Cobot" is reseller and community framing.** The word "collaborative" appears essentially + nowhere in the manual's specification or safety sections `[DS]`. Do not let the label do + risk-assessment work. +- **The vendor assigns application-level safety to you.** From the F/T sensor manual: the + robot, sensor and all other equipment "must be evaluated with a risk assessment. The robot + integrator must ensure that all local safety measures and regulations are respected." + `[DS]` In a lab, the integrator is the person reading this. + +### The lab-specific interlocks no datasheet covers + +- **No autonomous motion while a person is inside the envelope.** Physical guarding or a + presence interlock. Not a policy behavior, not a convention, not a sign. +- **Nothing moves over an open reagent, a sample rack, or a scarce input** unless a dropped or + swept plate landing there is acceptable. This is a layout decision made once. +- **Payload reported to the controller on every state change**, per the collision-detection + failure modes above. +- **An interlock whose miss rate nobody has measured is not an interlock.** This is + `autonomous_lab/vision.py`'s rule verbatim: a detector whose sensitivity has not been + measured on held-out real failures "is not a safety device. It is a source of false + confidence, and installing one makes a lab less safe than having none, because the operator + stops looking." That applies to an F/T threshold and a camera-based guard identically. Set + the threshold, then **test it against the actual failure** -- drive the arm into the + condition deliberately, at reduced speed, and record whether the interlock fired. +- **The fixture is inside the safety argument.** A printed fixture that shifts mid-run puts + the instrument at risk and is silent by construction -- nothing in it reports its own + position, and the first evidence is the collision `[repo: PRINTED_FIXTURES 5]`. + +--- + +## 8. The order to do it in + +A ladder, in the shape `autonomous_lab/printed.py` already uses for printed parts, and with +the same rule stated first because it is the whole content: + +> **NO RUNG IMPLIES THE ONE ABOVE IT.** Each rung tests a different physical claim and the +> claims are independent. Simulation success says nothing about the arm. The arm moving +> safely over an empty deck says nothing about contact. Contact with a dry fixture says +> nothing about material. And material moving says nothing about the experiment working, +> because nobody measured the outcome. + +### Rung 1 -- simulation + +*What it establishes.* The code path runs end to end. Observation and action shapes are what +you think. The training loop converges to something. Gross failures -- inverted axes, +mis-scaled actions, a gripper command in the wrong units -- surface here for free. + +*What it does not establish.* Real performance. A visual-matching simulator correlated with +real-world policy ranking at r = 0.924 for specific embodiments and tasks `[EXP: +arXiv:2405.05941]`, which is a genuinely encouraging number and is not a licence to +generalize: a controlled re-evaluation states that "common simulation benchmarks are not a +reliable proxy for real world performance" `[EXP: CoRL 2023, arXiv:2310.09289]`, and +sim-screen-then-hardware-confirm on a new setup is an engineering bet rather than an +established result `[PRACTICE]`. + +*The specific trap at this rung.* Simulation numbers get quoted as hardware numbers. DP3's +"10 demonstrations" and Diffusion Policy's 200/300 counts are both simulation figures that +circulate as real-robot capability claims (5). + +### Rung 2 -- empty deck, reduced speed + +*What it establishes.* The policy's commanded trajectory stays inside the envelope, does not +drive into the deck, and does not saturate a joint. This is `Rung.DRY_RUN` from the printed +guide, applied to a policy instead of a part. + +*Preconditions.* Safety Boundary configured to the task box. Reduced Mode on. Collision +detection at a sensitivity you chose deliberately and not at 0. Payload configured for what is +actually on the flange. E-stop in reach of a hand that is on it. All of 7 in place before the +first autonomous step, because this is the first time nobody is driving. + +*What it does not establish.* Anything about contact. Nothing is there to contact. + +### Rung 3 -- fixture and labware, no material + +*What it establishes.* The policy's trajectory against real geometry. This is where the +printed fixture earns or loses its `FITTED` and `DRY_RUN` rungs from +`docs/PRINTED_FIXTURES.md` 4, and where a set of failures becomes available that rung 2 could +not reach: + +- the plate is not the nominal plate (SLAS footprint tolerance is +/-0.25 mm only near the + corners and +/-0.5 mm elsewhere; corner radius spans 1.58 to 4.78 mm); +- the lip does not retain the plate at angle, if a wedge is involved; +- the stack height puts the plate outside the Z envelope; +- the part shifts under arm contact -- the failure that puts the instrument at risk and that + no rung below this one can see. + +*And a policy-specific one:* if the fixture was not in the demonstrations, its appearance in +frame is now out of distribution, and the policy has never seen this image. That is a rung-3 +discovery and it is cheap here and expensive later. + +*What it does not establish.* Anything about material. A dry plate weighs differently and +spills nothing. + +### Rung 4 -- material, with an assay reading the outcome + +*What it establishes.* Whether the task was accomplished, as opposed to whether the arm moved. +The assay is the external criterion from 6, and it is what turns a run into a +`teaching.MachineObservation` -- a value with a metric, units, conditions, and evidence you can +go back to. Without it you have a recording of motion. + +*Requirements, which are 6 restated as a checklist.* Criterion defined and pre-registered. +Scored by someone blind to which policy ran, where that is possible. Reported as k out of n, +with the held-out factors named. Trial count chosen from the interval table, against the +decision you actually need to make. Checkpoint frozen before the first trial. Batch +pre-committed or a sequential procedure used. Judged on the worst observation, not the mean. + +*What it still does not establish.* That the result transfers to a different plate type, a +different liquid, a different volume, a different aspiration height, or a reprinted fixture. +Each of those is its own condition, in the sense `teaching.Envelope` refuses to pool: an +envelope is a range over **one** experiment, and pooling two widens it past the failure it +existed to catch. + +### The demotion rule + +Any of these sends the system back down, and the rung it goes back to is the one that first +tested the thing that changed: + +| event | back to | +| --- | --- | +| camera moved, remounted, or refocused | rung 2, and re-evaluate at rung 4 -- the policy did not read the calibration (2) | +| fixture reprinted, or filament/spool/lot changed | rung 3 -- it is a different part until measured again (4) | +| gripper, wrist camera, or payload changed | rung 2 -- collision detection depends on the dynamic model (7) | +| arm remounted, re-zeroed, or controller replaced | rung 2 -- friction parameters and mounting direction feed the same model | +| plate vendor or labware lot changed | rung 3 -- tolerances are per-product, not per-standard | +| new checkpoint | rung 4, with a fresh frozen checkpoint and a fresh trial batch | + +None of these is paranoia. Each names a variable that is inside the training distribution and +is not recorded anywhere unless you record it. + +--- + +## What this guide refuses to tell you + +Kept as a list rather than buried, because the gaps are the part most likely to get filled in +by somebody's search results. + +- **A success rate for any lab manipulation task.** None has been published. Every figure + quoted above is bound to the paper, task, hardware and trial count it came from, and none of + those is a plate on a deck. +- **A number of demonstrations for a new task.** Not knowable in advance. The published + real-robot range spans about 400x, the only controlled study finds diversity dominates count, + and the most-quoted paper ran no ablation on the question at all. +- **A demonstration-count multiplier for a physical fixture.** No study uses a physical fixture + as the manipulated variable against an unfixtured control. The measurement that would produce + one is a paired collection with and without the nest, on the same task, evaluated per 6. +- **A controlled measurement of the third-person viewpoint penalty specifically.** Searched + for; not found. Most successful video-to-action results use egocentric capture, depth, or + multi-view triangulation, so the third-person monocular case is the one with the least + evidence and the most enthusiasm. +- **A millimeter figure for human-to-robot-hand retargeting error.** One paper defines exactly + the right metrics and reports them only as bar charts with no numeric values in the text + reached. +- **A mapping from N mm of hand-eye calibration error to M mm of end-effector error on a + manipulation task.** Searched for specifically; not found. Papers analyze reprojection error, + or robot kinematic error, or policy success rate, and nothing closes that loop. The + small-angle amplification arithmetic in 2 is this guide's own calculation, offered as a check + to run rather than a published result. +- **How fast a hand-eye calibration drifts out of tolerance in a working cell.** No published + study. Every recalibration-frequency figure located is vendor guidance. +- **Whether recalibrating a moved camera restores a policy.** No study runs that experiment. + The inference that it does not, for pixel-space policies, rests on those architectures taking + no extrinsics input -- which is confirmed, but is not the same as the experiment. +- **Any independently measured accuracy, repeatability or payload validation for xArm 5/6/7.** + Everything traces to UFACTORY documentation or its own storefront. +- **A price for the arm.** Two of the vendor's own US web properties disagree. Get a quote. +- **Whether the arm's safety functions carry a Performance Level or Category rating**, and + whether any power-and-force-limiting validation exists. Zero hits for 15066, PL d, Category 3 + or safety-rated in a 254-page manual. +- **Per-task demonstration counts for pi-0 and pi-0.5.** They do not exist in the papers, which + report hours. Any per-task number attributed to them is from outside them. +- **Whether any published policy trained with zero action information anywhere in the pipeline + performs a real manipulation task at useful rates.** Searched; not found. Absence of evidence + from a targeted search, not proof of impossibility. +- **Whether any of the published success rates quoted here are statistically distinguishable + from each other.** Real-robot trial counts are small throughout -- 25, 20, and in one case 5 + -- and none of these papers reports confidence intervals on real-robot success except one + standard deviation across tasks in DP3. + +--- + +## Sources + +Capture, retargeting, and what video does not carry +- *Phantom: Training Robots Without Robots Using Only Human Videos*, CoRL 2025, + arXiv:2503.00779 -- teleop-vs-video exchange rate, per-task rates, and the paper's own scope + limits (pinch grasps only, quasi-static only, capped by the hand pose estimator). +- *What Matters When Cotraining Robot Manipulation Policies on Everyday Human Videos?*, + arXiv:2606.06627 (4 Jun 2026) -- cotraining gain versus robot-data scale; hand-pose noise + injection; monocular estimator substitution. Dated seven weeks before this writing and not + independently replicated. +- *Human2Sim2Robot*, CoRL 2025, arXiv:2504.12609 -- 20-30 cm hand pose errors; morphological + argument against direct retargeting. +- *LAPA: Latent Action Pretraining from Videos*, ICLR 2025, arXiv:2410.11758 -- action-free + pretraining, action-labeled finetuning, and the fine-motor deficit. +- *Latent Action Learning Requires Supervision in the Presence of Distractors*, ICML 2025, + arXiv:2502.00379 -- 2.5% action supervision, 4.2x downstream improvement. +- *UMI: Universal Manipulation Interface*, RSS 2024, arXiv:2402.10329 -- proprioception + ablation, capture precision, collection-rate comparison against bare hands. +- *DexMachina*, arXiv:2505.24853 and *SPIDER*, arXiv:2511.09484 -- kinematic retargeting is not + dynamically feasible, and the machinery built to repair it. +- Learning-from-observation theory: NeurIPS 2019, arXiv:1910.04417 (inverse-dynamics + disagreement bound); arXiv:2102.10769 (injectivity requirement). +- *Open X-Embodiment*, arXiv:2310.08864 -- cross-embodiment cotraining as a low-data rescue + that can be a tax on data-rich embodiments. + +Calibration and viewpoint +- IEEE Transactions on Robotics and Automation, 1989 -- AX=XB formulation and the + non-parallel-axis requirement. Accuracy figures from the original were not verified. +- PLOS ONE 2022, *Accuracy evaluation of hand-eye calibration techniques for vision-guided + robots* -- noise-dependent solver ranking; explicit statement that real-data ground truth is + unavailable. +- Frontiers in Robotics and AI 7:65 (2020) -- static versus motion-mode residuals and the + timestamp-synchronization attribution. +- arXiv:2303.06766 -- next-best-view pose selection; 4.74 to 2.95 mm with the algorithm fixed. +- Sensors 24(1):113 (2024) -- solver comparison and reprojection-error optimization. +- Sensors 22(24):9997 (2022) -- 1.49 px maximum thermal image drift, 5 to 45 C. +- *Decomposing the Generalization Gap in Imitation Learning for Visual Robotic Manipulation*, + ICRA 2024, arXiv:2307.03659 -- real-robot factor ranking, camera pose hardest. +- *LIBERO-Plus*, arXiv:2510.13626 -- camera perturbation collapse across nine VLAs (simulation). +- *THE COLOSSEUM*, RSS 2024, arXiv:2402.08191 -- 2D versus 3D model sensitivity to camera pose. + Its perturbation units could not be confirmed and are not reproduced here. +- *Do You Know Where Your Camera Is? View-Invariant Policy Learning with Camera Conditioning*, + ICRA 2026, arXiv:2510.02268 -- the shortcut mechanism, and confirmation that the standard + baselines take no extrinsics input. + +Demonstration counts +- ACT / ALOHA, RSS 2023, arXiv:2304.13705. Verified by direct PDF text extraction (counts, + Tables I and II). +- Diffusion Policy, RSS 2023 and arXiv:2303.04137v5 (IJRR extended). Verified by direct PDF + text extraction (Table 3, Section 7, Appendix C). Table 3 and Appendix C.2.1 contradict each + other on the sauce tasks. +- Mobile ALOHA, CoRL 2024, arXiv:2401.02117. Verified by direct PDF text extraction (Section + 6.1, Figure 4, Tables 3 and 4). +- DP3 / 3D Diffusion Policy, RSS 2024, arXiv:2403.03954. +- ALOHA Unleashed, CoRL 2024, arXiv:2410.13126 -- **single HTML fetch, not re-verified against + the PDF.** +- *Data Scaling Laws in Imitation Learning for Robotic Manipulation*, ICLR 2025, + arXiv:2410.18647. +- pi-0, arXiv:2410.24164; pi-0.5, arXiv:2504.16054 -- both report hours, not per-task counts. + +Fixtures and pose variance +- *ManiBox*, arXiv:2411.01850 -- spatial scaling law and its five measured points (simulation). +- *MimicGen*, CoRL 2023, arXiv:2310.17596 -- D0/D1/D2 reset distributions at fixed budget; + generation rates; the 10-demonstration source-agent results including 0% on real hardware. +- *Learning from 10 Demos: Generalisable and Sample-Efficient Policy Learning with Oriented + Affordance Frames*, CoRL 2025, arXiv:2410.12124. +- *Equivariant Diffusion Policy*, IJRR 2026, DOI 10.1177/02783649261424445. +- Discrete analogue only: NeurIPS 2020, *Toward the Fundamental Limits of Imitation Learning* + -- tabular BC bound linear in state-space size. Cited as an analogue, not as support for a + continuous-pose claim. + +Evaluation +- *Robot Learning as an Empirical Science: Best Practices for Policy Evaluation*, + arXiv:2409.09491 (Sept 2024). +- *A Careful Examination of Large Behavior Models for Multitask Dexterous Manipulation*, + arXiv:2507.05331 (July 2025; also Science Robotics) -- blind protocol, roughly 1,800 real + trials, the 27% QA audit. +- *Is Your Imitation Learning Policy Better than Mine? Policy Comparison with Near-Optimal + Stopping*, RSS 2025, arXiv:2503.10966 -- the p-hacking warning; STEP. +- *Beyond Binary Success*, arXiv:2603.13616 (March 2026) -- anytime-valid inference; graded + scores separate policies faster. +- *Evaluating Real-World Robot Manipulation Policies in Simulation* (SIMPLER), CoRL 2024, + arXiv:2405.05941 -- validation MSE at Pearson r = 0.308. +- *Deep Reinforcement Learning that Matters*, AAAI 2018, arXiv:1709.06560 -- seed variance; + explicit refusal to name a minimum trial count. +- *RoboArena*, CoRL 2025, arXiv:2506.18123 -- distributed double-blind pairwise evaluation as a + check on single-lab distributions. +- *PhAIL*, arXiv:2605.29710 -- the 13-paper survey of trial counts. Single-author preprint, + convenience sample. +- CoRL 2023, arXiv:2310.09289 -- pretraining image distribution matters more than scale, and + simulation benchmarks are not a reliable proxy for real-world performance. Full text was not + reachable; abstract- and summary-level claims only. +- A GPU vendor's engineering blog, 2026-07-11, on evaluating general-purpose robot policies -- + used only for its Clopper-Pearson figures, which were independently reproduced here. Its + benchmark details were not verified. + +Arm documentation, read directly +- xArm User Manual V2.0.0 (254 pp), text extracted locally. Appendix 2 Technical Specifications + pp. 217-221; E-stop p. 23; Dedicated Safety I/O p. 52; Maximum Payload p. 237; Appendix 7 + Applied Standards pp. 231-233. +- UFACTORY 6 Axis Force Torque Sensor manual V2.5.2. Section 5 Technical Specifications p. 12; + Section 3.1 SDK control p. 9; Section 1.2 overload values and risk assessment p. 3. +- UFACTORY support article on current- and dynamic-model-based collision detection. +- `xArm-Developer/xArm-Python-SDK` -- BSD 3-Clause LICENSE and `doc/api/xarm_api.md`. + +Repo cross-references +- `docs/PRINTED_FIXTURES.md` -- 1 the contact boundary, 4 the qualification ladder, 5 shrinkage + and the failure modes, 6 the tilt module and its acceptance test. +- `hardware/README.md`, `hardware/tilt_module.scad` -- the worked fixture, its parameters, and + the replicate design this guide's rung 4 reuses. +- `autonomous_lab/printed.py` -- `Rung`, `DimensionState`, `Contact`, and the refusal table. +- `autonomous_lab/teaching.py` -- `MIN_DEMONSTRATIONS`, `Envelope`, `Attainment`, and the rule + that one good machine run is not parity. +- `autonomous_lab/vision.py` -- `VisionRequirement.VALIDATION`, and why an unmeasured detector + makes a lab less safe than none. +- `autonomous_lab/qc.py` -- `Basis`, and the broken absorbance read that disqualifies an optical + endpoint here. diff --git a/tests/test_feedback.py b/tests/test_feedback.py index 8ee1a82..5b69656 100644 --- a/tests/test_feedback.py +++ b/tests/test_feedback.py @@ -332,7 +332,7 @@ def test_measurement_availability_is_resolved_from_the_ledger_not_asserted(): def test_a_broken_instrument_and_an_unwired_one_refuse_alike_and_report_differently(): - """Wiring plr-tested in changes the reason and not the verdict. + """Wiring a plr-tested checkout in changes the reason and not the verdict. Both are UNAVAILABLE_MEASUREMENT because in both cases no number arrives. The reason is carried through from the ledger because the next action differs completely: one is a @@ -343,7 +343,7 @@ def test_a_broken_instrument_and_an_unwired_one_refuse_alike_and_report_differen unwired = build_ledger(protocol) wired_wc = Workcell.default() - wired_wc.plr_tested_root = "/nonexistent/plr-tested" + wired_wc.plr_tested_root = "/nonexistent/run-cards" wired = build_ledger(protocol, wired_wc) reader = next(i for i, s in enumerate(protocol.steps) if s.op == "read_absorbance") diff --git a/tests/test_imitation.py b/tests/test_imitation.py new file mode 100644 index 0000000..8e35114 --- /dev/null +++ b/tests/test_imitation.py @@ -0,0 +1,756 @@ +"""Device-free tests for the imitation layer: capture, coupling, and trust. + +The tests that matter here try to get a policy reported as trainable, equivalent to +teleoperation, or trusted with material when it is none of those, and the three richest +sources of that are the three the module was written against. + +The first is the capture that looks like a demonstration and is not. A video with no action +stream. A retargeting pipeline declared but never measured, which reads in a manifest exactly +like one that was. A capture with nothing set at all, where every default that means "not +checked" has to keep meaning that rather than quietly meaning "fine". + +The second is the fixture. A policy learned against a specific printed part, and the part is +the easiest thing in the workcell to change without telling anybody: reprint it with a tuned +clearance and every demonstration behind the policy describes a world that no longer exists. +A revision bump must decouple. A matching revision string on two parts nobody measured must +not couple, because a revision names a model file and `printed` refuses to estimate shrinkage +at all. + +The third is the success rate that is not one. Scored on the conditions the demonstrations +came from. Scored by the policy's own confidence. Computed off five trials, where one trial +is twenty percentage points. Each of those has to come back refused rather than reported with +a caveat beside it, because the number is what gets quoted and the caveat is not. + +The interval arithmetic is checked against published Clopper-Pearson values, since it is the +one thing in the module that can be wrong silently. +""" + +from __future__ import annotations + +import pathlib + +import pytest + +from autonomous_lab import imitation +from autonomous_lab.imitation import ( + Interlock, + DETERMINANTS, + INTERLOCK_NOTE, + MAX_TRIAL_WEIGHT_PP, + MIN_HELDOUT_TRIALS, + MODALITY_EVIDENCE, + NOT_ESTIMATED, + PUBLISHED_COUNTS, + Capture, + Coupling, + Evaluation, + FixtureCoupling, + Modality, + Policy, + Retargeting, + SuccessCriterion, + Trainable, + Trust, + demonstrations_needed, + equivalent_to_teleop, + fixture_coupled, + may_run, + published_span, + refusals, + success_interval, + success_rate, + summary, + trusted_with_material, + usable_for_training, +) +from autonomous_lab.printed import ( + Basis, + CriticalDimension, + DimensionState, + Fixture, + MATERIALS, +) + + +def _measured_part(name="nest", value=128.4): + return Fixture( + name=name, + material=MATERIALS["pla_fdm"], + dimensions=( + CriticalDimension( + name="pocket_length", + nominal_mm=128.5, + state=DimensionState.MEASURED, + measured_mm=value, + measured_after_processing=True, + instrument="digital caliper, three positions", + ), + ), + ) + + +def _designed_part(name="nest"): + return Fixture( + name=name, + material=MATERIALS["pla_fdm"], + dimensions=(CriticalDimension(name="pocket_length", nominal_mm=128.5),), + ) + + +def _coupling(revision="rev_a", measured=True, fixture_name="plate_nest"): + return FixtureCoupling( + fixture_name=fixture_name, + revision=revision, + fixture=_measured_part() if measured else _designed_part(), + constrains=("x", "y", "yaw"), + ) + + +def _teleop_capture(**kw): + """A capture with everything a lab controls already done, so a test can vary one thing.""" + defaults = dict( + name="teleop_session", + modality=Modality.TELEOPERATION, + episodes=50, + timestamps_synchronized=True, + camera_calibrated=True, + gripper_state_recorded=True, + coupling=_coupling(), + ) + defaults.update(kw) + return Capture(**defaults) + + +def _validated_retargeting(): + return Retargeting( + method="hand pose estimation, then ICP against the depth cloud", + depth_or_multiview=True, + error_measured_mm=9.4, + basis=Basis.IN_HOUSE, + ) + + +def _video_capture(modality=Modality.INSTRUMENTED_HUMAN_VIDEO, retargeting=None, **kw): + defaults = dict( + name="video_session", + modality=modality, + episodes=300, + retargeting=retargeting if retargeting is not None else _validated_retargeting(), + timestamps_synchronized=True, + camera_calibrated=True, + gripper_state_recorded=True, + coupling=_coupling(), + ) + defaults.update(kw) + return Capture(**defaults) + + +def _measured_evaluation(successes=17, trials=20, **kw): + defaults = dict( + trials=trials, + successes=successes, + held_out=True, + criterion=SuccessCriterion.EXTERNAL_MEASUREMENT, + blinded=True, + ) + defaults.update(kw) + return Evaluation(**defaults) + + +def _policy(**kw): + defaults = dict( + name="pick_and_nest", + method="action chunking", + captures=(_teleop_capture(),), + evaluation=_measured_evaluation(), + ) + defaults.update(kw) + return Policy(**defaults) + + +# -- vacuous passes ----------------------------------------------------------------- + + +def _interlock(**kw): + """A complete interlock. Every field must be set, so the helper states them explicitly + rather than relying on defaults that would quietly re-open the hole this leg closed.""" + base = dict( + workspace_bounded=True, + speed_bounded=True, + force_bounded=True, + enforced_by="controller soft limits and a hardware E-stop", + miss_rate_measured=True, + ) + base.update(kw) + return Interlock(**base) + + +def test_a_capture_with_nothing_set_is_not_usable(): + """The single most important test in the file. + + A dataclass of booleans and empty defaults is the easiest thing here to get wrong, because + every default that reads as True or as 'nothing to check' turns a recording nobody looked + at into a training set. A bare capture has to come back refused, and it has to name every + one of the things it lacks rather than only the first. + """ + standing = usable_for_training(Capture()) + assert not standing.usable + assert standing.verdict is not Trainable.USABLE + for reason in ( + Trainable.MODALITY_UNDECLARED, + Trainable.NO_EPISODES, + Trainable.NO_ACTION_STREAM, + Trainable.TIMESTAMPS_UNSYNCHRONIZED, + Trainable.CAMERA_UNCALIBRATED, + Trainable.GRIPPER_STATE_MISSING, + ): + assert reason in standing.reasons() + + +def test_a_policy_with_no_captures_does_not_read_as_trained_on_flawless_data(): + """`all()` over an empty tuple is True; a policy nobody fed must not inherit that.""" + bare = Policy() + assert not bare.trainable() + assert bare.episodes() == 0 + assert bare.training_coupling() is None + assert not may_run(bare, _coupling()).permitted + + +def test_a_bare_capture_is_not_equivalent_to_anything(): + ok, why = equivalent_to_teleop(Capture()) + assert not ok + assert "no modality is declared" in why + + +# -- a video is not a demonstration -------------------------------------------------- + + +def test_a_capture_with_no_action_stream_is_refused(): + """Pixels are the observation half. A capture that carries no actions trains nothing.""" + pixels_only = _video_capture( + modality=Modality.MONOCULAR_HUMAN_VIDEO, retargeting=Retargeting() + ) + assert not pixels_only.carries_actions + + standing = usable_for_training(pixels_only) + assert not standing.usable + assert Trainable.NO_ACTION_STREAM in standing.reasons() + detail = standing.details()[standing.reasons().index(Trainable.NO_ACTION_STREAM)] + assert "no action stream" in detail + assert "injective" in detail # the case where the action is unrecoverable in principle + assert Trainable.NO_ACTION_STREAM.needs_recapture + + +def test_retargeted_video_does_not_report_as_equivalent_to_teleop(): + """The refusal that keeps an exchange rate from becoming a silent multiplier. + + It has to hold for the BEST version of the video pipeline, not only the worst: depth, + multi-view, a measured retargeting error, in-house basis. That capture is usable for + training and is still not a teleoperated episode. + """ + best_video = _video_capture() + assert usable_for_training(best_video).usable + + ok, why = equivalent_to_teleop(best_video) + assert not ok + assert "no measured action stream" in why + assert "a measured estimator is still an estimator" in why + + worst_video = _video_capture( + modality=Modality.MONOCULAR_HUMAN_VIDEO, + retargeting=Retargeting(method="monocular hand pose", depth_or_multiview=False), + ) + assert not equivalent_to_teleop(worst_video)[0] + + +def test_a_handheld_gripper_is_not_teleoperation_either(): + """The nearest miss in the enum, and the one most likely to be filed as equivalent.""" + assert Modality.HANDHELD_GRIPPER.measured_action_stream + assert not Modality.HANDHELD_GRIPPER.on_the_robots_own_kinematics + + ok, why = equivalent_to_teleop(_teleop_capture(modality=Modality.HANDHELD_GRIPPER)) + assert not ok + assert "proprioception" in why + + +def test_only_captures_on_the_robots_own_kinematics_report_equivalent(): + for modality in Modality: + ok, _ = equivalent_to_teleop(_teleop_capture(modality=modality)) + assert ok is modality.on_the_robots_own_kinematics + assert sum(1 for m in Modality if m.on_the_robots_own_kinematics) == 2 + assert sum(1 for m in Modality if m.measured_action_stream) == 3 + + +def test_a_declared_retargeting_that_nobody_measured_is_refused(): + """A pipeline that exists is not a pipeline that was validated.""" + unmeasured = _video_capture( + retargeting=Retargeting(method="hand pose then wrist mapping", depth_or_multiview=True) + ) + standing = usable_for_training(unmeasured) + assert not standing.usable + assert Trainable.RETARGETING_UNVALIDATED in standing.reasons() + assert not Trainable.RETARGETING_UNVALIDATED.needs_recapture # measure it; do not recollect + + +def test_somebody_elses_retargeting_error_does_not_validate_this_one(): + borrowed = Retargeting( + method="published pipeline", + depth_or_multiview=True, + error_measured_mm=9.4, + basis=Basis.LITERATURE, + ) + assert not borrowed.validated + standing = usable_for_training(_video_capture(retargeting=borrowed)) + assert Trainable.RETARGETING_UNVALIDATED in standing.reasons() + + +def test_a_capture_cannot_both_measure_and_estimate_its_actions(): + with pytest.raises(ValueError) as excinfo: + Capture( + name="incoherent", + modality=Modality.TELEOPERATION, + episodes=10, + retargeting=_validated_retargeting(), + ) + assert "cannot both" in str(excinfo.value) + + +def test_each_refusal_names_what_specifically_is_missing(): + """One field off at a time, so no refusal can be hiding behind another.""" + cases = { + "timestamps_synchronized": (Trainable.TIMESTAMPS_UNSYNCHRONIZED, "common clock"), + "camera_calibrated": (Trainable.CAMERA_UNCALIBRATED, "camera pose"), + "gripper_state_recorded": (Trainable.GRIPPER_STATE_MISSING, "gripper state"), + } + for field_name, (reason, phrase) in cases.items(): + standing = usable_for_training(_teleop_capture(**{field_name: False})) + assert standing.reasons() == (reason,), field_name + assert phrase in standing.details()[0] + + +def test_the_camera_refusal_resolves_the_requirement_vision_already_owns(): + standing = usable_for_training(_teleop_capture(camera_calibrated=False)) + assert "pose" in standing.details()[0] + + +def test_a_fully_specified_teleop_capture_is_usable_and_claims_nothing_more(): + standing = usable_for_training(_teleop_capture()) + assert standing.usable + assert standing.verdict is Trainable.USABLE + assert "not a trained policy" in standing.reason + + +def test_every_modality_carries_its_evidence(): + assert set(MODALITY_EVIDENCE) == set(Modality) + assert all(len(v) > 100 for v in MODALITY_EVIDENCE.values()) + + +# -- the fixture is part of the policy's world ---------------------------------------- + + +def test_a_fixture_revision_change_decouples_an_existing_policy(): + """The finding the module exists for: nothing in the run reports this on its own.""" + policy = _policy() + check = fixture_coupled(policy, _coupling(revision="rev_b")) + assert not check.holds + assert check.state is Coupling.REVISED + assert "rev_b" in check.reason + assert "shrinkage" in check.reason + + +def test_a_matching_revision_on_two_parts_nobody_measured_does_not_couple(): + """A revision names a model file. The policy was trained against a part.""" + policy = _policy(captures=(_teleop_capture(coupling=_coupling(measured=False)),)) + check = fixture_coupled(policy, _coupling(measured=False)) + assert not check.holds + assert check.state is Coupling.UNVERIFIED + assert "fact about a model file" in check.reason + + +def test_a_measured_pair_at_the_same_revision_couples(): + check = fixture_coupled(_policy(), _coupling()) + assert check.holds + assert check.state is Coupling.COUPLED + assert "says nothing about whether it learned it" in check.reason + + +def test_a_different_fixture_is_not_a_degraded_coupling(): + check = fixture_coupled(_policy(), _coupling(fixture_name="tube_rack")) + assert check.state is Coupling.DIFFERENT_FIXTURE + + +def test_captures_against_two_revisions_couple_to_neither(): + """The same refusal `teaching` makes about pooling two sets of conditions into one range.""" + policy = _policy( + captures=( + _teleop_capture(name="monday", coupling=_coupling(revision="rev_a")), + _teleop_capture(name="tuesday", coupling=_coupling(revision="rev_b")), + ) + ) + assert policy.training_coupling() is None + assert len(policy.distinct_couplings()) == 2 + check = fixture_coupled(policy, _coupling(revision="rev_a")) + assert not check.holds + assert check.state is Coupling.MIXED + + +def test_a_policy_that_names_no_fixture_cannot_be_shown_to_have_changed_one(): + policy = _policy(captures=(_teleop_capture(coupling=None),)) + check = fixture_coupled(policy, _coupling()) + assert check.state is Coupling.UNIDENTIFIED + assert "cannot be reproduced" in check.reason + + +def test_a_proposed_run_naming_no_fixture_does_not_couple(): + assert fixture_coupled(_policy(), None).state is Coupling.UNIDENTIFIED + unnamed = FixtureCoupling(fixture_name="plate_nest") # no revision + assert not unnamed.identified + assert fixture_coupled(_policy(), unnamed).state is Coupling.UNIDENTIFIED + + +def test_coupling_anchors_on_printed_rather_than_on_a_boolean_here(): + """The measured-ness of the part is `printed`'s question and is resolved there.""" + anchored = _coupling() + assert anchored.anchored + assert not _coupling(measured=False).anchored + assert not FixtureCoupling(fixture_name="a", revision="b").anchored + # A fixture that declares no dimensions at all must not read as measured either. + bare = FixtureCoupling(fixture_name="a", revision="b", fixture=Fixture()) + assert not bare.anchored + + +def test_only_one_coupling_state_holds(): + assert sum(1 for c in Coupling if c.holds) == 1 + + +# -- a policy is untrusted until measured ---------------------------------------------- + + +def test_a_training_set_success_rate_is_not_a_success_rate(): + """The number that gets quoted, refused at both layers that could report it.""" + policy = _policy(evaluation=_measured_evaluation(successes=48, trials=50, held_out=False)) + report = success_rate(policy.evaluation) + assert not report.stated + assert report.point is None + assert "not a success rate" in report.reason + + standing = trusted_with_material(policy) + assert not standing.trusted + assert standing.verdict is Trust.TRAINING_SET_ONLY + + +def test_a_policy_scored_by_its_own_confidence_is_not_trusted(): + for criterion in (SuccessCriterion.POLICY_CONFIDENCE, SuccessCriterion.TRAINING_LOSS): + policy = _policy( + evaluation=_measured_evaluation(successes=95, trials=100, criterion=criterion) + ) + standing = trusted_with_material(policy) + assert not standing.trusted + assert standing.verdict is Trust.SELF_SCORED + assert standing.rate.point is None + assert not criterion.external + + +def test_a_thousand_self_scored_trials_do_not_substitute_for_an_external_criterion(): + policy = _policy( + evaluation=_measured_evaluation( + successes=900, trials=1000, criterion=SuccessCriterion.POLICY_CONFIDENCE + ) + ) + assert trusted_with_material(policy).verdict is Trust.SELF_SCORED + + +def test_too_few_trials_is_refused_rather_than_given_a_rate(): + """At 5 trials one trial is 20 points, which is the published row where 40 percent is 2.""" + policy = _policy(evaluation=_measured_evaluation(successes=2, trials=5)) + report = success_rate(policy.evaluation) + assert not report.stated + assert report.point is None and report.low is None and report.high is None + assert "20.0 percentage points" in report.reason + assert "convention" in report.reason # the floor is declared, not discovered + + standing = trusted_with_material(policy) + assert not standing.trusted + assert standing.verdict is Trust.UNDERPOWERED + + +def test_an_unevaluated_policy_is_untrusted_by_default(): + assert trusted_with_material(_policy(evaluation=None)).verdict is Trust.UNEVALUATED + assert trusted_with_material(_policy(evaluation=Evaluation())).verdict is Trust.UNEVALUATED + + +def test_a_measured_policy_reports_an_interval_rather_than_a_point_alone(): + policy = _policy(evaluation=_measured_evaluation(successes=17, trials=20)) + standing = trusted_with_material(policy, acceptable_rate=0.50) + assert standing.trusted + assert standing.verdict is Trust.MEASURED + assert standing.rate.stated + assert standing.rate.point == pytest.approx(0.85) + assert "interval" in standing.rate.describe() + assert "17 of 20" in standing.rate.describe() + assert standing.rate.width_pp is not None and standing.rate.width_pp > 30.0 + + +def test_only_one_trust_state_is_trusted(): + assert sum(1 for t in Trust if t.evidence_complete) == 1 + + +def test_blinding_is_reported_and_not_required(): + """Requiring it would be inventing a threshold out of hygiene nobody has measured.""" + unblinded = _policy(evaluation=_measured_evaluation(blinded=False)) + standing = trusted_with_material(unblinded, acceptable_rate=0.50) + assert standing.trusted + assert "unblinded" in standing.rate.reason + + +def test_an_evaluation_cannot_report_more_successes_than_trials(): + with pytest.raises(ValueError) as excinfo: + Evaluation(trials=10, successes=11, held_out=True) + assert "not an observation" in str(excinfo.value) + with pytest.raises(ValueError): + Evaluation(trials=-1) + + +def test_the_interval_matches_published_clopper_pearson_values(): + """The one thing here that could be wrong silently, checked against known values.""" + expected = { + (9, 10): (55.5, 99.7), + (18, 20): (68.3, 98.8), + (45, 50): (78.2, 96.7), + (90, 100): (82.4, 95.1), + (180, 200): (85.0, 93.8), + (900, 1000): (88.0, 91.8), + (5, 10): (18.7, 81.3), + (10, 20): (27.2, 72.8), + (25, 50): (35.5, 64.5), + (50, 100): (39.8, 60.2), + } + for (successes, trials), (low, high) in expected.items(): + got_low, got_high = success_interval(successes, trials) + assert round(100 * got_low, 1) == pytest.approx(low, abs=0.05) + assert round(100 * got_high, 1) == pytest.approx(high, abs=0.05) + + +def test_the_interval_refuses_impossible_observations(): + with pytest.raises(ValueError): + success_interval(1, 0) + with pytest.raises(ValueError): + success_interval(11, 10) + assert success_interval(0, 20)[0] == 0.0 + assert success_interval(20, 20)[1] == 1.0 + + +def test_the_trial_floor_is_derived_from_the_declared_ceiling(): + """The constant is computed from a stated convention rather than picked and defended.""" + assert MIN_HELDOUT_TRIALS == int(round(100.0 / MAX_TRIAL_WEIGHT_PP)) + assert 100.0 / MIN_HELDOUT_TRIALS <= MAX_TRIAL_WEIGHT_PP + + +# -- how many demonstrations ------------------------------------------------------------- + + +def test_demonstrations_needed_returns_no_number(): + estimate = demonstrations_needed("place a plate in a printed nest") + assert estimate.estimate is None + assert isinstance(estimate.refusal.why, str) + assert "not knowable in advance" in estimate.refusal.why + assert estimate.refusal.what_it_would_take + assert estimate.determinants == DETERMINANTS and len(DETERMINANTS) >= 4 + + +def test_the_published_table_names_its_tasks_and_methods_and_spans_two_orders_of_magnitude(): + low, high = published_span() + assert (low, high) == (20, 8658) + assert high / low > 100 + assert all(row.system and row.task and row.source for row in PUBLISHED_COUNTS) + described = demonstrations_needed("anything").describe() + assert "20 to 8658" in described + + +def test_the_table_keeps_the_row_that_contradicts_the_usual_reading(): + """Twice the demonstrations of any other task in its paper, and the worst result in it.""" + velcro = [r for r in PUBLISHED_COUNTS if r.task == "Thread Velcro"] + assert velcro and velcro[0].demonstrations == 100 and velcro[0].success == 0.20 + # And the row whose own paper contradicts itself keeps a null count rather than a guess. + unresolved = [r for r in PUBLISHED_COUNTS if r.demonstrations is None] + assert unresolved and all("contradict" in r.note or "HOURS" in r.note for r in unresolved) + + +def test_every_published_success_rate_carries_the_trial_count_where_the_paper_gave_one(): + """A rate off 5 rollouts and a rate off 25 are not the same kind of number.""" + cheap = [r for r in PUBLISHED_COUNTS if r.trials is not None and r.trials < 10] + assert cheap and all(r.note for r in cheap) + + +# -- composition and refusals -------------------------------------------------------------- + + +def test_a_trusted_policy_on_a_revised_fixture_may_not_run(): + """Each leg is answerable alone and none of them is the question a lab asks.""" + policy = _policy() + assert may_run(policy, _coupling(), acceptable_rate=0.50, interlock=_interlock()).permitted + + revised = may_run( + policy, _coupling(revision="rev_b"), acceptable_rate=0.50, interlock=_interlock() + ) + assert not revised.permitted + assert any("fixture coupling" in b for b in revised.blockers()) + assert revised.trust.trusted # the policy is fine; the world changed + + +def test_a_measured_policy_trained_on_an_unusable_capture_may_not_run(): + policy = _policy(captures=(_teleop_capture(gripper_state_recorded=False),)) + standing = may_run(policy, _coupling(), acceptable_rate=0.50, interlock=_interlock()) + assert not standing.permitted + assert any("not usable for training" in b for b in standing.blockers()) + + +def test_may_run_reports_every_blocker_rather_than_the_first(): + policy = Policy(name="nothing", captures=(Capture(),)) + standing = may_run(policy, None) + assert not standing.permitted + assert len(standing.blockers()) >= 3 + + +def test_the_refusals_are_returned_as_data(): + assert refusals() is NOT_ESTIMATED + assert len(NOT_ESTIMATED) >= 5 + for r in NOT_ESTIMATED: + assert r.quantity and r.why and r.what_it_would_take + quantities = " ".join(r.quantity for r in NOT_ESTIMATED) + assert "demonstrations required for a new task" in quantities + assert "exchange rate" in quantities + + +def test_the_module_exposes_no_helper_that_returns_a_demonstration_count(): + """A module that refuses a number in prose and returns one elsewhere has no discipline.""" + for name in imitation.__all__: + obj = getattr(imitation, name) + assert not name.startswith("estimate_") + assert getattr(obj, "__name__", "") not in ("demos_for", "how_many_demos") + + +def test_the_authority_boundary_is_referenced_in_prose_and_not_imported(): + assert "interlock" in INTERLOCK_NOTE.lower() + assert "proposes" in INTERLOCK_NOTE.lower() and "disposes" in INTERLOCK_NOTE.lower() + source = open(imitation.__file__, encoding="ascii").read() + assert "import authority" not in source + assert "from .authority" not in source + + +def test_the_module_is_ascii(): + """House rule, enforced here as well as in CI, because it is cheap to check and easy to break.""" + with open(imitation.__file__, "rb") as handle: + body = handle.read() + body.decode("ascii") + + +def test_the_summary_counts_what_it_says_it_counts(): + counts = summary() + assert counts["modalities"] == len(Modality) + assert counts["modalities_with_a_measured_action_stream"] == 3 + assert counts["published_task_rows_with_a_count"] < counts["published_task_rows"] + assert counts["min_heldout_trials"] == MIN_HELDOUT_TRIALS + + +# -- what the adversarial audit found, and what must stay found ------------------ + + +def test_a_policy_that_failed_every_trial_is_not_trusted(): + """The audit's first blocker, and the worst kind of bug this package can have. + + A policy evaluated correctly -- held out, externally scored, twenty trials -- that + succeeded in NONE of them came back trusted, because `trusted` was `verdict is MEASURED` + and the evaluation had indeed been conducted properly. Every word of the Trust docstring + said MEASURED does not mean good; the boolean said the opposite to anyone who read it + instead of the prose, and it is the boolean that gets branched on. + """ + failed = _policy(evaluation=_measured_evaluation(successes=0, trials=20)) + standing = trusted_with_material(failed, acceptable_rate=0.50) + assert standing.verdict is Trust.MEASURED # the evaluation WAS done properly + assert standing.evidence_complete + assert not standing.trusted + assert not may_run(failed, _coupling(), acceptable_rate=0.50, interlock=_interlock()).permitted + + +def test_without_a_declared_acceptance_rate_the_question_is_refused(): + """What rate is acceptable is a property of the task, not of this module. A tolerable + failure rate for a tip pickup a retry fixes is not one for a transfer that consumes the + last of a sample. Defaulting it would answer, invisibly, the only question here that + requires knowing what the robot is holding.""" + good = _policy(evaluation=_measured_evaluation(successes=98, trials=100)) + standing = trusted_with_material(good) + assert standing.evidence_complete + assert not standing.trusted + assert "no acceptable rate was declared" in standing.reason + + +def test_acceptance_is_judged_on_the_interval_not_the_point_estimate(): + """A point estimate is what the policy did on those trials. The lower bound is what the + trials support saying about the next one, and the gap is why the trial count is reported.""" + wide = _policy(evaluation=_measured_evaluation(successes=18, trials=20)) # point 90 percent + narrow = _policy(evaluation=_measured_evaluation(successes=90, trials=100)) # point 90 percent + assert not trusted_with_material(wide, acceptable_rate=0.80).trusted + assert trusted_with_material(narrow, acceptable_rate=0.80).trusted + + +def test_a_run_needs_a_bound_that_lives_outside_the_policy(): + """A learned policy carries no guarantee about an input it has not seen, so the bound on + what it can reach, how fast and how hard cannot come from the policy or from a model + checking the policy.""" + policy = _policy() + assert not may_run(policy, _coupling(), acceptable_rate=0.50).permitted + assert any("interlock" in b for b in may_run(policy, _coupling(), acceptable_rate=0.50).blockers()) + + # An interlock whose own miss rate nobody measured is a belief, not a safety device. + unmeasured = _interlock(miss_rate_measured=False) + standing = may_run(policy, _coupling(), acceptable_rate=0.50, interlock=unmeasured) + assert not standing.permitted + assert any("miss rate" in b for b in standing.blockers()) + + +def test_an_unrecorded_fixture_is_not_agreement_with_a_recorded_one(): + """The audit's fourth finding. `distinct_couplings` skipped a capture whose coupling was + None, so one fixtured session plus two unrecorded ones reported a single clean coupling. + Two thirds of the training data came from a world nobody wrote down.""" + fixtured = _teleop_capture(coupling=_coupling()) + unrecorded = _teleop_capture(coupling=None) + policy = _policy(captures=(fixtured, unrecorded, unrecorded)) + assert policy.training_coupling() is None + assert not fixture_coupled(policy, _coupling()).holds + + # Two captures against the SAME recorded revision still couple. + assert _policy(captures=(fixtured, fixtured)).training_coupling() is not None + + +def test_a_retargeting_error_must_be_a_positive_distance(): + """Zero and negative are the signature of a defaulted or sign-flipped value rather than a + measurement, in the one field the module says is almost never filled in.""" + for bad in (0.0, -5.0): + with pytest.raises(ValueError) as excinfo: + Retargeting(method="x", error_measured_mm=bad, basis=Basis.IN_HOUSE) + assert "not a measurement" in str(excinfo.value) + assert Retargeting(method="x", error_measured_mm=12.5, basis=Basis.IN_HOUSE).validated + + +def test_no_success_figure_is_quoted_as_a_bare_percentage(): + """The module quoted 91.7 percent against 45.8 percent five times as established. Those + are 11 of 12 and 11 of 24 -- trial counts this module's own `success_rate` refuses, since + one trial at n=12 moves the rate by 8.3 points. A module that refuses the caller's + 12-trial rate while quoting its own has no discipline.""" + import autonomous_lab.imitation as module + + source = pathlib.Path(module.__file__).read_text() + for stripped in ("91.7", "45.8", "52.8", "76.5"): + assert stripped not in source, f"{stripped} is a point estimate quoted without its count" + assert "11 of 12" in source and "11 of 24" in source + + +def test_the_simulation_result_is_marked_as_simulation(): + """It was welded onto a real-robot sentence with an 'and', cherry-picked to the worst of + nine models. The guide names that exact move as the trap it warns about.""" + import autonomous_lab.imitation as module + + source = pathlib.Path(module.__file__).read_text() + assert "SIMULATION" in source + assert "nine models" in source