Skip to content

QuantityRecord: PyUnitWizard's native inert interchange form (verified quantity records) #82

Description

@dprada

Resolution (2026-10-04). Delivered homogeneous QuantityRecord/Bundle MVP and design record resolved. API and qrec/0.3 formats remain provisional; extensions have independent owners #101–#106. Published receiving cases: Sabueso #32 and TopoMT sealed DFND inputs. MolSysMT #240 is legacy unit validation, not codec adoption.

Record. completed_proposals/quantity_record.md


What. Proposed: a canonical, backend-independent serialization of quantities, owned by PyUnitWizard. devguide/interop_future_directions.md §4 lists it as a post-1.0 candidate ("Evaluate a canonical serialization/deserialization API for quantities (value, unit, dimensionality, form) for JSON/YAML/data pipelines"). This issue brings a concrete consumer use case and asks to raise its priority.

Use case. MOLI needs every component to write quantities with their unit and to read them back safely (uibcdf/moli, quantity interchange contract, filed alongside this issue). Sabueso cards are JSON/SQLite documents that must persist concentrations, masses, areas and distances (uibcdf/sabueso#32). MolSysMT writes units as H5MSM dataset attributes, and MolSysViewer sends bare numbers to its frontend. With no shared codec, each component invents its own shape. Some already drop the unit, and the reader then assumes one. In the MolSysViewer case that assumption is wrong by 10× under a non-default policy (reported in uibcdf/molsysviewer).

Sketch (names and shape are PyUnitWizard's decision):

record = puw.to_record(q)          # {"value": 3.0, "unit": "nanomolar"}
                                   # arrays -> nested lists; unit in canonical long form
q = puw.from_record(record, dimensionality={"[N]": 1, "[L]": -3})
                                   # parses the unit, checks dimensionality, fails loudly

Properties worth guaranteeing:

  • Lossless round trip for scalars and n-D arrays, integer and float dtypes. Today the string form does not round-trip arrays (reported alongside this issue).
  • Independent of the session. It never goes through standard units or the default form on write; from_record returns the session's default form.
  • Canonical unit spelling. Unambiguous long names rather than case-sensitive symbols (nM/nm), and stable across pint versions.
  • Optional expected dimensionality on read, so a unit error surfaces at the boundary.
  • JSON-safe. Plain dicts, lists and numbers, with no backend objects and no pickling.

Why PyUnitWizard. It already owns parsing, forms and the shared kernel. A codec written in each consumer would diverge in exactly the details (spelling, arrays, precision) that make values misread.

Activity

  1. dprada commented on Sep 24, 2026

    @dprada
    CollaboratorAuthor

    Platform context: uibcdf/moli#12 (quantity interchange contract). Concrete misread this would prevent: uibcdf/molsysviewer#96.

  2. changed the title [-]Canonical value-unit serialization of quantities for cross-component data[/-] [+]Quantity serialization codec: negotiated containers with an integrity digest[/+] on Sep 24, 2026
  3. dprada commented on Sep 24, 2026

    @dprada
    CollaboratorAuthor

    Design updated. This issue is now the implementation project for the design recorded in #83.

    The per-value to_record/from_record sketch above is superseded. The chosen design, with the measurements and the alternatives discarded, is in #83:

    • Negotiated containers. The unit is written once, inside the same serialized object as the values.
    • Integrity digest. blake2b-128 over the canonical manifest plus the values' little-endian bytes. Any change made outside the codec fails on read.
    • The codec is the only writer and reader. It accepts quantities only, converts with an explicit to_unit, and has no default unit.
    • Reader handshake. The reader declares the field and the expected unit or dimensionality.
    • TaggedQuantities (values plus a per-value code into a local unit table) for genuinely heterogeneous collections, under the same digest.

    Scope of this project:

    1. The codec API (names are this repository's decision): write, read with handshake, digest verification, block digests for growing data, an explicit reseal tool.
    2. TaggedQuantities.
    3. Canonical unit spelling owned by PyUnitWizard, stable across pint versions.
    4. A PyUnitWizard devguide page that turns Design record: serializing and exchanging quantities across MOLI and MolSysSuite #83 into durable documentation, so that the question "how do we serialize quantities?" is answered by this library's docs.
    5. The open questions of Design record: serializing and exchanging quantities across MOLI and MolSysSuite #83 §5, decided and recorded there.
    6. Property tests: exact round trips (scalars, n-D, dtypes, compound units), the slip catalogue of Design record: serializing and exchanging quantities across MOLI and MolSysSuite #83 (9 deliberate slips), and conformance under a non-default session policy.

    Related: #81 (array string round trip), uibcdf/moli#12, uibcdf/sabueso#32 (first consumer).

  4. dprada commented on Sep 24, 2026

    @dprada
    CollaboratorAuthor

    Documentation scope: the devguide page this project produces should include the alternatives review from #83 (MCO, ChEMBL's transcription-error flags, OpenFF, pint, ASDF, CF/UDUNITS, HDF5/Parquet/BagIt checksums, UCUM, QUDT, UO, D-SI), with its sources. It explains to future readers why this design exists and what it borrows. Consider also a user-facing docs page, 'Serializing quantities safely', built from the same material. Interoperability follow-ups: #85 (forms and unit dialects), #84 (cross-registry pint).

  5. changed the title [-]Quantity serialization codec: negotiated containers with an integrity digest[/-] [+]Native inert interchange form: verified quantity records (negotiated containers, integrity digest)[/+] on Sep 24, 2026
  6. dprada commented on Sep 24, 2026

    @dprada
    CollaboratorAuthor

    Reoriented (Diego, 2026-09-24). This project now implements PyUnitWizard's native inert interchange form (#83, 'Design direction'). The codec is the form:

    • write with convert(q, to_form=<native>);
    • read with convert(record, to_form=<backend>);
    • TaggedQuantities is the heterogeneous variant.

    The form does not compute. The descriptor carries the canonical name, UCUM, the SI factor/offset/exponents and an optional kind.

    Added to the scope:

    A live, computing form is explicitly out of scope; its re-evaluation is tracked in #87.

  7. dprada commented on Sep 24, 2026

    @dprada
    CollaboratorAuthor

    Draft qrec/0.2: format, prototype results and validated CF binding

    The prototype and its 60 tests exist and can be committed when implementation starts. Below is the draft specification they implement.


    Working name: qrec; the final name is still to be chosen. Implementation issue:
    #82. Design record: #83.

    Prototype: qrec.py (0.1), qrec2.py (0.2), minimal_reader.py (a reader for 0.1 records
    that uses only the Python standard library). Tests: 60 passing (test_qrec.py,
    test_qrec2.py).

    What it is

    An inert form, like PyUnitWizard's "string" form: it represents quantities and
    converts to and from every other form, but it does not compute. To compute, convert it to
    pint, openmm.unit, astropy.units or unyt.

    record = to_native(q, field="bioactivity.ic50", kind=QUDT_CONCENTRATION)  # write
    q      = from_native(record, field="bioactivity.ic50", to_form="openmm.unit",
                         unit="uM", kind=QUDT_CONCENTRATION)                   # read: handshake

    Unit descriptor

    key required meaning
    unit yes PyUnitWizard's canonical name ("nanomolar", "angstrom ** 2")
    si yes {"factor", "offset", "exponents"}, where value_SI = value × factor + offset. Exponents use the keys m, kg, s, A, K, mol, cd, with zero exponents omitted. The semantics match QUDT's conversion multiplier and dimension vector.
    ucum when known UCUM code for third parties ("nmol/L"; note that UCUM has no M for molar)
    kind optional quantity-kind IRI (QUDT or OBO UO). Distinguishes Hz/Bq, J/N·m and absolute temperature from a difference.
    cf_units HDF5, NetCDF and Zarr bindings UDUNITS string, written as the CF units attribute and covered by the digest

    Readers cross-check every spelling they understand: name ↔ si ↔ ucum ↔ cf_units. A
    disagreement is an error. This catches a renamed unit even when someone recomputes the
    digest; the prototype tests that case for both the SI factor and the UCUM code.

    Tolerance. The relative tolerance for SI factors is 1e-6. Measured: pint (CODATA 2022)
    and UDUNITS (u = 1.6605402e-27 kg, an older constant) differ by about 7e-7 for the
    atomic mass unit. That is inside the tolerance, but close, so the value needs a documented
    rationale rather than a guess.

    Integrity

    • Canonical manifest bytes: a tagged little-endian binary encoding, which is
      language-independent (Python writes 1e-06, JavaScript 0.000001).
    • Values: C-order, little-endian bytes, with NaN canonicalized.
    • Digest: a blake2b-128 digest per block, and a record digest over the manifest plus the
      block digests. Only the codec produces them. A write that bypasses the codec (raw
      h5py, a hand edit, a raw append) fails on read.
    • Threat model: mistakes, not adversaries.

    Bindings

    Binding Values Manifest Measured
    JSON text nested lists manifest key 100k float64: 2.4× raw size; read 40 ms, of which 34 ms is json.loads
    JSON base64 base64 of the little-endian bytes manifest key 100k: 1.33× raw; write 5 ms, read 4 ms
    HDF5 the dataset itself, native dtype, no base64 attributes qrec_manifest, qrec_block_digests, qrec_digest, plus the CF units attribute 48 MB float32: write 246 ms (raw h5py 150 ms), read + verify 101 ms (raw 14 ms), file size +0.01 MB
    Bundle many small scalars in one object one descriptor per entry, one digest for the bundle 50 scalars: 166 B/scalar (separate records: 347 B); write 5.8 ms, read 2.8 ms
    • HDF5 appends go through the codec: the incoming quantity is converted to the record's
      unit and sealed as a new block.
    • CF compatibility: every CF tool reads units as usual. A CF tool that rewrites units
      breaks the seal, and the reader refuses the record.
    • All 15 CF strings in the prototype's table were validated with cf-units 3.3.1
      (UDUNITS-2): each one parses and converts to the expected SI value. UDUNITS does not
      know nM, Da, dalton or M.

    Slips refused (tests)

    • JSON records:
      • a raw append;
      • a hand edit of a value;
      • a renamed unit;
      • a respelled unit;
      • an edited SI factor;
      • a deleted manifest;
      • reordered values;
      • truncated values;
      • pasted values of another dimension;
      • a removed digest;
      • a record copied into another field;
      • a UCUM code or SI factor changed and then resealed.
    • HDF5: values written with h5py, the CF units attribute edited, the manifest unit
      edited, the manifest deleted, the digest deleted, a raw resize-and-write.
    • Bundles: a scalar edited, a unit renamed, an entry removed, two entries swapped.
    • Handshake: a wrong dimension, and a wrong kind (Hz vs Bq, which share SI dimensions).

    Not detectable by any format: a writer that is wrong and consistent (it meant nM and
    wrote pM everywhere). That case is covered by ArgDigest contracts, cross-component canaries,
    conformance tests under a non-default policy, and domain heuristics (ChEMBL's
    3-or-6-orders-of-magnitude flag).

    Still open

    1. The name of the form.
    2. Bundles repeat the descriptor per entry. A local unit table (as in TaggedQuantities)
      would shrink them further.
    3. UCUM and CF spellings are tables in the prototype. The implementation must derive
      them with real parsers (UCUM, UDUNITS via cf-units, Interoperability forms and unit dialects: openff-units, ASDF, CF/UDUNITS, UCUM, QUDT #85), and fail when a unit has no
      spelling rather than invent one.
    4. Broadcast codes (TaggedQuantities), logarithmic kinds (pIC50, pKa), and verifying lazily
      on first access for very large data.
    5. Arrow/Parquet and Zarr bindings (field metadata and array attributes; same rules).
    6. Whether the form becomes the translation hub between forms (Design record: serializing and exchanging quantities across MOLI and MolSysSuite #83, "Design direction").
  8. changed the title [-]Native inert interchange form: verified quantity records (negotiated containers, integrity digest)[/-] [+]QuantityRecord: PyUnitWizard's native inert interchange form (verified quantity records)[/+] on Sep 24, 2026
  9. dprada commented on Sep 24, 2026

    @dprada
    CollaboratorAuthor

    Name decided (Diego, 2026-09-24): QuantityRecord.

    • Class: QuantityRecord.
    • Form name: "record", alongside "string", "pint", "openmm.unit"… Usage: puw.convert(q, to_form="record"), and back with puw.convert(rec, to_form="pint").
    • Format identifier: qrec/<version> (quantity record); the current draft is qrec/0.2.
    • The heterogeneous variant keeps the working name TaggedQuantities for now. Whether it becomes TaggedRecord, so that the two read as a family, is open.

    Names avoided: "passport", because standards/PYUNITWIZARD_GUIDE.md lists passports as an anti-pattern and this form is verified on every read, not remembered; and "canonical", which would be confused with standard units.

  10. dprada commented on Sep 24, 2026

    @dprada
    CollaboratorAuthor

    Decision (Diego, 2026-09-24): TaggedQuantities is not a separate class. It is a layout of QuantityRecord.

    Layout Unit Use
    homogeneous one descriptor the normal case: coordinates, normalized data, most card fields
    tagged a local table of descriptors plus one code per value collections whose units really differ: per-source bioactivity units such as nM, µM and %, or mixed-unit Arrow and pandas columns

    Why one class:

    Consequences:

    • A tagged record cannot become a single-unit backend quantity implicitly. Reading requires either an explicit target unit (a vectorized conversion) or an explicit split by unit.
    • The name question (TaggedQuantities vs TaggedRecord) disappears.

    Implementation order. Define the tagged layout in the specification now, so that the format never has to break for it. Implement it when a measured consumer case exists, following the value-certification lesson in #83. In Sabueso, source values keep their verbatim unit in asserted_value and normalized values can be homogeneous, so the need may be smaller than first assumed.

  11. dprada commented on Sep 24, 2026

    @dprada
    CollaboratorAuthor

    MVP implemented in PR #88 (provisional API):

    • the record form and QuantityRecordBundle;
    • the seal, the handshake and no default unit;
    • strict JSON and base64 encodings;
    • RecordError;
    • a standard-library reference reader and frozen test vectors;
    • user docs and a section of the canonical guide;
    • a devguide report, the old serialization draft marked superseded, and the 1.0 checklist updated.

    Evidence: 626 tests pass locally, and the full hosted matrix (Linux and macOS × 3.11–3.14) is green. Deferred items are listed in the PR and in devguide/pending_proposals/quantity_record.md.

  12. dprada commented on Sep 24, 2026

    @dprada
    CollaboratorAuthor

    Released in PyUnitWizard 0.27.0 (provisional API): https://github.com/uibcdf/pyunitwizard/releases/tag/0.27.0

    • Staged route: exact-commit gates, 8 clean staged installs, and promotion of the same file (SHA-256 2102b1bd…c3aaac), verified on the public channel.
    • A clean install from the public channel on Python 3.14 passes a QuantityRecord smoke test.
    • The five published vectors are reproduced byte for byte from the staged install.

    The issue stays open for the deferred scope: the HDF5/CF binding, the tagged layout, UCUM/CF via parsers, Arrow/Zarr, appended blocks, and promotion out of provisional status.

  13. dprada commented on Oct 5, 2026

    @dprada
    CollaboratorAuthor

    Delivered homogeneous QuantityRecord/Bundle MVP and design record resolved. API and qrec/0.3 formats remain provisional; extensions have independent owners #101–#106. Published receiving cases: Sabueso #32 and TopoMT sealed DFND inputs. MolSysMT #240 is legacy unit validation, not codec adoption.

    Guard: tests/test_quantity_record.py::test_every_change_outside_the_codec_is_refused. Record: completed_proposals/quantity_record.md.

    Runtime fc06259 qualified on Linux/macOS Python 3.11–3.14 (full matrix 37237982579, release gates 37237984869). Normally installed wheel outside checkout: 63 record cases pass, one explicit JSON-NaN skip. Archival head 426e2fd additionally passes local full suite (706 passed/12 skipped) and CI 37238940812 / policy 37238941450. Exact source/artifact identities and limits: receipt. No new public package or stable format promotion.

  14. LMMV commented on Oct 5, 2026

    @LMMV
    Contributor

    Central reconciliation is published at 04480983260804113022bc20c214875a5eb5611e. PyUnitWizard #82/#83/#100 are owner-closed for the delivered homogeneous QuantityRecord/Bundle MVP, design and write-reactivation repair. Native matrix 37237982579 at original runtime fc062598fb968988acb493cd2e6ef2f2537c3423 independently reports eight Linux/macOS Python 3.11–3.14 cells, 701 passed/14 skips each. Release gates 37237984869 have six successful jobs; closure source 426e2fd409d22adca757df163da8d67b01185499 has separate successful routine CI 37238940812. Local 706/12, installed development wheel 63/1 and three later consumer canaries retain their distinct scopes and mapped-dependency limitations; no new public package/fresh Conda solve is inferred.

    The pinned TopoMT DFND input SHA-256 agrees with original source and exactly matches the provider's coordinates/atom_radii/epsilon fixtures. The relevant owning canary explicitly reads under metre policy, converts to nm and rejects wrong fields; this is inspected compatibility evidence, not TopoMT algorithm qualification. Closed Sabueso #32 retains owner-reported negotiated-column/bundle evidence. MolSysMT #240 is corrected to owner-closed legacy H5MSM unit validation, not codec adoption; future provider HDF5 binding belongs to PyUnitWizard #101, with other independent extensions #102–#106. The central policy's obsolete pending-record hyperlink now targets the stable owning issue.

    QuantityRecord/Bundle and qrec/0.3 remain provisional. #46 stays partial for actual member adoption and explicit provider promotion. The principal maintainer confirms OpenFF integration is currently active: central review waits for its team's completed handoff. Historical OpenFF notices are not current adoption certification. No OpenFF repair, consumer code change, component scientific test/installation, guide rollout or artifact/publication mutation follows.

    Independent review receipt; maintained inventory; maintained adoption record. Generated queue index and offline governance pass. This documentation-only commit uses the accepted internal [skip ci] route; no new hosted test/coverage success is claimed. Central Codecov/GPG recovery remains separately MolSysSuite #69.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    needs-triageAwaiting maintainer triageproposalDesign or governance proposal

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions