Product Operations · Evaluation · AI-Native Delivery · Queens, NY · English & Spanish
I build the checks, not just the thing. Most of what is here is a working artifact with its own way of failing: a graded eval harness, a deterministic gate, a method written down clearly enough that someone else could run it.
rolefit — reads job postings the way a careful reader would. Residency exclusions and years floors live in the body, not the structured fields. Deterministic gates, a graded model judgment layer, and an eval harness that fails the build.
ai-landscape — a capture-to-publish pipeline where the summarising step is graded. An invented number is caught deterministically; a dropped hedge and an inverted finding go to a model, which is then graded against fixtures.
ai-coach-design — a design record: every touchpoint in a 90-day cohort classified human-essential or AI opportunity, including the five things the AI must not do, because a model can imitate all of them convincingly.
roomread — a workshop instrument. Put researched options in front of a room, make each person back a few and name the assumption that could kill one, then read it back live. METHOD.md is the point.
How I work day to day: aguedaschwartz.com/practice
